python / python/cpython

time/datetime.fromisoformat: C accelerator accepts a fractional seconds component with no decimal mark

オープン
#155,175 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

extension-modules type-bug
主要言語
Python
スター
77.2k
フォーク
35.9k
PR マージ指標
PR 指標を取得中

説明

Bug report

Bug description:

In the basic ISO 8601 time format, the C implementation of
time.fromisoformat() / datetime.fromisoformat() accepts digits following
HHMMSS and silently interprets them as a fractional seconds component, even
though no decimal mark (. or ,) is present. ISO 8601 requires a decimal
sign to introduce a fraction, and the pure-Python implementation correctly
rejects these strings.

>>> from datetime import time, datetime
>>> time.fromisoformat('12345678')
datetime.time(12, 34, 56, 780000)          # expected ValueError
>>> time.fromisoformat('123456789')
datetime.time(12, 34, 56, 789000)          # expected ValueError
>>> datetime.fromisoformat('2020-01-01T12345678')
datetime.datetime(2020, 1, 1, 12, 34, 56, 780000)   # expected ValueError

The pure-Python implementation rejects all of the above:

>>> import _pydatetime
>>> _pydatetime.time.fromisoformat('12345678')
Traceback (most recent call last):
  ...
ValueError: Invalid isoformat string: '12345678'

The same parser is used for the UTC offset, so a malformed offset is accepted
too, with the extra digits silently discarded:

>>> time.fromisoformat('12:34:56+00000000')
datetime.time(12, 34, 56, tzinfo=datetime.timezone.utc)   # expected ValueError

Two further details show the leniency is unintended rather than deliberate:

  • The equivalent extended-format string is correctly rejected —
    time.fromisoformat('12:34:5678') raises ValueError.
  • The behaviour is not even self-consistent across lengths. '1234567'
    (7 digits) raises ValueError, while '12345678' (8 digits) is accepted.

This matters because the failure is silent and produces a wrong value rather
than an error: a truncated or corrupted timestamp string parses successfully
into a time/datetime that differs from the input.

Cause

In parse_hh_mm_ss_ff() (Modules/_datetimemodule.c), the HH/MM/SS loop
breaks when it sees a decimal mark, but in the basic format it can also fall
out of the loop normally after SS with characters still unconsumed (the
!has_separator branch does --p). The code following the loop then parses
whatever remains as a fraction, without ever checking that a decimal mark
introduced it.

The pure-Python _parse_hh_mm_ss_ff() in Lib/_pydatetime.py performs that
check explicitly:

if pos < len_str:
    if tstr[pos] not in '.,':
        raise ValueError("Invalid microsecond separator")
Documentation

fromisoformat() is documented to accept "any valid ISO 8601 format" subject to
an explicit list of exceptions; a fraction without a decimal sign is not among
them.

This is the same class of C/pure-Python divergence as gh-152157 (empty fraction
before a timezone designator) and gh-152079.

Affected versions

Reproduced directly on 3.12.3 and on main (3.16.0a0). I did not have 3.13,
3.14 or 3.15 interpreters to hand, but comparing parse_hh_mm_ss_ff() across
the branches shows the cause is present in all of them: the
else if (!has_separator) { --p; } branch and the unguarded "Parse fractional
components" block that follows the loop are unchanged on every active branch.
The 3.14, 3.15 and main copies of the function are identical; the 3.12 and
3.13 copies differ only in other validation paths (decimal mark on
hour/minute, empty fraction, microsecond separator).

CPython versions tested on:

3.12, 3.16, CPython main branch

Operating systems tested on:

Linux

Linked PRs
  • gh-155177

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

調査の方向性

Modules/_datetimemodule.c の parse_hh_mm_ss_ff() から始め、Modules/_datetimemodule.c の検証を Lib/_pydatetime.py の _parse_hh_mm_ss_ff() と比較してください。報告にある基本形式と不正なオフセットの例を再現し、その後、十進区切り記号のない入力が C 実装で一貫して拒否されるよう、リグレッションテストを追加してください。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
c, python
領域
backend
issue の種類
バグ
難易度
2/5
見積もり時間
1〜3時間
活発さ
静か
明瞭さ
明確に書かれている
初心者へのやさしさ
25/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。