python / python/cpython

time/datetime.fromisoformat: C accelerator accepts a fractional seconds component with no decimal mark

Open
#155,175 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

extension-modules type-bug
Dominant language
Python
Stars
77.2k
Forks
35.9k
PR merge metrics
PR metrics pending

Description

Bug report

Bug description:

In the basic ISO 8601 time format, the C implementation of
time.fromisoformat() / datetime.fromisoformat() accepts digits following
HHMMSS and silently interprets them as a fractional seconds component, even
though no decimal mark (. or ,) is present. ISO 8601 requires a decimal
sign to introduce a fraction, and the pure-Python implementation correctly
rejects these strings.

>>> from datetime import time, datetime
>>> time.fromisoformat('12345678')
datetime.time(12, 34, 56, 780000)          # expected ValueError
>>> time.fromisoformat('123456789')
datetime.time(12, 34, 56, 789000)          # expected ValueError
>>> datetime.fromisoformat('2020-01-01T12345678')
datetime.datetime(2020, 1, 1, 12, 34, 56, 780000)   # expected ValueError

The pure-Python implementation rejects all of the above:

>>> import _pydatetime
>>> _pydatetime.time.fromisoformat('12345678')
Traceback (most recent call last):
  ...
ValueError: Invalid isoformat string: '12345678'

The same parser is used for the UTC offset, so a malformed offset is accepted
too, with the extra digits silently discarded:

>>> time.fromisoformat('12:34:56+00000000')
datetime.time(12, 34, 56, tzinfo=datetime.timezone.utc)   # expected ValueError

Two further details show the leniency is unintended rather than deliberate:

  • The equivalent extended-format string is correctly rejected —
    time.fromisoformat('12:34:5678') raises ValueError.
  • The behaviour is not even self-consistent across lengths. '1234567'
    (7 digits) raises ValueError, while '12345678' (8 digits) is accepted.

This matters because the failure is silent and produces a wrong value rather
than an error: a truncated or corrupted timestamp string parses successfully
into a time/datetime that differs from the input.

Cause

In parse_hh_mm_ss_ff() (Modules/_datetimemodule.c), the HH/MM/SS loop
breaks when it sees a decimal mark, but in the basic format it can also fall
out of the loop normally after SS with characters still unconsumed (the
!has_separator branch does --p). The code following the loop then parses
whatever remains as a fraction, without ever checking that a decimal mark
introduced it.

The pure-Python _parse_hh_mm_ss_ff() in Lib/_pydatetime.py performs that
check explicitly:

if pos < len_str:
    if tstr[pos] not in '.,':
        raise ValueError("Invalid microsecond separator")
Documentation

fromisoformat() is documented to accept "any valid ISO 8601 format" subject to
an explicit list of exceptions; a fraction without a decimal sign is not among
them.

This is the same class of C/pure-Python divergence as gh-152157 (empty fraction
before a timezone designator) and gh-152079.

Affected versions

Reproduced directly on 3.12.3 and on main (3.16.0a0). I did not have 3.13,
3.14 or 3.15 interpreters to hand, but comparing parse_hh_mm_ss_ff() across
the branches shows the cause is present in all of them: the
else if (!has_separator) { --p; } branch and the unguarded "Parse fractional
components" block that follows the loop are unchanged on every active branch.
The 3.14, 3.15 and main copies of the function are identical; the 3.12 and
3.13 copies differ only in other validation paths (decimal mark on
hour/minute, empty fraction, microsecond separator).

CPython versions tested on:

3.12, 3.16, CPython main branch

Operating systems tested on:

Linux

Linked PRs
  • gh-155177

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with parse_hh_mm_ss_ff() in Modules/_datetimemodule.c and compare its validation with _parse_hh_mm_ss_ff() in Lib/_pydatetime.py. Reproduce the basic-format and malformed-offset examples from the report, then add regression coverage so inputs without a decimal mark are rejected consistently by the C implementation.

Written by the indexing model from the issue text.

Assessment

Tech stack
c, python
Domain
backend
Issue type
Bug
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Quiet
Clarity
Clearly specified
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.