python / python/cpython

time/datetime.fromisoformat: C accelerator accepts a fractional seconds component with no decimal mark

Ouverte
#155,175 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub

Personne n'a encore pris cette issue.

extension-modules type-bug
Langage dominant
Python
Étoiles
77.2k
Forks
35.9k
Métriques de merge des PR
Métriques de PR en attente

Description

Bug report

Bug description:

In the basic ISO 8601 time format, the C implementation of
time.fromisoformat() / datetime.fromisoformat() accepts digits following
HHMMSS and silently interprets them as a fractional seconds component, even
though no decimal mark (. or ,) is present. ISO 8601 requires a decimal
sign to introduce a fraction, and the pure-Python implementation correctly
rejects these strings.

>>> from datetime import time, datetime
>>> time.fromisoformat('12345678')
datetime.time(12, 34, 56, 780000)          # expected ValueError
>>> time.fromisoformat('123456789')
datetime.time(12, 34, 56, 789000)          # expected ValueError
>>> datetime.fromisoformat('2020-01-01T12345678')
datetime.datetime(2020, 1, 1, 12, 34, 56, 780000)   # expected ValueError

The pure-Python implementation rejects all of the above:

>>> import _pydatetime
>>> _pydatetime.time.fromisoformat('12345678')
Traceback (most recent call last):
  ...
ValueError: Invalid isoformat string: '12345678'

The same parser is used for the UTC offset, so a malformed offset is accepted
too, with the extra digits silently discarded:

>>> time.fromisoformat('12:34:56+00000000')
datetime.time(12, 34, 56, tzinfo=datetime.timezone.utc)   # expected ValueError

Two further details show the leniency is unintended rather than deliberate:

  • The equivalent extended-format string is correctly rejected —
    time.fromisoformat('12:34:5678') raises ValueError.
  • The behaviour is not even self-consistent across lengths. '1234567'
    (7 digits) raises ValueError, while '12345678' (8 digits) is accepted.

This matters because the failure is silent and produces a wrong value rather
than an error: a truncated or corrupted timestamp string parses successfully
into a time/datetime that differs from the input.

Cause

In parse_hh_mm_ss_ff() (Modules/_datetimemodule.c), the HH/MM/SS loop
breaks when it sees a decimal mark, but in the basic format it can also fall
out of the loop normally after SS with characters still unconsumed (the
!has_separator branch does --p). The code following the loop then parses
whatever remains as a fraction, without ever checking that a decimal mark
introduced it.

The pure-Python _parse_hh_mm_ss_ff() in Lib/_pydatetime.py performs that
check explicitly:

if pos < len_str:
    if tstr[pos] not in '.,':
        raise ValueError("Invalid microsecond separator")
Documentation

fromisoformat() is documented to accept "any valid ISO 8601 format" subject to
an explicit list of exceptions; a fraction without a decimal sign is not among
them.

This is the same class of C/pure-Python divergence as gh-152157 (empty fraction
before a timezone designator) and gh-152079.

Affected versions

Reproduced directly on 3.12.3 and on main (3.16.0a0). I did not have 3.13,
3.14 or 3.15 interpreters to hand, but comparing parse_hh_mm_ss_ff() across
the branches shows the cause is present in all of them: the
else if (!has_separator) { --p; } branch and the unguarded "Parse fractional
components" block that follows the loop are unchanged on every active branch.
The 3.14, 3.15 and main copies of the function are identical; the 3.12 and
3.13 copies differ only in other validation paths (decimal mark on
hour/minute, empty fraction, microsecond separator).

CPython versions tested on:

3.12, 3.16, CPython main branch

Operating systems tested on:

Linux

Linked PRs
  • gh-155177

Guide de contribution

Ouvrir le guide de contribution

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Piste de recherche

Commencez par parse_hh_mm_ss_ff() dans Modules/_datetimemodule.c et comparez sa validation avec celle de _parse_hh_mm_ss_ff() dans Lib/_pydatetime.py. Reproduisez les exemples de format de base et d’offset mal formé du rapport, puis ajoutez une couverture de régression afin que les entrées sans séparateur décimal soient systématiquement rejetées par l’implémentation C.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
c, python
Domaine
backend
Type d'issue
Bug
Difficulté
2/5
Temps estimé
1-3 heures
Activité
Calme
Clarté
Clairement spécifiée
Accessibilité débutants
25/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.