python / python/cpython

Wrong `SyntaxError.offset` for "Non-UTF-8 code starting with ..." when a non-ASCII character precedes the invalid byte

Ouverte
#157,378 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub

Personne n'a encore pris cette issue.

interpreter-core topic-parser type-bug
Langage dominant
Python
Étoiles
77.2k
Forks
35.9k
Métriques de merge des PR
Métriques de PR en attente

Description

Bug report

Bug description:

SyntaxError.offset and SyntaxError.end_offset for the error
"Non-UTF-8 code starting with ..." are too small when a valid multi-byte
character precedes the invalid byte on the same line.

Reproduction

compile(b'\xc3\xa9X\x80', '<bug>', 'exec')

Same via a file:

$ printf '\xc3\xa9X\x80' > bug.py
$ ./python bug.py

The input is "éX" followed by the invalid byte 0x80, with no encoding
cookie. The invalid byte is the 3rd character of line 1.

Expected

  File "<bug>", line 1
    éX
      ^
SyntaxError: Non-UTF-8 code starting with '\x80' on line 1, but no encoding declared; ...

offset == 3, end_offset == 3 (1-based character column).

Actual

  File "<bug>", line 1
    éX
     ^
SyntaxError: Non-UTF-8 code starting with '\x80' on line 1, but no encoding declared; ...

offset == 2, end_offset == 2.

Observed values:

build offset end_offset
main (575fe3914f4) 2 2
3.14.7 3 3
3.15.0rc2+dev (e325fae3578) 3 3

More inputs on main, all one column short per preceding multi-byte character:

compile(b'\xc3\xa9abc\x80', ...)  # offset 4, expected 5
compile(b'\t\xc3\xa9X\x80', ...)  # offset 3, expected 4
compile(b'a\xc3\xa9b\x80c', ...)  # offset 3, expected 4

Cause

_PyTokenizer_ensure_utf8() (Parser/tokenizer/helpers.c) computes a 1-based
character column and passes it to _PyTokenizer_syntaxerror_known_range():

_PyTokenizer_syntaxerror_known_range(tok,
        col_offset + 1, col_offset + 1, ...)

Since 59a691361cb (gh-156894, PR #156901), _syntaxerror_range() converts its
col_offset/end_col_offset arguments from bytes to characters with
byte_col_to_char_col(), so the character column from ensure_utf8() is
converted a second time.

The same function is unchanged on current main (b1e7554ac1e).

Suggested fix

Pass byte columns from ensure_utf8(), e.g. (int)(badchar - line_start) + 1
for both arguments, and let _syntaxerror_range() do the byte-to-character
conversion.

Tests

Lib/test/test_source_encoding.py only checks the message text for this error;
no test checks offset/end_offset, so this case is not covered.

CPython versions tested on:

CPython main branch

Operating systems tested on:

Linux

Linked PRs
  • gh-157412
  • gh-157489

Guide de contribution

Ouvrir le guide de contribution

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Piste de recherche

Commencez dans Parser/tokenizer/helpers.c, au niveau de _PyTokenizer_ensure_utf8(), et examinez _PyTokenizer_syntaxerror_known_range(), puis lisez la conversion des octets en caractères dans le code de gestion des erreurs de syntaxe. Ajoutez une couverture de régression dans Lib/test/test_source_encoding.py pour offset et end_offset lorsqu’un caractère multioctet précède l’octet invalide, puis exécutez les tests pertinents d’encodage du code source.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
python
Domaine
compilers
Type d'issue
Bug
Difficulté
2/5
Temps estimé
1-3 heures
Activité
À l'abandon
Clarté
Clairement spécifiée
Accessibilité débutants
35/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.