Wrong `SyntaxError.offset` for "Non-UTF-8 code starting with ..." when a non-ASCII character precedes the invalid byte
Dieses Issue hat noch niemand übernommen.
- Vorherrschende Sprache
- Python
- Sterne
- 77.2k
- Forks
- 35.9k
- PR-Merge-Kennzahlen
- PR-Kennzahlen ausstehend
Beschreibung
Bug report
Bug description:
SyntaxError.offset and SyntaxError.end_offset for the error
"Non-UTF-8 code starting with ..." are too small when a valid multi-byte
character precedes the invalid byte on the same line.
Reproduction
compile(b'\xc3\xa9X\x80', '<bug>', 'exec')
Same via a file:
$ printf '\xc3\xa9X\x80' > bug.py
$ ./python bug.py
The input is "éX" followed by the invalid byte 0x80, with no encoding
cookie. The invalid byte is the 3rd character of line 1.
Expected
File "<bug>", line 1
éX
^
SyntaxError: Non-UTF-8 code starting with '\x80' on line 1, but no encoding declared; ...
offset == 3, end_offset == 3 (1-based character column).
Actual
File "<bug>", line 1
éX
^
SyntaxError: Non-UTF-8 code starting with '\x80' on line 1, but no encoding declared; ...
offset == 2, end_offset == 2.
Observed values:
| build | offset | end_offset |
|---|---|---|
| main (575fe3914f4) | 2 | 2 |
| 3.14.7 | 3 | 3 |
| 3.15.0rc2+dev (e325fae3578) | 3 | 3 |
More inputs on main, all one column short per preceding multi-byte character:
compile(b'\xc3\xa9abc\x80', ...) # offset 4, expected 5
compile(b'\t\xc3\xa9X\x80', ...) # offset 3, expected 4
compile(b'a\xc3\xa9b\x80c', ...) # offset 3, expected 4
Cause
_PyTokenizer_ensure_utf8() (Parser/tokenizer/helpers.c) computes a 1-based
character column and passes it to _PyTokenizer_syntaxerror_known_range():
_PyTokenizer_syntaxerror_known_range(tok,
col_offset + 1, col_offset + 1, ...)
Since 59a691361cb (gh-156894, PR #156901), _syntaxerror_range() converts its
col_offset/end_col_offset arguments from bytes to characters with
byte_col_to_char_col(), so the character column from ensure_utf8() is
converted a second time.
The same function is unchanged on current main (b1e7554ac1e).
Suggested fix
Pass byte columns from ensure_utf8(), e.g. (int)(badchar - line_start) + 1
for both arguments, and let _syntaxerror_range() do the byte-to-character
conversion.
Tests
Lib/test/test_source_encoding.py only checks the message text for this error;
no test checks offset/end_offset, so this case is not covered.
CPython versions tested on:
CPython main branch
Operating systems tested on:
Linux
Linked PRs
- gh-157412
- gh-157489
Beitragsleitfaden
Erste Schritte
- Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
- Forke das Repository und arbeite in einem Branch.
- Öffne einen Pull Request, der die Issue-Nummer nennt.
Rechercherichtung
Beginne in Parser/tokenizer/helpers.c bei _PyTokenizer_ensure_utf8() und überprüfe _PyTokenizer_syntaxerror_known_range(); lies anschließend die Byte-zu-Zeichen-Konvertierung im Code zur Behandlung von Syntaxfehlern. Füge in Lib/test/test_source_encoding.py eine Regressionstestabdeckung für offset und end_offset mit einem Mehrbytezeichen vor dem ungültigen Byte hinzu und führe die relevanten Tests zur Quellkodierung aus.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- python
- Bereich
- compilers
- Issue-Typ
- Bug
- Schwierigkeit
- 2/5
- Geschätzter Aufwand
- 1-3 Stunden
- Aktivitätsstatus
- Veraltet
- Klarheit
- Klar beschrieben
- Anfängerfreundlichkeit
- 35/100