Wrong `SyntaxError.offset` for "Non-UTF-8 code starting with ..." when a non-ASCII character precedes the invalid byte
Nessuno ha ancora preso questa issue.
- Lingua principale
- Python
- Stelle
- 77.2k
- Fork
- 35.9k
- Metriche di merge delle PR
- Metriche PR in attesa
Descrizione
Bug report
Bug description:
SyntaxError.offset and SyntaxError.end_offset for the error
"Non-UTF-8 code starting with ..." are too small when a valid multi-byte
character precedes the invalid byte on the same line.
Reproduction
compile(b'\xc3\xa9X\x80', '<bug>', 'exec')
Same via a file:
$ printf '\xc3\xa9X\x80' > bug.py
$ ./python bug.py
The input is "éX" followed by the invalid byte 0x80, with no encoding
cookie. The invalid byte is the 3rd character of line 1.
Expected
File "<bug>", line 1
éX
^
SyntaxError: Non-UTF-8 code starting with '\x80' on line 1, but no encoding declared; ...
offset == 3, end_offset == 3 (1-based character column).
Actual
File "<bug>", line 1
éX
^
SyntaxError: Non-UTF-8 code starting with '\x80' on line 1, but no encoding declared; ...
offset == 2, end_offset == 2.
Observed values:
| build | offset | end_offset |
|---|---|---|
| main (575fe3914f4) | 2 | 2 |
| 3.14.7 | 3 | 3 |
| 3.15.0rc2+dev (e325fae3578) | 3 | 3 |
More inputs on main, all one column short per preceding multi-byte character:
compile(b'\xc3\xa9abc\x80', ...) # offset 4, expected 5
compile(b'\t\xc3\xa9X\x80', ...) # offset 3, expected 4
compile(b'a\xc3\xa9b\x80c', ...) # offset 3, expected 4
Cause
_PyTokenizer_ensure_utf8() (Parser/tokenizer/helpers.c) computes a 1-based
character column and passes it to _PyTokenizer_syntaxerror_known_range():
_PyTokenizer_syntaxerror_known_range(tok,
col_offset + 1, col_offset + 1, ...)
Since 59a691361cb (gh-156894, PR #156901), _syntaxerror_range() converts its
col_offset/end_col_offset arguments from bytes to characters with
byte_col_to_char_col(), so the character column from ensure_utf8() is
converted a second time.
The same function is unchanged on current main (b1e7554ac1e).
Suggested fix
Pass byte columns from ensure_utf8(), e.g. (int)(badchar - line_start) + 1
for both arguments, and let _syntaxerror_range() do the byte-to-character
conversion.
Tests
Lib/test/test_source_encoding.py only checks the message text for this error;
no test checks offset/end_offset, so this case is not covered.
CPython versions tested on:
CPython main branch
Operating systems tested on:
Linux
Linked PRs
- gh-157412
- gh-157489
Guida per i contributori
Apri la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Direzione di ricerca
Inizia in Parser/tokenizer/helpers.c da _PyTokenizer_ensure_utf8() e analizza _PyTokenizer_syntaxerror_known_range(), quindi leggi la conversione da byte a carattere nel codice di gestione degli errori di sintassi. Aggiungi la copertura di regressione in Lib/test/test_source_encoding.py per offset e end_offset con un carattere multibyte prima del byte non valido, quindi esegui i test pertinenti di codifica del codice sorgente.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- python
- Ambito
- compilers
- Tipo di issue
- Bug
- Difficoltà
- 2/5
- Tempo stimato
- 1-3 ore
- Stato di attività
- Ferma
- Chiarezza
- Specificata chiaramente
- Idoneità per principianti
- 35/100