python / python/cpython

Wrong `SyntaxError.offset` for "Non-UTF-8 code starting with ..." when a non-ASCII character precedes the invalid byte

Aberta
#157,378 0 comentários 0 reações 0 responsáveis Ver no GitHub

Ninguém assumiu esta issue ainda.

interpreter-core topic-parser type-bug
Linguagem predominante
Python
Estrelas
77.2k
Forks
35.9k
Métricas de merge de PRs
Métricas de PR pendentes

Descrição

Bug report

Bug description:

SyntaxError.offset and SyntaxError.end_offset for the error
"Non-UTF-8 code starting with ..." are too small when a valid multi-byte
character precedes the invalid byte on the same line.

Reproduction

compile(b'\xc3\xa9X\x80', '<bug>', 'exec')

Same via a file:

$ printf '\xc3\xa9X\x80' > bug.py
$ ./python bug.py

The input is "éX" followed by the invalid byte 0x80, with no encoding
cookie. The invalid byte is the 3rd character of line 1.

Expected

  File "<bug>", line 1
    éX
      ^
SyntaxError: Non-UTF-8 code starting with '\x80' on line 1, but no encoding declared; ...

offset == 3, end_offset == 3 (1-based character column).

Actual

  File "<bug>", line 1
    éX
     ^
SyntaxError: Non-UTF-8 code starting with '\x80' on line 1, but no encoding declared; ...

offset == 2, end_offset == 2.

Observed values:

build offset end_offset
main (575fe3914f4) 2 2
3.14.7 3 3
3.15.0rc2+dev (e325fae3578) 3 3

More inputs on main, all one column short per preceding multi-byte character:

compile(b'\xc3\xa9abc\x80', ...)  # offset 4, expected 5
compile(b'\t\xc3\xa9X\x80', ...)  # offset 3, expected 4
compile(b'a\xc3\xa9b\x80c', ...)  # offset 3, expected 4

Cause

_PyTokenizer_ensure_utf8() (Parser/tokenizer/helpers.c) computes a 1-based
character column and passes it to _PyTokenizer_syntaxerror_known_range():

_PyTokenizer_syntaxerror_known_range(tok,
        col_offset + 1, col_offset + 1, ...)

Since 59a691361cb (gh-156894, PR #156901), _syntaxerror_range() converts its
col_offset/end_col_offset arguments from bytes to characters with
byte_col_to_char_col(), so the character column from ensure_utf8() is
converted a second time.

The same function is unchanged on current main (b1e7554ac1e).

Suggested fix

Pass byte columns from ensure_utf8(), e.g. (int)(badchar - line_start) + 1
for both arguments, and let _syntaxerror_range() do the byte-to-character
conversion.

Tests

Lib/test/test_source_encoding.py only checks the message text for this error;
no test checks offset/end_offset, so this case is not covered.

CPython versions tested on:

CPython main branch

Operating systems tested on:

Linux

Linked PRs
  • gh-157412
  • gh-157489

Guia de contribuição

Abrir o guia de contribuição

Primeiros passos

  1. Leia a issue inteira e depois o guia de contribuição do projeto.
  2. Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
  3. Faça um fork do repositório e trabalhe em uma branch.
  4. Abra um pull request que referencie o número da issue.

Direção de pesquisa

Comece em Parser/tokenizer/helpers.c, em _PyTokenizer_ensure_utf8(), e revise _PyTokenizer_syntaxerror_known_range(); depois, leia a conversão de bytes para caracteres no código de tratamento de erros de sintaxe. Adicione cobertura de regressão em Lib/test/test_source_encoding.py para offset e end_offset com um caractere multibyte antes do byte inválido e execute os testes relevantes de codificação do código-fonte.

Escrita pelo modelo de indexação a partir do texto da issue.

Avaliação

Stack de tecnologia
python
Domínio
compilers
Tipo de issue
Bug
Dificuldade
2/5
Tempo estimado
1-3 horas
Status de atividade
Estagnada
Clareza
Claramente especificada
Facilidade para iniciantes
35/100

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.