python / python/cpython

Wrong `SyntaxError.offset` for "Non-UTF-8 code starting with ..." when a non-ASCII character precedes the invalid byte

オープン
#157,378 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

interpreter-core topic-parser type-bug
主要言語
Python
スター
77.2k
フォーク
35.9k
PR マージ指標
PR 指標を取得中

説明

Bug report

Bug description:

SyntaxError.offset and SyntaxError.end_offset for the error
"Non-UTF-8 code starting with ..." are too small when a valid multi-byte
character precedes the invalid byte on the same line.

Reproduction

compile(b'\xc3\xa9X\x80', '<bug>', 'exec')

Same via a file:

$ printf '\xc3\xa9X\x80' > bug.py
$ ./python bug.py

The input is "éX" followed by the invalid byte 0x80, with no encoding
cookie. The invalid byte is the 3rd character of line 1.

Expected

  File "<bug>", line 1
    éX
      ^
SyntaxError: Non-UTF-8 code starting with '\x80' on line 1, but no encoding declared; ...

offset == 3, end_offset == 3 (1-based character column).

Actual

  File "<bug>", line 1
    éX
     ^
SyntaxError: Non-UTF-8 code starting with '\x80' on line 1, but no encoding declared; ...

offset == 2, end_offset == 2.

Observed values:

build offset end_offset
main (575fe3914f4) 2 2
3.14.7 3 3
3.15.0rc2+dev (e325fae3578) 3 3

More inputs on main, all one column short per preceding multi-byte character:

compile(b'\xc3\xa9abc\x80', ...)  # offset 4, expected 5
compile(b'\t\xc3\xa9X\x80', ...)  # offset 3, expected 4
compile(b'a\xc3\xa9b\x80c', ...)  # offset 3, expected 4

Cause

_PyTokenizer_ensure_utf8() (Parser/tokenizer/helpers.c) computes a 1-based
character column and passes it to _PyTokenizer_syntaxerror_known_range():

_PyTokenizer_syntaxerror_known_range(tok,
        col_offset + 1, col_offset + 1, ...)

Since 59a691361cb (gh-156894, PR #156901), _syntaxerror_range() converts its
col_offset/end_col_offset arguments from bytes to characters with
byte_col_to_char_col(), so the character column from ensure_utf8() is
converted a second time.

The same function is unchanged on current main (b1e7554ac1e).

Suggested fix

Pass byte columns from ensure_utf8(), e.g. (int)(badchar - line_start) + 1
for both arguments, and let _syntaxerror_range() do the byte-to-character
conversion.

Tests

Lib/test/test_source_encoding.py only checks the message text for this error;
no test checks offset/end_offset, so this case is not covered.

CPython versions tested on:

CPython main branch

Operating systems tested on:

Linux

Linked PRs
  • gh-157412
  • gh-157489

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

調査の方向性

Parser/tokenizer/helpers.c の _PyTokenizer_ensure_utf8() から始め、_PyTokenizer_syntaxerror_known_range() を確認し、その後、構文エラー処理コード内のバイトから文字への変換を読みます。無効なバイトの前にマルチバイト文字がある場合の offset と end_offset について、Lib/test/test_source_encoding.py にリグレッションテストのカバレッジを追加し、関連するソースエンコーディングテストを実行します。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
python
領域
compilers
issue の種類
バグ
難易度
2/5
見積もり時間
1〜3時間
活発さ
停滞
明瞭さ
明確に書かれている
初心者へのやさしさ
35/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。