Wrong `SyntaxError.offset` for "Non-UTF-8 code starting with ..." when a non-ASCII character precedes the invalid byte
还没有人认领这个 Issue。
- 主要语言
- Python
- 星标
- 77.2k
- 派生
- 35.9k
- PR 合并指标
- PR 指标待抓取
描述
Bug report
Bug description:
SyntaxError.offset and SyntaxError.end_offset for the error
"Non-UTF-8 code starting with ..." are too small when a valid multi-byte
character precedes the invalid byte on the same line.
Reproduction
compile(b'\xc3\xa9X\x80', '<bug>', 'exec')
Same via a file:
$ printf '\xc3\xa9X\x80' > bug.py
$ ./python bug.py
The input is "éX" followed by the invalid byte 0x80, with no encoding
cookie. The invalid byte is the 3rd character of line 1.
Expected
File "<bug>", line 1
éX
^
SyntaxError: Non-UTF-8 code starting with '\x80' on line 1, but no encoding declared; ...
offset == 3, end_offset == 3 (1-based character column).
Actual
File "<bug>", line 1
éX
^
SyntaxError: Non-UTF-8 code starting with '\x80' on line 1, but no encoding declared; ...
offset == 2, end_offset == 2.
Observed values:
| build | offset | end_offset |
|---|---|---|
| main (575fe3914f4) | 2 | 2 |
| 3.14.7 | 3 | 3 |
| 3.15.0rc2+dev (e325fae3578) | 3 | 3 |
More inputs on main, all one column short per preceding multi-byte character:
compile(b'\xc3\xa9abc\x80', ...) # offset 4, expected 5
compile(b'\t\xc3\xa9X\x80', ...) # offset 3, expected 4
compile(b'a\xc3\xa9b\x80c', ...) # offset 3, expected 4
Cause
_PyTokenizer_ensure_utf8() (Parser/tokenizer/helpers.c) computes a 1-based
character column and passes it to _PyTokenizer_syntaxerror_known_range():
_PyTokenizer_syntaxerror_known_range(tok,
col_offset + 1, col_offset + 1, ...)
Since 59a691361cb (gh-156894, PR #156901), _syntaxerror_range() converts its
col_offset/end_col_offset arguments from bytes to characters with
byte_col_to_char_col(), so the character column from ensure_utf8() is
converted a second time.
The same function is unchanged on current main (b1e7554ac1e).
Suggested fix
Pass byte columns from ensure_utf8(), e.g. (int)(badchar - line_start) + 1
for both arguments, and let _syntaxerror_range() do the byte-to-character
conversion.
Tests
Lib/test/test_source_encoding.py only checks the message text for this error;
no test checks offset/end_offset, so this case is not covered.
CPython versions tested on:
CPython main branch
Operating systems tested on:
Linux
Linked PRs
- gh-157412
- gh-157489
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
调研方向
从 Parser/tokenizer/helpers.c 中的 _PyTokenizer_ensure_utf8() 开始,检查 _PyTokenizer_syntaxerror_known_range(),然后阅读语法错误处理代码中的字节到字符转换。在 Lib/test/test_source_encoding.py 中,为无效字节前存在多字节字符时的 offset 和 end_offset 添加回归测试覆盖,并运行相关的源代码编码测试。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- python
- 领域
- compilers
- Issue 类型
- 缺陷
- 难度
- 2/5
- 预计耗时
- 1-3 小时
- 活跃度
- 停滞
- 描述清晰度
- 描述清楚
- 新手友好度
- 35/100