tokenize: `e.offset` on `IndentationError` is not set correctly
還沒有人認領這個 Issue。
- 主要語言
- Python
- 星號
- 77.2k
- 分支
- 36k
- PR 合併指標
- PR 指標待擷取
描述
Bug report
Bug description:
import io, tokenize
input_s = """
def foo():
x = 1
x = "some_very_long_expression"
"""
try:
list(tokenize.generate_tokens(io.StringIO(input_s).readline))
except IndentationError as e:
print(f"{e.offset = }")
The line x = "some_very_long_expression" is indented with 3 spaces instead of 4, so we get an error.
We would expect the offset to be around 3, since that is where the issue is.
On python 3.12+, the offset is instead 35, equal to the length of the line + 1.
This appears to be caused by the refactor of tokenize done in 3.12 for PEP 701
Potentially related to https://github.com/python/cpython/issues/84515
Funnily enough, the exact same thing seems to have happened around 3.8 https://github.com/python/cpython/issues/75676
Tested versions:
| Version | e.offset |
|---|---|
| 3.11.15 | 3 |
| 3.12.13 | 35 |
| 3.13.13 | 35 |
| 3.14.4 | 35 |
| 3.15.0b1 | 35 |
Testing command:
nix-shell -p python311 --run "python --version && python tmp.py"
(and replace 311 by 312, 313, etc)
CPython versions tested on:
3.15
Operating systems tested on:
Linux
Linked PRs
- gh-154281
貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
研究方向
從 tokenize.generate_tokens 開始,在列出的 Python 版本中重現所提供的 IndentationError 範例。追蹤 tokenizer 如何建構 IndentationError,並將 offset 與 Python 3.11 的行為進行比較。當縮排錯誤的行回報的 offset 接近 3,且該範例具備回歸測試覆蓋時,即表示完成。
由索引模型根據 Issue 內容生成。
評估
- 技術堆疊
- python
- 領域
- compilers
- Issue 類型
- 缺陷
- 難度
- 3/5
- 預估耗時
- 1-2 天
- 活躍度
- 停滯
- 描述清晰度
- 基本清楚
- 新手友好度
- 35/100