Regression in tokenizer handling of `\r`
未关闭
@serhiy-storchaka 已经在做这个了。
开始于 2024年12月30日。
topic-parser
type-bug
- 主要语言
- Python
- 星标
- 77.2k
- 派生
- 35.9k
- PR 合并指标
- PR 指标待抓取
描述
Bug report
Bug description:
Python 3.12 onwards we get a weird \r} token when trying to parse a file just containing '{\r}':
$ printf '{\r}' | python3.11 -m tokenize
1,0-1,1: OP '{'
1,1-1,2: ERRORTOKEN '\r'
1,2-1,3: OP '}'
1,3-1,4: NEWLINE ''
2,0-2,0: ENDMARKER ''
$ printf '{\r}' | python3.12 -m tokenize
1,0-1,1: OP '{'
1,1-1,3: OP '\r}'
1,3-1,4: NEWLINE ''
2,0-2,0: ENDMARKER ''
Weirdly, AST generation passes just fine in both cases:
$ printf '{\r}' | python3.11 -m ast
Module(
body=[
Expr(
value=Dict(keys=[], values=[]))],
type_ignores=[])
$ printf '{\r}' | python3.12 -m ast
Module(
body=[
Expr(
value=Dict(keys=[], values=[]))],
type_ignores=[])
Expected behaviour
I'd expect the \r to yield a NL instead, and we get a } OP as expected.
CPython versions tested on:
3.11, 3.12, 3.13, 3.14
Operating systems tested on:
macOS
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
评估
这个 Issue 还没有评估数据。