Regression in tokenizer handling of `\r`
Open
@serhiy-storchaka is already working on this.
Since Dec 30, 2024.
topic-parser
type-bug
- Dominant language
- Python
- Stars
- 77.2k
- Forks
- 35.9k
- PR merge metrics
- PR metrics pending
Description
Bug report
Bug description:
Python 3.12 onwards we get a weird \r} token when trying to parse a file just containing '{\r}':
$ printf '{\r}' | python3.11 -m tokenize
1,0-1,1: OP '{'
1,1-1,2: ERRORTOKEN '\r'
1,2-1,3: OP '}'
1,3-1,4: NEWLINE ''
2,0-2,0: ENDMARKER ''
$ printf '{\r}' | python3.12 -m tokenize
1,0-1,1: OP '{'
1,1-1,3: OP '\r}'
1,3-1,4: NEWLINE ''
2,0-2,0: ENDMARKER ''
Weirdly, AST generation passes just fine in both cases:
$ printf '{\r}' | python3.11 -m ast
Module(
body=[
Expr(
value=Dict(keys=[], values=[]))],
type_ignores=[])
$ printf '{\r}' | python3.12 -m ast
Module(
body=[
Expr(
value=Dict(keys=[], values=[]))],
type_ignores=[])
Expected behaviour
I'd expect the \r to yield a NL instead, and we get a } OP as expected.
CPython versions tested on:
3.11, 3.12, 3.13, 3.14
Operating systems tested on:
macOS
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.