t-string tokeniser reports `MemoryError` on invalid input
Chưa có ai nhận issue này.
- Ngôn ngữ chính
- Python
- Star
- 77.2k
- Fork
- 35.9k
- Chỉ số merge pull request
- Chỉ số pull request đang chờ
Mô tả
Bug report
In short
I found a bug in the t-string tokeniser by fuzzing it, specifically in set_ftstring_expr (from Parser/lexer/lexer.c).
The following is a reproducer which consistently reports a MemoryError, which isn't the expected error (I would expect TokenError).
import tokenize
import io
list(tokenize.tokenize(io.BytesIO(b't"{!\n!x').readline))
results in the following:
Traceback (most recent call last):
File "<python-input-3>", line 1, in <module>
list(tokenize.tokenize(io.BytesIO(b't"{!\n!x').readline))
~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/bits/python315/Lib/tokenize.py", line 499, in tokenize
yield from _generate_tokens_from_c_tokenizer(rl_gen.__next__, encoding, extra_tokens=True)
File "/home/bits/python315/Lib/tokenize.py", line 634, in _generate_tokens_from_c_tokenizer
for info in it:
^^
MemoryError
Longer version
set_ftstring_expr in Parser/lexer/lexer.c, extracts the source text of the expression inside {...} so it can be attached as metadata to the token.
The function tracks two fields on the tokenizer mode struct:
last_expr_sizeis set when{is seen, with valuestrlen(tok->cur)(bytes fromtok->curto end of buffer)last_expr_endis set when!,}, or:is seen with valuestrlen(tok->start)(bytes from the delimiter to end of buffer)
The intended expression length is last_expr_size - last_expr_end (number of characters between the two positions).
However, this does not work when there are two ! across two lines, for example:
t"{expr!conv1
n!conv2
The sequence of events (LLM-analysed):
-
{on line 1:_PyLexer_update_ftstring_expr(tok, '{')setslast_expr_size = strlen(tok->cur)(bytes remaining on line 1 after{) andlast_expr_end = -1. -
First
!on line 1:_PyLexer_update_ftstring_expr(tok, '!')setslast_expr_end = strlen(tok->start)(a smaller value from the same line 1 buffer).set_ftstring_exprruns, computeslast_expr_size - last_expr_end > 0, stores result intoken->metadata. Crucially,last_expr_endis now ≥ 0. -
Newline:
_PyLexer_update_ftstring_expr(tok, 0)would normally append the next line's content and growlast_expr_size, keeping the measurements in sync. But thecase 0branch has a guard: it skips the append whenlast_expr_end >= 0. Because the first!already setlast_expr_end, the append is skipped andlast_expr_sizeis locked at its small line-1 value. -
Second
!on line 2:_PyLexer_update_ftstring_expr(tok, '!')setslast_expr_end = strlen(tok->start)measured in the new line 2 buffer. If line 2 has more content after!than line 1 had after{, this newlast_expr_end > last_expr_size. A newtokenstruct is active (the previous one was emitted), sotoken->metadata == NULLandset_ftstring_exprruns the full computation -- producing a negativePy_ssize_t.
That negative value is then used in three places:
PyMem_Malloc((last_expr_size - last_expr_end + 1) * sizeof(char))cast tosize_t,-N+1becomes huge.PyUnicode_DecodeUTF8(buf, last_expr_size - last_expr_end, NULL)The length argument isPy_ssize_tbutunicodeobject.cimmediately checksif (size > PY_SSIZE_T_MAX)after casting -- a negative value cast tosize_tis huge and trips the overflow guard, raisingPyErr_NoMemory.- Loop bounds (
for i < ...; while i < ...) -- a negative bound means the loops never execute, so the comment-stripping pass is silently skipped for inputs that reach thehash_detectedbranch.
Proposed fix
Compute expr_len once at the top of set_ftstring_expr and return -1 immediately if it is negative. Returning -1 signals a tokenizer error, which surfaces to Python callers as TokenError -- the correct outcome for malformed source. All five downstream uses of the subtraction are replaced with expr_len.
The fix (with a regression test) is already implemented on my branch.
CPython versions tested on:
CPython main branch
Operating systems tested on:
Linux
Linked PRs
- gh-149445
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Hướng nghiên cứu
Bắt đầu trong Parser/lexer/lexer.c tại set_ftstring_expr và theo dõi cách kết quả của nó đi đến đường dẫn tokenize.py. Sử dụng trình tái hiện tokenize được cung cấp và xem lại bài kiểm thử hồi quy hiện có trên phần triển khai được liên kết; hoàn thành khi đầu vào không hợp lệ gây ra TokenError thay vì MemoryError.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Đánh giá
- Công nghệ
- c, python
- Lĩnh vực
- backend, compilers
- Loại issue
- Lỗi
- Độ khó
- 3/5
- Thời gian dự kiến
- 1-2 ngày
- Mức độ hoạt động
- Đình trệ
- Độ rõ ràng
- Đặc tả rõ ràng
- Mức phù hợp với người mới
- 25/100