python / python/cpython

Keyword typo suggestions differ between file and string tokenizers

オープン
#156,047 コメント 4 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

3.14 3.15 3.16 interpreter-core topic-parser type-bug
主要言語
Python
スター
77.2k
フォーク
35.9k
PR マージ指標
PR 指標を取得中

説明

Bug report

Bug description:

Keyword-typo SyntaxError suggestions depend on whether CPython processes the same source using the file tokenizer or the string tokenizer.

Create t.py containing:

a=b=c=d=e=f=g=h=i=0
def fn():
    retrun True

Running the file directly:

$ python t.py

uses the file-tokenizer path and produces the helpful diagnostic:

  File ".../t.py", line 3
    retrun True
    ^^^^^^
SyntaxError: invalid syntax. Did you mean 'return'?

Running the same file as a module:

$ python -m t

uses the string-tokenizer path and produces:

  File ".../t.py", line 3
    retrun True
           ^^^^
SyntaxError: invalid syntax

The same discrepancy is visible in the default PyREPL-based REPL.

Pasting the complete block as one input:

a=b=c=d=e=f=g=h=i=0
def fn():
    retrun True

produces the generic error:

  File "<python-input-0>", line 3
    retrun True
           ^^^^
SyntaxError: invalid syntax

Entering and submitting the assignment separately, followed by the function:

a=b=c=d=e=f=g=h=i=0
def fn():
    retrun True

produces:

  File "<python-input-2>", line 2
    retrun True
    ^^^^^^
SyntaxError: invalid syntax. Did you mean 'return'?

I would expect the same source code to produce the same keyword suggestion regardless of whether it is:

  • executed as a file with python t.py;
  • loaded as a module with python -m t;
  • pasted into PyREPL as one block; or
  • entered into PyREPL statement by statement.

This appears related to the keyword-typo suggestions added in #132449.

traceback.TracebackException._find_keyword_typos() treats the two source paths differently:

  • When the source is loaded from the filename, tokens outside the reported error line are skipped before consuming the token-search budget.
  • When the source is included directly in the SyntaxError metadata, all non-keyword NAME tokens in the extracted source fragment consume that budget.

The search is limited to ten such tokens. In this example, a through i account for nine tokens and fn is the tenth, so the string-source path stops before examining retrun. The file-source path skips those earlier lines and finds the typo.

One possible approach would be to prioritize tokens on the reported error line for both source paths, and only then search the surrounding source fragment. That would preserve the existing work limit while avoiding this tokenizer-dependent result.

CPython versions tested on:

CPython main branch, 3.15, 3.14

Operating systems tested on:

macOS

Linked PRs
  • gh-156087

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

調査の方向性

traceback.TracebackException._find_keyword_typos() から始め、filename と SyntaxError メタデータのパス、およびここで説明されている 10 トークンの予算を比較します。上記の 4 つの実行モードを再現します。同じソースが、ファイル実行、モジュールの読み込み、そして両方の PyREPL 入力形式で同じキーワード候補を生成すれば完了です。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
python
領域
compilers
issue の種類
バグ
難易度
3/5
見積もり時間
1〜2日
活発さ
停滞
明瞭さ
明確に書かれている
初心者へのやさしさ
25/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。