python / python/cpython

Keyword typo suggestions differ between file and string tokenizers

Aperta
#156,047 4 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

3.14 3.15 3.16 interpreter-core topic-parser type-bug
Lingua principale
Python
Stelle
77.2k
Fork
35.9k
Metriche di merge delle PR
Metriche PR in attesa

Descrizione

Bug report

Bug description:

Keyword-typo SyntaxError suggestions depend on whether CPython processes the same source using the file tokenizer or the string tokenizer.

Create t.py containing:

a=b=c=d=e=f=g=h=i=0
def fn():
    retrun True

Running the file directly:

$ python t.py

uses the file-tokenizer path and produces the helpful diagnostic:

  File ".../t.py", line 3
    retrun True
    ^^^^^^
SyntaxError: invalid syntax. Did you mean 'return'?

Running the same file as a module:

$ python -m t

uses the string-tokenizer path and produces:

  File ".../t.py", line 3
    retrun True
           ^^^^
SyntaxError: invalid syntax

The same discrepancy is visible in the default PyREPL-based REPL.

Pasting the complete block as one input:

a=b=c=d=e=f=g=h=i=0
def fn():
    retrun True

produces the generic error:

  File "<python-input-0>", line 3
    retrun True
           ^^^^
SyntaxError: invalid syntax

Entering and submitting the assignment separately, followed by the function:

a=b=c=d=e=f=g=h=i=0
def fn():
    retrun True

produces:

  File "<python-input-2>", line 2
    retrun True
    ^^^^^^
SyntaxError: invalid syntax. Did you mean 'return'?

I would expect the same source code to produce the same keyword suggestion regardless of whether it is:

  • executed as a file with python t.py;
  • loaded as a module with python -m t;
  • pasted into PyREPL as one block; or
  • entered into PyREPL statement by statement.

This appears related to the keyword-typo suggestions added in #132449.

traceback.TracebackException._find_keyword_typos() treats the two source paths differently:

  • When the source is loaded from the filename, tokens outside the reported error line are skipped before consuming the token-search budget.
  • When the source is included directly in the SyntaxError metadata, all non-keyword NAME tokens in the extracted source fragment consume that budget.

The search is limited to ten such tokens. In this example, a through i account for nine tokens and fn is the tenth, so the string-source path stops before examining retrun. The file-source path skips those earlier lines and finds the typo.

One possible approach would be to prioritize tokens on the reported error line for both source paths, and only then search the surrounding source fragment. That would preserve the existing work limit while avoiding this tokenizer-dependent result.

CPython versions tested on:

CPython main branch, 3.15, 3.14

Operating systems tested on:

macOS

Linked PRs
  • gh-156087

Guida per i contributori

Apri la guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Direzione di ricerca

Inizia da traceback.TracebackException._find_keyword_typos(), confrontando i suoi percorsi per filename e metadati di SyntaxError e il budget di dieci token descritto qui. Riproduci le quattro modalità di esecuzione sopra indicate; il lavoro è completo quando la stessa sorgente produce lo stesso suggerimento di parola chiave durante l’esecuzione di un file, il caricamento di un modulo e entrambe le forme di input di PyREPL.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
python
Ambito
compilers
Tipo di issue
Bug
Difficoltà
3/5
Tempo stimato
1-2 giorni
Stato di attività
Ferma
Chiarezza
Specificata chiaramente
Idoneità per principianti
25/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.