Escaped quotes cause source comments to leak into debug f-string output and t-string metadata
Nessuno ha ancora preso questa issue.
- Lingua principale
- Python
- Stelle
- 77.2k
- Fork
- 35.9k
- Metriche di merge delle PR
- Metriche PR in attesa
Descrizione
Bug report
Bug description:
When an f-string or t-string replacement expression contains a backslash-escaped quote, a following source comment can be copied into runtime output or t-string metadata.
Observed effects:
- A debug f-string includes the source comment in the resulting str.
- A debug t-string includes the comment in Template.strings.
- A debug t-string includes both the debug = and the comment inInterpolation.expression.
- A non-debug t-string also retains the comment inInterpolation.expression.
The replacement expression itself evaluates to the correct value. The problemappears to be in reconstruction of the replacement-field source text.
Reproducer
f_result = f"{'\'' = # c
}"
t_debug = t"{'a\'b' = # c
}"
t_plain = t"{'a\'b' # c
}"
print("debug f-string result:", repr(f_result))
print("\ndebug t-string strings:", repr(t_debug.strings))
print("\ndebug t-string expression:", repr(t_debug.interpolations[0].expression))
print("\nnon-debug t-string expression:", repr(t_plain.interpolations[0].expression))
Actual behavior
On CPython main commit b86a41c, the output contains the source text # c in all four observed values.
debug f-string result: '\'\\\'\' = # c\n"\'"'
debug t-string strings: ("'a\\'b' = # c\n", '')
debug t-string expression: "'a\\'b' = # c"
non-debug t-string expression: "'a\\'b' # c"
Possible cause
The behavior appears to be caused by inconsistent escaped-character handling between two scans in _PyLexer_set_ftstring_expr() in Parser/lexer/string.c.
The initial scan that detects a comment outside a string literal explicitly skips a backslash and the character following it:
if (ch == '\\') {
i++;
continue;
}
However, after a comment has been detected, the subsequent reconstruction pass that removes comments appears not to perform the equivalent skip. It changes in_string whenever it encounters a quote:
if (ch == '"' || ch == '\'') {
if (!in_string) {
in_string = 1;
quote_char = ch;
}
else if (ch == quote_char) {
in_string = 0;
}
}
else if (ch == '#' && !in_string) {
/* skip the comment */
}
For:
'a\'b' = # c
the second pass may interpret the escaped quote as the end of the string and the real closing quote as the beginning of another string:
' begin string
\ treated as an ordinary character
' incorrectly treated as end of string
b
' incorrectly treated as beginning of another string
# incorrectly considered to be inside a string
Consequently, the # is not recognized by the comment-removal branch and the comment is copied into the reconstructed text.
This also explains why the same input pattern affects both f-strings and t-strings: both use the same replacement-expression reconstruction logic.
This is a proposed root cause based on the observed behavior and source review, not a confirmed diagnosis.
CPython versions tested on:
CPython main branch, 3.16
Operating systems tested on:
Linux
Linked PRs
- gh-154914
Guida per i contributori
Apri la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Direzione di ricerca
Inizia in Parser/lexer/string.c, in _PyLexer_set_ftstring_expr(), confrontando la scansione iniziale degli escaped characters con il successivo passaggio di ricostruzione per la rimozione dei commenti. Esegui il reproducer fornito per f-string e t-string su CPython main. Il lavoro è completato quando le escaped quotes non fanno più comparire # c nell’output di debug f-string, nelle stringhe t-string o nei metadati di interpolazione.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- python
- Ambito
- compilers
- Tipo di issue
- Bug
- Difficoltà
- 3/5
- Tempo stimato
- 1-2 giorni
- Stato di attività
- Ferma
- Chiarezza
- Specificata chiaramente
- Idoneità per principianti
- 35/100