python / python/cpython

Escaped quotes cause source comments to leak into debug f-string output and t-string metadata

Open
#154,711 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

interpreter-core topic-parser type-bug
Dominant language
Python
Stars
77.2k
Forks
35.9k
PR merge metrics
PR metrics pending

Description

Bug report

Bug description:

When an f-string or t-string replacement expression contains a backslash-escaped quote, a following source comment can be copied into runtime output or t-string metadata.

Observed effects:

  • A debug f-string includes the source comment in the resulting str.
  • A debug t-string includes the comment in Template.strings.
  • A debug t-string includes both the debug = and the comment inInterpolation.expression.
  • A non-debug t-string also retains the comment inInterpolation.expression.

The replacement expression itself evaluates to the correct value. The problemappears to be in reconstruction of the replacement-field source text.

Reproducer

f_result = f"{'\'' = # c
}"

t_debug = t"{'a\'b' = # c
}"
t_plain = t"{'a\'b' # c
}"

print("debug f-string result:", repr(f_result))
print("\ndebug t-string strings:", repr(t_debug.strings))
print("\ndebug t-string expression:", repr(t_debug.interpolations[0].expression))
print("\nnon-debug t-string expression:", repr(t_plain.interpolations[0].expression))

Actual behavior

On CPython main commit b86a41c, the output contains the source text # c in all four observed values.

debug f-string result: '\'\\\'\' = # c\n"\'"'

debug t-string strings: ("'a\\'b' = # c\n", '')

debug t-string expression: "'a\\'b' = # c"

non-debug t-string expression: "'a\\'b' # c"

Possible cause

The behavior appears to be caused by inconsistent escaped-character handling between two scans in _PyLexer_set_ftstring_expr() in Parser/lexer/string.c.

The initial scan that detects a comment outside a string literal explicitly skips a backslash and the character following it:

if (ch == '\\') {
    i++;
    continue;
}

However, after a comment has been detected, the subsequent reconstruction pass that removes comments appears not to perform the equivalent skip. It changes in_string whenever it encounters a quote:

if (ch == '"' || ch == '\'') {
    if (!in_string) {
        in_string = 1;
        quote_char = ch;
    }
    else if (ch == quote_char) {
        in_string = 0;
    }
}
else if (ch == '#' && !in_string) {
    /* skip the comment */
}

For:

'a\'b' = # c

the second pass may interpret the escaped quote as the end of the string and the real closing quote as the beginning of another string:

'       begin string
\       treated as an ordinary character
'       incorrectly treated as end of string
b
'       incorrectly treated as beginning of another string
#       incorrectly considered to be inside a string

Consequently, the # is not recognized by the comment-removal branch and the comment is copied into the reconstructed text.
This also explains why the same input pattern affects both f-strings and t-strings: both use the same replacement-expression reconstruction logic.
This is a proposed root cause based on the observed behavior and source review, not a confirmed diagnosis.

CPython versions tested on:

CPython main branch, 3.16

Operating systems tested on:

Linux

Linked PRs
  • gh-154914

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in Parser/lexer/string.c at _PyLexer_set_ftstring_expr(), comparing the initial escaped-character scan with the later comment-removal reconstruction pass. Run the provided f-string and t-string reproducer on CPython main. Done means escaped quotes no longer cause # c to appear in debug f-string output, t-string strings, or interpolation metadata.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
compilers
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Clearly specified
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.