python / python/cpython

Escaped quotes cause source comments to leak into debug f-string output and t-string metadata

オープン
#154,711 コメント 3 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

interpreter-core topic-parser type-bug
主要言語
Python
スター
77.2k
フォーク
35.9k
PR マージ指標
PR 指標を取得中

説明

Bug report

Bug description:

When an f-string or t-string replacement expression contains a backslash-escaped quote, a following source comment can be copied into runtime output or t-string metadata.

Observed effects:

  • A debug f-string includes the source comment in the resulting str.
  • A debug t-string includes the comment in Template.strings.
  • A debug t-string includes both the debug = and the comment inInterpolation.expression.
  • A non-debug t-string also retains the comment inInterpolation.expression.

The replacement expression itself evaluates to the correct value. The problemappears to be in reconstruction of the replacement-field source text.

Reproducer

f_result = f"{'\'' = # c
}"

t_debug = t"{'a\'b' = # c
}"
t_plain = t"{'a\'b' # c
}"

print("debug f-string result:", repr(f_result))
print("\ndebug t-string strings:", repr(t_debug.strings))
print("\ndebug t-string expression:", repr(t_debug.interpolations[0].expression))
print("\nnon-debug t-string expression:", repr(t_plain.interpolations[0].expression))

Actual behavior

On CPython main commit b86a41c, the output contains the source text # c in all four observed values.

debug f-string result: '\'\\\'\' = # c\n"\'"'

debug t-string strings: ("'a\\'b' = # c\n", '')

debug t-string expression: "'a\\'b' = # c"

non-debug t-string expression: "'a\\'b' # c"

Possible cause

The behavior appears to be caused by inconsistent escaped-character handling between two scans in _PyLexer_set_ftstring_expr() in Parser/lexer/string.c.

The initial scan that detects a comment outside a string literal explicitly skips a backslash and the character following it:

if (ch == '\\') {
    i++;
    continue;
}

However, after a comment has been detected, the subsequent reconstruction pass that removes comments appears not to perform the equivalent skip. It changes in_string whenever it encounters a quote:

if (ch == '"' || ch == '\'') {
    if (!in_string) {
        in_string = 1;
        quote_char = ch;
    }
    else if (ch == quote_char) {
        in_string = 0;
    }
}
else if (ch == '#' && !in_string) {
    /* skip the comment */
}

For:

'a\'b' = # c

the second pass may interpret the escaped quote as the end of the string and the real closing quote as the beginning of another string:

'       begin string
\       treated as an ordinary character
'       incorrectly treated as end of string
b
'       incorrectly treated as beginning of another string
#       incorrectly considered to be inside a string

Consequently, the # is not recognized by the comment-removal branch and the comment is copied into the reconstructed text.
This also explains why the same input pattern affects both f-strings and t-strings: both use the same replacement-expression reconstruction logic.
This is a proposed root cause based on the observed behavior and source review, not a confirmed diagnosis.

CPython versions tested on:

CPython main branch, 3.16

Operating systems tested on:

Linux

Linked PRs
  • gh-154914

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

調査の方向性

Parser/lexer/string.c の _PyLexer_set_ftstring_expr() から始め、最初の escaped characters のスキャンと、その後のコメント削除の再構築パスを比較します。提供されている f-string と t-string の再現プログラムを CPython main で実行します。escaped quotes によってデバッグ f-string 出力、t-string の文字列、または補間メタデータに # c が表示されなくなれば完了です。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
python
領域
compilers
issue の種類
バグ
難易度
3/5
見積もり時間
1〜2日
活発さ
停滞
明瞭さ
明確に書かれている
初心者へのやさしさ
35/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。