A backreference does not match characters which are matched case-insensitively as literals
還沒有人認領這個 Issue。
- 主要語言
- Python
- 星號
- 77.2k
- 分支
- 36k
- PR 合併指標
- PR 指標待擷取
描述
import re
print(re.fullmatch('(?i)ς', 'σ'))
print(re.fullmatch(r'(?i)(.)\1', 'ςσ'))
print(re.fullmatch('(?i)Σ', 'σ'))
print(re.fullmatch(r'(?i)(.)\1', 'Σσ'))
<re.Match object; span=(0, 1), match='σ'>
None
<re.Match object; span=(0, 1), match='σ'>
<re.Match object; span=(0, 2), match='Σσ'>
'ς' (GREEK SMALL LETTER FINAL SIGMA) matches 'σ' (GREEK SMALL LETTER SIGMA) when it is a literal in the pattern, but not when it is matched by a backreferenced group. 'Σ' (GREEK CAPITAL LETTER SIGMA) matches it in both cases.
This is because a literal is expanded at compile time into the alternatives listed in _casefix._EXTRA_CASES, which groups the characters having the same uppercase, so (?i)ς is compiled to a set containing both sigmas. A backreference has no compiled set to expand, and GROUPREF_UNI_IGNORE compares sre_lower_unicode() of the two characters, which lowercases 'Σ' to 'σ' but leaves 'ς' unchanged.
28 pairs are affected, among them 'µ' (MICRO SIGN) and 'μ' (GREEK SMALL LETTER MU), 'ſ' (LATIN SMALL LETTER LONG S) and 's' (LATIN SMALL LETTER S), the Greek symbol variants, and the Cyrillic historic letters added in Unicode 9.0.
sre_lower_unicode() can return a key which is the same for all characters matched case-insensitively, and then a backreference matches whatever a literal matches. This takes three steps:
-
Compare the simple case folding instead of the lowercase. This unifies most of the pairs,
'µ'with'μ'among them. -
Lowercase the result. A character whose full case folding is longer than one character is not unified with its case partners by the folding -- both SHARP S characters are folded to
"ss", and every letter with ypogegrammeni to two characters -- so such a character is lowercased instead, which is what keeps it with them, since they share the lowercase. The folding is lowercased in turn, since it is not always downwards: Cherokee letters are folded to their uppercase. -
Hardcode the four pairs which no case mapping unifies: LATIN SMALL LETTER DOTLESS I (
'i'and'ı'), the two pairs of Greek letters with tonos and with oxia ('ΐ'and'ΐ','ΰ'and'ΰ'), and the two ST ligatures ('ſt'and'st'). All eight characters are Unicode 1.1, and nothing added since has joined them.
_casefix._EXTRA_CASES is then unused and can be removed, together with the alternatives which the compiler expanded a literal into. Its generator keeps computing the groups of characters which have to be matched case-insensitively, and fails if sre_lower_unicode() does not fold such a group to a single code.
Linked PRs
- gh-156514
貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
研究方向
從正規表示式編譯器的 _casefix._EXTRA_CASES 產生以及呼叫 sre_lower_unicode() 的 GROUPREF_UNI_IGNORE 路徑開始。檢查現有的正規表示式測試,確認不區分大小寫的字面值和反向參照。完成的標準是:受影響的 Unicode 對透過字面值和反向參照都能一致比對,而且產生器的 folding 驗證仍然通過。
由索引模型根據 Issue 內容生成。
評估
- 技術堆疊
- python
- 領域
- compilers
- Issue 類型
- 缺陷
- 難度
- 4/5
- 預估耗時
- 3-5 天
- 活躍度
- 停滯
- 描述清晰度
- 描述清楚
- 新手友好度
- 30/100