`\p{...}` gives wrong results when mixed with letters under `re.IGNORECASE`
未關閉
@serhiy-storchaka 已經在處理了。
開始於 2026年8月7日。
extension-modules
topic-regex
type-bug
- 主要語言
- Python
- 星號
- 77.2k
- 分支
- 36k
- PR 合併指標
- PR 指標待擷取
描述
Bug description:
\p{Lu} matches uppercase Unicode letters. It breaks when mixed with a standard letter with re.IGNORECASE:
import re
property_only = re.compile(r"[\p{Lu}]", re.IGNORECASE)
property_plus_one = re.compile(r"[\p{Lu}1]", re.IGNORECASE)
property_plus_a = re.compile(r"[\p{Lu}a]", re.IGNORECASE)
assert property_only.fullmatch("B")
assert property_plus_one.fullmatch("B")
assert property_plus_a.fullmatch("B"), "adding 'a' broke the match for 'B'"
Expected: adding "a" to the character class shouldn't break the existing match
Actual: AssertionError: adding 'a' broke the match for 'B'
CPython versions tested on:
CPython main branch
Operating systems tested on:
macOS
Linked PRs
- gh-155299
貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
評估
這個 Issue 還沒有評估資料。