Speed up matching of case-insensitive character sets
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 35/100
- Loại issue
- Tính năng
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Đình trệ
- Công nghệ
- python
- Lĩnh vực
- performance
Hướng nghiên cứu
Trước tiên, hãy kiểm tra PR được liên kết gh-152055, sau đó kiểm tra entry point SRE(count) và cách xử lý các opcode IN, IN_IGNORE, IN_UNI_IGNORE và IN_LOC_IGNORE. Chạy benchmark pyperf được cung cấp với các bản build chưa áp dụng patch và đã áp dụng patch, đồng thời xác nhận rằng các benchmark lặp của tập ký tự không phân biệt chữ hoa chữ thường được cải thiện mà không làm thay đổi trường hợp tập phủ định vốn không thay đổi.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Feature or enhancement
Proposal:
A REPEAT_ONE over a case-insensitive character set — e.g. [a-z]+ with re.IGNORECASE — does not use the fast SRE(count) path. The compiled inner opcode is IN_IGNORE / IN_UNI_IGNORE / IN_LOC_IGNORE, none of which has a case in SRE(count). The case-sensitive SRE_OP_IN already has a fast case.
Adding the three IN_*_IGNORE cases to SRE(count) lets them scan inline.
Benchmark
| Benchmark | before | after | |
|---|---|---|---|
[a-z]+ re.I|re.A (IN_IGNORE) |
1.28 us | 538 ns | 2.38x |
[a-z]+ re.I (IN_UNI_IGNORE) |
1.41 us | 711 ns | 1.98x |
[aeiou]+ re.I |
1.35 us | 706 ns | 1.91x |
[a-z0-9]+ re.I |
1.31 us | 696 ns | 1.88x |
[a-z0-9_]+ re.I |
1.31 us | 703 ns | 1.86x |
[a-z]+ re.L|re.I bytes (IN_LOC_IGNORE) |
2.09 us | 1.38 us | 1.52x |
findall [a-z]+ re.I |
109 us | 88.8 us | 1.22x |
findall [a-z_][a-z0-9_]* re.I |
103 us | 89.2 us | 1.15x |
[^0-9]+ re.I is unchanged — it has no cased members, so it stays a plain
IN (already fast).
benchmark script (pyperf)
"""Benchmark: SRE(count) fast path for case-insensitive set repeats."""
import re
import pyperf
N = 100
MIXED = ("aBcDeFgHiJkLmNoPqRsTuVwX" * N)[:N]
ALNUM = ("aB3dE6gH9kLmN0pQrStUvWx1" * N)[:N]
WORD = ("aB_dE_gH_kLmN_pQrStUvW_1" * N)[:N]
NODIGIT = ("aBcDeF gHiJkL!mNoPqR.sT?" * N)[:N]
BYTES = MIXED.encode("latin1")
SCANS = [
("scan_alpha_uni", re.compile(r"[a-z]+", re.I), MIXED),
("scan_alpha_asc", re.compile(r"[a-z]+", re.I | re.A), MIXED),
("scan_alnum_uni", re.compile(r"[a-z0-9]+", re.I), ALNUM),
("scan_word_uni", re.compile(r"[a-z0-9_]+",re.I), WORD),
("scan_neg_uni", re.compile(r"[^0-9]+", re.I), NODIGIT),
("scan_vowels_uni", re.compile(r"[aeiou]+", re.I), "aAeEiIoOuU" * (N // 10)),
("scan_alpha_loc", re.compile(rb"[a-z]+", re.L | re.I), BYTES),
]
DOC = ("The Quick Brown Fox jumps over 12 Lazy Dogs near IP 10_0_0_1 and Node7. " * 50)
FINDS = [
("find_words_ci", re.compile(r"[a-z]+", re.I), DOC),
("find_ident_ci", re.compile(r"[a-z_][a-z0-9_]*", re.I), DOC),
]
def make_scan(p, s):
def run():
assert p.match(s) is not None
return run
runner = pyperf.Runner()
for name, p, s in SCANS:
runner.bench_func(name, make_scan(p, s))
for name, p, s in FINDS:
runner.bench_func(name, (lambda p, s: lambda: p.findall(s))(p, s))
Run under the unpatched and patched builds, then
python -m pyperf compare_to before.json after.json --table.
Has this already been discussed elsewhere?
No response given
Links to previous discussion of this feature:
No response
Linked PRs
- gh-152055
- Ngôn ngữ chính
- Python
- Star
- 77.2k
- Fork
- 36k
- Merge trung bình
- 1 ngày 9 giờ
- Pull request đã merge (30 ngày)
- 558
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của python/cpython
-
docs pending
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
-
stdlib type-feature
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
-
stdlib type-feature
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
-
build type-bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
-
stdlib topic-email type-feature
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100