`<regex>`: Consider special-casing some simple loop structures
Nobody has claimed this yet.
- Dominant language
- C++
- Stars
- 11.2k
- Forks
- 1.7k
- Avg merge
- 4d 15h
- Merged PRs (30d)
- 22
Description
Currently, loops like a* or \d* go through the whole NFA state transition processing for each repetition, which slows down the handling of such loops a lot.
I think it's probably worth it to recognize and special-case such simple and comonly used loop structures in the matcher. (But we should not add such special-case handling to non-simple loops and instead resolve #5957 to effectively extend the simple loop handling to branchless loops.)
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Begin with the regex matcher and its NFA state-transition processing, the entry points named in the issue. Compare simple loops such as a* and \d* with non-simple loops; done means special-casing only the simple structures while leaving non-simple loops to the approach described alongside #5957.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- performance
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100