Search is extremely slow or freezes on certain unicode strings
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 8.9k
- Forks
- 1.6k
- PR merge metrics
- No merged PRs in 30d
Description
The following search has been running for over 20 minutes on my computer and it still has not ended:
pattern.search.search("{CD+ and? CD? CD?} *? *? INFECT|AFFLICT",
pattern.en.parsetree(u"2012 — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — the coordination of Epidemiological surveillance,"))
The search pattern I'm using usually works on full documents within seconds. This string in the example is a snippet from a document that was causing a bulk processing job to freeze.
Replacing the unicode —s with -s brings the processing time back to being nearly instantaneous.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reproducing the exact pattern.search.search and pattern.en.parsetree example from the issue, comparing the em-dash input with hyphens. Trace the search path to identify why this Unicode string takes over 20 minutes, then verify that the problematic input completes promptly without regressing normal full-document searches.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- performance, search
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 25/100