clips / clips/pattern

Search is extremely slow or freezes on certain unicode strings

Open
#104 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
8.9k
Forks
1.6k
PR merge metrics
No merged PRs in 30d

Description

The following search has been running for over 20 minutes on my computer and it still has not ended:

pattern.search.search("{CD+ and? CD? CD?} *? *? INFECT|AFFLICT",
  pattern.en.parsetree(u"2012 — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — the coordination of Epidemiological surveillance,"))

The search pattern I'm using usually works on full documents within seconds. This string in the example is a snippet from a document that was causing a bulk processing job to freeze.

Replacing the unicode s with -s brings the processing time back to being nearly instantaneous.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the exact pattern.search.search and pattern.en.parsetree example from the issue, comparing the em-dash input with hyphens. Trace the search path to identify why this Unicode string takes over 20 minutes, then verify that the problematic input completes promptly without regressing normal full-document searches.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
performance, search
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.