apache / apache/lucene

A Stateful Filter That Works Across Index Segments [LUCENE-2506]

Open
#3,580 10 comments 0 reactions 0 assignees View on GitHub
affects-version:3.0.2 legacy-jira-label:filter legacy-jira-label:lucene legacy-jira-label:search legacy-jira-priority:Major module:core/index type:enhancement
Dominant language
Java
Stars
3.6k
Forks
1.4k
Avg merge
2d 11h
Merged PRs (30d)
88

Description

By design, Lucene's Filter abstraction is applied once for every segment in the index during searching. In particular, the reader provided to its #getDocIdSet method does not represent the whole underlying index. In other words, if the index has more than one segment the given reader only represents a single segment. As a result, that definition of the filter suffers the limitation of not having the ability to permit/prohibit documents in the search results based on the terms that reside in segments that precede the current one.

To address this limitation, we introduce here a StatefulFilter which specifically builds on the Filter class so as to make it capable of remembering terms in segments spanning the whole underlying index. To reiterate, the need for making filters stateful stems from the fact that some, although not most, filters care about the terms that they may have come across in prior segments. It does so by keeping track of the past terms from prior segments in a cache that is maintained in a StatefulTermsEnum instance on a per-thread basis.

Additionally, to address the case where a filter might want to accept the last matching term, we keep track of the TermsEnum#docFreq of the terms in the segments filtered thus far. By comparing the sum of such TermsEnum#docFreq with that of the top-level reader, we can tell if the current segment is the last segment in which the current term appears. Ideally, for this to work correctly, we require the user to explicitly set the top-level reader on the StatefulFilter. Knowing what the top-level reader is also helps the StatefulFilter to clean up after itself once the search has concluded.

Note that we leave it up to each concrete sub-class of the stateful filter to decide what to remember in its state and what not to. In other words, it can choose to remember as much or as little from prior segments as it deems necessary. In keeping with the TermsEnum interface, which the StatefulTermsEnum class extends, the filter must decide which terms to accept or not, based on the holistic state of the search.

---
Migrated from [LUCENE-2506](https://issues.apache.org/jira/browse/LUCENE-2506) by Karthick Sankarachary, 1 vote, updated Nov 26 2010
Attachments: [LUCENE-2506.patch](https://apache.github.io/lucene-jira-archive/attachments/LUCENE-2506/LUCENE-2506.patch)
Linked issues:
- #3424

Contributor guide

Open the contributing guide

Research direction

Start with the Filter#getDocIdSet and TermsEnum entry points, then read the attached LUCENE-2506.patch and the linked issue #3424 to understand the proposed StatefulFilter and StatefulTermsEnum design. Done means agreeing on and implementing state across index segments, including top-level reader setup and cleanup, with validation for terms spanning segments and last-term handling.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
backend, search
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.