apache / apache/lucene

Allow users to create their own DirectoryTaxonomyReaders with empty taxoArrays instead of letting the taxoEpoch decide [LUCENE-10482]

Open
#11,518 12 comments 0 reactions 0 assignees View on GitHub
affects-version:9.1 legacy-jira-priority:Minor module:facet type:enhancement
Dominant language
Java
Stars
3.6k
Forks
1.4k
Avg merge
2d 11h
Merged PRs (30d)
88

Description

I was experimenting with the taxonomy index and `DirectoryTaxonomyReaders` in my day job where we were trying to replace the index underneath a reader asynchronously and then call the `doOpenIfChanged` call on it.

It turns out that the taxonomy index uses its own index based counter (the `taxonomyIndexEpoch`}) to determine if the index was opened in write mode after the last time it was written and if not, it directly tries to reuse the previous `taxoArrays` it had created. This logic fails in a scenario where both the old and new index were opened just once but the index itself is completely different in both the cases.

In such a case, it would be good to give the user the flexibility to inform the DTR to recreate its `taxoArrays`}, `ordinalCache` and `categoryCache`} (not refreshing these arrays causes it to fail in various ways). Luckily, such a constructor already exists! But it is private today! The idea here is to allow subclasses of DTR to use this constructor.

Curious to see what other folks think about this idea.

---
Migrated from [LUCENE-10482](https://issues.apache.org/jira/browse/LUCENE-10482) by Gautam Worah (@gautamworah96), updated Apr 19 2022
Pull requests: https://github.com/apache/lucene/pull/762

Contributor guide

Open the contributing guide

Research direction

Start with the DirectoryTaxonomyReader constructors and the existing private constructor that accepts empty taxoArrays. Confirm how subclasses would use it when replacing a taxonomy index, then make the constructor available as requested and verify that taxoArrays, ordinalCache, and categoryCache can be recreated for the new index.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
search
Issue type
Feature
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.