Lucene: Avoid GPU OOMs and also Java-heap OOMs
Nobody has claimed this yet.
- Dominant language
- Cuda
- Stars
- 854
- Forks
- 236
- Avg merge
- 3d 3h
- Merged PRs (30d)
- 62
Description
This fixes two issues:
-
We were eagerly loading the full dataset into GPU instead of letting CAGRA manage streaming it to GPU.
-
We were passing full datasets to createMultiLayerHnswGraph which crashed my java heap space when benchmarking 10M 1536d (~60GB) dataset; now we can stream manageable subsets instead.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reviewing PR #141 and the Lucene integration described in the issue. Trace how CAGRA receives the dataset and how createMultiLayerHnswGraph is called, then use the 10M 1536d benchmark scenario to verify that GPU memory and Java heap usage remain manageable without loading the full dataset at once.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java
- Domain
- machine-learning, search
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 18/100