NVIDIA / NVIDIA/cuvs

Lucene: Avoid GPU OOMs and also Java-heap OOMs

Open
#2,459 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Lucene
Dominant language
Cuda
Stars
854
Forks
236
Avg merge
3d 3h
Merged PRs (30d)
62

Description

This fixes two issues:

  1. We were eagerly loading the full dataset into GPU instead of letting CAGRA manage streaming it to GPU.

  2. We were passing full datasets to createMultiLayerHnswGraph which crashed my java heap space when benchmarking 10M 1536d (~60GB) dataset; now we can stream manageable subsets instead.

PR: https://github.com/rapidsai/cuvs-lucene/pull/141

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reviewing PR #141 and the Lucene integration described in the issue. Trace how CAGRA receives the dataset and how createMultiLayerHnswGraph is called, then use the 10M 1536d benchmark scenario to verify that GPU memory and Java heap usage remain manageable without loading the full dataset at once.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
machine-learning, search
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
18/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.