NVIDIA / NVIDIA/cuvs

Lucene: look into observed weak ivf-pq indexing-time performance

Open
#2,462 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Lucene
Dominant language
Cuda
Stars
854
Forks
236
Avg merge
3d 3h
Merged PRs (30d)
62

Description

Try to reproduce the following performance on 26.02 and see if problems disappear on latest branches/PRs

Image
Vector Dimensions: 64

  Vector Similarity: DOT_PRODUCT  (vectors are assumed pre-normalized (per the code comment))

  GPU (cuVS CAGRA) Settings — from StandaloneIndexer.createIndexWriterConfig():

  - CagraGraphBuildAlgo: IVF_PQ (the enum value, no sub-parameters like nlist/nprobe/PQ bits are set — they use cuVS defaults)
  - graphDegree: 32 (default is 64)
  - intermediateGraphDegree: 48 (default is 128)
  - maxConn (HNSW M): 16
  - beamWidth: 100
  - numMergeWorkers: 16
  - Merge executor: a fixed thread pool of 16

  CPU (Lucene HNSW) Settings:

  - Lucene99HnswVectorsFormat(maxConn=16, beamWidth=100, mergeWorkers=16, mergePool)
  - Codec: Lucene101Codec with BEST_SPEED mode

  IndexWriter Config (shared):

  - RAMBufferSizeMB: 30 GB (30 * 1024)
  - maxBufferedDocs: disabled (auto-flush off)
  - useCompoundFile: false
  - openMode: CREATE

  Flush / Commit Behavior:

  - GPU mode: commits are skipped during loading (!useGpu guard on line ~180 of StandaloneIndexer). All docs accumulate in RAM, then close() triggers a single commit() → forceMerge(4) → commit().
  - CPU mode: commits every 200,000 docs (COMMIT_BATCH_SIZE = 200000)

  Indexing Threads:

  - GPU: 2 threads (code says useGpu ? 2 : MERGE_WORKERS), though the Quip doc notes 1 thread was optimal and 2 was slower
  - CPU: 16 threads

  Queue size: 150,000 documents

  IVF_PQ sub-parameters: The code uses IVF_PQ as the CagraGraphBuildAlgo enum but does not set any IVF_PQ-specific parameters (like nlist, nprobe, PQ sub-quantizers, PQ bits). These all use cuVS defaults. The AcceleratedHNSWParams.Builder only exposes
  withCagraGraphBuildAlgo, withGraphDegree, withIntermediateGraphDegree, withMaxConn, withBeamWidth, withNumMergeWorkers, and withMergeExecutorService.

My hypothesis is that most of these issues are resolved by using better IVF_PQ settings and switching to the latest main branch for cuvs and cuvs-lucene (with my PRs for cuvs-lucene: https://github.com/rapidsai/cuvs-lucene/pull/141, https://github.com/rapidsai/cuvs-lucene/pull/140), but there are also some timing issues that we need to update in the searchscale benchmarking repo (https://github.com/SearchScale/vectorsearch-benchmarks/pull/14, https://github.com/SearchScale/vectorsearch-benchmarks/pull/16). Will verify this (https://github.com/rapidsai/cuvs-lucene/issues/143) or continue investigating after resolving https://github.com/SearchScale/vectorsearch-benchmarks/issues/17

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the reported indexing performance on 26.02, then compare it with the latest cuVS and cuVS-Lucene branches and the referenced pull requests. Inspect StandaloneIndexer.createIndexWriterConfig() and the loading/commit behavior around line 180, then check vectorsearch-benchmarks issue 17. Done means determining whether the weak IVF-PQ performance and timing issues still occur and recording the findings.

Written by the indexing model from the issue text.

Assessment

Domain
machine-learning, performance, search
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.