NVIDIA / NVIDIA/cuvs

[DOC] Stratifying vectors in CAGRA's `extend()` across clusters to keep them distant improves recall

Open
#1,727 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

doc
Dominant language
Cuda
Stars
854
Forks
236
Avg merge
3d 3h
Merged PRs (30d)
62

Description

@jinsolp recently noted that clustering a series of vectors and then ordering multiple batches of the stratified vectors when calling extend() seemed to maintain recall while ordering the batches with similar and adjacent vectors had a very large negative impact on the recall post-extend. We should definitely document this behavior, because this could be a way to improve the streamability of CAGRA index construction.

cc @anaruse @tfeher @enp1s0 @yan-zaretskiy for awareness.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by locating the documentation for CAGRA's extend() entry point and any existing guidance on index construction. Document that stratifying vectors across clusters and ordering multiple batches can preserve post-extend recall, while batches of similar or adjacent vectors can significantly reduce it, and explain how this may improve streamability.

Written by the indexing model from the issue text.

Assessment

Domain
documentation, search
Issue type
Documentation
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.