[DOC] Stratifying vectors in CAGRA's `extend()` across clusters to keep them distant improves recall
Nobody has claimed this yet.
- Dominant language
- Cuda
- Stars
- 854
- Forks
- 236
- Avg merge
- 3d 3h
- Merged PRs (30d)
- 62
Description
@jinsolp recently noted that clustering a series of vectors and then ordering multiple batches of the stratified vectors when calling extend() seemed to maintain recall while ordering the batches with similar and adjacent vectors had a very large negative impact on the recall post-extend. We should definitely document this behavior, because this could be a way to improve the streamability of CAGRA index construction.
cc @anaruse @tfeher @enp1s0 @yan-zaretskiy for awareness.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by locating the documentation for CAGRA's extend() entry point and any existing guidance on index construction. Document that stratifying vectors across clusters and ordering multiple batches can preserve post-extend recall, while batches of similar or adjacent vectors can significantly reduce it, and explain how this may improve streamability.
Written by the indexing model from the issue text.
Assessment
- Domain
- documentation, search
- Issue type
- Documentation
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100