[FEA] BSDG: Add batch option
Open
@jinsolp is already working on this.
Since Jul 27, 2026.
feature request
- Dominant language
- Cuda
- Stars
- 854
- Forks
- 236
- Avg merge
- 3d 3h
- Merged PRs (30d)
- 62
Description
BSDG (Billion Scale Data Generator)
Open-sourced version in cuvs bench currently flushes the full dataset to disk on generation.
For very large datasets, users may want to generate in partitions.
If user gives a batch_size and a batch_id we should just return/save that specific batch.
# returns rows between batch_size * batch_id ~ batch_size * (batch_id+1)
batch = gen_batch(total_rows, batch_size, batch_id)
This is straightforward because we know the number of points per cluster, so we should be only able to generate the necessary clusters that have vectors belong to that partition of the full dataset.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.