intel / intel/ScalableVectorSearch

New index building APIs for simpler and flexible usage

Open
#189 1 comment 0 reactions 0 assignees View on GitHub
enhancement
Dominant language
C++
Stars
236
Forks
48
Avg merge
4d 22h
Merged PRs (30d)
10

Description

In SVS, we separate the concept of dataset and index. This allows an index to accept different datasets, such as FP32, Scalar Quantization (`SQDataset`), LVQ, and LeanVec.
Currently, using `SQDataset` as an example, users must call `compress` to get a `SQDataset`, then pass it to the index. However, most use cases don't care about the dataset itself - this two-step process is often unnecessary for most users. Additionally, this approach makes runtime fallback impossible to implement in SVS, as each dataset (i.e., type) is determined at compile time.
```cpp
auto loaded =
svs::VectorDataLoader(std::filesystem::path(SVS_DATA_DIR) / "data_f32.svs").load();
auto data =
svs::scalar::SQDataset::compress(loaded, threadpool); // SQDataset is determined at compile time
auto parameters = svs::index::vamana::VamanaBuildParameters{};
svs::Vamana index = svs::Vamana::build(
parameters, data, svs::distance::DistanceL2(), num_threads
);
```

I propose adding index building APIs that directly accept uncompressed data and take the dataset type as a parameter to determine which dataset format to use internally.
```cpp
auto loaded =
svs::VectorDataLoader(std::filesystem::path(SVS_DATA_DIR) / "data_f32.svs").load();
auto parameters = svs::index::vamana::VamanaBuildParameters{};
svs::Vamana index = svs::Vamana::build(
parameters, data, svs::distance::DistanceL2(), num_threads, svs::SQ8
); // internally, SVS fallbacks to uncompressed data if scalar quantization failed
```
Another advantage of this is that we could optionally utilize uncompressed data to build the graph rather than compressed data, which typically gives worse quality graphs as compression introduces approximation errors that can degrade the graph structure during construction.
Note that when calling `build`, we can compress the data and build the graph in parallel, as the two things are independent.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.