dotnet / dotnet/extensions

Hybrid Retrieval: Missing Sparse Embedding Abstraction in .NET

Open
#7,563 1 comment 1 reaction 0 assignees View on GitHub
api-suggestion area-ai
Dominant language
C#
Stars
3.2k
Forks
894
Avg merge
1d 12h
Merged PRs (30d)
23

Description

### Background and motivation

While implementing hybrid retrieval (dense + sparse + RRF) on top of Qdrant,
we ran into a fundamental gap in the current .NET ecosystem.

`Microsoft.Extensions.AI` provides `IEmbeddingGenerator`
for dense embeddings, but there is no corresponding abstraction for sparse embeddings.

This is not a niche requirement. Sparse retrieval (BM25 / SPLADE) is a
**first-class citizen in production RAG pipelines**, especially for:

- **CJK languages**: Dense-only retrieval degrades significantly for Chinese/Japanese/Korean,
where sparse + dense hybrid is considered best practice, not an optimization.
- **Exact keyword matching**: Product codes, proper nouns, technical terms —
cases where semantic similarity alone consistently underperforms.
- **Enterprise search**: Most production RAG systems use hybrid retrieval as the baseline.

Without this abstraction, the .NET ecosystem faces a concrete fragmentation risk:

1. Developers bypass `IVectorStore` entirely and use `Qdrant.Client` / `Milvus.Client` directly
2. Community packages build their own incompatible sparse abstractions
3. Microsoft's abstraction layer loses relevance for real-world RAG workloads

We have already confirmed this gap through decompilation of
`CommunityToolkit.VectorData.Qdrant 1.0.0-preview.3`:
the current `HybridSearchAsync` implementation uses **payload keyword Match filter**,
not sparse vector search. The official docs also confirm:
> "Sparse vector based hybrid search is not currently supported."

Additionally, `VectorStoreVectorAttribute` and `VectorStoreVectorProperty` currently
only model dense vectors (`Dimensions`, `IndexKind`, `DistanceFunction`),
with no mechanism to distinguish or define sparse vector fields.
This means the gap exists at both the **generation layer** (no `ISparseEmbeddingGenerator`)
and the **storage layer** (no sparse vector property support in MEVD).

The most viable workaround today is routing through TEI `/embed-sparse`,
but that still requires a custom abstraction that doesn't exist anywhere
in the ecosystem. FastEmbed, the de facto solution, is Python-only.

### API Proposal

```csharp
namespace Microsoft.Extensions.AI;

// Mirrors IEmbeddingGenerator for sparse vectors
public interface ISparseEmbeddingGenerator
{
Task GenerateAsync(
TInput input,
CancellationToken cancellationToken = default);

Task> GenerateBatchAsync(
IEnumerable inputs,
CancellationToken cancellationToken = default);
}

public readonly struct SparseEmbedding
{
public ReadOnlyMemory Indices { get; init; }
public ReadOnlyMemory Values { get; init; }
}
```

### API Usage

```csharp
// At index time: generate both dense and sparse vectors per chunk
ISparseEmbeddingGenerator sparseGen = new TeiSparseEmbeddingGenerator(endpoint);
IEmbeddingGenerator> denseGen = ...;

var dense = await denseGen.GenerateAsync(chunk);
var sparse = await sparseGen.GenerateAsync(chunk);

// Write to Qdrant with both vectors (dense + sparse)
await collection.UpsertAsync(new Point
{
DenseVector = dense.Vector,
SparseVector = new SparseVector(sparse.Indices, sparse.Values)
});

// At query time: same abstraction, enabling true dense + sparse + RRF hybrid
var queryDense = await denseGen.GenerateAsync(query);
var querySparse = await sparseGen.GenerateAsync(query);
```

### Alternative Designs

A broader fix would also require extending `VectorStoreVectorAttribute` and
`VectorStoreVectorProperty` in `Microsoft.Extensions.VectorData` to support
sparse vector fields — potentially with a new `[VectorStoreSparseVector]` attribute.
However, `ISparseEmbeddingGenerator` in MEAI is the foundational piece that
needs to land first before the storage layer can follow.

### Risks

Without this abstraction:
- CJK RAG workloads have no viable path in the .NET ecosystem
- Community implementations will diverge, fragmenting the ecosystem
- `IVectorStore` / `IKeywordHybridSearchable` remain insufficient
for production hybrid retrieval, undermining the value of MEVD as a whole

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.