Azure / Azure/azure-sdk-for-python

[Cosmos] [Embedding V0] Wire resolver into _run_hybrid_search, plumb generator, and add diagnostics span

Aperta
#46,733 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Cosmos feature-request
Lingua principale
Python
Stelle
5.6k
Fork
3.4k
Merge medio
2g
PR unite (30g)
217

Descrizione

# Wire `_resolve_embeddings` into `_run_hybrid_search` (sync + async), plumb generator through dispatcher, and add diagnostics span

Parent: 46729
Depends on: 46732

## Goal

Make the resolver actually run on real queries, fail-fast when embeddings are needed but no generator is configured, and surface embedding-generation latency as a first-class OpenTelemetry span. (Diagnostics scope merged from #46734.)

## Scope

### Pipeline plumbing

1. zure/cosmos/_execution_context/execution_dispatcher.py :: _ProxyQueryExecutionContext:
- Accept mbedding_generator in its options dict.
- Pass it through to the hybrid-search aggregator.

2. hybrid_search_aggregator.py (sync) and io/hybrid_search_aggregator.py (async):
- In _run_hybrid_search / async sibling, after the plan is obtained and **before** per-partition fan-out, call _resolve_embeddings (from #46732) and replace the working SqlQuerySpec with the augmented one.
- **Fast-fail**: if the plan's mbeddingParameterMap is non-empty but mbedding_generator is None, raise a ValueError (or CosmosHttpResponseError 400) with:
> "Query requires embedding generation but no embedding_generator was provided to query_items."

3. Thread the generator from container.py / io/_container.py → xecution_dispatcher → aggregator (already started in #46730; confirm the option key name is mbedding_generator).

4. The resolver is called exactly **once per top-level query attempt**, not on every continuation page. Assert this with a counting mock in unit tests.

### Diagnostics (merged from #46734)

Inside _resolve_embeddings (and async sibling), wrap the generator call in an OpenTelemetry span:

`python
from azure.core.settings import settings
import time, logging

_logger = logging.getLogger("azure.cosmos")

tracer = settings.tracing_implementation()
if tracer is not None:
with tracer.span(name="cosmos.embedding_generation") as span:
span.add_attribute("cosmos.embedding.count", len(texts))
span.add_attribute("cosmos.embedding.generator_type", type(generator).__name__)
start = time.perf_counter()
vectors = generator.generate_embeddings(list(texts))
span.add_attribute("cosmos.embedding.latency_ms",
int((time.perf_counter() - start) * 1000))
else:
start = time.perf_counter()
vectors = generator.generate_embeddings(list(texts))
_logger.info(
"embedding_generation count=%d generator_type=%s latency_ms=%d",
len(texts), type(generator).__name__,
int((time.perf_counter() - start) * 1000),
)
`

Mirror in async aggregator (await the call inside the span).

**Privacy:**
- **Never** log raw input strings.
- **Never** log returned vectors.
- Parameter key names (e.g. @documentdb-hybridsearchquery-embedding-0) may be logged — they are not customer data.

If a tracer is configured, record a span event when the generator raises (so failures appear in traces, not just logs).

## Acceptance criteria

- End-to-end happy path works with a mocked plan + mocked generator (sync + async).
- Fast-fail exception raised with the right message when generator is missing.
- Generator called exactly once per top-level query attempt (assert with a counting mock).
- With an OTel tracer, a cosmos.embedding_generation span is produced with count, latency_ms, and generator_type attributes.
- Without a tracer, an INFO-level log line is produced.
- Raw text / vectors do not appear in span attributes or log output (assert in tests).

## Files likely touched

- sdk/cosmos/azure-cosmos/azure/cosmos/_execution_context/execution_dispatcher.py
- sdk/cosmos/azure-cosmos/azure/cosmos/_execution_context/aio/execution_dispatcher.py
- sdk/cosmos/azure-cosmos/azure/cosmos/_execution_context/hybrid_search_aggregator.py
- sdk/cosmos/azure-cosmos/azure/cosmos/_execution_context/aio/hybrid_search_aggregator.py
- sdk/cosmos/azure-cosmos/azure/cosmos/container.py (option threading)
- sdk/cosmos/azure-cosmos/azure/cosmos/aio/_container.py (option threading)

## Dependencies

- `_resolve_embeddings` helper (#46732).

Guida per i contributori

Apri la guida per i contributori

Direzione di ricerca

Inizia da _resolve_embeddings di #46732 e dagli entry point sincroni e asincroni di _run_hybrid_search, quindi segui embedding_generator attraverso i dispatcher di esecuzione e i file dei container elencati nell’issue. Aggiungi o esegui unit test per il caso positivo, il fallimento rapido, l’invocazione singola, il tracing, il logging e i criteri di privacy; il lavoro è completo quando tutti i criteri di accettazione sincroni e asincroni sono soddisfatti.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
python
Ambito
backend-api-design, databases
Tipo di issue
Funzionalità
Difficoltà
4/5
Tempo stimato
3-5 giorni
Stato di attività
Tranquilla
Chiarezza
Specificata chiaramente
Idoneità per principianti
48/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.