api_server: unbounded in-memory trace retention (ApiServerSpanExporter.trace_dict / InMemoryExporter._spans) with web=False — ~200 KB leaked per LLM turn, OOM under min-instances >= 1
- Langage dominant
- Python
- Étoiles
- 21.5k
- Forks
- 4k
- Merge moyen
- 1 j 14 h
- PR mergées (30 j)
- 37
Description
## Summary
`get_fast_api_app()` unconditionally installs two internal span exporters that retain trace data in process memory **without any bound and with no way to disable them** — including with `web=False`, i.e. the production serving path. In a long-lived process (Cloud Run with `min-instances >= 1`, GKE, any always-on deployment) this is a memory leak that grows linearly with traffic until the container is OOM-killed.
We measured **~196 KB retained per LLM turn** in a production chatbot, causing an OOM kill (`Out-of-memory event detected in container`) every ~6 days, with a ~10-minute window of user-facing 503/504 errors before each kill.
## Environment
- google-adk **2.7.0** (also verified present in earlier 2.x; the same classes exist since 1.x `fast_api.py`)
- Python 3.13 · Cloud Run gen2 · `min-instances: 1` · 2 GiB
- App built via `get_fast_api_app(agents_dir=..., web=False, session_service_uri="postgresql+asyncpg://...", otel_to_cloud=True)`
## Root cause
`src/google/adk/cli/api_server.py` (line numbers from 2.7.0):
1. **`ApiServerSpanExporter.export()` — line 472:**
```python
self.trace_dict[attributes["gcp.vertex.agent.event_id"]] = attributes
```
For every `call_llm` / `execute_tool` span, the full attribute dict is stored — including `gcp.vertex.agent.llm_request` and `gcp.vertex.agent.llm_response`, i.e. the **entire serialized prompt, history, and tool context**. Entries are never evicted. With RAG-sized system prompts this is easily 50–200 KB per model call.
2. **`InMemoryExporter.export()` — line 500:**
```python
self._spans.extend(spans)
```
**Every span in the process** is appended to an unbounded list. Its `clear()` is only called from the eval flow, which never runs in a production server.
3. Both are installed **unconditionally** in `AdkWebServer.get_app()` — line 1127:
```python
_setup_telemetry(
otel_to_cloud=otel_to_cloud,
internal_exporters=[
export_lib.SimpleSpanProcessor(ApiServerSpanExporter(trace_dict)),
export_lib.SimpleSpanProcessor(memory_exporter),
],
)
```
`get_fast_api_app()` exposes no flag to disable them.
## Why this looks like an oversight rather than a design choice
In 2.7 the debug endpoints that *read* these buffers (`/debug/trace/{event_id}`, `/debug/trace/session/{session_id}`) were moved to `dev_server.py` and are only mounted with `web=True`. The exporters that *feed* them, however, are still installed unconditionally in `api_server.py`. So with `web=False` a production process pays unbounded memory for buffers **whose only readers are not even mounted**. The reader was gated; the writer was not.
(Related report of the same "dev default leaks into the prod path" class: #3251, `InMemoryMemoryService` always instantiated in `api_server`.)
## Production measurements
| Signal | Measurement | Rules out |
|---|---|---|
| Hours with <20 requests | **+0.0 MiB/h** (n=29) | background thread / slow GC |
| Day 1 · 1,882 LLM turns | +351 MiB ⇒ **~196 KB/turn** | — |
| Day 2 · 775 LLM turns | +148 MiB ⇒ **~200 KB/turn** | same per-turn cost at very different volume ⇒ linear retention |
| Container | OOM-killed 4× in 60 days | — |
Retention is purely traffic-proportional. Memory is perfectly flat during idle hours and never shrinks.
## Minimal reproduction
```python
from google.adk.cli.fast_api import get_fast_api_app
from opentelemetry import trace
app = get_fast_api_app(agents_dir="agents", web=False)
tp = trace.get_tracer_provider()
procs = tp._active_span_processor._span_processors
print([type(p.span_exporter).__name__ for p in procs])
# ['ApiServerSpanExporter', 'InMemoryExporter'] <- installed with web=False
# Any traffic now grows ApiServerSpanExporter.trace_dict and
# InMemoryExporter._spans for the lifetime of the process.
```
## Expected behavior
One of (in order of preference):
1. **Don't install the internal exporters when `web=False`** — symmetric with the 2.7 gating of the debug endpoints that consume them.
2. Bound the containers (`deque(maxlen=...)` / LRU eviction on `trace_dict`), so retention is finite by construction even in dev.
3. Expose a flag on `get_fast_api_app()` (e.g. `enable_debug_trace_buffers: bool`) so servers can opt out.
## Workaround we're using
We register an extra `SpanProcessor` after `get_fast_api_app()` that trims those same containers in place (FIFO) after ADK's processors have written. It works, but it depends on private structure (`_active_span_processor._span_processors`, `trace_dict`, `_spans`), so it will break silently-adjacent on any rename — hence this report.
Guide de contribution
Ouvrir le guide de contribution
Évaluation
Cette issue n'a pas encore été évaluée.