google / google/adk-python

api_server: unbounded in-memory trace retention (ApiServerSpanExporter.trace_dict / InMemoryExporter._spans) with web=False — ~200 KB leaked per LLM turn, OOM under min-instances >= 1

Đang mở
#6,915 2 bình luận 0 reaction 1 người được giao Được @sanketpatil06 nhận Xem trên GitHub
request clarification tracing
Ngôn ngữ chính
Python
Star
21.5k
Fork
4k
Merge trung bình
1 ngày 14 giờ
Pull request đã merge (30 ngày)
37

Mô tả

## Summary

`get_fast_api_app()` unconditionally installs two internal span exporters that retain trace data in process memory **without any bound and with no way to disable them** — including with `web=False`, i.e. the production serving path. In a long-lived process (Cloud Run with `min-instances >= 1`, GKE, any always-on deployment) this is a memory leak that grows linearly with traffic until the container is OOM-killed.

We measured **~196 KB retained per LLM turn** in a production chatbot, causing an OOM kill (`Out-of-memory event detected in container`) every ~6 days, with a ~10-minute window of user-facing 503/504 errors before each kill.

## Environment

- google-adk **2.7.0** (also verified present in earlier 2.x; the same classes exist since 1.x `fast_api.py`)
- Python 3.13 · Cloud Run gen2 · `min-instances: 1` · 2 GiB
- App built via `get_fast_api_app(agents_dir=..., web=False, session_service_uri="postgresql+asyncpg://...", otel_to_cloud=True)`

## Root cause

`src/google/adk/cli/api_server.py` (line numbers from 2.7.0):

1. **`ApiServerSpanExporter.export()` — line 472:**

```python
self.trace_dict[attributes["gcp.vertex.agent.event_id"]] = attributes
```

For every `call_llm` / `execute_tool` span, the full attribute dict is stored — including `gcp.vertex.agent.llm_request` and `gcp.vertex.agent.llm_response`, i.e. the **entire serialized prompt, history, and tool context**. Entries are never evicted. With RAG-sized system prompts this is easily 50–200 KB per model call.

2. **`InMemoryExporter.export()` — line 500:**

```python
self._spans.extend(spans)
```

**Every span in the process** is appended to an unbounded list. Its `clear()` is only called from the eval flow, which never runs in a production server.

3. Both are installed **unconditionally** in `AdkWebServer.get_app()` — line 1127:

```python
_setup_telemetry(
otel_to_cloud=otel_to_cloud,
internal_exporters=[
export_lib.SimpleSpanProcessor(ApiServerSpanExporter(trace_dict)),
export_lib.SimpleSpanProcessor(memory_exporter),
],
)
```

`get_fast_api_app()` exposes no flag to disable them.

## Why this looks like an oversight rather than a design choice

In 2.7 the debug endpoints that *read* these buffers (`/debug/trace/{event_id}`, `/debug/trace/session/{session_id}`) were moved to `dev_server.py` and are only mounted with `web=True`. The exporters that *feed* them, however, are still installed unconditionally in `api_server.py`. So with `web=False` a production process pays unbounded memory for buffers **whose only readers are not even mounted**. The reader was gated; the writer was not.

(Related report of the same "dev default leaks into the prod path" class: #3251, `InMemoryMemoryService` always instantiated in `api_server`.)

## Production measurements

| Signal | Measurement | Rules out |
|---|---|---|
| Hours with <20 requests | **+0.0 MiB/h** (n=29) | background thread / slow GC |
| Day 1 · 1,882 LLM turns | +351 MiB ⇒ **~196 KB/turn** | — |
| Day 2 · 775 LLM turns | +148 MiB ⇒ **~200 KB/turn** | same per-turn cost at very different volume ⇒ linear retention |
| Container | OOM-killed 4× in 60 days | — |

Retention is purely traffic-proportional. Memory is perfectly flat during idle hours and never shrinks.

## Minimal reproduction

```python
from google.adk.cli.fast_api import get_fast_api_app
from opentelemetry import trace

app = get_fast_api_app(agents_dir="agents", web=False)

tp = trace.get_tracer_provider()
procs = tp._active_span_processor._span_processors
print([type(p.span_exporter).__name__ for p in procs])
# ['ApiServerSpanExporter', 'InMemoryExporter'] <- installed with web=False

# Any traffic now grows ApiServerSpanExporter.trace_dict and
# InMemoryExporter._spans for the lifetime of the process.
```

## Expected behavior

One of (in order of preference):

1. **Don't install the internal exporters when `web=False`** — symmetric with the 2.7 gating of the debug endpoints that consume them.
2. Bound the containers (`deque(maxlen=...)` / LRU eviction on `trace_dict`), so retention is finite by construction even in dev.
3. Expose a flag on `get_fast_api_app()` (e.g. `enable_debug_trace_buffers: bool`) so servers can opt out.

## Workaround we're using

We register an extra `SpanProcessor` after `get_fast_api_app()` that trims those same containers in place (FIFO) after ADK's processors have written. It works, but it depends on private structure (`_active_span_processor._span_processors`, `trace_dict`, `_spans`), so it will break silently-adjacent on any rename — hence this report.

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.