anthropics / anthropics/claude-agent-sdk-python

Per-message trace-context propagation for persistent sessions — TRACEPARENT is injected only at connect(), so all MCP/tool spans attribute to the first turn's trace

Open
#1,126 0 comments 0 reactions 0 assignees View on GitHub
documentation enhancement
Dominant language
Python
Stars
8.1k
Forks
1.3k
Avg merge
2d 31m
Merged PRs (30d)
1

Description

## Feature request

A per-message mechanism to update the CLI subprocess's W3C trace context for long-lived `ClaudeSDKClient` sessions — today `TRACEPARENT`/`TRACESTATE` are injected into the child process environment **once, at `connect()`**, and never again.

## Current behavior

`_internal/transport/subprocess_cli.py` (v0.2.121, ~L504–528) reads the ambient OpenTelemetry span context at transport spawn and injects `TRACEPARENT`/`TRACESTATE` into the CLI child-process env. The CLI honors it (with `CLAUDE_CODE_ENABLE_TELEMETRY=1` + `CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1`, plus `CLAUDE_CODE_PROPAGATE_TRACEPARENT=1` when a custom base URL is in use) and stamps outbound MCP HTTP tool calls with `traceparent` — with **fresh span IDs per tool call**, but always the **spawn-time trace ID**, for the lifetime of the process.

For a persistent session serving many user turns, the host application typically starts a new trace per turn (the turn-per-trace model used by LLM observability backends such as Langfuse). The host-side spans land in each turn's trace correctly, but every MCP/tool span from every turn attributes to the **first turn's** trace. Effects:

- The first turn's trace accumulates the whole session's tool activity; its trace-window duration reads minutes for a seconds-long turn.
- Turns ≥2 appear to make no tool calls at all in their own traces.

We audited the control-protocol message subtypes in v0.2.121 and found no field that carries trace context on the streaming input envelope, so there is no host-side workaround: `options.env` is spawn-time-only, and the context cannot cross the SDK boundary afterward.

## Reproduction sketch

1. Streamable-HTTP MCP server (e.g. FastMCP) with one trivial tool, wrapped in ASGI middleware that logs the inbound `traceparent` header per request.
2. Host: OTel `TracerProvider` + any instrumentation that opens a span around each `query()` (or open one manually).
3. `ClaudeSDKClient` with telemetry env: `CLAUDE_CODE_ENABLE_TELEMETRY=1`, `CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1`, `CLAUDE_CODE_PROPAGATE_TRACEPARENT=1`, `OTEL_TRACES_EXPORTER=otlp`, `OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf`, `OTEL_EXPORTER_OTLP_ENDPOINT=`.
4. Send two prompts on the same client, each inside a **different** host span (different trace IDs), each forcing a tool call.
5. Observe: both tool calls' `traceparent` headers carry the trace ID that was ambient at `connect()`; the second turn's new trace ID never reaches the wire.

## Ask

Any of these would solve it (in our order of preference):

1. **Re-read the ambient OTel context per `query()`** and forward it to the CLI (control protocol or input envelope), so the CLI parents its per-interaction spans — and stamps outbound MCP calls — under the caller's current trace.
2. **Accept trace context on the streaming input message** (an explicit field the host sets per message).
3. Expose a documented control-protocol request to update `TRACEPARENT` mid-session.

Related work, for context: anthropics/claude-code#31039 (the spawn-time injection feature as shipped), anthropics/claude-code#76391 (SEP-414 `_meta` trace-context propagation for MCP — complementary: it addresses the MCP hop, but the CLI-side parent context would still be spawn-frozen), and #952 in this repo (adjacent).

## Side observation (worth a docs note even if the feature is declined)

`OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf` is load-bearing for the CLI's tracing: with an HTTP OTLP endpoint configured but the protocol var unset, the exporter apparently defaults to gRPC, fails to construct, and the entire telemetry subsystem — span creation included, hence also traceparent stamping on MCP calls — dies silently. A one-line note in the observability docs ("set the protocol explicitly") would save others the bisection.

## Environment

- `claude-agent-sdk==0.2.121` (Python), macOS + Linux (containerized)
- Bedrock mode (`CLAUDE_CODE_USE_BEDROCK=1`, custom `ANTHROPIC_BEDROCK_BASE_URL`)
- Streamable-HTTP MCP server (FastMCP)
- Verified against 0.1.48 as well: that CLI generation does not stamp `traceparent` on MCP calls at all, even with telemetry fully active — the behavior described above is the current (0.2.x) one.

## Our mitigation meanwhile

Server-side: MCP tool spans are detached into their own traces and joined to the calling turn via the `claudecode/toolUseId` the CLI already sends in `_meta` (thank you for that — it's the only per-call key that crosses the boundary today, and it works). The feature above would let those spans nest under the correct turn natively.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reading _internal/transport/subprocess_cli.py around the spawn-time TRACEPARENT/TRACESTATE injection noted in the issue, then reproduce with a persistent ClaudeSDKClient and a FastMCP server logging inbound traceparent headers. The issue is done only when a maintainer-approved per-message trace-context path exists and two queries in different host traces produce MCP/tool spans under their respective trace IDs.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
cli, observability
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
28/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.