modelcontextprotocol / modelcontextprotocol/python-sdk

[v2] JSON-RPC error responses leave client OpenTelemetry spans UNSET

未关闭
#3,174 9 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

bug needs confirmation P2 v2
主要语言
Python
星标
24.3k
派生
4k
平均合并
1 天 1 小时
30 天内合并 PR
31

描述

Initial Checks
Description

At current main (11934c90aeff5e1e68aee223edd00c5d1fce1d5c), JSON-RPC error responses handled by JSONRPCDispatcher leave the corresponding client OpenTelemetry span looking like a normal exit:

  • span status: UNSET
  • error.type: absent
  • rpc.response.status_code: absent
  • exception events: none

The request still fails correctly with MCPError; this report is limited to client-side telemetry correctness.

The ordering appears to explain the result: send_raw_request() receives outcome inside the client span, exits the span, and only afterward converts ErrorData into MCPError at lines 432–433. The span context manager therefore observes no exception.

This differs from the current server OTel middleware, which records the error attributes and status, and from the design assumption documented in #2381 that a propagated MCPError would be recorded by the client span.

I ran the same in-memory success/error pair five times. Every run produced:

METHOD=resources/list CONTROL=UNSET RED_CODE=-32602 RED_MESSAGE='forced failure' RED_STATUS=UNSET ERROR_TYPE=None RPC_STATUS=None EVENTS=0

Expected: the failed client span should be distinguishable from the successful control and carry the JSON-RPC error code/status consistently with the SDK's server-side instrumentation.

Scope and claim boundary: this concerns JSON-RPC/stream-backed dispatcher calls, not direct in-process dispatch generally. The reproducer is entirely in-memory. I have not tested or claimed anything about a deployed transport, collector, or production frequency.

Closest adjacent work found: #2132 targets the removed BaseSession architecture and predates the current dispatcher; #2854 adds GenAI attributes but does not handle JSON-RPC errors. I did not find an issue or current PR covering this ordering on the present dispatcher.

Disclosure: I used AI assistance to help reproduce and analyze this issue. I reviewed the source, reran the evidence independently, and take responsibility for the report.

Example Code
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter
from opentelemetry.trace import SpanKind

exporter = InMemorySpanExporter()
provider = TracerProvider()
provider.add_span_processor(SimpleSpanProcessor(exporter))
trace.set_tracer_provider(provider)

import anyio
from mcp.client.client import Client
from mcp.server.lowlevel.server import Server
from mcp.shared.exceptions import MCPError
from mcp_types import INVALID_PARAMS, ListResourcesResult

state = {"fail": False}

async def list_resources(ctx, params):
    if state["fail"]:
        raise MCPError(INVALID_PARAMS, "forced failure")
    return ListResourcesResult(resources=[])

async def main():
    server = Server(name="otel-audit", version="0.0.0", on_list_resources=list_resources)
    async with Client(server, mode="legacy", cache=None) as client:
        state["fail"] = False
        exporter.clear()
        await client.session.list_resources()
        [control] = [s for s in exporter.get_finished_spans() if s.kind == SpanKind.CLIENT]

        state["fail"] = True
        exporter.clear()
        try:
            await client.session.list_resources()
        except MCPError as exc:
            assert exc.error.code == INVALID_PARAMS
        else:
            raise AssertionError("expected MCPError")

        [failed] = [s for s in exporter.get_finished_spans() if s.kind == SpanKind.CLIENT]
        print(
            "control=", control.status.status_code.name,
            "failed=", failed.status.status_code.name,
            "error.type=", failed.attributes.get("error.type"),
            "rpc.response.status_code=", failed.attributes.get("rpc.response.status_code"),
            "events=", len(failed.events),
        )
        assert control.name == failed.name == "MCP send resources/list"
        assert failed.status.status_code.name == "UNSET"
        assert failed.attributes.get("error.type") is None
        assert failed.attributes.get("rpc.response.status_code") is None
        assert len(failed.events) == 0

anyio.run(main)
Python & MCP Python SDK
Python: 3.12.10
mcp: 2.0.0b2.dev33+11934c90
opentelemetry-sdk: 1.39.1
OS: Windows 11

Clean checkout command:
uv run --python 3.12 --frozen --no-default-groups --with opentelemetry-sdk==1.39.1 python repro.py

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

调研方向

从 src/mcp/shared/jsonrpc_dispatcher.py 中的 send_raw_request() 开始,重点检查 span 作用域以及从 ErrorData 到 MCPError 的转换,然后比较 src/mcp/server/_otel.py 中的属性和状态处理。运行提供的内存中成功/错误复现程序,并验证失败的客户端 span 能够与成功的 span 区分开来,且包含 JSON-RPC 错误信息。

由索引模型根据 Issue 内容生成。

评估

技术栈
python
领域
backend-api-design, observability
Issue 类型
缺陷
难度
3/5
预计耗时
1-2 天
活跃度
活跃
描述清晰度
描述清楚
新手友好度
76/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。