modelcontextprotocol / modelcontextprotocol/python-sdk

streamable-http client: no size bound before JSONRPCMessage.model_validate_json — one large server message can OOM the client

Offen
#3,330 2 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Dieses Issue hat noch niemand übernommen.

v1 v2
Vorherrschende Sprache
Python
Sterne
24.3k
Forks
4k
Ø Merge
1 T. 1 Std.
Gemergte PRs (30 T.)
31

Beschreibung

Summary

StreamableHTTPTransport parses every inbound SSE event and JSON response body with JSONRPCMessage.model_validate_json / model_validate_json(content) and no size bound. Because pydantic validation of a large JSON document allocates several times the wire size in live Python objects, a single oversized message from a server can exhaust the client's memory. There is no hook to inspect or reject a message before it is parsed.

We hit this in production: a client connected to ~30 MCP servers, one of which answered tools/list with a 21.9 MB single SSE event. Measured with tracemalloc at the parse site:

#1  663.4 MB in 9,957,044 blocks
    mcp/client/streamable_http.py:217
      message = JSONRPCMessage.model_validate_json(sse.data)
#2  46.9 MB in 3 blocks
    httpx_sse/_decoders.py:61   (whole event buffered as one string)
#3  46.9 MB in 2 blocks
    httpx_sse/_decoders.py:120  (copy from slicing)

Roughly a 7× amplification of the wire size in live objects, transient but concurrent with other parses. A trivial request that triggered only catalogue loading peaked at 2.35 GB RSS from a 56 MB baseline; heavier concurrent work reached 4.2 GB. Removing that one server dropped the same request's peak to 456 MB. Versions: mcp 1.27.1, Python 3.14.

Why a client-side bound is needed

The client cannot know in advance that a server will return a huge payload, and a misbehaving or misconfigured server should not be able to OOM its client. Today the only outcome is process death with nothing identifying the responsible server — the failure surfaces as an unexplained kill rather than an actionable error.

What we did as a workaround, and why it was awkward

We wrapped StreamableHTTPTransport._handle_sse_event and _handle_json_response to measure sse.data / the response body and reject anything over a configured cap before parsing. Two things made this harder than expected, and both seem worth addressing upstream:

  1. Raising from the handler is not viable. Every call site wraps it in except Exception: logger.debug(...) and then reconnects with Last-Event-ID, which replays the same oversized payload. The raise is swallowed and the pending request hangs.
  2. Delivering a bare Exception on the read stream does not fail the request either. In BaseSession._receive_loop, isinstance(message, Exception) routes to _handle_incoming, which for the default message handler is a no-op; only a JSONRPCResponse/JSONRPCError matching _response_streams[request_id] completes send_request, and that awaits with timeout=None. To fail the request we had to synthesize a JSONRPCError and recover the request id from the raw payload with a regex, since parsing it is exactly what we were trying to avoid.
Suggested improvements
  • An optional client-side maximum message size (constructor argument and/or environment variable) enforced before model_validate_json, on both the SSE and application/json paths.
  • When it trips, fail the corresponding pending request with a distinct error identifying the endpoint and the observed size, rather than letting the reconnect path replay the payload.
  • Failing that, a documented hook to inspect a raw message before parsing would let clients implement this without patching private methods.
  • Independently: _handle_sse_event returning False for a parse failure causes the caller to treat the stream as ended and reconnect with Last-Event-ID; for a deterministic failure (such as a payload that will always be too large, or malformed JSON) this retries something that cannot succeed.

Happy to open a PR if a maintainer indicates the preferred shape (constructor arg vs. env var vs. pre-parse hook).

Beitragsleitfaden

Beitragsleitfaden öffnen

Erste Schritte

  1. Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
  3. Forke das Repository und arbeite in einem Branch.
  4. Öffne einen Pull Request, der die Issue-Nummer nennt.

Rechercherichtung

Beginne in mcp/client/streamable_http.py an der SSE-Parsing-Stelle um Zeile 217 und untersuche _handle_sse_event und _handle_json_response für beide eingehenden Pfade. Lies dann BaseSession._receive_loop und _handle_incoming, um die Fehlerbehandlung für ausstehende Anfragen zu verstehen. Fertig ist die Aufgabe, wenn übergroße Nachrichten vor model_validate_json abgewiesen werden, die zugehörige Anfrage mit einem verwertbaren Fehler fehlschlägt und die Verbindung beim Wiederherstellen die Nutzlast nicht erneut abspielt.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
python
Bereich
api, networking
Issue-Typ
Bug
Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Aktivitätsstatus
Aktiv
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
48/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.