modelcontextprotocol / modelcontextprotocol/python-sdk

streamable-http client: no size bound before JSONRPCMessage.model_validate_json — one large server message can OOM the client

Aperta
#3,330 2 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

v1 v2
Lingua principale
Python
Stelle
24.3k
Fork
4k
Merge medio
1g 1h
PR unite (30g)
31

Descrizione

Summary

StreamableHTTPTransport parses every inbound SSE event and JSON response body with JSONRPCMessage.model_validate_json / model_validate_json(content) and no size bound. Because pydantic validation of a large JSON document allocates several times the wire size in live Python objects, a single oversized message from a server can exhaust the client's memory. There is no hook to inspect or reject a message before it is parsed.

We hit this in production: a client connected to ~30 MCP servers, one of which answered tools/list with a 21.9 MB single SSE event. Measured with tracemalloc at the parse site:

#1  663.4 MB in 9,957,044 blocks
    mcp/client/streamable_http.py:217
      message = JSONRPCMessage.model_validate_json(sse.data)
#2  46.9 MB in 3 blocks
    httpx_sse/_decoders.py:61   (whole event buffered as one string)
#3  46.9 MB in 2 blocks
    httpx_sse/_decoders.py:120  (copy from slicing)

Roughly a 7× amplification of the wire size in live objects, transient but concurrent with other parses. A trivial request that triggered only catalogue loading peaked at 2.35 GB RSS from a 56 MB baseline; heavier concurrent work reached 4.2 GB. Removing that one server dropped the same request's peak to 456 MB. Versions: mcp 1.27.1, Python 3.14.

Why a client-side bound is needed

The client cannot know in advance that a server will return a huge payload, and a misbehaving or misconfigured server should not be able to OOM its client. Today the only outcome is process death with nothing identifying the responsible server — the failure surfaces as an unexplained kill rather than an actionable error.

What we did as a workaround, and why it was awkward

We wrapped StreamableHTTPTransport._handle_sse_event and _handle_json_response to measure sse.data / the response body and reject anything over a configured cap before parsing. Two things made this harder than expected, and both seem worth addressing upstream:

  1. Raising from the handler is not viable. Every call site wraps it in except Exception: logger.debug(...) and then reconnects with Last-Event-ID, which replays the same oversized payload. The raise is swallowed and the pending request hangs.
  2. Delivering a bare Exception on the read stream does not fail the request either. In BaseSession._receive_loop, isinstance(message, Exception) routes to _handle_incoming, which for the default message handler is a no-op; only a JSONRPCResponse/JSONRPCError matching _response_streams[request_id] completes send_request, and that awaits with timeout=None. To fail the request we had to synthesize a JSONRPCError and recover the request id from the raw payload with a regex, since parsing it is exactly what we were trying to avoid.
Suggested improvements
  • An optional client-side maximum message size (constructor argument and/or environment variable) enforced before model_validate_json, on both the SSE and application/json paths.
  • When it trips, fail the corresponding pending request with a distinct error identifying the endpoint and the observed size, rather than letting the reconnect path replay the payload.
  • Failing that, a documented hook to inspect a raw message before parsing would let clients implement this without patching private methods.
  • Independently: _handle_sse_event returning False for a parse failure causes the caller to treat the stream as ended and reconnect with Last-Event-ID; for a deterministic failure (such as a payload that will always be too large, or malformed JSON) this retries something that cannot succeed.

Happy to open a PR if a maintainer indicates the preferred shape (constructor arg vs. env var vs. pre-parse hook).

Guida per i contributori

Apri la guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Direzione di ricerca

Inizia in mcp/client/streamable_http.py, nel punto di analisi SSE intorno alla riga 217, e ispeziona _handle_sse_event e _handle_json_response per entrambi i percorsi in ingresso. Poi leggi BaseSession._receive_loop e _handle_incoming per comprendere la gestione dei fallimenti delle richieste in sospeso. Il lavoro è completato quando i messaggi troppo grandi vengono rifiutati prima di model_validate_json, la richiesta corrispondente fallisce con un errore utile e la riconnessione non riproduce nuovamente il payload.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
python
Ambito
api, networking
Tipo di issue
Bug
Difficoltà
4/5
Tempo stimato
3-5 giorni
Stato di attività
Attiva
Chiarezza
Abbastanza chiara
Idoneità per principianti
48/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.