modelcontextprotocol / modelcontextprotocol/python-sdk
streamable-http client: no size bound before JSONRPCMessage.model_validate_json — one large server message can OOM the client
Nessuno ha ancora preso questa issue.
- Lingua principale
- Python
- Stelle
- 24.3k
- Fork
- 4k
- Merge medio
- 1g 1h
- PR unite (30g)
- 31
Descrizione
Summary
StreamableHTTPTransport parses every inbound SSE event and JSON response body with JSONRPCMessage.model_validate_json / model_validate_json(content) and no size bound. Because pydantic validation of a large JSON document allocates several times the wire size in live Python objects, a single oversized message from a server can exhaust the client's memory. There is no hook to inspect or reject a message before it is parsed.
We hit this in production: a client connected to ~30 MCP servers, one of which answered tools/list with a 21.9 MB single SSE event. Measured with tracemalloc at the parse site:
#1 663.4 MB in 9,957,044 blocks
mcp/client/streamable_http.py:217
message = JSONRPCMessage.model_validate_json(sse.data)
#2 46.9 MB in 3 blocks
httpx_sse/_decoders.py:61 (whole event buffered as one string)
#3 46.9 MB in 2 blocks
httpx_sse/_decoders.py:120 (copy from slicing)
Roughly a 7× amplification of the wire size in live objects, transient but concurrent with other parses. A trivial request that triggered only catalogue loading peaked at 2.35 GB RSS from a 56 MB baseline; heavier concurrent work reached 4.2 GB. Removing that one server dropped the same request's peak to 456 MB. Versions: mcp 1.27.1, Python 3.14.
Why a client-side bound is needed
The client cannot know in advance that a server will return a huge payload, and a misbehaving or misconfigured server should not be able to OOM its client. Today the only outcome is process death with nothing identifying the responsible server — the failure surfaces as an unexplained kill rather than an actionable error.
What we did as a workaround, and why it was awkward
We wrapped StreamableHTTPTransport._handle_sse_event and _handle_json_response to measure sse.data / the response body and reject anything over a configured cap before parsing. Two things made this harder than expected, and both seem worth addressing upstream:
- Raising from the handler is not viable. Every call site wraps it in
except Exception: logger.debug(...)and then reconnects withLast-Event-ID, which replays the same oversized payload. The raise is swallowed and the pending request hangs. - Delivering a bare
Exceptionon the read stream does not fail the request either. InBaseSession._receive_loop,isinstance(message, Exception)routes to_handle_incoming, which for the default message handler is a no-op; only aJSONRPCResponse/JSONRPCErrormatching_response_streams[request_id]completessend_request, and that awaits withtimeout=None. To fail the request we had to synthesize aJSONRPCErrorand recover the request id from the raw payload with a regex, since parsing it is exactly what we were trying to avoid.
Suggested improvements
- An optional client-side maximum message size (constructor argument and/or environment variable) enforced before
model_validate_json, on both the SSE andapplication/jsonpaths. - When it trips, fail the corresponding pending request with a distinct error identifying the endpoint and the observed size, rather than letting the reconnect path replay the payload.
- Failing that, a documented hook to inspect a raw message before parsing would let clients implement this without patching private methods.
- Independently:
_handle_sse_eventreturningFalsefor a parse failure causes the caller to treat the stream as ended and reconnect withLast-Event-ID; for a deterministic failure (such as a payload that will always be too large, or malformed JSON) this retries something that cannot succeed.
Happy to open a PR if a maintainer indicates the preferred shape (constructor arg vs. env var vs. pre-parse hook).
Guida per i contributori
Apri la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Direzione di ricerca
Inizia in mcp/client/streamable_http.py, nel punto di analisi SSE intorno alla riga 217, e ispeziona _handle_sse_event e _handle_json_response per entrambi i percorsi in ingresso. Poi leggi BaseSession._receive_loop e _handle_incoming per comprendere la gestione dei fallimenti delle richieste in sospeso. Il lavoro è completato quando i messaggi troppo grandi vengono rifiutati prima di model_validate_json, la richiesta corrispondente fallisce con un errore utile e la riconnessione non riproduce nuovamente il payload.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- python
- Ambito
- api, networking
- Tipo di issue
- Bug
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Stato di attività
- Attiva
- Chiarezza
- Abbastanza chiara
- Idoneità per principianti
- 48/100