modelcontextprotocol / modelcontextprotocol/python-sdk

streamable-http client: no size bound before JSONRPCMessage.model_validate_json — one large server message can OOM the client

Ouverte
#3,330 2 commentaires 0 réactions 0 personnes assignées Voir sur GitHub

Personne n'a encore pris cette issue.

v1 v2
Langage dominant
Python
Étoiles
24.3k
Forks
4k
Merge moyen
1 j 1 h
PR mergées (30 j)
31

Description

Summary

StreamableHTTPTransport parses every inbound SSE event and JSON response body with JSONRPCMessage.model_validate_json / model_validate_json(content) and no size bound. Because pydantic validation of a large JSON document allocates several times the wire size in live Python objects, a single oversized message from a server can exhaust the client's memory. There is no hook to inspect or reject a message before it is parsed.

We hit this in production: a client connected to ~30 MCP servers, one of which answered tools/list with a 21.9 MB single SSE event. Measured with tracemalloc at the parse site:

#1  663.4 MB in 9,957,044 blocks
    mcp/client/streamable_http.py:217
      message = JSONRPCMessage.model_validate_json(sse.data)
#2  46.9 MB in 3 blocks
    httpx_sse/_decoders.py:61   (whole event buffered as one string)
#3  46.9 MB in 2 blocks
    httpx_sse/_decoders.py:120  (copy from slicing)

Roughly a 7× amplification of the wire size in live objects, transient but concurrent with other parses. A trivial request that triggered only catalogue loading peaked at 2.35 GB RSS from a 56 MB baseline; heavier concurrent work reached 4.2 GB. Removing that one server dropped the same request's peak to 456 MB. Versions: mcp 1.27.1, Python 3.14.

Why a client-side bound is needed

The client cannot know in advance that a server will return a huge payload, and a misbehaving or misconfigured server should not be able to OOM its client. Today the only outcome is process death with nothing identifying the responsible server — the failure surfaces as an unexplained kill rather than an actionable error.

What we did as a workaround, and why it was awkward

We wrapped StreamableHTTPTransport._handle_sse_event and _handle_json_response to measure sse.data / the response body and reject anything over a configured cap before parsing. Two things made this harder than expected, and both seem worth addressing upstream:

  1. Raising from the handler is not viable. Every call site wraps it in except Exception: logger.debug(...) and then reconnects with Last-Event-ID, which replays the same oversized payload. The raise is swallowed and the pending request hangs.
  2. Delivering a bare Exception on the read stream does not fail the request either. In BaseSession._receive_loop, isinstance(message, Exception) routes to _handle_incoming, which for the default message handler is a no-op; only a JSONRPCResponse/JSONRPCError matching _response_streams[request_id] completes send_request, and that awaits with timeout=None. To fail the request we had to synthesize a JSONRPCError and recover the request id from the raw payload with a regex, since parsing it is exactly what we were trying to avoid.
Suggested improvements
  • An optional client-side maximum message size (constructor argument and/or environment variable) enforced before model_validate_json, on both the SSE and application/json paths.
  • When it trips, fail the corresponding pending request with a distinct error identifying the endpoint and the observed size, rather than letting the reconnect path replay the payload.
  • Failing that, a documented hook to inspect a raw message before parsing would let clients implement this without patching private methods.
  • Independently: _handle_sse_event returning False for a parse failure causes the caller to treat the stream as ended and reconnect with Last-Event-ID; for a deterministic failure (such as a payload that will always be too large, or malformed JSON) this retries something that cannot succeed.

Happy to open a PR if a maintainer indicates the preferred shape (constructor arg vs. env var vs. pre-parse hook).

Guide de contribution

Ouvrir le guide de contribution

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Piste de recherche

Commencez dans mcp/client/streamable_http.py, au niveau du point d’analyse SSE vers la ligne 217, et examinez _handle_sse_event et _handle_json_response pour les deux chemins entrants. Lisez ensuite BaseSession._receive_loop et _handle_incoming afin de comprendre la gestion des échecs des requêtes en attente. C’est terminé lorsque les messages trop volumineux sont rejetés avant model_validate_json, que la requête correspondante échoue avec une erreur exploitable et que la reconnexion ne rejoue pas la charge utile.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
python
Domaine
api, networking
Type d'issue
Bug
Difficulté
4/5
Temps estimé
3-5 jours
Activité
Active
Clarté
Plutôt claire
Accessibilité débutants
48/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.