modelcontextprotocol / modelcontextprotocol/python-sdk

StreamableHTTP client: _handle_reconnection resets attempt counter to 0, causing infinite retry loop

未關閉 適合新手
#2,393 2 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

還沒有人認領這個 Issue。

bug P2
主要語言
Python
星號
24.3k
分支
4k
平均合併
1 天 1 小時
30 天內合併 PR
31

描述

Bug Description

_handle_reconnection() in streamable_http.py resets the attempt counter to 0 on line 494 when a reconnection succeeds but the stream ends without delivering a complete response. This makes MAX_RECONNECTION_ATTEMPTS ineffective — the counter only applies to consecutive exceptions, not total reconnection attempts. If the server accepts the connection but the stream drops repeatedly, the client retries forever.

Reproduction

MCP version: 1.26.0 (also confirmed unpatched in 1.27.0 main branch)

SSCE — Minimal reproducer

Server (server.py): A server that accepts SSE connections but closes them before sending a complete response, with a last-event-id header to trigger the reconnection path.

"""Minimal MCP server that drops SSE streams to trigger infinite reconnect."""
import asyncio
from starlette.applications import Starlette
from starlette.routing import Route
from starlette.responses import Response
from sse_starlette.sse import EventSourceResponse
import uvicorn, json, uuid

session_id = str(uuid.uuid4())

async def handle_mcp(request):
    if request.method == "POST":
        body = await request.json()
        method = body.get("method")

        if method == "initialize":
            resp = {
                "jsonrpc": "2.0",
                "id": body["id"],
                "result": {
                    "protocolVersion": "2025-06-18",
                    "capabilities": {"tools": {}},
                    "serverInfo": {"name": "drop-server", "version": "0.1.0"},
                },
            }
            return Response(
                json.dumps(resp),
                media_type="application/json",
                headers={"mcp-session-id": session_id},
            )

        if method == "notifications/initialized":
            return Response(status_code=202, headers={"mcp-session-id": session_id})

        if method == "tools/call":
            # Return SSE that sends a priming event with an ID, then drops
            async def event_generator():
                yield {"event": "message", "id": "evt-1", "data": ""}
                # Close without sending the actual response
                return

            return EventSourceResponse(
                event_generator(),
                headers={"mcp-session-id": session_id},
            )

        return Response(status_code=404)

    if request.method == "GET":
        # GET stream for server-initiated messages — also drop immediately
        async def get_stream():
            yield {"event": "message", "id": "evt-get-1", "data": ""}
            return

        return EventSourceResponse(
            get_stream(),
            headers={"mcp-session-id": session_id},
        )

    if request.method == "DELETE":
        return Response(status_code=200)

app = Starlette(routes=[Route("/mcp", handle_mcp, methods=["POST", "GET", "DELETE"])])

if __name__ == "__main__":
    uvicorn.run(app, host="0.0.0.0", port=9999)

Client (client.py): Connects and calls a tool, demonstrating the infinite loop.

"""Client that demonstrates infinite reconnection loop."""
import asyncio, logging
logging.basicConfig(level=logging.INFO)

from mcp import ClientSession
from mcp.client.streamable_http import streamable_http_client

async def main():
    async with streamable_http_client("http://localhost:9999/mcp") as (read, write, _):
        async with ClientSession(read, write) as session:
            await session.initialize()
            print("Initialized. Calling tool (will hang forever)...")
            # This call will trigger _handle_reconnection with attempt=0 reset
            try:
                result = await asyncio.wait_for(
                    session.call_tool("any_tool", {}),
                    timeout=30,
                )
            except asyncio.TimeoutError:
                print("CONFIRMED: call_tool hung for 30s (infinite reconnect loop)")

asyncio.run(main())
Steps
  1. pip install mcp[cli] sse-starlette uvicorn
  2. Run server: python server.py
  3. Run client: python client.py
  4. Observe: client logs show repeated "GET stream disconnected, reconnecting in 1000ms..." messages. The call_tool never returns. After 30s the wait_for timeout fires, confirming the hang.

Root Cause

In streamable_http.py, _handle_reconnection (line 437):

async def _handle_reconnection(self, ctx, last_event_id, retry_interval_ms, attempt=0):
    if attempt >= MAX_RECONNECTION_ATTEMPTS:  # Only 2
        return

    # ... reconnects, iterates SSE ...

    # Line 494: Stream ended without response — resets attempt to 0!
    await self._handle_reconnection(ctx, reconnect_last_event_id, reconnect_retry_ms, 0)

When the reconnection succeeds (HTTP 200) but the stream ends without a complete JSONRPCResponse, line 494 recurses with attempt=0, restarting the counter. Only the exception path (line 498) increments the counter. A server that accepts connections but drops streams causes infinite recursion at 1-second intervals.

Expected Behavior

After MAX_RECONNECTION_ATTEMPTS total reconnection attempts (regardless of whether they succeeded at the HTTP level), the client should give up and propagate an error to the caller.

Suggested Fix

Track total attempts across the recursion rather than resetting on successful connect:

# Line 494: increment instead of reset
await self._handle_reconnection(ctx, reconnect_last_event_id, reconnect_retry_ms, attempt + 1)

Or add a separate max_total_reconnection_attempts counter that is never reset.

Impact

In production, this causes MCP client coroutines to hang forever when a server experiences transient stream drops. The calling application has no way to recover without wrapping every MCP call in asyncio.wait_for(). We discovered this when an agentquant research job hung for 5+ hours in a reconnection loop after an MCP server's SSE stream dropped.

貢獻指南

開啟貢獻指南

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

研究方向

從 streamable_http.py 中的 _handle_reconnection 開始,追蹤成功重新連線路徑和例外路徑,重點關注 attempt 如何透過遞迴傳遞。執行提供的 server.py 和 client.py 重現程式;當重複出現的不完整串流在 MAX_RECONNECTION_ATTEMPTS 次嘗試後停止,且呼叫端收到錯誤而不是無限期掛起時,即表示完成。

由索引模型根據 Issue 內容生成。

評估

技術堆疊
python
領域
api
Issue 類型
缺陷
難度
2/5
預估耗時
1-3 小時
活躍度
冷清
描述清晰度
描述清楚
新手友好度
78/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。