modelcontextprotocol / modelcontextprotocol/python-sdk
mcp.server.fastmcp.FastMCP HTTP transport deadlocks after ~5 sequential client sessions on Windows (companion to PrefectHQ/fastmcp#4192)
还没有人认领这个 Issue。
- 主要语言
- Python
- 星标
- 24.3k
- 派生
- 4k
- 平均合并
- 1 天 1 小时
- 30 天内合并 PR
- 31
描述
mcp.server.fastmcp.FastMCP HTTP transport deadlocks after ~5 sequential client sessions on Windows (companion to PrefectHQ/fastmcp#4192)
Summary
Filing here as a companion to PrefectHQ/fastmcp#4192, which contains the full investigation, cross-control matrix, stack-dump evidence, and minimal reproducers. The reason for the companion issue: the bug reproduces on Anthropic's mcp.server.fastmcp.FastMCP shipped inside this SDK, faster than on PrefectHQ's downstream fork, and the SDK maintainers should see the signal directly rather than having to chase a downstream cross-reference.
The one-line result that's yours
A 30-sequential-session driver running against an HTTP server built with mcp.server.fastmcp.FastMCP 1.27.1 on Windows + Python 3.14 hangs at session #5, with gradual slowdown visible from session #4 onwards. The same driver against PrefectHQ fastmcp.FastMCP 3.3.1 hangs at session #12; against raw mcp.server.lowlevel.Server it passes 30/30; against the Node.js TypeScript reference server (@modelcontextprotocol/server-everything) it passes 30/30.
That is: this SDK's FastMCP wrapper exhibits the defect more aggressively than the downstream fork, and the SDK's own lower-level surface is unaffected.
Cross-control matrix (excerpt; full matrix in PrefectHQ/fastmcp#4192)
| Server | Python | Sessions before hang |
|---|---|---|
mcp.server.fastmcp.FastMCP 1.27.1 |
3.14 | 5 |
mcp.server.fastmcp.FastMCP 1.27.1 |
3.13.7 | 12 |
fastmcp.FastMCP 3.3.1 (PrefectHQ) |
3.14 | 12 |
mcp.server.lowlevel.Server + StreamableHTTPSessionManager |
3.14 | 30/30 ✅ |
@modelcontextprotocol/server-everything (Node.js streamableHttp) |
n/a | 30/30 ✅ |
Python-version variance rules out a 3.14 regression. The Node.js control rules out a protocol-level defect. The raw-SDK control rules out the SDK's lower-level transport.
Mechanism (per stack-dump evidence in #4192)
python -m asyncio ps <PID> on the hung server process shows multiple Task entries (one per "closed" session) parked in TaskGroup.__aexit__, all with child tasks named sse_starlette.sse.EventSourceResponse.__call__.<locals>.cancel_on_finish blocked in Event.wait / sleep on _ping, _listen_for_disconnect, _stream_response, and the SDK's standalone_sse_writer.
mcp.server.fastmcp.FastMCP.create_streamable_http_app instantiates EventSourceResponse from sse_starlette as part of its ASGI app construction. The raw mcp.server.lowlevel.Server path writes text/event-stream directly via its own ASGI write loop without ever instantiating EventSourceResponse — which is why it isn't affected. The defect is in the wrapper layer's interaction with sse_starlette, not the SDK's core transport.
Reproducers
Two minimal reproducers in the PrefectHQ issue: one for PrefectHQ FastMCP, one for raw SDK (the passing control). A third targeting mcp.server.fastmcp.FastMCP directly is trivially derivable — same shape, different import — and I'm happy to add it here if useful.
Why the cap is lower on this SDK's wrapper
Working hypothesis (not yet verified): mcp.server.fastmcp.FastMCP and PrefectHQ's fastmcp.FastMCP use slightly different sse_starlette integration shapes. Either version-pin differences, or an extra layer of session-lifecycle handling in one wrapper that creates more orphan tasks per session, would explain why the threshold lands at 5 here vs 12 downstream. Confirming this would need a side-by-side diff of the two wrappers' create_streamable_http_app paths and sse_starlette versions, which I haven't done.
Platform / environment
- Windows 11 (
ProactorEventLoop) - Python 3.14 and 3.13.7
mcp1.27.1 (the SDK that shipsmcp.server.fastmcp.FastMCP)- Linux not validated — PrefectHQ's public CI runs Linux + Python 3.12 where
epoll-based loop semantics may drain orphan tasks faster. Maintainer Linux validation requested in #4192; same request stands here.
The question for SDK maintainers
Is mcp.server.fastmcp.FastMCP considered an actively-maintained surface that the SDK team will fix, or is it a thin reference wrapper where PrefectHQ's fork is the canonical production target? The answer changes where the fix should land and whether this issue stays open here or gets closed in favour of #4192.
If it's in-scope here, the cross-controls in #4192 already narrow the surface to "EventSourceResponse integration in the FastMCP wrapper layer." The raw SDK path is the working reference.
Workaround for users of this SDK
Migrate from mcp.server.fastmcp.FastMCP to mcp.server.lowlevel.Server + manually-constructed StreamableHTTPSessionManager + Starlette + uvicorn. ~1 day's mechanical rewrite for a server with ~7 tools. Empirically 30/30 sessions passing on the same workload.
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
调研方向
从 mcp.server.fastmcp.FastMCP.create_streamable_http_app 开始,检查其与 EventSourceResponse 的集成,并将原始的 mcp.server.lowlevel.Server 路径作为正常工作的对照。复现 Windows 上顺序会话挂起的问题,将 wrapper 的行为与 PrefectHQ/fastmcp#4192 中的证据进行比较;当关闭的会话在重复运行多个会话后不再留下孤儿任务或导致死锁时,即可认为问题已解决。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- python
- 领域
- api, backend, networking
- Issue 类型
- 缺陷
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 活跃度
- 冷清
- 描述清晰度
- 基本清楚
- 新手友好度
- 42/100