modelcontextprotocol / modelcontextprotocol/python-sdk
OAuthClientProvider auth lock is permanently poisoned when httpx closes async_auth_flow from a different task (RuntimeError: The current task is not holding this lock)
還沒有人認領這個 Issue。
- 主要語言
- Python
- 星號
- 24.3k
- 分支
- 4k
- 平均合併
- 1 天 1 小時
- 30 天內合併 PR
- 31
描述
Summary
OAuthClientProvider.async_auth_flow holds self.context.lock (an anyio.Lock) across the entire httpx auth-flow generator, including every yield (src/mcp/client/auth/oauth2.py:582 in v2.0.0; same code is present on v2.1.0 and main). anyio.Lock.release() is bound to the acquiring task. When httpx closes the auth generator from a different task than the one that advanced it — which happens routinely when a request is cancelled mid-flight (network drop, timeout, task-group teardown) — the async with __aexit__ runs in the closing task and raises:
RuntimeError: The current task is not holding this lock
The exception escapes into a fire-and-forget teardown task ("Task exception was never retrieved"), and the lock is left permanently held. Every subsequent request through the same OAuthClientProvider then blocks forever at async with self.context.lock: without sending any HTTP. For a long-lived client that reuses the provider across reconnects, that server is dead until the whole process restarts.
Environment
mcp2.0.0 (reproduced; the sameasync with self.context.lock:pattern is unchanged in 2.1.0 and on main)- Python 3.11.15, macOS (darwin), anyio 4.14.2, asyncio backend
- Long-lived agent process (Nous Research Hermes) with OAuth-backed streamable-HTTP MCP servers; surfaced downstream as NousResearch/hermes-agent#81051
Observed traceback
Two captures from the same host, one per yield point (authorization-code exchange and the main yield request):
ERROR asyncio: Task exception was never retrieved
future: <Task finished name='Task-598' coro=<<async_generator_athrow without __name__>()> exception=RuntimeError('The current task is not holding this lock')>
Traceback (most recent call last):
File ".../site-packages/mcp/client/auth/oauth2.py", line 754, in async_auth_flow
yield request
GeneratorExit
During handling of the above exception, another exception occurred:
Traceback (most recent call last):
File ".../site-packages/mcp/client/auth/oauth2.py", line 582, in async_auth_flow
async with self.context.lock:
File ".../site-packages/anyio/_core/_synchronization.py", line 173, in __aexit__
self.release()
File ".../site-packages/anyio/_backends/_asyncio.py", line 1935, in release
raise RuntimeError("The current task is not holding this lock")
RuntimeError: The current task is not holding this lock
(The other capture is identical except the GeneratorExit lands at line 746, token_response = yield await self._perform_authorization().)
After this fires, every reconnect attempt for that server times out with no HTTP traffic — the flow generator never gets past line 582.
Minimal reproduction (no network needed)
import asyncio
import httpx2
from mcp.client.auth.oauth2 import OAuthClientProvider
from mcp.shared.auth import OAuthClientMetadata
class MemStorage:
async def get_tokens(self): return None
async def set_tokens(self, tokens): pass
async def get_client_info(self): return None
async def set_client_info(self, client_info): pass
async def main():
provider = OAuthClientProvider(
server_url="https://example.invalid/mcp",
client_metadata=OAuthClientMetadata(redirect_uris=["http://localhost:1/callback"]),
storage=MemStorage(),
)
# Task A: advance the flow. With no stored tokens it suspends at
# `response = yield request`, still inside `async with context.lock`.
flow = provider.async_auth_flow(httpx2.Request("POST", "https://example.invalid/mcp"))
await flow.asend(None)
# Task B: httpx tears the stream down from a different task on
# cancellation, closing the generator there.
try:
await asyncio.get_running_loop().create_task(flow.aclose())
except RuntimeError as exc:
print(f"aclose raised: {exc!r}") # <- fires
print("lock still held:", provider.context.lock.locked()) # True
# The "reconnect" — blocks forever at line 582:
flow2 = provider.async_auth_flow(httpx2.Request("POST", "https://example.invalid/mcp"))
try:
await asyncio.wait_for(flow2.asend(None), timeout=2)
print("second flow proceeds")
except asyncio.TimeoutError:
print("second flow BLOCKED on poisoned lock") # <- happens
finally:
await flow2.aclose()
asyncio.run(main())
Output on 2.0.0:
aclose raised: RuntimeError('The current task is not holding this lock')
lock still held: True
second flow BLOCKED on poisoned lock
Root cause
anyio.Lock is task-bound by design; httpx makes no guarantee that the auth-flow generator is closed from the task that advanced it (cancellation/teardown commonly runs aclose() from a sibling task). Holding a task-bound lock across the generator's yields therefore poisons the lock on any cross-task close: the release both raises and never happens.
Suggested fix
Serializing the auth flow per-context is still needed (token refresh must not race), but the primitive must allow release from the closing task. Options:
- Use
anyio.Semaphore(1)instead ofanyio.LockforOAuthContext.lock. anyio semaphores are not owner-bound, so release from the closing task is legal on both asyncio and trio backends. One-line change in theOAuthContextdataclass plus theasync withkeeps working. - Alternatively, wrap the generator body so
GeneratorExit/cross-task unwind releases via a tolerant path (catch the ownershipRuntimeErrorand force the lock back to a released state), though anyio has no public API for that today.
Happy to send a PR for option 1 if that direction is acceptable.
貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
研究方向
從 src/mcp/client/auth/oauth2.py 中的 OAuthContext 和 OAuthClientProvider.async_auth_flow 開始,然後執行 issue 中的最小重現。確認從另一個 task 關閉 flow 不會引發例外或讓 context 維持鎖定狀態,且後續的 flow 會繼續執行,而不是遭到阻塞。
由索引模型根據 Issue 內容生成。
評估
- 技術堆疊
- python
- 領域
- api, authentication
- Issue 類型
- 缺陷
- 難度
- 2/5
- 預估耗時
- 1-3 小時
- 活躍度
- 活躍
- 描述清晰度
- 描述清楚
- 新手友好度
- 82/100