modelcontextprotocol / modelcontextprotocol/python-sdk
OAuthClientProvider auth lock is permanently poisoned when httpx closes async_auth_flow from a different task (RuntimeError: The current task is not holding this lock)
Dieses Issue hat noch niemand übernommen.
- Vorherrschende Sprache
- Python
- Sterne
- 24.3k
- Forks
- 4k
- Ø Merge
- 1 T. 1 Std.
- Gemergte PRs (30 T.)
- 31
Beschreibung
Summary
OAuthClientProvider.async_auth_flow holds self.context.lock (an anyio.Lock) across the entire httpx auth-flow generator, including every yield (src/mcp/client/auth/oauth2.py:582 in v2.0.0; same code is present on v2.1.0 and main). anyio.Lock.release() is bound to the acquiring task. When httpx closes the auth generator from a different task than the one that advanced it — which happens routinely when a request is cancelled mid-flight (network drop, timeout, task-group teardown) — the async with __aexit__ runs in the closing task and raises:
RuntimeError: The current task is not holding this lock
The exception escapes into a fire-and-forget teardown task ("Task exception was never retrieved"), and the lock is left permanently held. Every subsequent request through the same OAuthClientProvider then blocks forever at async with self.context.lock: without sending any HTTP. For a long-lived client that reuses the provider across reconnects, that server is dead until the whole process restarts.
Environment
mcp2.0.0 (reproduced; the sameasync with self.context.lock:pattern is unchanged in 2.1.0 and on main)- Python 3.11.15, macOS (darwin), anyio 4.14.2, asyncio backend
- Long-lived agent process (Nous Research Hermes) with OAuth-backed streamable-HTTP MCP servers; surfaced downstream as NousResearch/hermes-agent#81051
Observed traceback
Two captures from the same host, one per yield point (authorization-code exchange and the main yield request):
ERROR asyncio: Task exception was never retrieved
future: <Task finished name='Task-598' coro=<<async_generator_athrow without __name__>()> exception=RuntimeError('The current task is not holding this lock')>
Traceback (most recent call last):
File ".../site-packages/mcp/client/auth/oauth2.py", line 754, in async_auth_flow
yield request
GeneratorExit
During handling of the above exception, another exception occurred:
Traceback (most recent call last):
File ".../site-packages/mcp/client/auth/oauth2.py", line 582, in async_auth_flow
async with self.context.lock:
File ".../site-packages/anyio/_core/_synchronization.py", line 173, in __aexit__
self.release()
File ".../site-packages/anyio/_backends/_asyncio.py", line 1935, in release
raise RuntimeError("The current task is not holding this lock")
RuntimeError: The current task is not holding this lock
(The other capture is identical except the GeneratorExit lands at line 746, token_response = yield await self._perform_authorization().)
After this fires, every reconnect attempt for that server times out with no HTTP traffic — the flow generator never gets past line 582.
Minimal reproduction (no network needed)
import asyncio
import httpx2
from mcp.client.auth.oauth2 import OAuthClientProvider
from mcp.shared.auth import OAuthClientMetadata
class MemStorage:
async def get_tokens(self): return None
async def set_tokens(self, tokens): pass
async def get_client_info(self): return None
async def set_client_info(self, client_info): pass
async def main():
provider = OAuthClientProvider(
server_url="https://example.invalid/mcp",
client_metadata=OAuthClientMetadata(redirect_uris=["http://localhost:1/callback"]),
storage=MemStorage(),
)
# Task A: advance the flow. With no stored tokens it suspends at
# `response = yield request`, still inside `async with context.lock`.
flow = provider.async_auth_flow(httpx2.Request("POST", "https://example.invalid/mcp"))
await flow.asend(None)
# Task B: httpx tears the stream down from a different task on
# cancellation, closing the generator there.
try:
await asyncio.get_running_loop().create_task(flow.aclose())
except RuntimeError as exc:
print(f"aclose raised: {exc!r}") # <- fires
print("lock still held:", provider.context.lock.locked()) # True
# The "reconnect" — blocks forever at line 582:
flow2 = provider.async_auth_flow(httpx2.Request("POST", "https://example.invalid/mcp"))
try:
await asyncio.wait_for(flow2.asend(None), timeout=2)
print("second flow proceeds")
except asyncio.TimeoutError:
print("second flow BLOCKED on poisoned lock") # <- happens
finally:
await flow2.aclose()
asyncio.run(main())
Output on 2.0.0:
aclose raised: RuntimeError('The current task is not holding this lock')
lock still held: True
second flow BLOCKED on poisoned lock
Root cause
anyio.Lock is task-bound by design; httpx makes no guarantee that the auth-flow generator is closed from the task that advanced it (cancellation/teardown commonly runs aclose() from a sibling task). Holding a task-bound lock across the generator's yields therefore poisons the lock on any cross-task close: the release both raises and never happens.
Suggested fix
Serializing the auth flow per-context is still needed (token refresh must not race), but the primitive must allow release from the closing task. Options:
- Use
anyio.Semaphore(1)instead ofanyio.LockforOAuthContext.lock. anyio semaphores are not owner-bound, so release from the closing task is legal on both asyncio and trio backends. One-line change in theOAuthContextdataclass plus theasync withkeeps working. - Alternatively, wrap the generator body so
GeneratorExit/cross-task unwind releases via a tolerant path (catch the ownershipRuntimeErrorand force the lock back to a released state), though anyio has no public API for that today.
Happy to send a PR for option 1 if that direction is acceptable.
Beitragsleitfaden
Erste Schritte
- Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
- Forke das Repository und arbeite in einem Branch.
- Öffne einen Pull Request, der die Issue-Nummer nennt.
Rechercherichtung
Beginne in src/mcp/client/auth/oauth2.py bei OAuthContext und OAuthClientProvider.async_auth_flow und führe dann die minimale Reproduktion aus dem Issue aus. Überprüfe, dass das Schließen des Flows aus einem anderen Task weder einen Fehler auslöst noch den Context gesperrt lässt und dass ein nachfolgender Flow fortgesetzt wird, statt zu blockieren.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- python
- Bereich
- api, authentication
- Issue-Typ
- Bug
- Schwierigkeit
- 2/5
- Geschätzter Aufwand
- 1-3 Stunden
- Aktivitätsstatus
- Aktiv
- Klarheit
- Klar beschrieben
- Anfängerfreundlichkeit
- 82/100