MCP session stale after server scale-to-zero - @retry_on_errors reuses dead cached session
- Langage dominant
- Python
- Étoiles
- 21.5k
- Forks
- 4k
- Merge moyen
- 1 j 14 h
- PR mergées (30 j)
- 37
Description
## Bug Description
When an MCP server (e.g. on Cloud Run) scales to zero and back, the server's in-memory sessions are lost. The agent's `@retry_on_errors` decorator retries the failed call, but `MCPSessionManager.create_session()` returns the **same cached dead session** because `_is_session_disconnected()` only checks transport-level stream closure - the transport is still alive, only the server-side session is gone.
This means every retry reuses the stale session and fails again, until all retries are exhausted.
## Steps to Reproduce
1. Deploy an MCP server on Cloud Run (or any auto-scaling platform) with scale-to-zero enabled
2. Connect an ADK agent using `McpToolset` with a Streamable HTTP connection
3. Make a successful tool call (session is created and cached)
4. Wait for the MCP server to scale to zero (~15 minutes idle)
5. Make another tool call - the server scales back up with fresh state
## Expected Behavior
The agent should detect the server-side session loss and create a fresh session on retry.
## Actual Behavior
- `@retry_on_errors` catches the error and retries
- `create_session()` finds the cached session
- `_is_session_disconnected()` returns `False` (transport streams are still open)
- The same dead session is reused → retry fails with the same error
- All retries exhausted → permanent failure
## Root Cause
`_is_session_disconnected()` only checks `_read_stream._closed` and `_write_stream._closed`. In a scale-to-zero scenario, the HTTP transport is fine - the server responds normally with
a 404/error for the unknown session ID. The check needs a way to detect **servtion, not just transport failure.
The existing `_MCP_GRACEFUL_ERROR_HANDLING` flag handles the case where the asyncio task dies, but not this scenario where the task is alive and the server responds normally.
## Environment
- ADK version: 2.7.1
- MCP server: Cloud Run with Streamable HTTP transport
- Python: 3.12+
## Suggested Fix
Add an `invalidate_session()` method to `MCPSessionManager` that marks a session key as stale. Have `_execute_with_session()` call it on any exception before re-raising, so the next `@retry_on_errors` attempt builds a fresh session. PR incoming.
Guide de contribution
Ouvrir le guide de contribution
Évaluation
Cette issue n'a pas encore été évaluée.