google / google/adk-python

MCP session stale after server scale-to-zero - @retry_on_errors reuses dead cached session

オープン
#7,060 コメント 6 件 リアクション 0 件 担当者 1 名 @llalitkumarrr が担当を希望しています GitHub で見る
mcp request clarification spam
主要言語
Python
スター
21.5k
フォーク
4k
平均マージ
1日 14時間
マージ済み PR(30日)
37

説明

## Bug Description

When an MCP server (e.g. on Cloud Run) scales to zero and back, the server's in-memory sessions are lost. The agent's `@retry_on_errors` decorator retries the failed call, but `MCPSessionManager.create_session()` returns the **same cached dead session** because `_is_session_disconnected()` only checks transport-level stream closure - the transport is still alive, only the server-side session is gone.

This means every retry reuses the stale session and fails again, until all retries are exhausted.

## Steps to Reproduce

1. Deploy an MCP server on Cloud Run (or any auto-scaling platform) with scale-to-zero enabled
2. Connect an ADK agent using `McpToolset` with a Streamable HTTP connection
3. Make a successful tool call (session is created and cached)
4. Wait for the MCP server to scale to zero (~15 minutes idle)
5. Make another tool call - the server scales back up with fresh state

## Expected Behavior

The agent should detect the server-side session loss and create a fresh session on retry.

## Actual Behavior

- `@retry_on_errors` catches the error and retries
- `create_session()` finds the cached session
- `_is_session_disconnected()` returns `False` (transport streams are still open)
- The same dead session is reused → retry fails with the same error
- All retries exhausted → permanent failure

## Root Cause

`_is_session_disconnected()` only checks `_read_stream._closed` and `_write_stream._closed`. In a scale-to-zero scenario, the HTTP transport is fine - the server responds normally with
a 404/error for the unknown session ID. The check needs a way to detect **servtion, not just transport failure.

The existing `_MCP_GRACEFUL_ERROR_HANDLING` flag handles the case where the asyncio task dies, but not this scenario where the task is alive and the server responds normally.

## Environment

- ADK version: 2.7.1
- MCP server: Cloud Run with Streamable HTTP transport
- Python: 3.12+

## Suggested Fix

Add an `invalidate_session()` method to `MCPSessionManager` that marks a session key as stale. Have `_execute_with_session()` call it on any exception before re-raising, so the next `@retry_on_errors` attempt builds a fresh session. PR incoming.

コントリビューションガイド

コントリビューションガイドを開く

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。