MagicStack / MagicStack/asyncpg
Connection.close(timeout=) waits forever on a pending cancel when the server never acknowledges it
まだ誰も着手していません。
- 主要言語
- Python
- スター
- 8.1k
- フォーク
- 468
- PR マージ指標
- 30日以内にマージされた PR はありません
説明
## Summary
When a statement times out (`command_timeout`) while the server, or a pooler in front of it, is frozen, asyncpg requests a cancel and then every later operation on that connection, including `close(timeout=...)`, awaits the cancel acknowledgement with no bound. The `timeout` argument of `close()` does not cover that wait, and a subsequent transport loss does not resolve it either, so the connection can never be closed gracefully and any caller that awaits `close()` hangs indefinitely.
## Versions
- asyncpg 0.30.0 and 0.31.0 (same code shape in both)
- Python 3.12.3, Linux
- Observed through SQLAlchemy 2.0.52's asyncpg dialect, which calls `Connection.close(timeout=2)` when invalidating a connection after a `TimeoutError`, but the behaviour is asyncpg's.
## Where in the source (0.31.0)
- `asyncpg/protocol/protocol.pyx`, `close(self, timeout)`: awaits `self.cancel_sent_waiter` and then `if self.cancel_waiter is not None: await self.cancel_waiter` before the part that is guarded by `timeout`.
- `_request_cancel()` (called from `_on_timeout()`) creates `cancel_waiter`; it is resolved only by a ReadyForQuery arriving on the original socket.
- `_handle_waiter_on_connection_lost()` and `_on_connection_lost()` resolve `self.waiter` only; `cancel_waiter` is left pending when the transport is lost.
- `abort()` returns early when `self.closing` is already set, so cancelling a stuck `close()` from outside and then calling `Connection._abort()` does not close the transport.
## Reproduction
1. Run PostgreSQL behind pgbouncer (transaction pooling), or plain PostgreSQL.
2. Open a connection with `command_timeout=5`, run `SELECT pg_sleep(40)`.
3. While it runs, freeze the server process (`podman pause` / `kill -STOP` on postgres, or on pgbouncer).
4. The statement raises `asyncio.TimeoutError` after 5 s and asyncpg starts a cancel task.
5. Now `await conn.close(timeout=2)`: it never returns while the freeze lasts. If the frozen side is later closed by a pooler timeout (pgbouncer `query_timeout` closes the client socket), `close()` still never returns because the transport loss resolves only the query waiter.
Observed with an asyncio task-stack watchdog: the caller sits in `Connection.close` → `protocol.close` → `await self.cancel_waiter`, and the `_cancel` task sits in `connect_utils` awaiting the cancel connection's `on_disconnect`, for as long as the server stays frozen (minutes; unbounded).
## Expected
- `close(timeout=t)` should bound the wait for the cancel acknowledgement by `t` (or by the connection's `command_timeout`) and fall back to aborting the transport.
- `_on_connection_lost()` should resolve `cancel_waiter` (with the same connection-lost exception it uses for `waiter`) so that a lost transport cannot leave a permanently pending cancel.
## Workaround we use
An application-level guard that, on `TimeoutError`/`CancelledError`, runs `asyncio.wait_for(conn.close(timeout=g), g)` and on expiry calls `conn.terminate()` followed by an explicit `conn._transport.abort()`, because `terminate()` alone leaves the socket open once `close()` has marked the protocol as closing.
コントリビューションガイド
このリポジトリのコントリビューションガイドは索引されていません
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
調査の方向性
asyncpg/protocol/protocol.pyx で close(self, timeout)、_request_cancel()、_handle_waiter_on_connection_lost()、_on_connection_lost() を読み始めます。issue に記載されているサーバーがフリーズしたケースを再現し、その後 close(timeout=2) が戻り、transport の喪失によって cancel_waiter が pending のまま残らないことを確認します。その間、接続は期待どおり abort にフォールバックします。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- postgresql, python
- 領域
- databases
- issue の種類
- バグ
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 活発さ
- 活発
- 明瞭さ
- 明確に書かれている
- 初心者へのやさしさ
- 68/100