MagicStack / MagicStack/asyncpg

Connection.close(timeout=) waits forever on a pending cancel when the server never acknowledges it

Đang mở
#1,356 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Ngôn ngữ chính
Python
Star
8.1k
Fork
468
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Mô tả

Summary

When a statement times out (command_timeout) while the server, or a pooler in front of it, is frozen, asyncpg requests a cancel and then every later operation on that connection, including close(timeout=...), awaits the cancel acknowledgement with no bound. The timeout argument of close() does not cover that wait, and a subsequent transport loss does not resolve it either, so the connection can never be closed gracefully and any caller that awaits close() hangs indefinitely.

Versions

  • asyncpg 0.30.0 and 0.31.0 (same code shape in both)
  • Python 3.12.3, Linux
  • Observed through SQLAlchemy 2.0.52's asyncpg dialect, which calls Connection.close(timeout=2) when invalidating a connection after a TimeoutError, but the behaviour is asyncpg's.

Where in the source (0.31.0)

  • asyncpg/protocol/protocol.pyx, close(self, timeout): awaits self.cancel_sent_waiter and then if self.cancel_waiter is not None: await self.cancel_waiter before the part that is guarded by timeout.
  • _request_cancel() (called from _on_timeout()) creates cancel_waiter; it is resolved only by a ReadyForQuery arriving on the original socket.
  • _handle_waiter_on_connection_lost() and _on_connection_lost() resolve self.waiter only; cancel_waiter is left pending when the transport is lost.
  • abort() returns early when self.closing is already set, so cancelling a stuck close() from outside and then calling Connection._abort() does not close the transport.

Reproduction

  1. Run PostgreSQL behind pgbouncer (transaction pooling), or plain PostgreSQL.
  2. Open a connection with command_timeout=5, run SELECT pg_sleep(40).
  3. While it runs, freeze the server process (podman pause / kill -STOP on postgres, or on pgbouncer).
  4. The statement raises asyncio.TimeoutError after 5 s and asyncpg starts a cancel task.
  5. Now await conn.close(timeout=2): it never returns while the freeze lasts. If the frozen side is later closed by a pooler timeout (pgbouncer query_timeout closes the client socket), close() still never returns because the transport loss resolves only the query waiter.

Observed with an asyncio task-stack watchdog: the caller sits in Connection.closeprotocol.closeawait self.cancel_waiter, and the _cancel task sits in connect_utils awaiting the cancel connection's on_disconnect, for as long as the server stays frozen (minutes; unbounded).

Expected

  • close(timeout=t) should bound the wait for the cancel acknowledgement by t (or by the connection's command_timeout) and fall back to aborting the transport.
  • _on_connection_lost() should resolve cancel_waiter (with the same connection-lost exception it uses for waiter) so that a lost transport cannot leave a permanently pending cancel.

Workaround we use

An application-level guard that, on TimeoutError/CancelledError, runs asyncio.wait_for(conn.close(timeout=g), g) and on expiry calls conn.terminate() followed by an explicit conn._transport.abort(), because terminate() alone leaves the socket open once close() has marked the protocol as closing.

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Hướng nghiên cứu

Bắt đầu trong asyncpg/protocol/protocol.pyx, đọc close(self, timeout), _request_cancel(), _handle_waiter_on_connection_lost() và _on_connection_lost(). Tái hiện trường hợp máy chủ bị treo được mô tả trong issue, sau đó xác minh rằng close(timeout=2) trả về và việc mất transport không để cancel_waiter ở trạng thái pending, trong khi kết nối chuyển sang abort như mong đợi.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
postgresql, python
Lĩnh vực
databases
Loại issue
Lỗi
Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức độ hoạt động
Sôi nổi
Độ rõ ràng
Đặc tả rõ ràng
Mức phù hợp với người mới
68/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.