MagicStack / MagicStack/asyncpg

Connection.close(timeout=) waits forever on a pending cancel when the server never acknowledges it

Abierto
#1,356 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Nadie ha tomado este issue todavía.

Lenguaje dominante
Python
Estrellas
8.1k
Forks
468
Métricas de merge de PR
Sin PR fusionados en 30 d

Descripción

Summary

When a statement times out (command_timeout) while the server, or a pooler in front of it, is frozen, asyncpg requests a cancel and then every later operation on that connection, including close(timeout=...), awaits the cancel acknowledgement with no bound. The timeout argument of close() does not cover that wait, and a subsequent transport loss does not resolve it either, so the connection can never be closed gracefully and any caller that awaits close() hangs indefinitely.

Versions

  • asyncpg 0.30.0 and 0.31.0 (same code shape in both)
  • Python 3.12.3, Linux
  • Observed through SQLAlchemy 2.0.52's asyncpg dialect, which calls Connection.close(timeout=2) when invalidating a connection after a TimeoutError, but the behaviour is asyncpg's.

Where in the source (0.31.0)

  • asyncpg/protocol/protocol.pyx, close(self, timeout): awaits self.cancel_sent_waiter and then if self.cancel_waiter is not None: await self.cancel_waiter before the part that is guarded by timeout.
  • _request_cancel() (called from _on_timeout()) creates cancel_waiter; it is resolved only by a ReadyForQuery arriving on the original socket.
  • _handle_waiter_on_connection_lost() and _on_connection_lost() resolve self.waiter only; cancel_waiter is left pending when the transport is lost.
  • abort() returns early when self.closing is already set, so cancelling a stuck close() from outside and then calling Connection._abort() does not close the transport.

Reproduction

  1. Run PostgreSQL behind pgbouncer (transaction pooling), or plain PostgreSQL.
  2. Open a connection with command_timeout=5, run SELECT pg_sleep(40).
  3. While it runs, freeze the server process (podman pause / kill -STOP on postgres, or on pgbouncer).
  4. The statement raises asyncio.TimeoutError after 5 s and asyncpg starts a cancel task.
  5. Now await conn.close(timeout=2): it never returns while the freeze lasts. If the frozen side is later closed by a pooler timeout (pgbouncer query_timeout closes the client socket), close() still never returns because the transport loss resolves only the query waiter.

Observed with an asyncio task-stack watchdog: the caller sits in Connection.closeprotocol.closeawait self.cancel_waiter, and the _cancel task sits in connect_utils awaiting the cancel connection's on_disconnect, for as long as the server stays frozen (minutes; unbounded).

Expected

  • close(timeout=t) should bound the wait for the cancel acknowledgement by t (or by the connection's command_timeout) and fall back to aborting the transport.
  • _on_connection_lost() should resolve cancel_waiter (with the same connection-lost exception it uses for waiter) so that a lost transport cannot leave a permanently pending cancel.

Workaround we use

An application-level guard that, on TimeoutError/CancelledError, runs asyncio.wait_for(conn.close(timeout=g), g) and on expiry calls conn.terminate() followed by an explicit conn._transport.abort(), because terminate() alone leaves the socket open once close() has marked the protocol as closing.

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Línea de trabajo

Comienza en asyncpg/protocol/protocol.pyx leyendo close(self, timeout), _request_cancel(), _handle_waiter_on_connection_lost() y _on_connection_lost(). Reproduce el caso de servidor congelado descrito en el issue y, después, verifica que close(timeout=2) retorna y que la pérdida del transporte no deja cancel_waiter en estado pending, mientras la conexión vuelve a abortar como se espera.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
postgresql, python
Área
databases
Tipo de issue
Error
Dificultad
4/5
Tiempo estimado
3-5 días
Estado de actividad
Activo
Claridad
Bien especificado
Aptitud para principiantes
68/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.