MagicStack / MagicStack/asyncpg

Upon postgres segfault, connections in `async for record in conn.cursor` fail to detect loss of connectivity and hang indefinitely

未关闭
#549 3 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

主要语言
Python
星标
8.1k
派生
468
PR 合并指标
30 天内没有已合并 PR

描述

* **asyncpg version**: 0.20.1
* **PostgreSQL version**: PostgreSQL 10.6 on x86_64-pc-linux-gnu, compiled by gcc (GCC) 4.8.3 20140911 (Red Hat 4.8.3-9), 64-bit
* **Do you use a PostgreSQL SaaS? If so, which? Can you reproduce
the issue with a local PostgreSQL install?**: unknown - this depends on being somewhere critically inside the iterator for `conn.cursor(...)`
* **Python version**: 3.7.7 (default, Mar 24 2020, 03:07:37) \n[GCC 9.2.0]
* **Platform**: python:3.7-alpine
* **Do you use pgbouncer?**: No.
* **Did you install asyncpg with pip?**: yes
* **If you built asyncpg locally, which version of Cython did you use?**: N/A
* **Can the issue be reproduced under both asyncio and
[uvloop](https://github.com/magicstack/uvloop)?**: uvloop is not used in this installation.

I made use of a `asyncpg.pool.Pool` created via `create_pool` with a command_timeout of None.

Midway through heavy load, the RDS instance crashed and rebooted, losing all connections.

However, the connections actively in use jammed and as of this time have been jammed for well over an hour. Using `awaitwhat` and a shell, I was able to determine that all the coroutines were jammed on the following:

```
async def _execute_iter(self, conn):
async for record in conn.cursor(self.query, *self.bindings.values):
yield record

async def execute(self):
conn = self.conn()
if conn is None:
raise ValueError("Connection has been garbage collected!")

cursor = self._execute_iter(conn)
if self.iterable:
return cursor
items = []
async for record in cursor: # <---- jammed in here indefinitely!
items.append(record)
return tuple(items)
```

Upon grabbing a stalled coroutine with my backdoor shell, I discovered a surprising behavior!

```
>>> tasks[0].get_stack()[0].f_locals['conn']
0x7fe02a56df10>
>>> tasks[0].get_stack()[0].f_locals['conn'].is_closed()
False
>>> tasks[0].get_stack()[0].f_locals['conn']
0x7fe02a56df10>
>>> tasks[0].get_stack()[0].f_locals['conn'].is_in_transaction()
True
>>> tasks[0].get_stack()[0].f_locals['conn']._con.get_server_pid()
22062
>>> await tasks[0].get_stack()[0].f_locals['conn']._con.fetch('SELECT 1')
Traceback (most recent call last):
File "/usr/local/lib/python3.7/asyncio/selector_events.py", line 814, in _read_ready__data_received
data = self._sock.recv(self.max_size)
ConnectionResetError: [Errno 104] Connection reset by peer

The above exception was the direct cause of the following exception:

Traceback (most recent call last):
File "/usr/local/lib/python3.7/site-packages/aioconsole/execute.py", line 95, in aexec
result, new_local = yield from coro
File "", line 2, in __corofn
File "/usr/local/lib/python3.7/site-packages/asyncpg/connection.py", line 420, in fetch
return await self._execute(query, args, 0, timeout)
File "/usr/local/lib/python3.7/site-packages/asyncpg/connection.py", line 1403, in _execute
query, args, limit, timeout, return_status=return_status)
File "/usr/local/lib/python3.7/site-packages/asyncpg/connection.py", line 1411, in __execute
return await self._do_execute(query, executor, timeout)
File "/usr/local/lib/python3.7/site-packages/asyncpg/connection.py", line 1423, in _do_execute
stmt = await self._get_statement(query, None)
File "/usr/local/lib/python3.7/site-packages/asyncpg/connection.py", line 328, in _get_statement
statement = await self._protocol.prepare(stmt_name, query, timeout)
File "asyncpg/protocol/protocol.pyx", line 163, in prepare
asyncpg.exceptions.ConnectionDoesNotExistError: connection was closed in the middle of operation
>>> tasks[0].get_stack()[0].f_locals['conn'].is_closed()
Traceback (most recent call last):
File "/usr/local/lib/python3.7/site-packages/aioconsole/execute.py", line 95, in aexec
result, new_local = yield from coro
File "", line 2, in __corofn
File "/usr/local/lib/python3.7/site-packages/asyncpg/pool.py", line 56, in call_con_method
meth_name))
asyncpg.exceptions._base.InterfaceError: cannot call Connection.is_closed(): connection has been released back to the pool
>>>
```

I submit to you that in the code for iterating through a cursor, something does not properly handle the case of a sudden connection reset. Notice how when I actually attempt to use the stalled connection which is blocked in an `async for record in cursor` generator, it suddenly realizes the connection is dead!

If I write b”x” in the socket itself, that also is enough to wake up asyncio to being disconnected!

I will experiment with a `command_timeout` to hopefully add in a worst-case timer for blocked queries, however this

贡献指南

这个仓库没有索引到贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

调研方向

从 _execute_iter 中的游标迭代以及 asyncpg/connection.py、asyncpg/pool.py 和 asyncpg/protocol/protocol.pyx 中展示的连接丢失路径开始。在执行 async for record in conn.cursor(...) 期间复现一次 reset,并跟踪被阻塞的操作如何被释放。完成的标准是迭代器检测到连接丢失并以适当的连接错误退出,而不是无限期挂起。

由索引模型根据 Issue 内容生成。

评估

技术栈
postgresql, python
领域
backend, databases
Issue 类型
缺陷
难度
4/5
预计耗时
3-5 天
活跃度
停滞
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。