aws / aws/aws-advanced-python-wrapper

[aio] aurora_connection_tracker closes its own connection on the first statement of every cluster-endpoint connection

オープン
#1,276 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る
bug
主要言語
Python
スター
98
フォーク
22
平均マージ
1日 9時間
マージ済み PR(30日)
5

説明

### Describe the bug

With the async wrapper (`aws_advanced_python_wrapper.aio`, via the `postgresql+aws_wrapper_psycopg` SQLAlchemy dialect) and the default plugin chain, every new connection to an Aurora PostgreSQL cluster writer endpoint fails on its first statement with `FailoverSuccessError`. The retries also open several connections per attempt (the failover's writer connection plus topology probes), which push a huge spike in connection count.

### Expected Behavior

A connection made through the cluster writer endpoint with the default plugins should run its first statement normally, as it does with `wrapper_plugins=failover,host_monitoring_v2` (the chain used in `docs/examples/PGSQLAlchemyAsyncFailover.py`).

### What plugins are used? What other connection properties were set?

Default chain (`wrapper_plugins` not set, so `initial_connection,aurora_connection_tracker,failover_v2,host_monitoring_v2`), `wrapper_dialect=aurora-pg`. Also reproduced with `wrapper_plugins=aurora_connection_tracker` alone. Does not reproduce with `wrapper_plugins=failover,host_monitoring_v2` or `failover_v2` alone.

### Current Behavior

Every connection, on its first statement:

```
aws_advanced_python_wrapper.aio.aurora_connection_tracker: [OpenedConnectionTracker] Invalidating opened connections to host: .cluster-.us-east-1.rds.amazonaws.com:5432/
aws_advanced_python_wrapper.aio.aurora_connection_tracker: [OpenedConnectionTracker] Opened Connections Tracked:
aws_advanced_python_wrapper.aio.aurora_connection_tracker: [AsyncAuroraConnectionTrackerPlugin] failover handler: pre=.cluster-.us-east-1.rds.amazonaws.com:5432/ post=tf-2026061217...chnc4uhow0rs.us-east-1.rds.amazonaws.com:5432/ pinned=tf-2026061217...chnc4uhow0rs.us-east-1.rds.amazonaws.com:5432/
aws_advanced_python_wrapper.aio.aurora_connection_tracker: [OpenedConnectionTracker] Invalidating opened connections to host: .cluster-.us-east-1.rds.amazonaws.com:5432/
iter 0: operational error (FailoverSuccessError)
```

### Reproduction Steps

```python
import asyncio, logging, os, sys
from sqlalchemy import text
from sqlalchemy.exc import OperationalError
from sqlalchemy.ext.asyncio import create_async_engine
from aws_advanced_python_wrapper.aio import release_resources_async

logging.basicConfig(level=logging.WARNING, stream=sys.stdout, format="%(name)s: %(message)s")
logging.getLogger("aws_advanced_python_wrapper.aio.aurora_connection_tracker").setLevel(logging.DEBUG)

CLUSTER_ENDPOINT, DB_NAME, USER, PASSWORD = (os.environ[k] for k in ("PGHOST", "PGDATABASE", "PGUSER", "PGPASSWORD"))
PLUGINS = "" if sys.argv[1] == "default" else "&wrapper_plugins=failover,host_monitoring_v2"

async def main():
engine = create_async_engine(
f"postgresql+aws_wrapper_psycopg://{USER}:{PASSWORD}@{CLUSTER_ENDPOINT}:5432/{DB_NAME}"
f"?wrapper_dialect=aurora-pg{PLUGINS}")
try:
for i in range(3):
try:
async with engine.connect() as conn:
row = await conn.execute(text("SELECT pg_catalog.aurora_db_instance_identifier()"))
print(f"iter {i}: connected to instance {row.scalar_one()}")
except OperationalError as exc:
print(f"iter {i}: operational error ({type(exc.orig).__name__})")
finally:
await engine.dispose()
await release_resources_async()

asyncio.run(main())
```

### Possible Solution

Claude output
> In `aws_advanced_python_wrapper/aio/aurora_connection_tracker.py`, `_pin_current_writer` first pins the writer from topology, which is the instance endpoint. Its stale-topology guard then calls `get_host_role(conn)`, gets WRITER, and replaces the pin with `plugin_service.current_host_info`, which is the URL host, i.e. the cluster endpoint. `_same_host` compares host strings, so the cluster endpoint never equals the instance endpoint. On the first `execute`, `_invalidate_writer_change` compares the cluster-endpoint pin with the instance-endpoint topology writer, reports a writer change, and `invalidate_all` closes every connection keyed under the cluster endpoint, which after `_fill_instance_alias` includes the connection about to execute. The sync tracker does not pin at connect time and only compares instance against instance, so it is unaffected.

### Additional Information/Context

_No response_

### The AWS Advanced Python Wrapper version used

3.1.0

### python version used

3.14.7

### Operating System and version

Debian GNU/Linux 12 (bookworm), aarch64, `python:3.14-slim` image

コントリビューションガイド

コントリビューションガイドを開く

調査の方向性

aws_advanced_python_wrapper/aio/aurora_connection_tracker.py、特に _pin_current_writer、_invalidate_writer_change、_same_host、_fill_instance_alias から始め、次にデフォルトのプラグインチェーンで提供された SQLAlchemy の非同期再現を実行します。固定されたホストがトポロジーの writer とどのように比較されるかを追跡し、最初のステートメントが自身の接続を無効化せずに成功することを確認します。一方、文書化されている failover,host_monitoring_v2 チェーンは対照として残します。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
aws, postgresql, python, sqlalchemy
領域
databases
issue の種類
バグ
難易度
4/5
見積もり時間
3〜5日
活発さ
活発
明瞭さ
明確に書かれている
初心者へのやさしさ
72/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。