aws / aws/aws-advanced-python-wrapper

[aio] aurora_connection_tracker closes its own connection on the first statement of every cluster-endpoint connection

Offen
#1,276 1 Kommentar 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
bug
Vorherrschende Sprache
Python
Sterne
98
Forks
22
Ø Merge
1 T. 9 Std.
Gemergte PRs (30 T.)
5

Beschreibung

### Describe the bug

With the async wrapper (`aws_advanced_python_wrapper.aio`, via the `postgresql+aws_wrapper_psycopg` SQLAlchemy dialect) and the default plugin chain, every new connection to an Aurora PostgreSQL cluster writer endpoint fails on its first statement with `FailoverSuccessError`. The retries also open several connections per attempt (the failover's writer connection plus topology probes), which push a huge spike in connection count.

### Expected Behavior

A connection made through the cluster writer endpoint with the default plugins should run its first statement normally, as it does with `wrapper_plugins=failover,host_monitoring_v2` (the chain used in `docs/examples/PGSQLAlchemyAsyncFailover.py`).

### What plugins are used? What other connection properties were set?

Default chain (`wrapper_plugins` not set, so `initial_connection,aurora_connection_tracker,failover_v2,host_monitoring_v2`), `wrapper_dialect=aurora-pg`. Also reproduced with `wrapper_plugins=aurora_connection_tracker` alone. Does not reproduce with `wrapper_plugins=failover,host_monitoring_v2` or `failover_v2` alone.

### Current Behavior

Every connection, on its first statement:

```
aws_advanced_python_wrapper.aio.aurora_connection_tracker: [OpenedConnectionTracker] Invalidating opened connections to host: .cluster-.us-east-1.rds.amazonaws.com:5432/
aws_advanced_python_wrapper.aio.aurora_connection_tracker: [OpenedConnectionTracker] Opened Connections Tracked:
aws_advanced_python_wrapper.aio.aurora_connection_tracker: [AsyncAuroraConnectionTrackerPlugin] failover handler: pre=.cluster-.us-east-1.rds.amazonaws.com:5432/ post=tf-2026061217...chnc4uhow0rs.us-east-1.rds.amazonaws.com:5432/ pinned=tf-2026061217...chnc4uhow0rs.us-east-1.rds.amazonaws.com:5432/
aws_advanced_python_wrapper.aio.aurora_connection_tracker: [OpenedConnectionTracker] Invalidating opened connections to host: .cluster-.us-east-1.rds.amazonaws.com:5432/
iter 0: operational error (FailoverSuccessError)
```

### Reproduction Steps

```python
import asyncio, logging, os, sys
from sqlalchemy import text
from sqlalchemy.exc import OperationalError
from sqlalchemy.ext.asyncio import create_async_engine
from aws_advanced_python_wrapper.aio import release_resources_async

logging.basicConfig(level=logging.WARNING, stream=sys.stdout, format="%(name)s: %(message)s")
logging.getLogger("aws_advanced_python_wrapper.aio.aurora_connection_tracker").setLevel(logging.DEBUG)

CLUSTER_ENDPOINT, DB_NAME, USER, PASSWORD = (os.environ[k] for k in ("PGHOST", "PGDATABASE", "PGUSER", "PGPASSWORD"))
PLUGINS = "" if sys.argv[1] == "default" else "&wrapper_plugins=failover,host_monitoring_v2"

async def main():
engine = create_async_engine(
f"postgresql+aws_wrapper_psycopg://{USER}:{PASSWORD}@{CLUSTER_ENDPOINT}:5432/{DB_NAME}"
f"?wrapper_dialect=aurora-pg{PLUGINS}")
try:
for i in range(3):
try:
async with engine.connect() as conn:
row = await conn.execute(text("SELECT pg_catalog.aurora_db_instance_identifier()"))
print(f"iter {i}: connected to instance {row.scalar_one()}")
except OperationalError as exc:
print(f"iter {i}: operational error ({type(exc.orig).__name__})")
finally:
await engine.dispose()
await release_resources_async()

asyncio.run(main())
```

### Possible Solution

Claude output
> In `aws_advanced_python_wrapper/aio/aurora_connection_tracker.py`, `_pin_current_writer` first pins the writer from topology, which is the instance endpoint. Its stale-topology guard then calls `get_host_role(conn)`, gets WRITER, and replaces the pin with `plugin_service.current_host_info`, which is the URL host, i.e. the cluster endpoint. `_same_host` compares host strings, so the cluster endpoint never equals the instance endpoint. On the first `execute`, `_invalidate_writer_change` compares the cluster-endpoint pin with the instance-endpoint topology writer, reports a writer change, and `invalidate_all` closes every connection keyed under the cluster endpoint, which after `_fill_instance_alias` includes the connection about to execute. The sync tracker does not pin at connect time and only compares instance against instance, so it is unaffected.

### Additional Information/Context

_No response_

### The AWS Advanced Python Wrapper version used

3.1.0

### python version used

3.14.7

### Operating System and version

Debian GNU/Linux 12 (bookworm), aarch64, `python:3.14-slim` image

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

Beginnen Sie mit aws_advanced_python_wrapper/aio/aurora_connection_tracker.py, insbesondere mit _pin_current_writer, _invalidate_writer_change, _same_host und _fill_instance_alias, und führen Sie anschließend die bereitgestellte asynchrone SQLAlchemy-Reproduktion mit der standardmäßigen Plugin-Kette aus. Verfolgen Sie, wie der angeheftete Host mit dem Writer der Topologie verglichen wird, und überprüfen Sie, dass die erste Anweisung erfolgreich ist, ohne ihre eigene Verbindung zu invalidieren, während die dokumentierte failover,host_monitoring_v2-Kette weiterhin als Kontrolle dient.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
aws, postgresql, python, sqlalchemy
Bereich
databases
Issue-Typ
Bug
Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Aktivitätsstatus
Aktiv
Klarheit
Klar beschrieben
Anfängerfreundlichkeit
72/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.