aws / aws/aws-advanced-python-wrapper

[aio] aurora_connection_tracker closes its own connection on the first statement of every cluster-endpoint connection

Open
#1,276 1 comment 0 reactions 0 assignees View on GitHub
bug
Dominant language
Python
Stars
98
Forks
22
Avg merge
1d 9h
Merged PRs (30d)
5

Description

### Describe the bug

With the async wrapper (`aws_advanced_python_wrapper.aio`, via the `postgresql+aws_wrapper_psycopg` SQLAlchemy dialect) and the default plugin chain, every new connection to an Aurora PostgreSQL cluster writer endpoint fails on its first statement with `FailoverSuccessError`. The retries also open several connections per attempt (the failover's writer connection plus topology probes), which push a huge spike in connection count.

### Expected Behavior

A connection made through the cluster writer endpoint with the default plugins should run its first statement normally, as it does with `wrapper_plugins=failover,host_monitoring_v2` (the chain used in `docs/examples/PGSQLAlchemyAsyncFailover.py`).

### What plugins are used? What other connection properties were set?

Default chain (`wrapper_plugins` not set, so `initial_connection,aurora_connection_tracker,failover_v2,host_monitoring_v2`), `wrapper_dialect=aurora-pg`. Also reproduced with `wrapper_plugins=aurora_connection_tracker` alone. Does not reproduce with `wrapper_plugins=failover,host_monitoring_v2` or `failover_v2` alone.

### Current Behavior

Every connection, on its first statement:

```
aws_advanced_python_wrapper.aio.aurora_connection_tracker: [OpenedConnectionTracker] Invalidating opened connections to host: .cluster-.us-east-1.rds.amazonaws.com:5432/
aws_advanced_python_wrapper.aio.aurora_connection_tracker: [OpenedConnectionTracker] Opened Connections Tracked:
aws_advanced_python_wrapper.aio.aurora_connection_tracker: [AsyncAuroraConnectionTrackerPlugin] failover handler: pre=.cluster-.us-east-1.rds.amazonaws.com:5432/ post=tf-2026061217...chnc4uhow0rs.us-east-1.rds.amazonaws.com:5432/ pinned=tf-2026061217...chnc4uhow0rs.us-east-1.rds.amazonaws.com:5432/
aws_advanced_python_wrapper.aio.aurora_connection_tracker: [OpenedConnectionTracker] Invalidating opened connections to host: .cluster-.us-east-1.rds.amazonaws.com:5432/
iter 0: operational error (FailoverSuccessError)
```

### Reproduction Steps

```python
import asyncio, logging, os, sys
from sqlalchemy import text
from sqlalchemy.exc import OperationalError
from sqlalchemy.ext.asyncio import create_async_engine
from aws_advanced_python_wrapper.aio import release_resources_async

logging.basicConfig(level=logging.WARNING, stream=sys.stdout, format="%(name)s: %(message)s")
logging.getLogger("aws_advanced_python_wrapper.aio.aurora_connection_tracker").setLevel(logging.DEBUG)

CLUSTER_ENDPOINT, DB_NAME, USER, PASSWORD = (os.environ[k] for k in ("PGHOST", "PGDATABASE", "PGUSER", "PGPASSWORD"))
PLUGINS = "" if sys.argv[1] == "default" else "&wrapper_plugins=failover,host_monitoring_v2"

async def main():
engine = create_async_engine(
f"postgresql+aws_wrapper_psycopg://{USER}:{PASSWORD}@{CLUSTER_ENDPOINT}:5432/{DB_NAME}"
f"?wrapper_dialect=aurora-pg{PLUGINS}")
try:
for i in range(3):
try:
async with engine.connect() as conn:
row = await conn.execute(text("SELECT pg_catalog.aurora_db_instance_identifier()"))
print(f"iter {i}: connected to instance {row.scalar_one()}")
except OperationalError as exc:
print(f"iter {i}: operational error ({type(exc.orig).__name__})")
finally:
await engine.dispose()
await release_resources_async()

asyncio.run(main())
```

### Possible Solution

Claude output
> In `aws_advanced_python_wrapper/aio/aurora_connection_tracker.py`, `_pin_current_writer` first pins the writer from topology, which is the instance endpoint. Its stale-topology guard then calls `get_host_role(conn)`, gets WRITER, and replaces the pin with `plugin_service.current_host_info`, which is the URL host, i.e. the cluster endpoint. `_same_host` compares host strings, so the cluster endpoint never equals the instance endpoint. On the first `execute`, `_invalidate_writer_change` compares the cluster-endpoint pin with the instance-endpoint topology writer, reports a writer change, and `invalidate_all` closes every connection keyed under the cluster endpoint, which after `_fill_instance_alias` includes the connection about to execute. The sync tracker does not pin at connect time and only compares instance against instance, so it is unaffected.

### Additional Information/Context

_No response_

### The AWS Advanced Python Wrapper version used

3.1.0

### python version used

3.14.7

### Operating System and version

Debian GNU/Linux 12 (bookworm), aarch64, `python:3.14-slim` image

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.