aws / aws/aws-advanced-python-wrapper

[aio] aurora_connection_tracker closes its own connection on the first statement of every cluster-endpoint connection

Open
#1,276 1 comment 0 reactions 0 assignees View on GitHub
bug
Dominant language
Python
Stars
98
Forks
22
Avg merge
1d 9h
Merged PRs (30d)
5

Description

### Describe the bug

With the async wrapper (`aws_advanced_python_wrapper.aio`, via the `postgresql+aws_wrapper_psycopg` SQLAlchemy dialect) and the default plugin chain, every new connection to an Aurora PostgreSQL cluster writer endpoint fails on its first statement with `FailoverSuccessError`. The retries also open several connections per attempt (the failover's writer connection plus topology probes), which push a huge spike in connection count.

### Expected Behavior

A connection made through the cluster writer endpoint with the default plugins should run its first statement normally, as it does with `wrapper_plugins=failover,host_monitoring_v2` (the chain used in `docs/examples/PGSQLAlchemyAsyncFailover.py`).

### What plugins are used? What other connection properties were set?

Default chain (`wrapper_plugins` not set, so `initial_connection,aurora_connection_tracker,failover_v2,host_monitoring_v2`), `wrapper_dialect=aurora-pg`. Also reproduced with `wrapper_plugins=aurora_connection_tracker` alone. Does not reproduce with `wrapper_plugins=failover,host_monitoring_v2` or `failover_v2` alone.

### Current Behavior

Every connection, on its first statement:

```
aws_advanced_python_wrapper.aio.aurora_connection_tracker: [OpenedConnectionTracker] Invalidating opened connections to host: .cluster-.us-east-1.rds.amazonaws.com:5432/
aws_advanced_python_wrapper.aio.aurora_connection_tracker: [OpenedConnectionTracker] Opened Connections Tracked:
aws_advanced_python_wrapper.aio.aurora_connection_tracker: [AsyncAuroraConnectionTrackerPlugin] failover handler: pre=.cluster-.us-east-1.rds.amazonaws.com:5432/ post=tf-2026061217...chnc4uhow0rs.us-east-1.rds.amazonaws.com:5432/ pinned=tf-2026061217...chnc4uhow0rs.us-east-1.rds.amazonaws.com:5432/
aws_advanced_python_wrapper.aio.aurora_connection_tracker: [OpenedConnectionTracker] Invalidating opened connections to host: .cluster-.us-east-1.rds.amazonaws.com:5432/
iter 0: operational error (FailoverSuccessError)
```

### Reproduction Steps

```python
import asyncio, logging, os, sys
from sqlalchemy import text
from sqlalchemy.exc import OperationalError
from sqlalchemy.ext.asyncio import create_async_engine
from aws_advanced_python_wrapper.aio import release_resources_async

logging.basicConfig(level=logging.WARNING, stream=sys.stdout, format="%(name)s: %(message)s")
logging.getLogger("aws_advanced_python_wrapper.aio.aurora_connection_tracker").setLevel(logging.DEBUG)

CLUSTER_ENDPOINT, DB_NAME, USER, PASSWORD = (os.environ[k] for k in ("PGHOST", "PGDATABASE", "PGUSER", "PGPASSWORD"))
PLUGINS = "" if sys.argv[1] == "default" else "&wrapper_plugins=failover,host_monitoring_v2"

async def main():
engine = create_async_engine(
f"postgresql+aws_wrapper_psycopg://{USER}:{PASSWORD}@{CLUSTER_ENDPOINT}:5432/{DB_NAME}"
f"?wrapper_dialect=aurora-pg{PLUGINS}")
try:
for i in range(3):
try:
async with engine.connect() as conn:
row = await conn.execute(text("SELECT pg_catalog.aurora_db_instance_identifier()"))
print(f"iter {i}: connected to instance {row.scalar_one()}")
except OperationalError as exc:
print(f"iter {i}: operational error ({type(exc.orig).__name__})")
finally:
await engine.dispose()
await release_resources_async()

asyncio.run(main())
```

### Possible Solution

Claude output
> In `aws_advanced_python_wrapper/aio/aurora_connection_tracker.py`, `_pin_current_writer` first pins the writer from topology, which is the instance endpoint. Its stale-topology guard then calls `get_host_role(conn)`, gets WRITER, and replaces the pin with `plugin_service.current_host_info`, which is the URL host, i.e. the cluster endpoint. `_same_host` compares host strings, so the cluster endpoint never equals the instance endpoint. On the first `execute`, `_invalidate_writer_change` compares the cluster-endpoint pin with the instance-endpoint topology writer, reports a writer change, and `invalidate_all` closes every connection keyed under the cluster endpoint, which after `_fill_instance_alias` includes the connection about to execute. The sync tracker does not pin at connect time and only compares instance against instance, so it is unaffected.

### Additional Information/Context

_No response_

### The AWS Advanced Python Wrapper version used

3.1.0

### python version used

3.14.7

### Operating System and version

Debian GNU/Linux 12 (bookworm), aarch64, `python:3.14-slim` image

Contributor guide

Open the contributing guide

Research direction

Start with aws_advanced_python_wrapper/aio/aurora_connection_tracker.py, especially _pin_current_writer, _invalidate_writer_change, _same_host, and _fill_instance_alias, then run the provided SQLAlchemy async reproduction with the default plugin chain. Trace how the pinned host compares with the topology writer and verify that the first statement succeeds without invalidating its own connection, while the documented failover,host_monitoring_v2 chain remains a control.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, postgresql, python, sqlalchemy
Domain
databases
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Clearly specified
Newbie friendliness
72/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.