aws / aws/aws-lambda-web-adapter

Freeze/thaw poisons pooled localhost connection on non-SnapStart functions → SendRequest/ConnectionReset with no retry; no config to disable pooling

Aperta
#830 2 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
Rust
Stelle
2.7k
Fork
162
Merge medio
2g 10h
PR unite (30g)
2

Descrizione

### Summary

On a regular (non-SnapStart) on-demand function, the execution environment freeze/thaw between invocations can leave the adapter's pooled keep-alive connection to the app server in a reset state. The next invocation reuses that dead connection, and because `fetch_response` forwards the request exactly once with no retry, the caller gets a `5xx`. There is currently no configuration to disable the adapter↔app connection pool outside of SnapStart.

### Setup

- Adapter: `public.ecr.aws/awsguru/aws-lambda-adapter:1.0.1` (x86_64), as a Lambda extension in a container image
- App: Next.js 16 standalone server on `:3000`
- Base image: `node:24-bookworm-slim`
- Front: API Gateway v2, payload format 2.0, `$default` route
- Region: `eu-central-1`
- Traffic: low, function scales to zero frequently
- `AWS_LWA_READINESS_CHECK_PATH=/api/health`, `AWS_LWA_READINESS_CHECK_PROTOCOL=tcp`

### Symptom

Intermittent `500 {"message":"Internal Server Error"}` to the client. Adapter log on the failing invocation:

```
ERROR Lambda runtime invoke{requestId="7ab2a21b-..."}:
hyper_util::client::legacy::Error(SendRequest,
hyper::Error(Io, Os { code: 104, kind: ConnectionReset, message: "Connection reset by peer" }))
```

The next request to the same execution environment succeeds. The app server does **not** crash — no error on its side, and it keeps serving.

### Timeline (one execution environment, from CloudWatch)

```
15:49:43.137 new exec env starts: app "Ready", "lambda-adapter State: Ready"
15:49:43.141 req A durationMs 760 initDurationMs 562 status success <- cold-start request, OK
15:49:52.105 req B durationMs 85 (warm, no init) status success <- reuses pooled conn, OK
15:49:55.823 req C durationMs 2486 (warm, no init) ERROR ConnectionReset <- 500 to caller
```

Req C is ~3.6s after req B, on a warm environment. The environment was frozen in between.

### Analysis

The adapter builds its client with a keep-alive pool:

https://github.com/aws/aws-lambda-web-adapter/blob/v1.0.1/src/lib.rs#L588-L600

```rust
if env::var("AWS_LAMBDA_INITIALIZATION_TYPE").as_deref() == Ok("snap-start") {
builder.pool_max_idle_per_host(0);
} else {
builder.pool_idle_timeout(Duration::from_secs(4));
}
```

`pool_idle_timeout` is measured with `Instant`/`CLOCK_MONOTONIC`, which does not advance while the execution environment is frozen. So a connection that has been idle far longer than 4s in wall-clock time still looks fresh to the pool and is reused — but the underlying socket was reset across the freeze/thaw. `fetch_response` then does:

https://github.com/aws/aws-lambda-web-adapter/blob/v1.0.1/src/lib.rs#L1012

```rust
let mut app_response = self.client.request(request).await?;
```

— a single attempt. The connection error propagates straight out as a runtime error → `5xx`.

This is the same failure mode the SnapStart branch above guards against with `pool_max_idle_per_host(0)`, but on-demand environments freeze/thaw as well and hit it too. Related prior reports of the `SendRequest` / `IncompleteMessage` family: #415, #295.

### Request

Either (ideally both):

1. **Expose the pool control as configuration** — e.g. `AWS_LWA_DISABLE_CONNECTION_POOL=true` (or `AWS_LWA_POOL_MAX_IDLE_PER_HOST=0`) so non-SnapStart functions can opt into `pool_max_idle_per_host(0)`. For localhost the cost of a fresh connection per request is negligible.
2. **Retry the forwarded request once on connection-level errors** (`ConnectionReset`, `BrokenPipe`, `IncompleteMessage` when zero bytes were written). These are safe to retry — no request bytes reached the app — and every mainstream HTTP client does this for pooled connections.

### Workarounds tried

- `AWS_LWA_READINESS_CHECK_PROTOCOL=tcp` — removes the connection the readiness probe would otherwise seed into the pool, but not request-to-request reuse (this report is with `tcp` already set).
- Client-side retry at the caller — works for service-to-service callers, not for browsers hitting the function directly.

Guida per i contributori

Apri la guida per i contributori

Direzione di ricerca

Inizia in src/lib.rs, nel client builder intorno alle righe 588-600, e in fetch_response intorno alla riga 1012; poi segui il modo in cui le richieste sottoposte a pool espongono gli errori di connessione. Esamina i test esistenti per la configurazione e il forwarding del client, se presenti. L’attività è completata quando il problema di freeze/thaw è stato risolto tramite il controllo del pool richiesto o un retry sicuro in caso di errore di connessione, con una copertura di regressione.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
aws, rust
Ambito
backend, cloud, networking
Tipo di issue
Bug
Difficoltà
4/5
Tempo stimato
3-5 giorni
Stato di attività
Attiva
Chiarezza
Abbastanza chiara
Idoneità per principianti
58/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.