nodejs / nodejs/node

Node.js 24 LTS HTTPS default Agent keep-alive behavior can cause long-running requests through AWS Global Accelerator to fail with read ETIMEDOUT

Aperta
#63,043 4 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

stale
Lingua principale
JavaScript
Stelle
122k
Fork
37.3k
Merge medio
4g 2h
PR unite (30g)
283

Descrizione

Summary

I may be missing an intended Agent configuration here, so I am opening this first as a question / interoperability report rather than a confirmed bug.

We are seeing a behavior difference between Node.js v22 and Node.js v24 LTS for a long-running HTTPS POST request routed through AWS Global Accelerator.

The same request:

  • succeeds with Node.js v22.22.2 through AWS Global Accelerator
  • succeeds with Node.js v24.15.0 when routed directly to the backend ALB hostname
  • fails with Node.js v24.15.0 when routed through AWS Global Accelerator
  • succeeds again with Node.js v24.15.0 through AWS Global Accelerator when using a custom HTTPS Agent with keepAlive: false and Connection: close

This issue is not about a proprietary API. The private endpoint used in our tests cannot be shared, but the behavior appears to be related to Node.js HTTPS Agent / TCP keepalive behavior for long-running requests where no HTTP response data is returned for about 60 seconds.

We are opening this issue to ask whether this Node.js v24 LTS behavior is expected, whether the current workaround is the recommended one, and whether Node.js should expose or document more precise controls for this case.

Environment

Client environment:

  • Node.js versions tested:
    • v22.22.2
    • v24.15.0
  • Platform:
    • macOS arm64
  • Protocol:
    • HTTPS over TCP
  • HTTP client:
    • built-in node:https
    • https.request()
  • Request type:
    • one long-running HTTPS POST request
    • backend intentionally delays the HTTP response for about 60 seconds

Network path variants:

  1. Client -> AWS Application Load Balancer directly
  2. Client -> AWS Global Accelerator -> AWS Application Load Balancer
Observed behavior
Node.js version Network path Result
v22.22.2 Direct ALB hostname Success after ~60s
v24.15.0 Direct ALB hostname Success after ~60s
v22.22.2 AWS Global Accelerator hostname Success after ~60s
v24.15.0 AWS Global Accelerator hostname Fails after ~39s with read ETIMEDOUT
v24.15.0 AWS Global Accelerator hostname + new https.Agent({ keepAlive: false }) + Connection: close Success after ~60s

The failing Node.js v24.15.0 request has an application-level request timeout configured much higher than the failure time, for example 180000 ms, so this is not caused by the application request timeout.

The observed error is:

read ETIMEDOUT

The error appears to come from the socket/TLS layer rather than from the HTTP request timeout callback.

Packet capture findings

AWS Support reviewed packet captures from the four test cases above.

Their analysis was:

  • Node.js v22 sends an initial TCP keepalive probe after a short idle period, but does not continue sending probes every 1 second afterwards.
  • Node.js v24 sends an initial TCP keepalive probe after a short idle period, then continues sending TCP keepalive probes at roughly 1 second intervals.
  • Through the direct ALB path, those probes are acknowledged and the request succeeds.
  • Through AWS Global Accelerator, the first probes are acknowledged, but after a while further zero-byte TCP keepalive probes are not acknowledged.
  • After 10 unacknowledged TCP keepalive probes, the client resets the connection.
  • This results in the observed failure at about 39 seconds.

In the failing Node.js v24 + Global Accelerator case, the packet capture did not show a remote FIN/RST from Global Accelerator before the client-side error. The visible reset was client-originated after the unacknowledged keepalive probes.

Minimal code shape

The failing version uses the default HTTPS agent behavior:

import https from "node:https";

const req = https.request(
  process.env.TARGET_URL,
  {
    method: "POST",
    timeout: 180_000,
    headers: {
      "Content-Type": "application/json",
      "Content-Length": "0",
    },
  },
  (res) => {
    res.resume();
    res.on("end", () => {
      console.log("status", res.statusCode);
    });
  },
);

req.on("timeout", () => {
  console.error("request timeout");
  req.destroy(new Error("request timeout"));
});

req.on("error", (err) => {
  console.error("request error", err);
});

req.end();

The workaround that succeeds is:

import https from "node:https";

const agent = new https.Agent({ keepAlive: false });

const req = https.request(
  process.env.TARGET_URL,
  {
    method: "POST",
    timeout: 180_000,
    agent,
    headers: {
      "Content-Type": "application/json",
      "Content-Length": "0",
      Connection: "close",
    },
  },
  (res) => {
    res.resume();
    res.on("end", () => {
      console.log("status", res.statusCode);
    });
  },
);

req.on("timeout", () => {
  console.error("request timeout");
  req.destroy(new Error("request timeout"));
});

req.on("error", (err) => {
  console.error("request error", err);
});

req.end();
Why this looks related to Node.js Agent / TCP keepalive behavior

The Node.js documentation says that:

  • Agent has keepAliveMsecs, which specifies the initial delay for TCP keepalive packets when keepAlive is used.
  • agent.keepSocketAlive(socket) defaults to calling socket.setKeepAlive(true, this.keepAliveMsecs).
  • socket.setKeepAlive() enables TCP keepalive and sets socket options including TCP_KEEPCNT=10 and TCP_KEEPINTVL=1.

This matches the packet capture pattern reported by AWS Support: approximately 1 second interval probes and failure after 10 unacknowledged probes.

Questions
  1. Is the observed Node.js v24 LTS behavior expected when using the default HTTPS global agent?

  2. Is new https.Agent({ keepAlive: false }) plus Connection: close the recommended application-level workaround for long-running single requests through TCP middleboxes that may not acknowledge zero-byte TCP keepalive probes?

  3. Is there a supported way in Node.js to configure the TCP keepalive interval and probe count per request or per Agent?

    For this case, being able to configure only the initial delay is not enough. AWS recommended settings equivalent to:

    • first probe after 60 seconds
    • probe interval 15 seconds
    • max unacknowledged probes 10
  4. If Node.js intentionally sets TCP_KEEPINTVL=1 and TCP_KEEPCNT=10 when Agent keepalive is enabled, should this be documented more prominently for long-running HTTPS requests?

  5. Would Node.js consider exposing per-socket or per-Agent options for TCP keepalive interval and probe count, where supported by the operating system?

Expected outcome

We are not necessarily claiming this is a Node.js bug. We would like confirmation on whether this is expected behavior and what the recommended Node.js-level mitigation should be.

If this is working as designed, a documentation clarification may be enough.

If the current API does not allow applications to tune the relevant TCP keepalive parameters without changing OS-level defaults, an API enhancement might be useful for long-running HTTP(S) clients running through TCP proxies, load balancers, accelerators, or other middleboxes.

Additional notes

AWS Global Accelerator documentation states that TCP keepalive packets with no payload should not be relied on to keep Global Accelerator connections active. In our case the failure happens much earlier than the documented Global Accelerator idle timeout because the Node.js v24 client appears to close/reset the connection after its own TCP keepalive probes go unacknowledged.

The key operational concern is that Node.js v24 is an LTS release, and this behavior can affect long-running HTTPS requests that were previously successful with Node.js v22.

Guida per i contributori

Apri la guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Direzione di ricerca

Inizia dai punti di ingresso integrati di node:https Agent e dal comportamento di socket.setKeepAlive() descritto nel report, quindi riproduci la richiesta di lunga durata attraverso i percorsi diretto e Global Accelerator su Node.js 22 e 24. Conferma se le impostazioni keepalive osservate e il workaround sono previste e documenta i controlli supportati o la necessità di un miglioramento dell’API.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
javascript, node.js
Ambito
backend, networking
Tipo di issue
Documentazione
Difficoltà
4/5
Tempo stimato
3-5 giorni
Stato di attività
Attiva
Chiarezza
Abbastanza chiara
Idoneità per principianti
48/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.