Node.js 24 LTS HTTPS default Agent keep-alive behavior can cause long-running requests through AWS Global Accelerator to fail with read ETIMEDOUT
まだ誰も着手していません。
- 主要言語
- JavaScript
- スター
- 122k
- フォーク
- 37.3k
- 平均マージ
- 4日 2時間
- マージ済み PR(30日)
- 283
説明
Summary
I may be missing an intended Agent configuration here, so I am opening this first as a question / interoperability report rather than a confirmed bug.
We are seeing a behavior difference between Node.js v22 and Node.js v24 LTS for a long-running HTTPS POST request routed through AWS Global Accelerator.
The same request:
- succeeds with Node.js v22.22.2 through AWS Global Accelerator
- succeeds with Node.js v24.15.0 when routed directly to the backend ALB hostname
- fails with Node.js v24.15.0 when routed through AWS Global Accelerator
- succeeds again with Node.js v24.15.0 through AWS Global Accelerator when using a custom HTTPS Agent with
keepAlive: falseandConnection: close
This issue is not about a proprietary API. The private endpoint used in our tests cannot be shared, but the behavior appears to be related to Node.js HTTPS Agent / TCP keepalive behavior for long-running requests where no HTTP response data is returned for about 60 seconds.
We are opening this issue to ask whether this Node.js v24 LTS behavior is expected, whether the current workaround is the recommended one, and whether Node.js should expose or document more precise controls for this case.
Environment
Client environment:
- Node.js versions tested:
- v22.22.2
- v24.15.0
- Platform:
- macOS arm64
- Protocol:
- HTTPS over TCP
- HTTP client:
- built-in
node:https https.request()
- built-in
- Request type:
- one long-running HTTPS POST request
- backend intentionally delays the HTTP response for about 60 seconds
Network path variants:
- Client -> AWS Application Load Balancer directly
- Client -> AWS Global Accelerator -> AWS Application Load Balancer
Observed behavior
| Node.js version | Network path | Result |
|---|---|---|
| v22.22.2 | Direct ALB hostname | Success after ~60s |
| v24.15.0 | Direct ALB hostname | Success after ~60s |
| v22.22.2 | AWS Global Accelerator hostname | Success after ~60s |
| v24.15.0 | AWS Global Accelerator hostname | Fails after ~39s with read ETIMEDOUT |
| v24.15.0 | AWS Global Accelerator hostname + new https.Agent({ keepAlive: false }) + Connection: close |
Success after ~60s |
The failing Node.js v24.15.0 request has an application-level request timeout configured much higher than the failure time, for example 180000 ms, so this is not caused by the application request timeout.
The observed error is:
read ETIMEDOUT
The error appears to come from the socket/TLS layer rather than from the HTTP request timeout callback.
Packet capture findings
AWS Support reviewed packet captures from the four test cases above.
Their analysis was:
- Node.js v22 sends an initial TCP keepalive probe after a short idle period, but does not continue sending probes every 1 second afterwards.
- Node.js v24 sends an initial TCP keepalive probe after a short idle period, then continues sending TCP keepalive probes at roughly 1 second intervals.
- Through the direct ALB path, those probes are acknowledged and the request succeeds.
- Through AWS Global Accelerator, the first probes are acknowledged, but after a while further zero-byte TCP keepalive probes are not acknowledged.
- After 10 unacknowledged TCP keepalive probes, the client resets the connection.
- This results in the observed failure at about 39 seconds.
In the failing Node.js v24 + Global Accelerator case, the packet capture did not show a remote FIN/RST from Global Accelerator before the client-side error. The visible reset was client-originated after the unacknowledged keepalive probes.
Minimal code shape
The failing version uses the default HTTPS agent behavior:
import https from "node:https";
const req = https.request(
process.env.TARGET_URL,
{
method: "POST",
timeout: 180_000,
headers: {
"Content-Type": "application/json",
"Content-Length": "0",
},
},
(res) => {
res.resume();
res.on("end", () => {
console.log("status", res.statusCode);
});
},
);
req.on("timeout", () => {
console.error("request timeout");
req.destroy(new Error("request timeout"));
});
req.on("error", (err) => {
console.error("request error", err);
});
req.end();
The workaround that succeeds is:
import https from "node:https";
const agent = new https.Agent({ keepAlive: false });
const req = https.request(
process.env.TARGET_URL,
{
method: "POST",
timeout: 180_000,
agent,
headers: {
"Content-Type": "application/json",
"Content-Length": "0",
Connection: "close",
},
},
(res) => {
res.resume();
res.on("end", () => {
console.log("status", res.statusCode);
});
},
);
req.on("timeout", () => {
console.error("request timeout");
req.destroy(new Error("request timeout"));
});
req.on("error", (err) => {
console.error("request error", err);
});
req.end();
Why this looks related to Node.js Agent / TCP keepalive behavior
The Node.js documentation says that:
AgenthaskeepAliveMsecs, which specifies the initial delay for TCP keepalive packets whenkeepAliveis used.agent.keepSocketAlive(socket)defaults to callingsocket.setKeepAlive(true, this.keepAliveMsecs).socket.setKeepAlive()enables TCP keepalive and sets socket options includingTCP_KEEPCNT=10andTCP_KEEPINTVL=1.
This matches the packet capture pattern reported by AWS Support: approximately 1 second interval probes and failure after 10 unacknowledged probes.
Questions
-
Is the observed Node.js v24 LTS behavior expected when using the default HTTPS global agent?
-
Is
new https.Agent({ keepAlive: false })plusConnection: closethe recommended application-level workaround for long-running single requests through TCP middleboxes that may not acknowledge zero-byte TCP keepalive probes? -
Is there a supported way in Node.js to configure the TCP keepalive interval and probe count per request or per Agent?
For this case, being able to configure only the initial delay is not enough. AWS recommended settings equivalent to:
- first probe after 60 seconds
- probe interval 15 seconds
- max unacknowledged probes 10
-
If Node.js intentionally sets
TCP_KEEPINTVL=1andTCP_KEEPCNT=10when Agent keepalive is enabled, should this be documented more prominently for long-running HTTPS requests? -
Would Node.js consider exposing per-socket or per-Agent options for TCP keepalive interval and probe count, where supported by the operating system?
Expected outcome
We are not necessarily claiming this is a Node.js bug. We would like confirmation on whether this is expected behavior and what the recommended Node.js-level mitigation should be.
If this is working as designed, a documentation clarification may be enough.
If the current API does not allow applications to tune the relevant TCP keepalive parameters without changing OS-level defaults, an API enhancement might be useful for long-running HTTP(S) clients running through TCP proxies, load balancers, accelerators, or other middleboxes.
Additional notes
AWS Global Accelerator documentation states that TCP keepalive packets with no payload should not be relied on to keep Global Accelerator connections active. In our case the failure happens much earlier than the documented Global Accelerator idle timeout because the Node.js v24 client appears to close/reset the connection after its own TCP keepalive probes go unacknowledged.
The key operational concern is that Node.js v24 is an LTS release, and this behavior can affect long-running HTTPS requests that were previously successful with Node.js v22.
コントリビューションガイド
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
調査の方向性
組み込みの node:https Agent エントリポイントと、レポートに記載された socket.setKeepAlive() の動作から始め、その後、Node.js 22 と 24 で direct パスおよび Global Accelerator パスに対する長時間実行リクエストを再現してください。観測された keepalive 設定と回避策が想定されたものかどうかを確認し、サポートされている制御方法、または API 拡張の必要性を文書化してください。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- javascript, node.js
- 領域
- backend, networking
- issue の種類
- ドキュメント
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 活発さ
- 活発
- 明瞭さ
- おおむね明確
- 初心者へのやさしさ
- 48/100