Node.js 24 LTS HTTPS default Agent keep-alive behavior can cause long-running requests through AWS Global Accelerator to fail with read ETIMEDOUT
还没有人认领这个 Issue。
- 主要语言
- JavaScript
- 星标
- 122k
- 派生
- 37.4k
- 平均合并
- 4 天 3 小时
- 30 天内合并 PR
- 272
描述
Summary
I may be missing an intended Agent configuration here, so I am opening this first as a question / interoperability report rather than a confirmed bug.
We are seeing a behavior difference between Node.js v22 and Node.js v24 LTS for a long-running HTTPS POST request routed through AWS Global Accelerator.
The same request:
- succeeds with Node.js v22.22.2 through AWS Global Accelerator
- succeeds with Node.js v24.15.0 when routed directly to the backend ALB hostname
- fails with Node.js v24.15.0 when routed through AWS Global Accelerator
- succeeds again with Node.js v24.15.0 through AWS Global Accelerator when using a custom HTTPS Agent with
keepAlive: falseandConnection: close
This issue is not about a proprietary API. The private endpoint used in our tests cannot be shared, but the behavior appears to be related to Node.js HTTPS Agent / TCP keepalive behavior for long-running requests where no HTTP response data is returned for about 60 seconds.
We are opening this issue to ask whether this Node.js v24 LTS behavior is expected, whether the current workaround is the recommended one, and whether Node.js should expose or document more precise controls for this case.
Environment
Client environment:
- Node.js versions tested:
- v22.22.2
- v24.15.0
- Platform:
- macOS arm64
- Protocol:
- HTTPS over TCP
- HTTP client:
- built-in
node:https https.request()
- built-in
- Request type:
- one long-running HTTPS POST request
- backend intentionally delays the HTTP response for about 60 seconds
Network path variants:
- Client -> AWS Application Load Balancer directly
- Client -> AWS Global Accelerator -> AWS Application Load Balancer
Observed behavior
| Node.js version | Network path | Result |
|---|---|---|
| v22.22.2 | Direct ALB hostname | Success after ~60s |
| v24.15.0 | Direct ALB hostname | Success after ~60s |
| v22.22.2 | AWS Global Accelerator hostname | Success after ~60s |
| v24.15.0 | AWS Global Accelerator hostname | Fails after ~39s with read ETIMEDOUT |
| v24.15.0 | AWS Global Accelerator hostname + new https.Agent({ keepAlive: false }) + Connection: close |
Success after ~60s |
The failing Node.js v24.15.0 request has an application-level request timeout configured much higher than the failure time, for example 180000 ms, so this is not caused by the application request timeout.
The observed error is:
read ETIMEDOUT
The error appears to come from the socket/TLS layer rather than from the HTTP request timeout callback.
Packet capture findings
AWS Support reviewed packet captures from the four test cases above.
Their analysis was:
- Node.js v22 sends an initial TCP keepalive probe after a short idle period, but does not continue sending probes every 1 second afterwards.
- Node.js v24 sends an initial TCP keepalive probe after a short idle period, then continues sending TCP keepalive probes at roughly 1 second intervals.
- Through the direct ALB path, those probes are acknowledged and the request succeeds.
- Through AWS Global Accelerator, the first probes are acknowledged, but after a while further zero-byte TCP keepalive probes are not acknowledged.
- After 10 unacknowledged TCP keepalive probes, the client resets the connection.
- This results in the observed failure at about 39 seconds.
In the failing Node.js v24 + Global Accelerator case, the packet capture did not show a remote FIN/RST from Global Accelerator before the client-side error. The visible reset was client-originated after the unacknowledged keepalive probes.
Minimal code shape
The failing version uses the default HTTPS agent behavior:
import https from "node:https";
const req = https.request(
process.env.TARGET_URL,
{
method: "POST",
timeout: 180_000,
headers: {
"Content-Type": "application/json",
"Content-Length": "0",
},
},
(res) => {
res.resume();
res.on("end", () => {
console.log("status", res.statusCode);
});
},
);
req.on("timeout", () => {
console.error("request timeout");
req.destroy(new Error("request timeout"));
});
req.on("error", (err) => {
console.error("request error", err);
});
req.end();
The workaround that succeeds is:
import https from "node:https";
const agent = new https.Agent({ keepAlive: false });
const req = https.request(
process.env.TARGET_URL,
{
method: "POST",
timeout: 180_000,
agent,
headers: {
"Content-Type": "application/json",
"Content-Length": "0",
Connection: "close",
},
},
(res) => {
res.resume();
res.on("end", () => {
console.log("status", res.statusCode);
});
},
);
req.on("timeout", () => {
console.error("request timeout");
req.destroy(new Error("request timeout"));
});
req.on("error", (err) => {
console.error("request error", err);
});
req.end();
Why this looks related to Node.js Agent / TCP keepalive behavior
The Node.js documentation says that:
AgenthaskeepAliveMsecs, which specifies the initial delay for TCP keepalive packets whenkeepAliveis used.agent.keepSocketAlive(socket)defaults to callingsocket.setKeepAlive(true, this.keepAliveMsecs).socket.setKeepAlive()enables TCP keepalive and sets socket options includingTCP_KEEPCNT=10andTCP_KEEPINTVL=1.
This matches the packet capture pattern reported by AWS Support: approximately 1 second interval probes and failure after 10 unacknowledged probes.
Questions
-
Is the observed Node.js v24 LTS behavior expected when using the default HTTPS global agent?
-
Is
new https.Agent({ keepAlive: false })plusConnection: closethe recommended application-level workaround for long-running single requests through TCP middleboxes that may not acknowledge zero-byte TCP keepalive probes? -
Is there a supported way in Node.js to configure the TCP keepalive interval and probe count per request or per Agent?
For this case, being able to configure only the initial delay is not enough. AWS recommended settings equivalent to:
- first probe after 60 seconds
- probe interval 15 seconds
- max unacknowledged probes 10
-
If Node.js intentionally sets
TCP_KEEPINTVL=1andTCP_KEEPCNT=10when Agent keepalive is enabled, should this be documented more prominently for long-running HTTPS requests? -
Would Node.js consider exposing per-socket or per-Agent options for TCP keepalive interval and probe count, where supported by the operating system?
Expected outcome
We are not necessarily claiming this is a Node.js bug. We would like confirmation on whether this is expected behavior and what the recommended Node.js-level mitigation should be.
If this is working as designed, a documentation clarification may be enough.
If the current API does not allow applications to tune the relevant TCP keepalive parameters without changing OS-level defaults, an API enhancement might be useful for long-running HTTP(S) clients running through TCP proxies, load balancers, accelerators, or other middleboxes.
Additional notes
AWS Global Accelerator documentation states that TCP keepalive packets with no payload should not be relied on to keep Global Accelerator connections active. In our case the failure happens much earlier than the documented Global Accelerator idle timeout because the Node.js v24 client appears to close/reset the connection after its own TCP keepalive probes go unacknowledged.
The key operational concern is that Node.js v24 is an LTS release, and this behavior can affect long-running HTTPS requests that were previously successful with Node.js v22.
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
调研方向
从内置的 node:https Agent 入口点以及报告中描述的 socket.setKeepAlive() 行为开始,然后在 Node.js 22 和 24 上通过 direct 和 Global Accelerator 路径复现长时间运行的请求。确认观察到的 keepalive 设置和变通方案是否符合预期,并记录受支持的控制方式或对 API 进行增强的需求。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- javascript, node.js
- 领域
- backend, networking
- Issue 类型
- 文档
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 活跃度
- 活跃
- 描述清晰度
- 基本清楚
- 新手友好度
- 48/100