nodejs / nodejs/node

zstd decoding complete frame in read stream halts stream

未关闭
#64,741 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

stream zlib
主要语言
JavaScript
星标
122k
派生
37.4k
平均合并
4 天 3 小时
30 天内合并 PR
272

描述

Version

v26.5.0

Platform
Darwin Mac.lan 24.6.0 Darwin Kernel Version 24.6.0: Tue Apr 21 20:19:12 PDT 2026; root:xnu-11417.140.69.710.16~1/RELEASE_ARM64_T6041 arm64
Subsystem

No response

What steps will reproduce the bug?
import { Buffer } from 'node:buffer';
import * as consumers from 'node:stream/consumers';
import * as zlib from 'node:zlib';

function consume(codec, writes) {
	const inflate = codec.createDecompress();
	const result = consumers.buffer(inflate);
	writes.forEach(write => inflate.write(write));
	inflate.end();
	return result;
}

function describe(buffer) {
	const aa = buffer.includes(0x41);
	const bb = buffer.includes(0x42);
	const which = function() {
		if (aa && bb) {
			return 'frameA + frameB';
		} else if (aa) {
			return 'frameA only';
		} else if (bb) {
			return 'frameB only';
		} else {
			return 'nothing';
		}
	}();
	return `${buffer.length} bytes (${which})`;
}

async function attempt(codec, writes) {
	try {
		return describe(await consume(codec, writes));
	} catch (error) {
		return `threw: ${error.message}`;
	}
}

const tests = [ {
	name: 'zstd',
	compress: buffer => zlib.zstdCompressSync(buffer),
	createDecompress: () => zlib.createZstdDecompress(),
}, {
	name: 'gzip',
	compress: buffer => zlib.gzipSync(buffer),
	createDecompress: () => zlib.createGunzip(),
} ];

console.log(`node ${process.version}`);
for (const codec of tests) {
	// 'A' / 'B'
	const frameA = codec.compress(Buffer.alloc(2000, 0x41));
	const frameB = codec.compress(Buffer.alloc(2000, 0x42));
	const concatenated = Buffer.concat([ frameA, frameB ]);

	// BUG: both frames in a single write — the boundary between them falls inside one input buffer.
	const bug = await attempt(codec, [ concatenated ]);
	// COUNTER-EXAMPLE: the same two frames, one per write — the boundary lands at a write boundary.
	const counter = await attempt(codec, [ frameA, frameB ]);

	console.log(`${codec.name}:`);
	console.log(`  stream, both frames in ONE write   -> ${bug}`);
	console.log(`  stream, one frame per write        -> ${counter}`);
}
How often does it reproduce? Is there a required condition?

This reproduces every time

What is the expected behavior? Why is that the expected behavior?

Decoding concatenated zstd payloads should decode correctly as one stream. Indeed, zstdcat will decode this correctly (omitted from example but you can take my word for it).

What do you see instead?
marcel[1:42:21PM] [~/xx] ~/Downloads/node-v26.5.0-darwin-arm64/bin/node a.mjs
node v26.5.0
zstd:
  stream, both frames in ONE write   -> 2000 bytes (frameA only)
  stream, one frame per write        -> 4000 bytes (frameA + frameB)
gzip:
  stream, both frames in ONE write   -> 4000 bytes (frameA + frameB)
  stream, one frame per write        -> 4000 bytes (frameA + frameB)
Additional information

This will cause spurious failures while decoding zstd streams. The odds of it happening are a function of stream read size vs zstd chunk size. In my case I ran into this while decoding a stream after about 3GB and tracked it down to this bug.

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

调研方向

首先针对 zlib 流式 API 运行复现脚本,比较在一次写入中拼接的 zstd 帧与每次写入一个帧的情况。跟踪 Node.js zstd 解压缩流的处理,并为单次写入的情况添加回归测试;当两个帧都产生 4000 字节、与 gzip 和 zstdcat 的行为一致时,即表示完成。

由索引模型根据 Issue 内容生成。

评估

技术栈
javascript, nodejs
领域
stream-processing
Issue 类型
缺陷
难度
3/5
预计耗时
1-2 天
活跃度
冷清
描述清晰度
基本清楚
新手友好度
68/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。