nodejs / nodejs/node

zstd decoding complete frame in read stream halts stream

Aperta
#64,741 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

stream zlib
Lingua principale
JavaScript
Stelle
122k
Fork
37.3k
Merge medio
4g 2h
PR unite (30g)
283

Descrizione

Version

v26.5.0

Platform
Darwin Mac.lan 24.6.0 Darwin Kernel Version 24.6.0: Tue Apr 21 20:19:12 PDT 2026; root:xnu-11417.140.69.710.16~1/RELEASE_ARM64_T6041 arm64
Subsystem

No response

What steps will reproduce the bug?
import { Buffer } from 'node:buffer';
import * as consumers from 'node:stream/consumers';
import * as zlib from 'node:zlib';

function consume(codec, writes) {
	const inflate = codec.createDecompress();
	const result = consumers.buffer(inflate);
	writes.forEach(write => inflate.write(write));
	inflate.end();
	return result;
}

function describe(buffer) {
	const aa = buffer.includes(0x41);
	const bb = buffer.includes(0x42);
	const which = function() {
		if (aa && bb) {
			return 'frameA + frameB';
		} else if (aa) {
			return 'frameA only';
		} else if (bb) {
			return 'frameB only';
		} else {
			return 'nothing';
		}
	}();
	return `${buffer.length} bytes (${which})`;
}

async function attempt(codec, writes) {
	try {
		return describe(await consume(codec, writes));
	} catch (error) {
		return `threw: ${error.message}`;
	}
}

const tests = [ {
	name: 'zstd',
	compress: buffer => zlib.zstdCompressSync(buffer),
	createDecompress: () => zlib.createZstdDecompress(),
}, {
	name: 'gzip',
	compress: buffer => zlib.gzipSync(buffer),
	createDecompress: () => zlib.createGunzip(),
} ];

console.log(`node ${process.version}`);
for (const codec of tests) {
	// 'A' / 'B'
	const frameA = codec.compress(Buffer.alloc(2000, 0x41));
	const frameB = codec.compress(Buffer.alloc(2000, 0x42));
	const concatenated = Buffer.concat([ frameA, frameB ]);

	// BUG: both frames in a single write — the boundary between them falls inside one input buffer.
	const bug = await attempt(codec, [ concatenated ]);
	// COUNTER-EXAMPLE: the same two frames, one per write — the boundary lands at a write boundary.
	const counter = await attempt(codec, [ frameA, frameB ]);

	console.log(`${codec.name}:`);
	console.log(`  stream, both frames in ONE write   -> ${bug}`);
	console.log(`  stream, one frame per write        -> ${counter}`);
}
How often does it reproduce? Is there a required condition?

This reproduces every time

What is the expected behavior? Why is that the expected behavior?

Decoding concatenated zstd payloads should decode correctly as one stream. Indeed, zstdcat will decode this correctly (omitted from example but you can take my word for it).

What do you see instead?
marcel[1:42:21PM] [~/xx] ~/Downloads/node-v26.5.0-darwin-arm64/bin/node a.mjs
node v26.5.0
zstd:
  stream, both frames in ONE write   -> 2000 bytes (frameA only)
  stream, one frame per write        -> 4000 bytes (frameA + frameB)
gzip:
  stream, both frames in ONE write   -> 4000 bytes (frameA + frameB)
  stream, one frame per write        -> 4000 bytes (frameA + frameB)
Additional information

This will cause spurious failures while decoding zstd streams. The odds of it happening are a function of stream read size vs zstd chunk size. In my case I ran into this while decoding a stream after about 3GB and tracked it down to this bug.

Guida per i contributori

Apri la guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Direzione di ricerca

Inizia eseguendo lo script di riproduzione sulle API di streaming di zlib, confrontando frame zstd concatenati in una singola scrittura con un frame per scrittura. Traccia la gestione dello stream di decompressione zstd di Node.js e aggiungi un test di regressione per il caso di una singola scrittura; il lavoro è completato quando entrambi i frame producono 4000 byte, come nel comportamento di gzip e zstdcat.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
javascript, nodejs
Ambito
stream-processing
Tipo di issue
Bug
Difficoltà
3/5
Tempo stimato
1-2 giorni
Stato di attività
Tranquilla
Chiarezza
Abbastanza chiara
Idoneità per principianti
68/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.