facebook / facebook/zstd

Multiframe ZSTD file: how to jump to and stream the second file?

Aperta
#4,569 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
C
Stelle
27.9k
Fork
2.6k
Merge medio
1g 3h
PR unite (30g)
8

Descrizione

TL;DR: This is not about bugs of ZSTD but about how to take advantage of its feature:

---

I compress two ndjson files into a [multiframe](https://github.com/facebook/zstd/blob/3bee41a70eaf343fbcae3637b3f6edbe52f35ed8/doc/zstd_compression_format.md#frames) ZST file where each ndjson is compressed into a frame. I have the following metadata `meta_data` (as a list) of the ZST file:

````python
import zstandard as zstd
from pathlib import Path

input_file = r"E:\Personal projects\tmp\test.zst"
input_file = Path(output_file)

meta_data = [{'name' : 'chunk_0.ndjson',
'uncompressed_size' : 2147473321,
'compressed_offset' : 0,
'uncompressed_offset' : 0,
'compressed_size' : 175631248},
{'name' : 'chunk_1.ndjson',
'uncompressed_size' : 2147473321,
'compressed_offset' : 175631248,
'uncompressed_offset' : 2147473321,
'compressed_size' : 175631248}]
````

In Python, how can we leverage the above `meta_data` to seek to `chunk_1.ndjson`, start decompressing, and stream it line-by-line? In this way, we don't need to
- decompress `chunk_0.ndjson`,
- load the whole compressed `chunk_1.ndjson` into the memory.

Thank your for your help.

Guida per i contributori

Apri la guida per i contributori

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.