facebook / facebook/zstd

Multiframe ZSTD file: how to jump to and stream the second file?

Abierto
#4,569 0 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
C
Estrellas
27.9k
Forks
2.6k
Merge medio
1 d 3 h
PR fusionados (30 d)
8

Descripción

TL;DR: This is not about bugs of ZSTD but about how to take advantage of its feature:

---

I compress two ndjson files into a [multiframe](https://github.com/facebook/zstd/blob/3bee41a70eaf343fbcae3637b3f6edbe52f35ed8/doc/zstd_compression_format.md#frames) ZST file where each ndjson is compressed into a frame. I have the following metadata `meta_data` (as a list) of the ZST file:

````python
import zstandard as zstd
from pathlib import Path

input_file = r"E:\Personal projects\tmp\test.zst"
input_file = Path(output_file)

meta_data = [{'name' : 'chunk_0.ndjson',
'uncompressed_size' : 2147473321,
'compressed_offset' : 0,
'uncompressed_offset' : 0,
'compressed_size' : 175631248},
{'name' : 'chunk_1.ndjson',
'uncompressed_size' : 2147473321,
'compressed_offset' : 175631248,
'uncompressed_offset' : 2147473321,
'compressed_size' : 175631248}]
````

In Python, how can we leverage the above `meta_data` to seek to `chunk_1.ndjson`, start decompressing, and stream it line-by-line? In this way, we don't need to
- decompress `chunk_0.ndjson`,
- load the whole compressed `chunk_1.ndjson` into the memory.

Thank your for your help.

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.