facebook / facebook/zstd

Multiframe ZSTD file: how to jump to and stream the second file?

オープン
#4,569 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る
主要言語
C
スター
27.9k
フォーク
2.6k
平均マージ
1日 3時間
マージ済み PR(30日)
8

説明

TL;DR: This is not about bugs of ZSTD but about how to take advantage of its feature:

---

I compress two ndjson files into a [multiframe](https://github.com/facebook/zstd/blob/3bee41a70eaf343fbcae3637b3f6edbe52f35ed8/doc/zstd_compression_format.md#frames) ZST file where each ndjson is compressed into a frame. I have the following metadata `meta_data` (as a list) of the ZST file:

````python
import zstandard as zstd
from pathlib import Path

input_file = r"E:\Personal projects\tmp\test.zst"
input_file = Path(output_file)

meta_data = [{'name' : 'chunk_0.ndjson',
'uncompressed_size' : 2147473321,
'compressed_offset' : 0,
'uncompressed_offset' : 0,
'compressed_size' : 175631248},
{'name' : 'chunk_1.ndjson',
'uncompressed_size' : 2147473321,
'compressed_offset' : 175631248,
'uncompressed_offset' : 2147473321,
'compressed_size' : 175631248}]
````

In Python, how can we leverage the above `meta_data` to seek to `chunk_1.ndjson`, start decompressing, and stream it line-by-line? In this way, we don't need to
- decompress `chunk_0.ndjson`,
- load the whole compressed `chunk_1.ndjson` into the memory.

Thank your for your help.

コントリビューションガイド

コントリビューションガイドを開く

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。