ZSTD compresion with dictionary causes odd errors
- 主要言語
- Python
- スター
- 211
- フォーク
- 58
- 平均マージ
- 1日 17時間
- マージ済み PR(30日)
- 6
説明
When activating the "use_dict" flag in an SChunk instance, storing data leads to errors.
The following code does not execute on my system:
```
import blosc2
import numpy as np
CHUNKSIZE = int(2**12)
NCHUNKS = 5
coptions = blosc2.cparams_dflts.copy()
coptions["codec"] = blosc2.Codec.ZSTD # this is already the default
coptions["use_dict"] = 1
_rng = np.random.default_rng()
def _make_data() -> bytes:
return _rng.random(CHUNKSIZE // 4, dtype=np.float32).tobytes()
data = [_make_data() for x in range(NCHUNKS)]
storage = blosc2.SChunk(
chunksize=CHUNKSIZE, cparams=coptions, dparams=blosc2.dparams_dflts
)
for x in data:
storage.append_data(x)
for index, x in enumerate(data):
assert storage.decompress_chunk(index) == x
```
Instead, it leads to the following `RuntimeError`:
```
Traceback (most recent call last):
File "/home/user/minimal_bug.py", line 26, in
storage.append_data(x)
File "/home/user/env/lib/python3.9/site-packages/blosc2/schunk.py", line 298, in append_data
return super(SChunk, self).append_data(data)
File "blosc2_ext.pyx", line 1105, in blosc2.blosc2_ext.SChunk.append_data
RuntimeError: Could not append the buffer
```
If the above code is run with `coptions["use_dict"] = 0`, it executes successfully.
Do specific flags need to be set for shared dictionary compression to be successful, or does the sizing of stored data have different requirements?
python-blosc2 version: `blosc2==2.3.2`
python version: `3.9.18`
platform: arch linux, conda based python install
コントリビューションガイド
調査の方向性
提供された最小スクリプトを python-blosc2 2.3.2 で再現し、schunk.py の SChunk.append_data と blosc2_ext.pyx の traceback パスを調査します。use_dict=1 がこれらのチャンクを拒否する理由を特定し、5 つすべてのバッファーを追加して、それぞれを辞書有効で正常に伸長できることを確認します。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- python
- 領域
- data
- issue の種類
- バグ
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 活発さ
- 停滞
- 明瞭さ
- おおむね明確
- 初心者へのやさしさ
- 28/100