Blosc / Blosc/python-blosc2

ZSTD compresion with dictionary causes odd errors

未关闭
#208 5 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Python
星标
211
派生
58
平均合并
1 天 17 小时
30 天内合并 PR
6

描述

When activating the "use_dict" flag in an SChunk instance, storing data leads to errors.

The following code does not execute on my system:

```
import blosc2
import numpy as np

CHUNKSIZE = int(2**12)
NCHUNKS = 5

coptions = blosc2.cparams_dflts.copy()
coptions["codec"] = blosc2.Codec.ZSTD # this is already the default
coptions["use_dict"] = 1

_rng = np.random.default_rng()

def _make_data() -> bytes:
return _rng.random(CHUNKSIZE // 4, dtype=np.float32).tobytes()

data = [_make_data() for x in range(NCHUNKS)]

storage = blosc2.SChunk(
chunksize=CHUNKSIZE, cparams=coptions, dparams=blosc2.dparams_dflts
)

for x in data:
storage.append_data(x)

for index, x in enumerate(data):
assert storage.decompress_chunk(index) == x
```

Instead, it leads to the following `RuntimeError`:

```
Traceback (most recent call last):
File "/home/user/minimal_bug.py", line 26, in
storage.append_data(x)
File "/home/user/env/lib/python3.9/site-packages/blosc2/schunk.py", line 298, in append_data
return super(SChunk, self).append_data(data)
File "blosc2_ext.pyx", line 1105, in blosc2.blosc2_ext.SChunk.append_data
RuntimeError: Could not append the buffer
```

If the above code is run with `coptions["use_dict"] = 0`, it executes successfully.

Do specific flags need to be set for shared dictionary compression to be successful, or does the sizing of stored data have different requirements?

python-blosc2 version: `blosc2==2.3.2`
python version: `3.9.18`
platform: arch linux, conda based python install

贡献指南

打开贡献指南

调研方向

使用 python-blosc2 2.3.2 重现所提供的最小脚本,并检查 schunk.py 中的 SChunk.append_data 以及 blosc2_ext.pyx 中的 traceback 路径。确定 use_dict=1 拒绝这些 chunk 的原因,然后验证在启用字典的情况下追加全部五个缓冲区并分别解压每个缓冲区都能成功。

由索引模型根据 Issue 内容生成。

评估

技术栈
python
领域
data
Issue 类型
缺陷
难度
4/5
预计耗时
3-5 天
活跃度
停滞
描述清晰度
基本清楚
新手友好度
28/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。