apache / apache/parquet-java

zstd compressor and decompressor use the same configuration

未关闭
#2,689 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
Component: Java Component: Parquet Priority: Major Type: bug
主要语言
Java
星标
3.1k
派生
1.6k
平均合并
3 天 12 小时
30 天内合并 PR
33

描述

I use spark to rewrite the parquet files that are compressed by zstd. And the parquet version is  1.12.2. I want to read the parquet files compressed by level 3 and compress them on another level. But the level can't be changed.
After I check the source, I found the problem was the codec was cached, and the configuration will not be updated:

I think the problem is important. I found it when I try to use a different level to compaction the files in the iceberg table. Asynchronous rewriting with a higher level can lead to higher compression ratio. This is important to save storage costs.

**Reporter**: [Peidian Li](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=lipeidian)

**Note**: *This issue was originally created as [PARQUET-2152](https://issues.apache.org/jira/browse/PARQUET-2152). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

贡献指南

这个仓库没有索引到贡献指南

调研方向

从 parquet-hadoop/src/main/java/org/apache/parquet/hadoop/CodecFactory.java 开始,重点查看第 144 行和第 226 行的缓存相关位置。跟踪压缩器和解压缩器配置的缓存方式,并确认重写可以使用不同的 Zstandard 压缩级别。完成的标准是,请求的级别不再受到之前已缓存的 codec 配置的阻塞;如果周边测试确定了合适的位置,则添加回归覆盖。

由索引模型根据 Issue 内容生成。

评估

技术栈
java, spark
领域
data-engineering
Issue 类型
缺陷
难度
3/5
预计耗时
1-2 天
活跃度
停滞
描述清晰度
基本清楚
新手友好度
45/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。