apache / apache/parquet-java

Improve the RLE encoding for Parquet Dictionary IDs

未关闭
#2,073 4 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
Component: Parquet Priority: Major Type: enhancement
主要语言
Java
星标
3.1k
派生
1.6k
平均合并
3 天 12 小时
30 天内合并 PR
33

描述

The IDs of Parquet Dictionary encoding is using `RunLengthBitPackingHybridEncoder`.
RunLengthBitPackingHybridEncoder handles encoding with `repeat` and `bitpacking`, we should improve it with the method likes `DeltaBinaryPackingWriter`

**Reporter**: [Dapeng Sun](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=dapengsun) / @sundapeng

**Note**: *This issue was originally created as [PARQUET-1059](https://issues.apache.org/jira/browse/PARQUET-1059). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

贡献指南

这个仓库没有索引到贡献指南

调研方向

Start by locating RunLengthBitPackingHybridEncoder and DeltaBinaryPackingWriter in the Java sources, then read their encoding and test coverage. Determine the intended improved handling for Parquet Dictionary IDs and define completion through encoding correctness and performance tests; the issue does not name specific files or tests.

由索引模型根据 Issue 内容生成。

评估

技术栈
java
领域
data-engineering
Issue 类型
重构
难度
5/5
预计耗时
一周以上
活跃度
停滞
描述清晰度
需要澄清
新手友好度
25/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。