apache / apache/parquet-java

Bulk skip in RunLengthBitPackingHybridDecoder / DictionaryValuesReader

未关闭
#3,772 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
Type: enhancement
主要语言
Java
星标
3.1k
派生
1.6k
平均合并
3 天 12 小时
30 天内合并 PR
33

描述

### Motivation

Make Hive leverage bulk skip when implementing probe decode for Parquet, similarly to https://issues.apache.org/jira/browse/HIVE-22731, which was about ORC.

### Problem

`ValuesReader.skip(int n)` ships with a naive default:

```java
public void skip(int n) {
for (int i = 0; i < n; i++) skip();
}
```

For dictionary-encoded columns (the common case), each `skip()` bottoms
out in `RunLengthBitPackingHybridDecoder.readInt()` — a mode switch,
array-index arithmetic, and a value the caller immediately discards.

Any filter-then-skip path (column-index row ranges, hash-join probe
filtering, runtime filters) pays this cost per skipped row.

### Proposal

1. Add `RunLengthBitPackingHybridDecoder.skipInts(int n)` — re-use
`readNext()` per run, then advance `currentCount` by
`min(n, currentCount)` instead of walking every value through
`readInt()`.
2. Override `skip(int)` on `DictionaryValuesReader` and
`RunLengthBitPackingHybridValuesReader` to call `decoder.skipInts(n)`.

### Component(s)

Core

贡献指南

这个仓库没有索引到贡献指南

调研方向

先阅读 RunLengthBitPackingHybridDecoder.readInt() 和 readNext(),然后检查 DictionaryValuesReader 和 RunLengthBitPackingHybridValuesReader 中的 skip(int)。确认字典编码值的解码方式,并找出相关的现有测试或测试入口点。完成标准是:三个 reader 中的批量跳过都使用 decoder 路径,同时保持跳过的值数量不变。

由索引模型根据 Issue 内容生成。

评估

技术栈
java
领域
data
Issue 类型
功能
难度
3/5
预计耗时
1-2 天
活跃度
活跃
描述清晰度
基本清楚
新手友好度
72/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。