apache / apache/parquet-java

Improve performance of InternalParquetRecordReader (1%)

未关闭
#3,226 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
Type: enhancement
主要语言
Java
星标
3.1k
派生
1.6k
平均合并
3 天 12 小时
30 天内合并 PR
33

描述

### Describe the enhancement requested

Profiling the load of a Parquet file with Java Mission Control, I've noticed that `InternalParquetRecordReader` [LongStream](https://github.com/apache/parquet-java/blob/1f1e07bbf750fba228851c2d63470c3da5726831/parquet-hadoop/src/main/java/org/apache/parquet/hadoop/InternalParquetRecordReader.java#L323) consumes relevant amount of time.

This `LongStream` can be replaced with a simpler Long Iterator that iterates from 0 to `pages.getRowCount()`.

To measure the overhead I've created a test project that overwrites `InternalParquetRecordReader` implementation with a Long Iterator: https://github.com/jerolba/parquet-rowindexiterator

The execution time is sensitive to the context of the JVM, but running the benchmark multiple times shows that LongStream is slower than LongIterator, between 1% and 4% depending on the run.

### Component(s)

_No response_

贡献指南

这个仓库没有索引到贡献指南

调研方向

从 parquet-hadoop/src/main/java/org/apache/parquet/hadoop/InternalParquetRecordReader.java 中第 323 行附近的 LongStream 开始,然后查看链接的基准测试项目 parquet-rowindexiterator。将 reader 当前的迭代方式与基准测试中的 Long Iterator 方法进行比较;完成的标准是 reader 避免报告的 LongStream 开销,同时保留行索引迭代行为。

由索引模型根据 Issue 内容生成。

评估

技术栈
java
领域
data-engineering, performance
Issue 类型
功能
难度
2/5
预计耗时
1-3 小时
活跃度
停滞
描述清晰度
描述清楚
新手友好度
55/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。