Improve performance of InternalParquetRecordReader (1%)
- 主要语言
- Java
- 星标
- 3.1k
- 派生
- 1.6k
- 平均合并
- 3 天 12 小时
- 30 天内合并 PR
- 33
描述
### Describe the enhancement requested
Profiling the load of a Parquet file with Java Mission Control, I've noticed that `InternalParquetRecordReader` [LongStream](https://github.com/apache/parquet-java/blob/1f1e07bbf750fba228851c2d63470c3da5726831/parquet-hadoop/src/main/java/org/apache/parquet/hadoop/InternalParquetRecordReader.java#L323) consumes relevant amount of time.
This `LongStream` can be replaced with a simpler Long Iterator that iterates from 0 to `pages.getRowCount()`.
To measure the overhead I've created a test project that overwrites `InternalParquetRecordReader` implementation with a Long Iterator: https://github.com/jerolba/parquet-rowindexiterator
The execution time is sensitive to the context of the JVM, but running the benchmark multiple times shows that LongStream is slower than LongIterator, between 1% and 4% depending on the run.
### Component(s)
_No response_
贡献指南
这个仓库没有索引到贡献指南
调研方向
从 parquet-hadoop/src/main/java/org/apache/parquet/hadoop/InternalParquetRecordReader.java 中第 323 行附近的 LongStream 开始,然后查看链接的基准测试项目 parquet-rowindexiterator。将 reader 当前的迭代方式与基准测试中的 Long Iterator 方法进行比较;完成的标准是 reader 避免报告的 LongStream 开销,同时保留行索引迭代行为。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- java
- 领域
- data-engineering, performance
- Issue 类型
- 功能
- 难度
- 2/5
- 预计耗时
- 1-3 小时
- 活跃度
- 停滞
- 描述清晰度
- 描述清楚
- 新手友好度
- 55/100