Implement async IO for Parquet file reader
- 主要语言
- Java
- 星标
- 3.1k
- 派生
- 1.6k
- 平均合并
- 3 天 12 小时
- 30 天内合并 PR
- 33
描述
ParquetFileReader's implementation has the following flow (simplified) -
- For every column -> Read from storage in 8MB blocks -> Read all uncompressed pages into output queue
- From output queues -> (downstream ) decompression + decoding
This flow is serialized, which means that downstream threads are blocked until the data has been read. Because a large part of the time spent is waiting for data from storage, threads are idle and CPU utilization is really low.
There is no reason why this cannot be made asynchronous _and_ parallel. So
For Column _i_ -> reading one chunk until end, from storage -> intermediate output queue -> read one uncompressed page until end -> output queue -> (downstream ) decompression + decoding
Note that this can be made completely self contained in ParquetFileReader and downstream implementations like Iceberg and Spark will automatically be able to take advantage without code change as long as the ParquetFileReader apis are not changed.
In past work with async io [Drill - async page reader ](https://github.com/apache/drill/blob/master/exec/java-exec/src/main/java/org/apache/drill/exec/store/parquet/columnreaders/AsyncPageReader.java) , I have seen 2x-3x improvement in reading speed for Parquet files.
**Reporter**: [Parth Chandra](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=parthc) / @parthchandra
#### Related issues:
- [Improve Parquet IO Performance within cloud datalakes](https://github.com/apache/parquet-java/issues/2912) (is depended upon by)
#### PRs and other links:
- [GitHub Pull Request #968](https://github.com/apache/parquet-java/pull/968)
**Note**: *This issue was originally created as [PARQUET-2149](https://issues.apache.org/jira/browse/PARQUET-2149). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*
贡献指南
这个仓库没有索引到贡献指南
调研方向
首先检查 ParquetFileReader 当前的串行流程和所引用的 Pull Request #968,然后将链接的 Drill AsyncPageReader 与之前的异步 I/O 工作进行比较。完成标准是:每列的存储读取和页面处理可以异步并行进行,同时不改变 ParquetFileReader APIs。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- java
- 领域
- data-engineering, performance
- Issue 类型
- 功能
- 难度
- 5/5
- 预计耗时
- 一周以上
- 活跃度
- 停滞
- 描述清晰度
- 基本清楚
- 新手友好度
- 25/100