apache / apache/parquet-java

Implement async IO for Parquet file reader

オープン
#2,686 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る
Component: Java Component: Parquet Priority: Major Type: enhancement
主要言語
Java
スター
3.1k
フォーク
1.6k
平均マージ
3日 12時間
マージ済み PR(30日)
33

説明

ParquetFileReader's implementation has the following flow (simplified) - 
      - For every column -> Read from storage in 8MB blocks -> Read all uncompressed pages into output queue 
      - From output queues -> (downstream ) decompression + decoding

This flow is serialized, which means that downstream threads are blocked until the data has been read. Because a large part of the time spent is waiting for data from storage, threads are idle and CPU utilization is really low.

There is no reason why this cannot be made asynchronous _and_ parallel. So 

For Column _i_ -> reading one chunk until end, from storage -> intermediate output queue -> read one uncompressed page until end -> output queue -> (downstream ) decompression + decoding

Note that this can be made completely self contained in ParquetFileReader and downstream implementations like Iceberg and Spark will automatically be able to take advantage without code change as long as the ParquetFileReader apis are not changed. 

In past work with async io  [Drill - async page reader ](https://github.com/apache/drill/blob/master/exec/java-exec/src/main/java/org/apache/drill/exec/store/parquet/columnreaders/AsyncPageReader.java) , I have seen 2x-3x improvement in reading speed for Parquet files.

**Reporter**: [Parth Chandra](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=parthc) / @parthchandra
#### Related issues:
- [Improve Parquet IO Performance within cloud datalakes](https://github.com/apache/parquet-java/issues/2912) (is depended upon by)
#### PRs and other links:
- [GitHub Pull Request #968](https://github.com/apache/parquet-java/pull/968)

**Note**: *This issue was originally created as [PARQUET-2149](https://issues.apache.org/jira/browse/PARQUET-2149). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

調査の方向性

まず、ParquetFileReader の現在の直列化されたフローと参照されている Pull Request #968 を確認し、次にリンクされている Drill AsyncPageReader を過去の非同期 I/O の取り組みと比較してください。ParquetFileReader APIs を変更せずに、列ごとのストレージ読み取りとページ処理を非同期かつ並列に進められるようになれば完了です。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
java
領域
data-engineering, performance
issue の種類
機能追加
難易度
5/5
見積もり時間
1週間以上
活発さ
停滞
明瞭さ
おおむね明確
初心者へのやさしさ
25/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。