Reuse hadoop file status and footer in ParquetRecordReader
- 主要言語
- Java
- スター
- 3.1k
- フォーク
- 1.6k
- 平均マージ
- 3日 12時間
- マージ済み PR(30日)
- 33
説明
## DESCRIPTION
Spark will send a listStatus RPC to get hadoop file status and read the parquet file footer before reading the parquet file. And send a same listStatus RPC to get the same hadoop file status and read the footer again in ParquetRecordReader. We can reuse the file status and the footer.
## PLANS
Save the hadoop file status in the ParquetMetadata and save the ParquetMetadata in the input split, so we can reuse them when init a new ParquetRecordReader.
**Reporter**: [Wan Kun](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=wankun)
#### PRs and other links:
- [GitHub Pull Request #1242](https://github.com/apache/parquet-mr/pull/1242)
**Note**: *This issue was originally created as [PARQUET-2415](https://issues.apache.org/jira/browse/PARQUET-2415). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*
コントリビューションガイド
このリポジトリのコントリビューションガイドは索引されていません
調査の方向性
まず ParquetRecordReader と、issue で説明されている input split のパスを確認し、次に ParquetMetadata が現在どのように footer を保持しているか、また Hadoop のファイルステータスがどのように取得されているかを調べます。既存のアプローチを GitHub Pull Request #1242 と比較します。新しい ParquetRecordReader の初期化時に、ステータスと footer の繰り返し取得が回避されれば完了です。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- hadoop, java
- 領域
- data-engineering
- issue の種類
- リファクタリング
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 活発さ
- 停滞
- 明瞭さ
- おおむね明確
- 初心者へのやさしさ
- 25/100