apache / apache/parquet-java

Controlling memory utilization by ParquetReader

オープン
#2,679 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る
Component: Parquet Priority: Major Type: enhancement
主要言語
Java
スター
3.1k
フォーク
1.6k
平均マージ
3日 12時間
マージ済み PR(30日)
33

説明

In Apache Druid, Parquet is one of the popular form of input source to ingest data into a druid cluster (https://druid.apache.org/docs/latest/development/extensions-core/parquet.html). We rely on the parquet-mr library to read the parquet files and then convert them into Druid native format row-for-row to ingest. A considerable amount of our usecases ingest the whole parquet files (ie all columns in a single shot) into the system.

A challenge that we face is that the parquet reader loads up an entire row group into memory as part of its normal operation. Row groups can be quite large (like, 1GB large) and sometimes it creates a pressure on our reader JVM leading to OOMs. Further, in some other cases it ends up creating GC pressure on the JVM leading to a decrease in the throughput of the ingestion tasks.

To mitigate this problem, we are considering that would it be better to have an option to download the Parquet rowgroup/file first and memory-map it for reading? The code which buffers the rowgroup works on the ByteBuffer interface already (https://github.com/apache/parquet-mr/blob/master/parquet-hadoop/src/main/java/org/apache/parquet/hadoop/ParquetFileReader.java#L1763), so it seems like it could compliment the MMappedByteBuffer implementation too. Such a thing would alleviate pressure off of our reader JVM there by heavily reducing the chances for OOMs.

We're very open to more ideas or already tried solutions around this problem. 

**Reporter**: [Rohan Garg](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=rohangarg)

**Note**: *This issue was originally created as [PARQUET-2141](https://issues.apache.org/jira/browse/PARQUET-2141). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

調査の方向性

parquet-hadoop/src/main/java/org/apache/parquet/hadoop/ParquetFileReader.java の 1763 行目付近から始め、行グループのバッファリングで ByteBuffer がどのように使用されているかを調査します。提案されている MMappedByteBuffer のアプローチを現在の動作と比較して評価し、その後、取り込み結果を変更せずに JVM メモリと GC の負荷を削減する、合意された解決策を定義して検証します。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
java
領域
performance
issue の種類
機能追加
難易度
5/5
見積もり時間
1週間以上
活発さ
停滞
明瞭さ
説明が足りない
初心者へのやさしさ
25/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。