apache / apache/parquet-java

Controlling memory utilization by ParquetReader

未关闭
#2,679 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
Component: Parquet Priority: Major Type: enhancement
主要语言
Java
星标
3.1k
派生
1.6k
平均合并
3 天 12 小时
30 天内合并 PR
33

描述

In Apache Druid, Parquet is one of the popular form of input source to ingest data into a druid cluster (https://druid.apache.org/docs/latest/development/extensions-core/parquet.html). We rely on the parquet-mr library to read the parquet files and then convert them into Druid native format row-for-row to ingest. A considerable amount of our usecases ingest the whole parquet files (ie all columns in a single shot) into the system.

A challenge that we face is that the parquet reader loads up an entire row group into memory as part of its normal operation. Row groups can be quite large (like, 1GB large) and sometimes it creates a pressure on our reader JVM leading to OOMs. Further, in some other cases it ends up creating GC pressure on the JVM leading to a decrease in the throughput of the ingestion tasks.

To mitigate this problem, we are considering that would it be better to have an option to download the Parquet rowgroup/file first and memory-map it for reading? The code which buffers the rowgroup works on the ByteBuffer interface already (https://github.com/apache/parquet-mr/blob/master/parquet-hadoop/src/main/java/org/apache/parquet/hadoop/ParquetFileReader.java#L1763), so it seems like it could compliment the MMappedByteBuffer implementation too. Such a thing would alleviate pressure off of our reader JVM there by heavily reducing the chances for OOMs.

We're very open to more ideas or already tried solutions around this problem. 

**Reporter**: [Rohan Garg](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=rohangarg)

**Note**: *This issue was originally created as [PARQUET-2141](https://issues.apache.org/jira/browse/PARQUET-2141). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

贡献指南

这个仓库没有索引到贡献指南

调研方向

从 parquet-hadoop/src/main/java/org/apache/parquet/hadoop/ParquetFileReader.java 第 1763 行附近开始,检查行组缓冲如何使用 ByteBuffer。将提议的 MMappedByteBuffer 方案与当前行为进行评估,然后定义并验证一个达成共识的解决方案,在不改变数据摄取结果的情况下减少 JVM 内存使用和 GC 压力。

由索引模型根据 Issue 内容生成。

评估

技术栈
java
领域
performance
Issue 类型
功能
难度
5/5
预计耗时
一周以上
活跃度
停滞
描述清晰度
需要澄清
新手友好度
25/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。