apache / apache/parquet-java

Provide option to use on-heap buffers for Snappy compression/decompression

Open
#1,408 9 comments 0 reactions 0 assignees View on GitHub
Component: Java Component: Parquet Priority: Major Type: enhancement
Dominant language
Java
Stars
3.1k
Forks
1.6k
Avg merge
3d 12h
Merged PRs (30d)
33

Description

The current code uses direct off-heap buffers for decompression. If many decompressors are instantiated across multiple threads, and/or the objects being decompressed are large, this can lead to a huge amount of off-heap allocation by the JVM. This can be exacerbated if overall, there is not heap contention, since no GC will be performed to reclaim the space used by these buffers.

It would be nice if there was a flag we cold use to simply allocate on-heap buffers here:

https://github.com/apache/incubator-parquet-mr/blob/master/parquet-hadoop/src/main/java/parquet/hadoop/codec/SnappyDecompressor.java#L28

We ran into an issue today where these buffers totaled a very large amount of storage and caused our Java processes (running within containers) to be terminated by the kernel OOM-killer.

**Reporter**: [Patrick Wendell](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=pwendell)
#### Related issues:
- [Parquet+Snappy can cause significant off-heap memory usage](https://issues.apache.org/jira/browse/SPARK-4073) (breaks)

**Note**: *This issue was originally created as [PARQUET-118](https://issues.apache.org/jira/browse/PARQUET-118). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with parquet-hadoop/src/main/java/parquet/hadoop/codec/SnappyDecompressor.java and inspect how decompression buffers are currently allocated. Define how a flag should select on-heap buffers, then verify that Snappy decompression still works and that the selected mode avoids the reported off-heap allocation growth.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
data-engineering
Issue type
Feature
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.