apache / apache/parquet-java

Support configurable for DirectByteBufferAllocator from Hadoop Configuration

Open
#2,443 0 comments 0 reactions 0 assignees View on GitHub
Component: Java Component: Parquet Priority: Minor Type: enhancement
Dominant language
Java
Stars
3.1k
Forks
1.6k
Avg merge
3d 12h
Merged PRs (30d)
33

Description

Now in [HadoopReadOptions\|([https://github.com/apache/parquet-mr/blob/master/parquet-hadoop/src/main/java/org/apache/parquet/HadoopReadOptions.java#L85]] class, we cannot change default allocator from Hadoop Configuration.

 

Add a config `parquet.allocator.direct.enabled` to enable `DirectByteBufferAllocator` in `ParuqetFileReader`. 

**Reporter**: [ShuMing Li](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=lishuming)
#### Related issues:
- [ColumnChunkPageWriter uses only heap memory.](https://github.com/apache/parquet-java/issues/2060) (relates to)
#### PRs and other links:
- [GitHub Pull Request #749](https://github.com/apache/parquet-mr/pull/749)

**Note**: *This issue was originally created as [PARQUET-1771](https://issues.apache.org/jira/browse/PARQUET-1771). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with parquet-hadoop/src/main/java/org/apache/parquet/HadoopReadOptions.java and the ParquetFileReader path that chooses the allocator. Trace how Hadoop Configuration is read and how DirectByteBufferAllocator is selected; done means parquet.allocator.direct.enabled controls that selection, with the existing behavior preserved when it is not enabled.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
data-engineering
Issue type
Feature
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Clearly specified
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.