Support configurable for DirectByteBufferAllocator from Hadoop Configuration
- Dominant language
- Java
- Stars
- 3.1k
- Forks
- 1.6k
- Avg merge
- 3d 12h
- Merged PRs (30d)
- 33
Description
Now in [HadoopReadOptions\|([https://github.com/apache/parquet-mr/blob/master/parquet-hadoop/src/main/java/org/apache/parquet/HadoopReadOptions.java#L85]] class, we cannot change default allocator from Hadoop Configuration.
Add a config `parquet.allocator.direct.enabled` to enable `DirectByteBufferAllocator` in `ParuqetFileReader`.
**Reporter**: [ShuMing Li](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=lishuming)
#### Related issues:
- [ColumnChunkPageWriter uses only heap memory.](https://github.com/apache/parquet-java/issues/2060) (relates to)
#### PRs and other links:
- [GitHub Pull Request #749](https://github.com/apache/parquet-mr/pull/749)
**Note**: *This issue was originally created as [PARQUET-1771](https://issues.apache.org/jira/browse/PARQUET-1771). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*
Contributor guide
No contributing guide indexed for this repository
Research direction
Start with parquet-hadoop/src/main/java/org/apache/parquet/HadoopReadOptions.java and the ParquetFileReader path that chooses the allocator. Trace how Hadoop Configuration is read and how DirectByteBufferAllocator is selected; done means parquet.allocator.direct.enabled controls that selection, with the existing behavior preserved when it is not enabled.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java
- Domain
- data-engineering
- Issue type
- Feature
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Stale
- Clarity
- Clearly specified
- Newbie friendliness
- 25/100