Add a configurableValueWriterFactory for parquet-hadoop module
- Dominant language
- Java
- Stars
- 3.1k
- Forks
- 1.6k
- Avg merge
- 3d 12h
- Merged PRs (30d)
- 33
Description
Currently, `ValuesWriterFactory` could be config at `ParquetProperties`, but `ParquetOutputFormat` which is used by Hadoop and Hive haven't support to custom a ValuesWriterFactory. We should support update `ParquetOutputFormat` to support it.
**Reporter**: [Dapeng Sun](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=dapengsun) / @sundapeng
#### Related issues:
- [Support enable/disable dictionary for column](https://github.com/apache/parquet-java/issues/2072) (relates to)
**Note**: *This issue was originally created as [PARQUET-1062](https://issues.apache.org/jira/browse/PARQUET-1062). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reading ParquetOutputFormat and ParquetProperties, then trace how ValuesWriterFactory is currently configured. Done means Hadoop and Hive users can supply a custom ValuesWriterFactory through ParquetOutputFormat; no specific test file is named in the issue.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java
- Domain
- data-engineering
- Issue type
- Feature
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100