Efficient storage for several INT_8 and INT_16
- Dominant language
- Java
- Stars
- 3.1k
- Forks
- 1.6k
- Avg merge
- 3d 12h
- Merged PRs (30d)
- 33
Description
In very large datasets, aggregating several INT8 into INT32 fields (or byte array) can make a big difference.
In parquet, efficient algorithms exist for INT32, so if the LogicalType is INT_8 the encoded int might take up only one byte.
However further optimizations could be made by allowing the user to better specify the types.
What about BYTE_ARRAY logical type, backed by FIXED_LEN_BYTE_ARRAY type (or eventually INT_32)?
**Reporter**: [Fernando Pereira](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=ferdonline)
**Note**: *This issue was originally created as [PARQUET-845](https://issues.apache.org/jira/browse/PARQUET-845). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reviewing how Parquet currently represents INT8 and INT16 logical types with INT32, then examine the existing encoding behavior for INT32 and BYTE_ARRAY values. Done means agreeing on a concrete type representation and implementing and validating the proposed storage optimization, but the issue does not name files or tests to run.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java
- Domain
- data-engineering
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100