apache / apache/parquet-java

Make the order of encodings in column metadata deterministic

未关闭
#3,215 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
Type: enhancement
主要语言
Java
星标
3.1k
派生
1.6k
平均合并
3 天 12 小时
30 天内合并 PR
33

描述

### Describe the enhancement requested

**Background**
The list of encodings used for a column is stored in the column metadata in the Parquet file footer.

The elements in this list are enumeration constants (`org.apache.parquet.format.Encoding`), which are collected in a `HashSet` when the records and fields are written. Later, when the footer is written, the set of enum constants are converted to a temporary list using the HashSet’s iteration order.

Since `Enum::hashCode` delegates to `Object::hashCode` (or, in later JDK versions, to `System::identityHashCode`), the order of the enum constants in this list can vary between runs in different processes, and thus files with identical encodings can represent this list differently (the elements may appear in different order).

**Rationale for changing this behaviour**
Two processes running the same version of parquet-java and having identical writer configurations can still produce files that are different at the binary level for the exact same written data.

For redundancy reasons, it is not uncommon to write data to Parquet files on two different machines. To verify that the same data has been written on both machines, it is currently not sufficient to compare the files ate the binary level. Instead, the files must be decoded and their actual data must be compared to ensure they are equal.

If the files can be made identical at the binary level, this verification process would be simplified.

**Suggested change**
The list of encodings is created in `org.apache.parquet.format.converter.ParquetMetadataConverter::toFormatEncodings`. A simple solution is to sort this list before it is returned, i.e. always return the `Encoding` enum constants in ascending ordinal order.

### Component(s)

Core

贡献指南

这个仓库没有索引到贡献指南

调研方向

从 org.apache.parquet.format.converter.ParquetMetadataConverter::toFormatEncodings 开始,检查 Encoding 值的 HashSet 如何变成返回的列表。当写入等效数据能够按确定性的升序 ordinal 顺序生成 Encoding 元数据,从而消除依赖进程的顺序时,工作即完成。

由索引模型根据 Issue 内容生成。

评估

技术栈
java
领域
data
Issue 类型
功能
难度
2/5
预计耗时
1-3 小时
活跃度
停滞
描述清晰度
描述清楚
新手友好度
58/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。