apache / apache/parquet-java

Incompatible behavior for ColumnChunk.file_offset between Parquet-mr and Impala

Ouverte
#1,516 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub
Component: Format Component: Java Component: Parquet Priority: Major Type: bug
Langage dominant
Java
Étoiles
3.1k
Forks
1.6k
Merge moyen
3 j 12 h
PR mergées (30 j)
33

Description

According to comments in [parquet.thrift](https://github.com/apache/incubator-parquet-format/blob/master/src/thrift/parquet.thrift#L479), this field is supposed to store offset of ColumnMetaData within the file column chunk is stored. My understanding is that this allows omitting ColumnMetaData within ColumnChunk (it is optional field, after all). Unfortunately, two major implementations, Parquet-mr and Impala, deviate from this definition when writing Parquet files. Impala implementation writes offset pointing to the ColumnChunk rather than ColumnMetaData, as can be found in [hdfs-parquet-table-reader.cc](https://github.com/cloudera/Impala/blob/24db37f4efdc493d218470dc045b61f5104c4fd0/be/src/exec/hdfs-parquet-table-writer.cc#L895). While this is still incorrect behavior according to the comments in parquet.thrift, this still allows access to the ColumnMetaData necessary for reading data.

Parquet-mr implementation can be found in [ParquetMetadataConverter](https://github.com/Parquet/parquet-mr/blob/fd8d18f26af9ad7813dda71352b5dcb0080306eb/parquet-hadoop/src/main/java/parquet/format/converter/ParquetMetadataConverter.java#L149), which writes the offset to the first data page. Not only this is incompatible behavior, but also it makes no sense because you cannot read the data with just data page offset. There is even a comment on that line saying "verify this is the right offset."

**Reporter**: [Eunsoo Roh](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=eunsoor)

**Note**: *This issue was originally created as [PARQUET-68](https://issues.apache.org/jira/browse/PARQUET-68). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Guide de contribution

Aucun guide de contribution indexé pour ce dépôt

Piste de recherche

Lisez src/thrift/parquet.thrift et ParquetMetadataConverter.java, puis comparez le ColumnChunk.file_offset documenté avec le comportement référencé de l’écrivain Impala. Déterminez le contrat d’offset compatible entre ces implémentations ; la tâche n’est terminée qu’une fois l’écart entre la spécification et l’implémentation résolu.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
java
Domaine
data-engineering
Type d'issue
Bug
Difficulté
5/5
Temps estimé
Plus d'une semaine
Activité
À l'abandon
Clarté
À clarifier
Accessibilité débutants
25/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.