apache / apache/parquet-java

Incompatible behavior for ColumnChunk.file_offset between Parquet-mr and Impala

Offen
#1,516 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Component: Format Component: Java Component: Parquet Priority: Major Type: bug
Vorherrschende Sprache
Java
Sterne
3.1k
Forks
1.6k
Ø Merge
3 T. 12 Std.
Gemergte PRs (30 T.)
33

Beschreibung

According to comments in [parquet.thrift](https://github.com/apache/incubator-parquet-format/blob/master/src/thrift/parquet.thrift#L479), this field is supposed to store offset of ColumnMetaData within the file column chunk is stored. My understanding is that this allows omitting ColumnMetaData within ColumnChunk (it is optional field, after all). Unfortunately, two major implementations, Parquet-mr and Impala, deviate from this definition when writing Parquet files. Impala implementation writes offset pointing to the ColumnChunk rather than ColumnMetaData, as can be found in [hdfs-parquet-table-reader.cc](https://github.com/cloudera/Impala/blob/24db37f4efdc493d218470dc045b61f5104c4fd0/be/src/exec/hdfs-parquet-table-writer.cc#L895). While this is still incorrect behavior according to the comments in parquet.thrift, this still allows access to the ColumnMetaData necessary for reading data.

Parquet-mr implementation can be found in [ParquetMetadataConverter](https://github.com/Parquet/parquet-mr/blob/fd8d18f26af9ad7813dda71352b5dcb0080306eb/parquet-hadoop/src/main/java/parquet/format/converter/ParquetMetadataConverter.java#L149), which writes the offset to the first data page. Not only this is incompatible behavior, but also it makes no sense because you cannot read the data with just data page offset. There is even a comment on that line saying "verify this is the right offset."

**Reporter**: [Eunsoo Roh](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=eunsoor)

**Note**: *This issue was originally created as [PARQUET-68](https://issues.apache.org/jira/browse/PARQUET-68). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Beitragsleitfaden

Für dieses Repository ist kein Beitragsleitfaden indexiert

Rechercherichtung

Lies src/thrift/parquet.thrift und ParquetMetadataConverter.java und vergleiche anschließend den dokumentierten ColumnChunk.file_offset mit dem referenzierten Verhalten des Impala-Schreibers. Ermittle den kompatiblen Offset-Vertrag zwischen diesen Implementierungen; die Aufgabe ist erst abgeschlossen, wenn die Diskrepanz zwischen Spezifikation und Implementierung geklärt ist.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
java
Bereich
data-engineering
Issue-Typ
Bug
Schwierigkeit
5/5
Geschätzter Aufwand
Über eine Woche
Aktivitätsstatus
Veraltet
Klarheit
Muss geklärt werden
Anfängerfreundlichkeit
25/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.