Incompatible behavior for ColumnChunk.file_offset between Parquet-mr and Impala
- Lenguaje dominante
- Java
- Estrellas
- 3.1k
- Forks
- 1.6k
- Merge medio
- 3 d 12 h
- PR fusionados (30 d)
- 33
Descripción
According to comments in [parquet.thrift](https://github.com/apache/incubator-parquet-format/blob/master/src/thrift/parquet.thrift#L479), this field is supposed to store offset of ColumnMetaData within the file column chunk is stored. My understanding is that this allows omitting ColumnMetaData within ColumnChunk (it is optional field, after all). Unfortunately, two major implementations, Parquet-mr and Impala, deviate from this definition when writing Parquet files. Impala implementation writes offset pointing to the ColumnChunk rather than ColumnMetaData, as can be found in [hdfs-parquet-table-reader.cc](https://github.com/cloudera/Impala/blob/24db37f4efdc493d218470dc045b61f5104c4fd0/be/src/exec/hdfs-parquet-table-writer.cc#L895). While this is still incorrect behavior according to the comments in parquet.thrift, this still allows access to the ColumnMetaData necessary for reading data.
Parquet-mr implementation can be found in [ParquetMetadataConverter](https://github.com/Parquet/parquet-mr/blob/fd8d18f26af9ad7813dda71352b5dcb0080306eb/parquet-hadoop/src/main/java/parquet/format/converter/ParquetMetadataConverter.java#L149), which writes the offset to the first data page. Not only this is incompatible behavior, but also it makes no sense because you cannot read the data with just data page offset. There is even a comment on that line saying "verify this is the right offset."
**Reporter**: [Eunsoo Roh](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=eunsoor)
**Note**: *This issue was originally created as [PARQUET-68](https://issues.apache.org/jira/browse/PARQUET-68). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*
Guía de contribución
No hay ninguna guía de contribución indexada para este repositorio
Línea de trabajo
Lee src/thrift/parquet.thrift y ParquetMetadataConverter.java y compara después el ColumnChunk.file_offset documentado con el comportamiento del escritor de Impala al que se hace referencia. Determina el contrato de offsets compatible entre estas implementaciones; el trabajo requiere resolver la discrepancia entre la especificación y la implementación.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- java
- Área
- data-engineering
- Tipo de issue
- Error
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Estado de actividad
- Estancado
- Claridad
- Necesita aclaración
- Aptitud para principiantes
- 25/100