apache / apache/parquet-java

Difference between parquet-mr implementation and parquet-format documentation

Open
#1,819 3 comments 0 reactions 0 assignees View on GitHub
Component: Format Component: Java Component: Parquet Priority: Major Type: bug
Dominant language
Java
Stars
3.1k
Forks
1.6k
Avg merge
3d 12h
Merged PRs (30d)
33

Description

Documentation at https://github.com/apache/parquet-format/blob/master/src/thrift/parquet.thrift

```
struct ColumnChunk {
/** File where column data is stored. If not set, assumed to be same file as
* metadata. This path is relative to the current file.
**/
1: optional string file_path

/** Byte offset in file_path to the ColumnMetaData **/
2: required i64 file_offset

...
```

and https://github.com/apache/parquet-format

```
4-byte magic number "PAR1"

...
```

suggests that ColumnChunk data should be followed by ColumnChunkMetaData.

However it looks like parquet-mr doesn't write ColumnMetaData after Columns at all and populates ColumnChunk.file_offset with an offset of the first data page:

from **ParquetMetadataConverter.java:153**:
```Java
for (ColumnChunkMetaData columnMetaData : columns) {
ColumnChunk columnChunk = new ColumnChunk(columnMetaData.getFirstDataPageOffset()); // verify this is the right offset
columnChunk.file_path = block.getPath(); // they are in the same file for now
```

Is it a bug in parquet-mr or in the documentation?

**Reporter**: [Konstantin Shaposhnikov](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=k.shaposhnikov@gmail.com) / @kostya-sh
#### PRs and other links:
- [parquet-format PR #56](https://github.com/apache/parquet-format/pull/56)

**Note**: *This issue was originally created as [PARQUET-291](https://issues.apache.org/jira/browse/PARQUET-291). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with ParquetMetadataConverter.java at line 153 and compare its ColumnChunk construction with the ColumnChunk definition in parquet.thrift and the parquet-format layout documentation. Review the linked parquet-format PR #56 and the migration documentation referenced in the issue. Done means the implementation and documentation discrepancy has a clearly recorded resolution.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
data-engineering
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.