apache / apache/parquet-java

Older `parquet-java` readers fail on files with unprojected VARIANT columns

Open
#3,633 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Java
Stars
3.1k
Forks
1.6k
Avg merge
3d 12h
Merged PRs (30d)
33

Description

### Summary

`parquet-java` 1.15.x fails when opening a Parquet file that contains a `VARIANT` column, even if the application only requests other columns and does not request the `VARIANT` column.

The failure happens while converting the footer schema, before projection is applied. The parquet-format compatibility guidance says that new logical types are intended to be forward compatible: https://github.com/apache/parquet-format/blob/master/CONTRIBUTING.md#compatibility-and-feature-enablement

Based on that, I think older readers should be able to tolerate an unknown logical type when that field is not part of the requested projection.

### Repro

This can be reproduced with the `VARIANT` fixture in `apache/parquet-testing`:

```bash
git clone https://github.com/apache/parquet-testing.git
PARQUET_FILE="$PWD/parquet-testing/shredded_variant/case-001.parquet"
```

The file contains an `id` column and a `var` column annotated with `VARIANT`.

Using `parquet-java` 1.15.x, try to read only `id` through the normal `GroupReadSupport` projection path:

```java
Path input = new Path(args[0]);

try (ParquetReader reader = ParquetReader.builder(new GroupReadSupport(), input)
.set(
ReadSupport.PARQUET_READ_SCHEMA,
"message root {\n" + "optional int32 id;\n" + "}")
.build()) {
Group row = reader.read();
System.out.println(row.getInteger("id", 0));
}
```

I also put up a draft repro PR with a failing unit test against `parquet-1.15.x`: https://github.com/kevinjqliu/parquet-java/pull/1

That test uses `apache/parquet-testing/shredded_variant/case-001.parquet` and requests only the `id` column. It still fails while reading footer metadata, before the projected read can happen.

### Actual behavior

The read fails during footer schema conversion:

```text
[ERROR] org.apache.parquet.hadoop.TestReadWithUnknownLogicalType.testReadProjectedColumnFromFileWithUnknownLogicalType -- Time elapsed: 0.283 s <<< ERROR!
java.lang.NullPointerException: Cannot invoke "org.apache.parquet.format.LogicalType$_Fields.ordinal()" because the return value of "org.apache.parquet.format.LogicalType.getSetField()" is null
at org.apache.parquet.format.converter.ParquetMetadataConverter.getLogicalTypeAnnotation(ParquetMetadataConverter.java:1174)
at org.apache.parquet.format.converter.ParquetMetadataConverter.buildChildren(ParquetMetadataConverter.java:1892)
at org.apache.parquet.format.converter.ParquetMetadataConverter.fromParquetSchema(ParquetMetadataConverter.java:1840)
at org.apache.parquet.format.converter.ParquetMetadataConverter.fromParquetMetadata(ParquetMetadataConverter.java:1670)
at org.apache.parquet.format.converter.ParquetMetadataConverter.readParquetMetadata(ParquetMetadataConverter.java:1630)
at org.apache.parquet.hadoop.ParquetFileReader.readFooter(ParquetFileReader.java:629)
at org.apache.parquet.hadoop.ParquetFileReader.(ParquetFileReader.java:934)
at org.apache.parquet.hadoop.ParquetFileReader.(ParquetFileReader.java:925)
```

The stack points to `ParquetMetadataConverter.getLogicalTypeAnnotation`.

### Expected behavior

Older readers should tolerate an unknown logical type in the footer when the field is not projected, preserving the physical schema or treating the field as having no known logical annotation.

They should only fail if the unsupported field is actually read or interpreted semantically.

### Notes

This is relevant for files written by newer or external writers that contain `VARIANT`, where older readers only need unrelated columns.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with ParquetMetadataConverter.getLogicalTypeAnnotation and the stack locations in ParquetMetadataConverter.java, then run the failing TestReadWithUnknownLogicalType test against parquet-testing/shredded_variant/case-001.parquet. Done means a projected read of id tolerates the unprojected VARIANT column while unsupported fields still fail only when semantically read.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
data
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Clearly specified
Newbie friendliness
55/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.