Older `parquet-java` readers fail on files with unprojected VARIANT columns
- Dominant language
- Java
- Stars
- 3.1k
- Forks
- 1.6k
- Avg merge
- 3d 12h
- Merged PRs (30d)
- 33
Description
### Summary
`parquet-java` 1.15.x fails when opening a Parquet file that contains a `VARIANT` column, even if the application only requests other columns and does not request the `VARIANT` column.
The failure happens while converting the footer schema, before projection is applied. The parquet-format compatibility guidance says that new logical types are intended to be forward compatible: https://github.com/apache/parquet-format/blob/master/CONTRIBUTING.md#compatibility-and-feature-enablement
Based on that, I think older readers should be able to tolerate an unknown logical type when that field is not part of the requested projection.
### Repro
This can be reproduced with the `VARIANT` fixture in `apache/parquet-testing`:
```bash
git clone https://github.com/apache/parquet-testing.git
PARQUET_FILE="$PWD/parquet-testing/shredded_variant/case-001.parquet"
```
The file contains an `id` column and a `var` column annotated with `VARIANT`.
Using `parquet-java` 1.15.x, try to read only `id` through the normal `GroupReadSupport` projection path:
```java
Path input = new Path(args[0]);
try (ParquetReader reader = ParquetReader.builder(new GroupReadSupport(), input)
.set(
ReadSupport.PARQUET_READ_SCHEMA,
"message root {\n" + "optional int32 id;\n" + "}")
.build()) {
Group row = reader.read();
System.out.println(row.getInteger("id", 0));
}
```
I also put up a draft repro PR with a failing unit test against `parquet-1.15.x`: https://github.com/kevinjqliu/parquet-java/pull/1
That test uses `apache/parquet-testing/shredded_variant/case-001.parquet` and requests only the `id` column. It still fails while reading footer metadata, before the projected read can happen.
### Actual behavior
The read fails during footer schema conversion:
```text
[ERROR] org.apache.parquet.hadoop.TestReadWithUnknownLogicalType.testReadProjectedColumnFromFileWithUnknownLogicalType -- Time elapsed: 0.283 s <<< ERROR!
java.lang.NullPointerException: Cannot invoke "org.apache.parquet.format.LogicalType$_Fields.ordinal()" because the return value of "org.apache.parquet.format.LogicalType.getSetField()" is null
at org.apache.parquet.format.converter.ParquetMetadataConverter.getLogicalTypeAnnotation(ParquetMetadataConverter.java:1174)
at org.apache.parquet.format.converter.ParquetMetadataConverter.buildChildren(ParquetMetadataConverter.java:1892)
at org.apache.parquet.format.converter.ParquetMetadataConverter.fromParquetSchema(ParquetMetadataConverter.java:1840)
at org.apache.parquet.format.converter.ParquetMetadataConverter.fromParquetMetadata(ParquetMetadataConverter.java:1670)
at org.apache.parquet.format.converter.ParquetMetadataConverter.readParquetMetadata(ParquetMetadataConverter.java:1630)
at org.apache.parquet.hadoop.ParquetFileReader.readFooter(ParquetFileReader.java:629)
at org.apache.parquet.hadoop.ParquetFileReader.(ParquetFileReader.java:934)
at org.apache.parquet.hadoop.ParquetFileReader.(ParquetFileReader.java:925)
```
The stack points to `ParquetMetadataConverter.getLogicalTypeAnnotation`.
### Expected behavior
Older readers should tolerate an unknown logical type in the footer when the field is not projected, preserving the physical schema or treating the field as having no known logical annotation.
They should only fail if the unsupported field is actually read or interpreted semantically.
### Notes
This is relevant for files written by newer or external writers that contain `VARIANT`, where older readers only need unrelated columns.
Contributor guide
No contributing guide indexed for this repository
Research direction
Start with ParquetMetadataConverter.getLogicalTypeAnnotation and the stack locations in ParquetMetadataConverter.java, then run the failing TestReadWithUnknownLogicalType test against parquet-testing/shredded_variant/case-001.parquet. Done means a projected read of id tolerates the unprojected VARIANT column while unsupported fields still fail only when semantically read.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java
- Domain
- data
- Issue type
- Bug
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Stale
- Clarity
- Clearly specified
- Newbie friendliness
- 55/100