apache / apache/parquet-java

Older `parquet-java` readers fail on files with unprojected VARIANT columns

Offen
#3,633 1 Kommentar 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Vorherrschende Sprache
Java
Sterne
3.1k
Forks
1.6k
Ø Merge
3 T. 12 Std.
Gemergte PRs (30 T.)
33

Beschreibung

### Summary

`parquet-java` 1.15.x fails when opening a Parquet file that contains a `VARIANT` column, even if the application only requests other columns and does not request the `VARIANT` column.

The failure happens while converting the footer schema, before projection is applied. The parquet-format compatibility guidance says that new logical types are intended to be forward compatible: https://github.com/apache/parquet-format/blob/master/CONTRIBUTING.md#compatibility-and-feature-enablement

Based on that, I think older readers should be able to tolerate an unknown logical type when that field is not part of the requested projection.

### Repro

This can be reproduced with the `VARIANT` fixture in `apache/parquet-testing`:

```bash
git clone https://github.com/apache/parquet-testing.git
PARQUET_FILE="$PWD/parquet-testing/shredded_variant/case-001.parquet"
```

The file contains an `id` column and a `var` column annotated with `VARIANT`.

Using `parquet-java` 1.15.x, try to read only `id` through the normal `GroupReadSupport` projection path:

```java
Path input = new Path(args[0]);

try (ParquetReader reader = ParquetReader.builder(new GroupReadSupport(), input)
.set(
ReadSupport.PARQUET_READ_SCHEMA,
"message root {\n" + "optional int32 id;\n" + "}")
.build()) {
Group row = reader.read();
System.out.println(row.getInteger("id", 0));
}
```

I also put up a draft repro PR with a failing unit test against `parquet-1.15.x`: https://github.com/kevinjqliu/parquet-java/pull/1

That test uses `apache/parquet-testing/shredded_variant/case-001.parquet` and requests only the `id` column. It still fails while reading footer metadata, before the projected read can happen.

### Actual behavior

The read fails during footer schema conversion:

```text
[ERROR] org.apache.parquet.hadoop.TestReadWithUnknownLogicalType.testReadProjectedColumnFromFileWithUnknownLogicalType -- Time elapsed: 0.283 s <<< ERROR!
java.lang.NullPointerException: Cannot invoke "org.apache.parquet.format.LogicalType$_Fields.ordinal()" because the return value of "org.apache.parquet.format.LogicalType.getSetField()" is null
at org.apache.parquet.format.converter.ParquetMetadataConverter.getLogicalTypeAnnotation(ParquetMetadataConverter.java:1174)
at org.apache.parquet.format.converter.ParquetMetadataConverter.buildChildren(ParquetMetadataConverter.java:1892)
at org.apache.parquet.format.converter.ParquetMetadataConverter.fromParquetSchema(ParquetMetadataConverter.java:1840)
at org.apache.parquet.format.converter.ParquetMetadataConverter.fromParquetMetadata(ParquetMetadataConverter.java:1670)
at org.apache.parquet.format.converter.ParquetMetadataConverter.readParquetMetadata(ParquetMetadataConverter.java:1630)
at org.apache.parquet.hadoop.ParquetFileReader.readFooter(ParquetFileReader.java:629)
at org.apache.parquet.hadoop.ParquetFileReader.(ParquetFileReader.java:934)
at org.apache.parquet.hadoop.ParquetFileReader.(ParquetFileReader.java:925)
```

The stack points to `ParquetMetadataConverter.getLogicalTypeAnnotation`.

### Expected behavior

Older readers should tolerate an unknown logical type in the footer when the field is not projected, preserving the physical schema or treating the field as having no known logical annotation.

They should only fail if the unsupported field is actually read or interpreted semantically.

### Notes

This is relevant for files written by newer or external writers that contain `VARIANT`, where older readers only need unrelated columns.

Beitragsleitfaden

Für dieses Repository ist kein Beitragsleitfaden indexiert

Rechercherichtung

Start with ParquetMetadataConverter.getLogicalTypeAnnotation and the stack locations in ParquetMetadataConverter.java, then run the failing TestReadWithUnknownLogicalType test against parquet-testing/shredded_variant/case-001.parquet. Done means a projected read of id tolerates the unprojected VARIANT column while unsupported fields still fail only when semantically read.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
java
Bereich
data
Issue-Typ
Bug
Schwierigkeit
3/5
Geschätzter Aufwand
1-2 Tage
Aktivitätsstatus
Veraltet
Klarheit
Klar beschrieben
Anfängerfreundlichkeit
55/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.