apache / apache/parquet-java

Isolate Avro usage in parquet-cli to commands that require Avro

Open
#3,672 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Java
Stars
3.1k
Forks
1.6k
Avg merge
3d 12h
Merged PRs (30d)
33

Description

## Motivation

`parquet-cli` still uses Avro as an internal row model in places where users are only trying to inspect Parquet files. Valid Parquet is not always valid Avro, so Parquet inspection should not require Parquet schemas or records to round-trip through Avro.

`cat/head` already has a group-reader fallback for one class of problem: Avro schema conversion failures such as invalid Avro field names. That fallback is useful, but it is still a special case. Other inspection paths, such as `scan`, still use the Avro-backed data path for Parquet input.

Concrete examples:

- Invalid Avro field names: `original-instruction` and `column with known type` are valid Parquet field names but invalid Avro names. This already required a `cat/head` group-reader fallback.
- Nested Parquet structures: `data/nested_lists.snappy.parquet` is a valid `parquet-testing` fixture, but `cat` and `scan` fail through the Avro-backed reader.
- Projection/type mismatch: UUID and INT96 Parquet columns can fail when Avro projection rewrites the requested physical type.
- Lost Parquet structure: shredded Variant and other newer logical types can have valid Parquet physical structures that Avro cannot represent without losing information.

Avro is still part of the CLI API for conversion workflows. The goal is not to remove Avro; it is to stop requiring Avro for Parquet inspection.

## Proposal

Use Avro only when the command or input/output format requires it. Use Parquet-native readers and metadata APIs for Parquet inspection and physical rewrites.

| Command | Avro needed? | Proposed internal path |
|---|---:|---|
| `help`, `version` | No | CLI metadata only |
| `meta`, `pages`, `dictionary`, `check-stats`, `column-index`, `column-size`, `footer`, `bloom-filter`, `size-stats`, `geospatial-stats` | No | `ParquetFileReader` and Parquet metadata/page/stat APIs |
| `prune`, `trans-compression`, `masking`, `rewrite` | No | Parquet rewrite utilities |
| `cat`, `head` | Only for non-Parquet inputs | Parquet -> `GroupReadSupport`; Avro/JSON -> Avro path |
| `scan` | Only for non-Parquet inputs | Parquet -> `GroupReadSupport`; Avro/JSON -> Avro path |
| `schema` | Yes by default | Keep Avro schema output; keep physical Parquet path for `--parquet` |
| `csv-schema` | Yes | Avro schema inference |
| `convert-csv` | Yes | CSV -> Avro records/schema -> Parquet |
| `convert` | Yes | Avro-backed conversion path |
| `to-avro` | Yes | Avro writer/schema path |

## Related issues

- https://github.com/apache/parquet-java/issues/2836 / https://github.com/apache/parquet-java/pull/3332: `cat` failed because Avro rejected a Parquet field name. The merged fix added a Parquet group-reader fallback.
- https://github.com/apache/parquet-java/issues/2835: `head` reported the same class of invalid Avro field-name failure.
- https://github.com/apache/parquet-java/issues/2708: `cat` failed on parquet-protobuf files with modern LIST encoding via `AvroRecordConverter`.
- https://github.com/apache/parquet-java/issues/1641: UUID values failed because Avro projection changed the requested physical type; skipping projection let the file read.
- https://github.com/apache/parquet-java/issues/3345: INT96 data failed through Avro requested-schema conversion.
- https://github.com/apache/parquet-java/issues/3477: Avro schema conversion dropped `typed_value` from shredded Variant schemas.
- https://github.com/apache/parquet-java/issues/2339: `convert-csv` hit Avro naming rules. This should stay in the Avro-backed conversion area, not drive Parquet inspection behavior.

Tangential: https://github.com/apache/parquet-java/issues/2657 and https://github.com/apache/parquet-java/issues/2630 show additional CLI coupling to Avro internals, but are mostly binary/API mismatch issues.

## Summary

Keep Avro for Avro-facing workflows: `schema` default output, `to-avro`, `convert`, CSV/JSON conversion, and Avro input handling.

Use Parquet-native readers for Parquet inspection and physical rewrites. This preserves the Avro API where intentional while making `parquet-cli` more robust for valid Parquet files that Avro cannot model.

### Appendix
#### Repro examples

Using `parquet-mr version 1.17.1` and fixtures from `apache/parquet-testing`:

```bash
git clone https://github.com/apache/parquet-testing.git
cd parquet-testing
```

These Parquet inspection commands currently fail:

```bash
parquet cat -n 1 data/nested_lists.snappy.parquet
parquet scan data/nested_lists.snappy.parquet

parquet cat -n 1 shredded_variant/case-001.parquet
parquet scan shredded_variant/case-001.parquet

parquet cat -n 1 data/int96_from_spark.parquet
parquet scan data/int96_from_spark.parquet
```

Observed failure classes:

- `data/nested_lists.snappy.parquet`: Avro schema conversion succeeds, but `cat` and `scan` fail during Avro-backed record reading with `ParquetDecodingException`, followed by `GroupColumnIO.getFirst` index failure.
- `shredded_variant/case-001.parquet`: Avro schema conversion succeeds, but `cat` and `scan` fail during Avro-backed record reading with `ParquetDecodingException`, followed by `GroupColumnIO.getLast` index failure.
- `data/int96_from_spark.parquet`: Avro schema conversion fails before reading records with `Argument error: INT96 is deprecated...`, blocking Parquet row inspection.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by tracing the parquet-cli entry points for cat, head, scan, schema, conversion, and rewrite commands, including the existing GroupReadSupport fallback. Reproduce the failures with data/nested_lists.snappy.parquet, shredded_variant/case-001.parquet, and data/int96_from_spark.parquet. Done means Parquet inspection and physical rewrites use native paths while Avro remains for conversion workflows and non-Parquet inputs.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
cli, data-engineering
Issue type
Refactor
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.