apache / apache/parquet-java

Decouple parquet-hadoop module from hadoop-mapreduce-client-core

Ouverte
#3,780 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub
Type: enhancement
Langage dominant
Java
Étoiles
3.1k
Forks
1.6k
Merge moyen
3 j 12 h
PR mergées (30 j)
33

Description

### Describe the enhancement requested

Decouple `ParquetReadOptions` from the legacy Hadoop `ParquetInputFormat`/`FileInputFormat` classes, which currently force pulling in the [`hadoop-mapreduce-client-core`](https://mvnrepository.com/artifact/org.apache.hadoop/hadoop-mapreduce-client-core) dependency (and its transitive JARs).

## Summary

To instantiate a `org.apache.parquet.hadoop.ParquetReader` we need to use `org.apache.parquet.ParquetReadOptions`, which references a set of keys located in `org.apache.parquet.hadoop.ParquetInputFormat` and calls a static `getFilter` method declared also on `ParquetInputFormat`, which `extends org.apache.hadoop.mapreduce.lib.input.FileInputFormat`.

`ParquetInputFormat` is part of the `parquet-hadoop` module, while `FileInputFormat` is declared in the `hadoop-mapreduce-client-core` JAR from the Hadoop project.

Because `ParquetReader` (via `ParquetReadOptions`) needs to call `ParquetInputFormat.getFilter(...)`, the JVM is forced to initialize `ParquetInputFormat`, and initializing a class triggers the loading and initialization of its superclass (`FileInputFormat`) along with its entire transitive dependency graph (`org.apache.hadoop.mapreduce.*`). For code that only needs to read a plain Parquet file (and not the MapReduce input-format machinery), this pulls an unwanted, heavy Hadoop-`mapreduce` dependency into the classpath and link set. `AvroParquetReader` and `ProtoParquetReader` extend from `ParquetReader` and have the same issue.

`ParquetInputFormat` has three responsibilities:
* define a set of property keys as constants
* deserialize the filter predicates from a configuration value
* support the integration of Parquet files into Hadoop MapReduce

This issue proposes to extract the first two responsibilities into two new Hadoop-agnostic classes, in the `org.apache.parquet.conf` package:
* `ParquetInputProperties`: the property-key constants only
* `ParquetInputFilters`: the filter deserialization logic
so that the configuration can be used without ever loading `FileInputFormat` and its transitive dependencies.

```
ParquetReader ← org.apache.parquet.hadoop (parquet-hadoop)
│ uses

ParquetReadOptions ← org.apache.parquet (parquet-hadoop)
│ (static import keys + getFilter call)

ParquetInputFormat.getFilter(...) ← org.apache.parquet.hadoop (parquet-hadoop)
│ extends

FileInputFormat ← org.apache.hadoop.mapreduce.lib.input (hadoop-mapreduce-client-core)
│ extends

InputFormat ← org.apache.hadoop.mapreduce.lib.input (hadoop-mapreduce-client-core)
│ (transitive)

{ InputSplit, JobContext, TaskAttemptContext,
RecordReader, ... } ← org.apache.hadoop.mapreduce* (hadoop-mapreduce-client-core)
```

With the new `ParquetInputProperties` / `ParquetInputFilters`, building `ParquetReadOptions` no longer touches any `org.apache.hadoop.mapreduce` type, so consumers not related to Hadoop avoid transitively including the MapReduce dependency.

## Public API / Behavioral change

No behavior change. This is a refactoring:

- New classes `org.apache.parquet.conf.ParquetInputProperties` and
`org.apache.parquet.conf.ParquetInputFilters` (in `parquet-hadoop`) hold the constants and the
`ParquetConfiguration`-based filter resolution respectively.
- The legacy `org.apache.hadoop.conf.Configuration`-based entry points remain on
`ParquetInputFormat` (they are used only via the legacy MapReduce path) and are kept for binary /
source compatibility.
- All pre-existing constants on `ParquetInputFormat` are now `@Deprecated` and delegate to the new
classes; source and binary compatibility are preserved.

## Acceptance criteria

- `ParquetReadOptions` (used for plain file reads) references only `ParquetInputProperties` / `ParquetInputFilters`, never `ParquetInputFormat` / `org.apache.hadoop.mapreduce.InputFormat`.
- Loading `ParquetInputProperties` / `ParquetInputFilters` / `ParquetReadOptions` does not initialize `FileInputFormat`.
- All existing tests still pass (`./mvnw test`).

### Component(s)

Core

Guide de contribution

Aucun guide de contribution indexé pour ce dépôt

Piste de recherche

Commencez par ParquetReadOptions et ParquetInputFormat dans le module parquet-hadoop, puis suivez les constantes existantes et la logique de résolution des filtres. Introduisez les classes org.apache.parquet.conf demandées tout en préservant les points d’entrée existants de ParquetInputFormat et la compatibilité. Vérifiez que les nouvelles classes et ParquetReadOptions n’initialisent pas FileInputFormat, puis exécutez ./mvnw test.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
hadoop, java
Domaine
data-engineering
Type d'issue
Refactorisation
Difficulté
4/5
Temps estimé
3-5 jours
Activité
Active
Clarté
Clairement spécifiée
Accessibilité débutants
68/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.