Decouple parquet-hadoop module from hadoop-mapreduce-client-core
- Lingua principale
- Java
- Stelle
- 3.1k
- Fork
- 1.6k
- Merge medio
- 3g 12h
- PR unite (30g)
- 33
Descrizione
### Describe the enhancement requested
Decouple `ParquetReadOptions` from the legacy Hadoop `ParquetInputFormat`/`FileInputFormat` classes, which currently force pulling in the [`hadoop-mapreduce-client-core`](https://mvnrepository.com/artifact/org.apache.hadoop/hadoop-mapreduce-client-core) dependency (and its transitive JARs).
## Summary
To instantiate a `org.apache.parquet.hadoop.ParquetReader` we need to use `org.apache.parquet.ParquetReadOptions`, which references a set of keys located in `org.apache.parquet.hadoop.ParquetInputFormat` and calls a static `getFilter` method declared also on `ParquetInputFormat`, which `extends org.apache.hadoop.mapreduce.lib.input.FileInputFormat`.
`ParquetInputFormat` is part of the `parquet-hadoop` module, while `FileInputFormat` is declared in the `hadoop-mapreduce-client-core` JAR from the Hadoop project.
Because `ParquetReader` (via `ParquetReadOptions`) needs to call `ParquetInputFormat.getFilter(...)`, the JVM is forced to initialize `ParquetInputFormat`, and initializing a class triggers the loading and initialization of its superclass (`FileInputFormat`) along with its entire transitive dependency graph (`org.apache.hadoop.mapreduce.*`). For code that only needs to read a plain Parquet file (and not the MapReduce input-format machinery), this pulls an unwanted, heavy Hadoop-`mapreduce` dependency into the classpath and link set. `AvroParquetReader` and `ProtoParquetReader` extend from `ParquetReader` and have the same issue.
`ParquetInputFormat` has three responsibilities:
* define a set of property keys as constants
* deserialize the filter predicates from a configuration value
* support the integration of Parquet files into Hadoop MapReduce
This issue proposes to extract the first two responsibilities into two new Hadoop-agnostic classes, in the `org.apache.parquet.conf` package:
* `ParquetInputProperties`: the property-key constants only
* `ParquetInputFilters`: the filter deserialization logic
so that the configuration can be used without ever loading `FileInputFormat` and its transitive dependencies.
```
ParquetReader ← org.apache.parquet.hadoop (parquet-hadoop)
│ uses
▼
ParquetReadOptions ← org.apache.parquet (parquet-hadoop)
│ (static import keys + getFilter call)
▼
ParquetInputFormat.getFilter(...) ← org.apache.parquet.hadoop (parquet-hadoop)
│ extends
▼
FileInputFormat ← org.apache.hadoop.mapreduce.lib.input (hadoop-mapreduce-client-core)
│ extends
▼
InputFormat ← org.apache.hadoop.mapreduce.lib.input (hadoop-mapreduce-client-core)
│ (transitive)
▼
{ InputSplit, JobContext, TaskAttemptContext,
RecordReader, ... } ← org.apache.hadoop.mapreduce* (hadoop-mapreduce-client-core)
```
With the new `ParquetInputProperties` / `ParquetInputFilters`, building `ParquetReadOptions` no longer touches any `org.apache.hadoop.mapreduce` type, so consumers not related to Hadoop avoid transitively including the MapReduce dependency.
## Public API / Behavioral change
No behavior change. This is a refactoring:
- New classes `org.apache.parquet.conf.ParquetInputProperties` and
`org.apache.parquet.conf.ParquetInputFilters` (in `parquet-hadoop`) hold the constants and the
`ParquetConfiguration`-based filter resolution respectively.
- The legacy `org.apache.hadoop.conf.Configuration`-based entry points remain on
`ParquetInputFormat` (they are used only via the legacy MapReduce path) and are kept for binary /
source compatibility.
- All pre-existing constants on `ParquetInputFormat` are now `@Deprecated` and delegate to the new
classes; source and binary compatibility are preserved.
## Acceptance criteria
- `ParquetReadOptions` (used for plain file reads) references only `ParquetInputProperties` / `ParquetInputFilters`, never `ParquetInputFormat` / `org.apache.hadoop.mapreduce.InputFormat`.
- Loading `ParquetInputProperties` / `ParquetInputFilters` / `ParquetReadOptions` does not initialize `FileInputFormat`.
- All existing tests still pass (`./mvnw test`).
### Component(s)
Core
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Direzione di ricerca
Inizia con ParquetReadOptions e ParquetInputFormat nel modulo parquet-hadoop, quindi segui le costanti esistenti e la logica di risoluzione dei filtri. Introduci le classi org.apache.parquet.conf richieste, preservando i punti di ingresso esistenti di ParquetInputFormat e la compatibilità. Verifica che le nuove classi e ParquetReadOptions non inizializzino FileInputFormat, quindi esegui ./mvnw test.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- hadoop, java
- Ambito
- data-engineering
- Tipo di issue
- Refactoring
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Stato di attività
- Attiva
- Chiarezza
- Specificata chiaramente
- Idoneità per principianti
- 68/100