apache / apache/parquet-java

Decouple parquet-hadoop module from hadoop-mapreduce-client-core

Open
#3,780 0 comments 0 reactions 0 assignees View on GitHub
Type: enhancement
Dominant language
Java
Stars
3.1k
Forks
1.6k
Avg merge
3d 12h
Merged PRs (30d)
33

Description

### Describe the enhancement requested

Decouple `ParquetReadOptions` from the legacy Hadoop `ParquetInputFormat`/`FileInputFormat` classes, which currently force pulling in the [`hadoop-mapreduce-client-core`](https://mvnrepository.com/artifact/org.apache.hadoop/hadoop-mapreduce-client-core) dependency (and its transitive JARs).

## Summary

To instantiate a `org.apache.parquet.hadoop.ParquetReader` we need to use `org.apache.parquet.ParquetReadOptions`, which references a set of keys located in `org.apache.parquet.hadoop.ParquetInputFormat` and calls a static `getFilter` method declared also on `ParquetInputFormat`, which `extends org.apache.hadoop.mapreduce.lib.input.FileInputFormat`.

`ParquetInputFormat` is part of the `parquet-hadoop` module, while `FileInputFormat` is declared in the `hadoop-mapreduce-client-core` JAR from the Hadoop project.

Because `ParquetReader` (via `ParquetReadOptions`) needs to call `ParquetInputFormat.getFilter(...)`, the JVM is forced to initialize `ParquetInputFormat`, and initializing a class triggers the loading and initialization of its superclass (`FileInputFormat`) along with its entire transitive dependency graph (`org.apache.hadoop.mapreduce.*`). For code that only needs to read a plain Parquet file (and not the MapReduce input-format machinery), this pulls an unwanted, heavy Hadoop-`mapreduce` dependency into the classpath and link set. `AvroParquetReader` and `ProtoParquetReader` extend from `ParquetReader` and have the same issue.

`ParquetInputFormat` has three responsibilities:
* define a set of property keys as constants
* deserialize the filter predicates from a configuration value
* support the integration of Parquet files into Hadoop MapReduce

This issue proposes to extract the first two responsibilities into two new Hadoop-agnostic classes, in the `org.apache.parquet.conf` package:
* `ParquetInputProperties`: the property-key constants only
* `ParquetInputFilters`: the filter deserialization logic
so that the configuration can be used without ever loading `FileInputFormat` and its transitive dependencies.

```
ParquetReader ← org.apache.parquet.hadoop (parquet-hadoop)
│ uses

ParquetReadOptions ← org.apache.parquet (parquet-hadoop)
│ (static import keys + getFilter call)

ParquetInputFormat.getFilter(...) ← org.apache.parquet.hadoop (parquet-hadoop)
│ extends

FileInputFormat ← org.apache.hadoop.mapreduce.lib.input (hadoop-mapreduce-client-core)
│ extends

InputFormat ← org.apache.hadoop.mapreduce.lib.input (hadoop-mapreduce-client-core)
│ (transitive)

{ InputSplit, JobContext, TaskAttemptContext,
RecordReader, ... } ← org.apache.hadoop.mapreduce* (hadoop-mapreduce-client-core)
```

With the new `ParquetInputProperties` / `ParquetInputFilters`, building `ParquetReadOptions` no longer touches any `org.apache.hadoop.mapreduce` type, so consumers not related to Hadoop avoid transitively including the MapReduce dependency.

## Public API / Behavioral change

No behavior change. This is a refactoring:

- New classes `org.apache.parquet.conf.ParquetInputProperties` and
`org.apache.parquet.conf.ParquetInputFilters` (in `parquet-hadoop`) hold the constants and the
`ParquetConfiguration`-based filter resolution respectively.
- The legacy `org.apache.hadoop.conf.Configuration`-based entry points remain on
`ParquetInputFormat` (they are used only via the legacy MapReduce path) and are kept for binary /
source compatibility.
- All pre-existing constants on `ParquetInputFormat` are now `@Deprecated` and delegate to the new
classes; source and binary compatibility are preserved.

## Acceptance criteria

- `ParquetReadOptions` (used for plain file reads) references only `ParquetInputProperties` / `ParquetInputFilters`, never `ParquetInputFormat` / `org.apache.hadoop.mapreduce.InputFormat`.
- Loading `ParquetInputProperties` / `ParquetInputFilters` / `ParquetReadOptions` does not initialize `FileInputFormat`.
- All existing tests still pass (`./mvnw test`).

### Component(s)

Core

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with ParquetReadOptions and ParquetInputFormat in the parquet-hadoop module, then trace the existing constants and filter-resolution logic. Introduce the requested org.apache.parquet.conf classes while preserving the legacy ParquetInputFormat entry points and compatibility. Verify that the new classes and ParquetReadOptions do not initialize FileInputFormat, then run ./mvnw test.

Written by the indexing model from the issue text.

Assessment

Tech stack
hadoop, java
Domain
data-engineering
Issue type
Refactor
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Clearly specified
Newbie friendliness
68/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.