Decouple parquet-hadoop module from hadoop-mapreduce-client-core
- 主要语言
- Java
- 星标
- 3.1k
- 派生
- 1.6k
- 平均合并
- 3 天 12 小时
- 30 天内合并 PR
- 33
描述
### Describe the enhancement requested
Decouple `ParquetReadOptions` from the legacy Hadoop `ParquetInputFormat`/`FileInputFormat` classes, which currently force pulling in the [`hadoop-mapreduce-client-core`](https://mvnrepository.com/artifact/org.apache.hadoop/hadoop-mapreduce-client-core) dependency (and its transitive JARs).
## Summary
To instantiate a `org.apache.parquet.hadoop.ParquetReader` we need to use `org.apache.parquet.ParquetReadOptions`, which references a set of keys located in `org.apache.parquet.hadoop.ParquetInputFormat` and calls a static `getFilter` method declared also on `ParquetInputFormat`, which `extends org.apache.hadoop.mapreduce.lib.input.FileInputFormat`.
`ParquetInputFormat` is part of the `parquet-hadoop` module, while `FileInputFormat` is declared in the `hadoop-mapreduce-client-core` JAR from the Hadoop project.
Because `ParquetReader` (via `ParquetReadOptions`) needs to call `ParquetInputFormat.getFilter(...)`, the JVM is forced to initialize `ParquetInputFormat`, and initializing a class triggers the loading and initialization of its superclass (`FileInputFormat`) along with its entire transitive dependency graph (`org.apache.hadoop.mapreduce.*`). For code that only needs to read a plain Parquet file (and not the MapReduce input-format machinery), this pulls an unwanted, heavy Hadoop-`mapreduce` dependency into the classpath and link set. `AvroParquetReader` and `ProtoParquetReader` extend from `ParquetReader` and have the same issue.
`ParquetInputFormat` has three responsibilities:
* define a set of property keys as constants
* deserialize the filter predicates from a configuration value
* support the integration of Parquet files into Hadoop MapReduce
This issue proposes to extract the first two responsibilities into two new Hadoop-agnostic classes, in the `org.apache.parquet.conf` package:
* `ParquetInputProperties`: the property-key constants only
* `ParquetInputFilters`: the filter deserialization logic
so that the configuration can be used without ever loading `FileInputFormat` and its transitive dependencies.
```
ParquetReader ← org.apache.parquet.hadoop (parquet-hadoop)
│ uses
▼
ParquetReadOptions ← org.apache.parquet (parquet-hadoop)
│ (static import keys + getFilter call)
▼
ParquetInputFormat.getFilter(...) ← org.apache.parquet.hadoop (parquet-hadoop)
│ extends
▼
FileInputFormat ← org.apache.hadoop.mapreduce.lib.input (hadoop-mapreduce-client-core)
│ extends
▼
InputFormat ← org.apache.hadoop.mapreduce.lib.input (hadoop-mapreduce-client-core)
│ (transitive)
▼
{ InputSplit, JobContext, TaskAttemptContext,
RecordReader, ... } ← org.apache.hadoop.mapreduce* (hadoop-mapreduce-client-core)
```
With the new `ParquetInputProperties` / `ParquetInputFilters`, building `ParquetReadOptions` no longer touches any `org.apache.hadoop.mapreduce` type, so consumers not related to Hadoop avoid transitively including the MapReduce dependency.
## Public API / Behavioral change
No behavior change. This is a refactoring:
- New classes `org.apache.parquet.conf.ParquetInputProperties` and
`org.apache.parquet.conf.ParquetInputFilters` (in `parquet-hadoop`) hold the constants and the
`ParquetConfiguration`-based filter resolution respectively.
- The legacy `org.apache.hadoop.conf.Configuration`-based entry points remain on
`ParquetInputFormat` (they are used only via the legacy MapReduce path) and are kept for binary /
source compatibility.
- All pre-existing constants on `ParquetInputFormat` are now `@Deprecated` and delegate to the new
classes; source and binary compatibility are preserved.
## Acceptance criteria
- `ParquetReadOptions` (used for plain file reads) references only `ParquetInputProperties` / `ParquetInputFilters`, never `ParquetInputFormat` / `org.apache.hadoop.mapreduce.InputFormat`.
- Loading `ParquetInputProperties` / `ParquetInputFilters` / `ParquetReadOptions` does not initialize `FileInputFormat`.
- All existing tests still pass (`./mvnw test`).
### Component(s)
Core
贡献指南
这个仓库没有索引到贡献指南
调研方向
从 parquet-hadoop 模块中的 ParquetReadOptions 和 ParquetInputFormat 开始,然后跟踪现有的常量和过滤器解析逻辑。引入所请求的 org.apache.parquet.conf 类,同时保留现有的 ParquetInputFormat 入口点和兼容性。验证新类和 ParquetReadOptions 不会初始化 FileInputFormat,然后运行 ./mvnw test。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- hadoop, java
- 领域
- data-engineering
- Issue 类型
- 重构
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 活跃度
- 活跃
- 描述清晰度
- 描述清楚
- 新手友好度
- 68/100