apache / apache/parquet-java

make it easy to read and write parquet files in java without depending on hadoop

未关闭
#1,497 6 条评论 3 个 reaction 已指派 0 人 在 GitHub 查看
Component: Parquet Priority: Major Type: enhancement
主要语言
Java
星标
3.1k
派生
1.6k
平均合并
3 天 12 小时
30 天内合并 PR
33

描述

I am happy to help with this but I'd love some guidance on:

1) likelihood of being accepted as a patch.
2) how critical it is to maintain backwards compatibility in APIs.

For instance, we probably want to introduce a new artifact that lives under the existing hadoop depending artifact, and move as much code as possible to that, keeping the hadoop apis in the old artifact.

Welcome comments on solving this issue.

**Reporter**: [Oscar Boykin](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=posco)
#### Related issues:
- [Make ParquetIO Read splittable](https://issues.apache.org/jira/browse/BEAM-4379) (blocks)
- [hadoop-common is not an optional dependency](https://github.com/apache/parquet-java/issues/2556) (is duplicated by)
- [Avoid leaking Hadoop API to downstream libraries](https://github.com/apache/parquet-java/issues/2097) (incorporates)
- [Add Java NIO Avro OutputFile InputFile](https://github.com/apache/parquet-java/issues/2447) (is related to)
#### PRs and other links:
- [GitHub Pull Request #1376](https://github.com/apache/parquet-java/pull/1376)

**Note**: *This issue was originally created as [PARQUET-1126](https://issues.apache.org/jira/browse/PARQUET-1126). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

贡献指南

这个仓库没有索引到贡献指南

调研方向

首先审查 pull request #1376,以及与 Hadoop 依赖、API 泄漏和 Java NIO 支持相关的 issue。该 issue 提议使用独立于 Hadoop 的 artifact,同时保留现有 artifact 中的兼容性,但没有定义已确定的设计或验收标准;开始之前请确认当前状态。

由索引模型根据 Issue 内容生成。

评估

技术栈
java
领域
data-engineering
Issue 类型
功能
难度
5/5
预计耗时
一周以上
活跃度
停滞
描述清晰度
需要澄清
新手友好度
20/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。