apache / apache/parquet-java

make it easy to read and write parquet files in java without depending on hadoop

オープン
#1,497 コメント 6 件 リアクション 3 件 担当者 0 名 GitHub で見る
Component: Parquet Priority: Major Type: enhancement
主要言語
Java
スター
3.1k
フォーク
1.6k
平均マージ
3日 12時間
マージ済み PR(30日)
33

説明

I am happy to help with this but I'd love some guidance on:

1) likelihood of being accepted as a patch.
2) how critical it is to maintain backwards compatibility in APIs.

For instance, we probably want to introduce a new artifact that lives under the existing hadoop depending artifact, and move as much code as possible to that, keeping the hadoop apis in the old artifact.

Welcome comments on solving this issue.

**Reporter**: [Oscar Boykin](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=posco)
#### Related issues:
- [Make ParquetIO Read splittable](https://issues.apache.org/jira/browse/BEAM-4379) (blocks)
- [hadoop-common is not an optional dependency](https://github.com/apache/parquet-java/issues/2556) (is duplicated by)
- [Avoid leaking Hadoop API to downstream libraries](https://github.com/apache/parquet-java/issues/2097) (incorporates)
- [Add Java NIO Avro OutputFile InputFile](https://github.com/apache/parquet-java/issues/2447) (is related to)
#### PRs and other links:
- [GitHub Pull Request #1376](https://github.com/apache/parquet-java/pull/1376)

**Note**: *This issue was originally created as [PARQUET-1126](https://issues.apache.org/jira/browse/PARQUET-1126). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

調査の方向性

まず pull request #1376 と、Hadoop 依存関係、API の漏洩、Java NIO のサポートに関する関連 issue を確認してください。この issue では、既存の artifact で互換性を維持しながら Hadoop に依存しない artifact を提案していますが、確定した設計や受け入れ基準は定義されていません。開始する前に現在の状況を確認してください。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
java
領域
data-engineering
issue の種類
機能追加
難易度
5/5
見積もり時間
1週間以上
活発さ
停滞
明瞭さ
説明が足りない
初心者へのやさしさ
20/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。