apache / apache/arrow-java

[JAVA][C++]Support Parquet Read and Write in Java

Đang mở
#279 8 bình luận 1 reaction 0 người được giao Xem trên GitHub
Type: enhancement
Ngôn ngữ chính
Java
Star
94
Fork
152
Merge trung bình
3 ngày 16 giờ
Pull request đã merge (30 ngày)
11

Mô tả

We added a new java interface to support parquet read and write from hdfs or local file.

The purpose of this implementation is that when we loading and dumping parquet data in Java, we can only use rowBased put and get methods. Since arrow already has C++ implementation to load and dump parquet, so we wrapped those codes as Java APIs.

After test, we noticed in our workload, performance improved more than 2x comparing with rowBased load and dump. So we want to contribute codes to arrow.

since this is a total independent change, there is no codes change to current arrow codes. We added two folders as listed:  java/adapter/parquet and cpp/src/jni/parquet

**Reporter**: [Chendi.Xue](https://issues.apache.org/jira/browse/ARROW-6720)
#### Related issues:
- [[Java][Dataset] Implement Datasets Java API ](https://github.com/apache/arrow/issues/17055) (incorporates)
- [[Java][Dataset] Support writing to files within dataset scanner via JNI](https://github.com/apache/arrow/issues/27628) (incorporates)
#### PRs and other links:
- [GitHub Pull Request apache/arrow#5522](https://github.com/apache/arrow/pull/5522)
- [GitHub Pull Request apache/arrow#5717](https://github.com/apache/arrow/pull/5717)
- [GitHub Pull Request apache/arrow#5719](https://github.com/apache/arrow/pull/5719)

**Note**: *This issue was originally created as [ARROW-6720](https://issues.apache.org/jira/browse/ARROW-6720). Please see the [migration documentation](https://github.com/apache/arrow/issues/14542) for further details.*

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Hướng nghiên cứu

Bắt đầu bằng cách xem xét java/adapter/parquet và cpp/src/jni/parquet, sau đó so sánh các pull request được liên kết #5522, #5717 và #5719 với các issue dataset liên quan. Kiểm tra các Java API để đọc/ghi Parquet trên HDFS và tệp cục bộ, cùng với các bài kiểm thử của chúng; được xem là hoàn thành khi đã có hỗ trợ cần thiết mà không trùng lặp công việc dataset đã được tích hợp.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
cpp, java
Lĩnh vực
backend, data
Loại issue
Tính năng
Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức độ hoạt động
Đình trệ
Độ rõ ràng
Khá rõ ràng
Mức phù hợp với người mới
20/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.