ParquetOutputFormat should support custom OutputCommitter
- Ngôn ngữ chính
- Java
- Star
- 3.1k
- Fork
- 1.6k
- Merge trung bình
- 3 ngày 12 giờ
- Pull request đã merge (30 ngày)
- 33
Mô tả
ParquetOutputFormat should support custom OutputCommitter.
There is a need to bypass current Hadoop functionality of writing output data under **_temporary** folder. Especially with AWS S3, there can be huge overhead of moving the files from **_temporary** folder to output folder.
**Reporter**: [Mikko Kupsu](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=mikkokupsu)
**Assignee**: [Steve Loughran](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=stevel@apache.org) / @steveloughran
#### Related issues:
- [Improve Parquet IO Performance within cloud datalakes](https://github.com/apache/parquet-java/issues/2912) (is depended upon by)
#### PRs and other links:
- [GitHub Pull Request #1361](https://github.com/apache/parquet-java/pull/1361)
**Note**: *This issue was originally created as [PARQUET-781](https://issues.apache.org/jira/browse/PARQUET-781). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*
Hướng dẫn đóng góp
Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này
Hướng nghiên cứu
Bắt đầu với ParquetOutputFormat và xem xét tích hợp OutputCommitter của Hadoop, sau đó đọc pull request #1361 về phần việc hiện có. Công việc được hoàn thành khi có thể hỗ trợ các custom committers mà vẫn tránh các lần ghi không cần thiết vào _temporary và việc di chuyển tệp trên S3.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Đánh giá
- Công nghệ
- aws, hadoop, java
- Lĩnh vực
- cloud, data-engineering
- Loại issue
- Tính năng
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức độ hoạt động
- Đình trệ
- Độ rõ ràng
- Khá rõ ràng
- Mức phù hợp với người mới
- 25/100