apache / apache/parquet-java

Implement async IO for Parquet file reader

Đang mở
#2,686 0 bình luận 0 reaction 0 người được giao Xem trên GitHub
Component: Java Component: Parquet Priority: Major Type: enhancement
Ngôn ngữ chính
Java
Star
3.1k
Fork
1.6k
Merge trung bình
3 ngày 12 giờ
Pull request đã merge (30 ngày)
33

Mô tả

ParquetFileReader's implementation has the following flow (simplified) - 
      - For every column -> Read from storage in 8MB blocks -> Read all uncompressed pages into output queue 
      - From output queues -> (downstream ) decompression + decoding

This flow is serialized, which means that downstream threads are blocked until the data has been read. Because a large part of the time spent is waiting for data from storage, threads are idle and CPU utilization is really low.

There is no reason why this cannot be made asynchronous _and_ parallel. So 

For Column _i_ -> reading one chunk until end, from storage -> intermediate output queue -> read one uncompressed page until end -> output queue -> (downstream ) decompression + decoding

Note that this can be made completely self contained in ParquetFileReader and downstream implementations like Iceberg and Spark will automatically be able to take advantage without code change as long as the ParquetFileReader apis are not changed. 

In past work with async io  [Drill - async page reader ](https://github.com/apache/drill/blob/master/exec/java-exec/src/main/java/org/apache/drill/exec/store/parquet/columnreaders/AsyncPageReader.java) , I have seen 2x-3x improvement in reading speed for Parquet files.

**Reporter**: [Parth Chandra](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=parthc) / @parthchandra
#### Related issues:
- [Improve Parquet IO Performance within cloud datalakes](https://github.com/apache/parquet-java/issues/2912) (is depended upon by)
#### PRs and other links:
- [GitHub Pull Request #968](https://github.com/apache/parquet-java/pull/968)

**Note**: *This issue was originally created as [PARQUET-2149](https://issues.apache.org/jira/browse/PARQUET-2149). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Hướng nghiên cứu

Bắt đầu bằng việc xem xét luồng tuần tự hiện tại của ParquetFileReader và Pull Request #968 được tham chiếu, sau đó so sánh Drill AsyncPageReader được liên kết với công việc I/O bất đồng bộ trước đây. Công việc được xem là hoàn tất khi việc đọc bộ nhớ lưu trữ và xử lý trang theo từng cột có thể tiến hành bất đồng bộ và song song mà không thay đổi các API của ParquetFileReader.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
java
Lĩnh vực
data-engineering, performance
Loại issue
Tính năng
Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức độ hoạt động
Đình trệ
Độ rõ ràng
Khá rõ ràng
Mức phù hợp với người mới
25/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.