apache / apache/parquet-java

Off heap memory leaks with large binary fields using Snappy

Đang mở
#2,114 0 bình luận 0 reaction 0 người được giao Xem trên GitHub
Component: Java Component: Parquet Priority: Major Type: bug
Ngôn ngữ chính
Java
Star
3.1k
Fork
1.6k
Merge trung bình
3 ngày 12 giờ
Pull request đã merge (30 ngày)
33

Mô tả

When I write a large pages (~100MB) that contains large binary fields (~1MB), the java application uses an unexpected amount of off-heap memory (1.2GB)

This problem was identified when using the `AvroParquetWriter` but its source lies in the parquet-hadoop submodule.

Diving a little bit deeper shows the following:
- writing fields into the ParquetWriter creates a SequenceBytesIn which is actually just a list of `BytesInput` for each field. When calling `bytes.writeAllTo(cos)` in the `CodecFactory`, it actually writes one `ByteInput` (which contains a single field) at a time.
- the `SnappyCompressor` receives the data in `setInput` one large field at a time. This calls `ByteBuffer.allocateDirect` each time with a growing size. But as the memory is actually allocated off-heap, this does not trigger the garbage collector which only sees small objects on the heap. The actual memory associated with the object is the size of all the fields added to the page until then, so off-heap the memory is growing quadratically.

I did not attach a pull request to this issue because I see multiple mitigation to the issue but I'm not really delighted by any of them:
- merge all the fields into one byte array before pushing them down to the `SnappyCompressor`. For instance we could replace the previous statement in the `CodecFactory` with `BytesInput.from(bytes.toByteArray()).writeAllTo(cos)`. But this generates an extra on-heap allocation the size of the whole page.
- force the `DirectBuffer` to be cleaned up with something like `((DirectBuffer)inputBuffer).cleaner().clean()` after having copied it to the new bigger buffer. The issue here would be that `DirectBuffer` is part of the internal API and is likely to be moved. Using reflexion could make the solution more resilient but is even "hackier" IMHO.

**Reporter**: [Remi Dettai](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=remi.dettai)

**Note**: *This issue was originally created as [PARQUET-1188](https://issues.apache.org/jira/browse/PARQUET-1188). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Hướng nghiên cứu

Bắt đầu trong submodule parquet-hadoop, lần theo ParquetWriter và SequenceBytesIn đến CodecFactory và SnappyCompressor. Tái hiện sự cố với các page có kích thước khoảng 100MB chứa các trường nhị phân 1MB, sau đó đánh giá các hướng giảm thiểu được đề xuất. Hoàn thành khi các thao tác ghi lớn không còn gây tăng trưởng bậc hai của bộ nhớ off-heap mà vẫn giữ nguyên hành vi của writer.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
java
Lĩnh vực
performance
Loại issue
Lỗi
Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức độ hoạt động
Đình trệ
Độ rõ ràng
Cần làm rõ
Mức phù hợp với người mới
35/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.