apache / apache/arrow-java

[Java] Parquet DatasetFileWriter not completely emptying allocator after finishing a write

Open
#1,255 0 comments 0 reactions 0 assignees View on GitHub
Type: usage
Dominant language
Java
Stars
94
Forks
152
Avg merge
3d 16h
Merged PRs (30d)
11

Description

### Describe the usage question you have. Please include as many useful details as possible.

I'm currently working on a somewhat simple project that processes Parquet files as both input and output:
- Main thread is a Parquet DatasetFactory that reads the file in batches. Here I have a child **readerAllocator**.
- The main thread schedules a fixed amount of worker threads that do the actual processing. They each have their own child **workerAllocator**.
- The main thread also schedules a single writer thread that reads (via a custom ArrowReader) a queue of worker thread futures from which to obtain the new VectorSchemaRoot with additional columns. The transfer is done just fine via VectorLoader/Unloader and sent to a child **writerArrowReaderAllocator**, which is then read by the actual writer's child **writerAllocator**.

I'll admit that this is my first time using the library and am not completely up to speed with the memory management, so my first reaction to seeing the allocators as an AutoCloseable, was to use Java's try-with-resource syntax to not have to faff about with closing the allocators manually.

Much, much, much more debugging later... I realize this is a terrible idea as the allocator's behavior is not always synced with the different objects (e.g.: DatasetFileWriter) that use them (i.e.: the allocators are not empty by the time the closeable objects are closed).

I would get my output parquet file (first one at least) but then the application would exit abruptly due to closing the allocators too early as they weren't empty.

Insanity set in and I wrote this (absolutely disgusting) piece of code out of desperation (putting it at the very end after the DatasetFileWriter.write call has finished):
````java
boolean childrenEmpty = false;
while (!childrenEmpty) {
childrenEmpty = true;
for (BufferAllocator childAllocator : rootAllocator.getChildAllocators()) {
if (allocator.getAllocatorMemory <= 0) allocator.close();
else childrenEmpty = false;
}
}
````

and... finally... everything started working... no memory errors anymore...

Doing a simple test with a parquet file of 1 row, it would take the **writerAllocator** roughly anywhere from 50 to 600 iterations until the allocated memory bytes would be down to 0 and then the allocator could be closed without errors.

I'm pretty I've made mistakes in my design given that this is my first time using the library. However, in every example I've seen (whether it be from random google results, the cookbook, or even Claude Opus), the writer's allocator is handled with a try-with-resource. Heck even the [class's test](https://github.com/apache/arrow-java/blob/main/dataset/src/test/java/org/apache/arrow/dataset/file/TestDatasetFileWriter.java) closes the allocator without checking if its empty...

So I would just like to know if there is some hidden known side effect that I haven't read about or if there is some usage guidelines I simply haven't followed.

I'm also open to any alternative suggestions on how to handle the closing of the allocators.

### Component(s)

OS: RHEL 9.4
Java version: OpenJDK 26.0.1 (compiling to version 25)
Arrow JAR versions: 19.0.0
Default Memory Allocation Manager: Netty
JVM options: `--sun-misc-unsafe-memory-access=allow --add-opens=java.base/java.nio=ALL-UNNAMED`

Contributor guide

Open the contributing guide

Research direction

Start with dataset/src/test/java/org/apache/arrow/dataset/file/TestDatasetFileWriter.java and the DatasetFileWriter.write call described in the report. Reproduce the one-row Parquet case with the writerAllocator and its child allocators, then trace when memory reaches zero relative to writer and allocator closure. Done means establishing whether this is expected usage or a library bug and documenting the correct lifecycle guidance.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
data
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.