apache / apache/amoro

[Improvement]: Solve High Memory Usage in Planning Phase Caused by Massive Deleted Files

Open
#4,255 2 comments 0 reactions 0 assignees View on GitHub
type:improvement
Dominant language
Java
Stars
1.2k
Forks
395
Avg merge
4d 10h
Merged PRs (30d)
33

Description

### Search before asking

- [x] I have searched in the [issues](https://github.com/apache/amoro/issues?q=is%3Aissue) and found no similar issues.

### What would you like to be improved?

Currently, the planning phase suffers from excessive memory consumption when dealing with tables containing massive deleted small files. The root cause lies in the inefficient storage of file relationships: a single DeleteFile is often associated with multiple DataFiles.
In the current implementation, these associations are likely stored as explicit lists or object references. When a table has a large volume of data files referencing the same delete files, the memory overhead for maintaining these references grows unboundedly. This redundancy causes the planning index to consume significantly more heap memory than necessary, leading to potential Out-Of-Memory (OOM) errors and degraded performance during query planning.

### How should we improve?

I propose optimizing the memory layout of the planning index by introducing RoaringBitmap to compress the association between DeleteFile and DataFile. Instead of storing explicit lists of file IDs or object references, we can use RoaringBitmaps to represent the set of DataFile IDs associated with each DeleteFile. RoaringBitmap provides highly efficient compression for integer sets (file IDs), significantly reducing the memory footprint required to store these many-to-many relationships.

### Are you willing to submit PR?

- [x] Yes I am willing to submit a PR!

### Subtasks

_No response_

### Code of Conduct

- [x] I agree to follow this project's [Code of Conduct](https://www.apache.org/foundation/policies/conduct)

Contributor guide

Open the contributing guide

Research direction

Start by locating the planning index and its DeleteFile-to-DataFile association, then trace how the planning phase builds and retains those relationships. Done means the association uses the proposed compressed representation while planning remains correct and the excessive memory consumption is reduced for tables with many deleted small files.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
data-engineering, performance
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.