Revisit Savepoint Design to handle OOM issues due to very large save points
Open
area:table-service
from-jira
priority:critical
type:improvement
- Dominant language
- Java
- Stars
- 6.2k
- Forks
- 2.5k
- Avg merge
- 2d 8h
- Merged PRs (30d)
- 111
Description
Github Issue - [https://github.com/apache/hudi/issues/9747]
## JIRA info
- Link: https://issues.apache.org/jira/browse/HUDI-6981
- Type: Improvement
- Fix version(s):
- 1.2.0
Contributor guide
No contributing guide indexed for this repository
Research direction
Start with the linked GitHub issue #9747 and JIRA item HUDI-6981 to recover the missing design context for Hudi savepoints. Identify the current savepoint design and the conditions that cause out-of-memory errors from very large savepoints. Done should mean an agreed and implemented design that handles those large savepoints without OOM failures, with validation appropriate to the changed behavior.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java
- Domain
- data-engineering, distributed-systems
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100