apache / apache/hudi

[SUPPORT] Inconsistent Checkpoint Size in Flink Applications with MoR

Open
#10,329 14 comments 0 reactions 0 assignees View on GitHub
engine:flink type:feature
Dominant language
Java
Stars
6.2k
Forks
2.5k
Avg merge
2d 8h
Merged PRs (30d)
111

Description

I have an Apache Flink Application that leverages RockDB and incremental checkpoint, but it seems that every time a compaction task occurs, the checkpoint size of the application during that interval increases dramatically. Is this due to when doing compaction, the streaming job has to load the entire table? if thats the case, Is there any way it can only compact and load the current partition?

Contributor guide

No contributing guide indexed for this repository

Research direction

No source files, tests, or reproducible configuration are identified. Begin by collecting the Flink/RocksDB incremental-checkpoint and compaction settings and comparing checkpoint metrics before and during compaction; done requires a confirmed explanation and an actionable scope for the current-partition behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
data-engineering, stream-processing
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.