apache / apache/fluss

Use Leader Epoch rather than High Watermark for Truncation

Open
#673 0 comments 0 reactions 1 assignee Claimed by @swuferhong View on GitHub
component=log component=server
Dominant language
Java
Stars
2.1k
Forks
625
Avg merge
3d 14h
Merged PRs (30d)
97

Description

### Search before asking

- [x] I searched in the [issues](https://github.com/alibaba/fluss/issues) and found nothing similar.

### Motivation

Use leader epoch rather than high watermark for log truncation to avoid inconsistent or lost data when the leader or follower goes offline abnormally.

Currently, the synchronization of fluss `highWatermark` involves first updating the leader and then updating the follower in the next round. `Log recovery` also depends on the `highWatermark` in the leader's local checkpoint cache to determine whether to truncate data. However, this approach poses a significant risk of data loss. In issue #674 , we will modify the `highWatermark synchronization` way to update the follower first and then the leader in the next round. While this change can solve the data loss problem caused by server abnormal exits, it may introduce new problem of data duplication. Therefore, in this issue, we aim to adopt the approach from Kafka [KIP-101](https://cwiki.apache.org/confluence/display/KAFKA/KIP-101+-+Alter+Replication+Protocol+to+use+Leader+Epoch+rather+than+High+Watermark+for+Truncation) to resolve this problem through supporting leader epoch cache to avoid Inconsistent or lost data when the leader or follower goes offline abnormally.

### Solution

_No response_

### Anything else?

_No response_

### Willingness to contribute

- [ ] I'm willing to submit a PR!

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.