coordinator: maintainer handoff can restart from a stale checkpoint
Nobody has claimed this yet.
- Dominant language
- Go
- Stars
- 56
- Forks
- 63
- Avg merge
- 2d 20h
- Merged PRs (30d)
- 34
Description
What did you do?
During a maintainer move, the origin’s final checkpoint can be discarded because it carries the previous ownership epoch.
Coordinator checkpoint = 100, current epoch = 1
|
| Start moving maintainer
v
Coordinator persists new epoch = 2
|
| Origin continues processing
v
Origin reports Stopped(epoch=1, checkpoint=200)
|
| Report is rejected as stale
v
Target starts with epoch=2, checkpoint=100
The stored checkpoint does not decrease, but the advancement from 100 to 200 is lost. The target unnecessarily replays that range.
The Coordinator should preserve the expected origin maintainer’s final checkpoint before creating the target maintainer, while continuing to reject unrelated stale-epoch reports.
What did you expect to see?
No response
What did you see instead?
As described above.
Versions of the cluster
Upstream TiDB cluster version (execute SELECT tidb_version(); in a MySQL client):
(paste TiDB cluster version here)
Upstream TiKV version (execute tikv-server --version):
(paste TiKV version here)
TiCDC version (execute cdc version):
(paste TiCDC version here)
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start at the Coordinator’s maintainer handoff and checkpoint handling, which are the entry points named in the report. Reproduce the epoch-2 transition where the origin reports Stopped(epoch=1, checkpoint=200), then verify that the expected origin’s final checkpoint is preserved while unrelated stale-epoch reports remain rejected.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- go
- Domain
- distributed-systems
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 52/100