pingcap / pingcap/ticdc

coordinator: maintainer handoff can restart from a stale checkpoint

Open
#6,283 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

affects-8.5 severity/major type/bug
Dominant language
Go
Stars
56
Forks
63
Avg merge
2d 20h
Merged PRs (30d)
34

Description

What did you do?

During a maintainer move, the origin’s final checkpoint can be discarded because it carries the previous ownership epoch.

Coordinator checkpoint = 100, current epoch = 1
       |
       | Start moving maintainer
       v
Coordinator persists new epoch = 2
       |
       | Origin continues processing
       v
Origin reports Stopped(epoch=1, checkpoint=200)
       |
       | Report is rejected as stale
       v
Target starts with epoch=2, checkpoint=100

The stored checkpoint does not decrease, but the advancement from 100 to 200 is lost. The target unnecessarily replays that range.

The Coordinator should preserve the expected origin maintainer’s final checkpoint before creating the target maintainer, while continuing to reject unrelated stale-epoch reports.

What did you expect to see?

No response

What did you see instead?

As described above.

Versions of the cluster

Upstream TiDB cluster version (execute SELECT tidb_version(); in a MySQL client):

(paste TiDB cluster version here)

Upstream TiKV version (execute tikv-server --version):

(paste TiKV version here)

TiCDC version (execute cdc version):

(paste TiCDC version here)

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start at the Coordinator’s maintainer handoff and checkpoint handling, which are the entry points named in the report. Reproduce the epoch-2 transition where the origin reports Stopped(epoch=1, checkpoint=200), then verify that the expected origin’s final checkpoint is preserved while unrelated stale-epoch reports remain rejected.

Written by the indexing model from the issue text.

Assessment

Tech stack
go
Domain
distributed-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
52/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.