pingcap / pingcap/tidb-operator

PD & TiDB only removes the failure members until all of the failed pods recover

Open
#2,160 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

priority:P2 status/discussion-wanted
Dominant language
Go
Stars
1.3k
Forks
540
Avg merge
3d 2h
Merged PRs (30d)
18

Description

Feature Request

Is your feature request related to a problem? Please describe:

PD & TiDB only removes the failure members in tc status until all of the failed pods recover, however, it's possible that only part of the failed pods recover, in which case, we can remove the failure members for the recover pods instead of waiting for all of the pods to recover.
We need to improve the failover logic.
Describe the feature you'd like:

Describe alternatives you've considered:

Teachability, Documentation, Adoption, Migration Strategy:

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by tracing the PD and TiDB failover logic that removes failed members and examine how it handles pod recovery. Reproduce a case where only some failed pods recover; done means the corresponding recovered members are removed without waiting for every failed pod to recover.

Written by the indexing model from the issue text.

Assessment

Tech stack
go, kubernetes
Domain
distributed-systems, infrastructure
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.