taosdata / taosdata/TDengine

if tdengine3 leader crash ,could not elect a new leader in K8s

Open
#20,399 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug
Dominant language
C
Stars
25.1k
Forks
5k
Avg merge
4d 59m
Merged PRs (30d)
7

Description

Bug Description
用官方提供k8s脚本部署tdengine3.0.2.5,配置3个mnode一旦leader机子宕机,选举会有问题导致一直无法选出leader,整个集群无法工作

To Reproduce

  1. 用官方k8s脚本(https://docs.taosdata.com/deployment/k8s/) 创建出replica为3的tdengine集群
  2. 进入集群通过 create mnode on dnode <dnode_id>; 的方式将3个replica都任命为mnode
  3. 然后在k8s下通过命令删除leader的replica(pod),后续此replica(pod)会自动重启
  4. 但即便自动重启完成,通过第二台replica(pod)的taos命令运行 show mnodes; 也会发现一直无法选出leader的现象

Expected Behavior
能选举成功

Screenshots
image

Environment (please complete the following information):

  • TDengine Version 3.0.2.5

Additional Context
原先的leader(mnode_id=1)的pod重启后的相关日志

03/10 02:04:12.273101 00000090 SYN vgId:1, begin election, sync:candidate, term:12, commit-index:10, first-ver:0, last-ver:26, min:-1, snap:10, snap-term:1, elect-times:9, as-leader-times:0, cfg-ch-times:0, hb-slow:0, hbr-slow:0, aq-items:-1, snaping:-1, replicas:3, last-cfg:-1, chging:0, restore:0, quorum:2, elect-lc-timer:10, hb:0, buffer:[10 10 26, 27), repl-mgrs:{0:0 [0 0, 0), 1:0 [0 0, 0), 2:0 [0 0, 0)}, members:{num:3, as:0, [tdengine-0.taosd.experimentb.svc.cluster.local:6030, tdengine-1.taosd.experimentb.svc.cluster.local:6030, tdengine-2.taosd.experimentb.svc.cluster.local:6030]}, hb:{0:1678413805684,1:1678413805684,2:1678413805684}, hb-reply:{0:1678413805684,1:1678413805684,2:1678413805684}
03/10 02:04:12.292439 00000090 SYN vgId:1, succeed to write raft store file:/var/lib/taos//mnode/sync/raft_store.json, term:13
03/10 02:04:12.311026 00000090 SYN vgId:1, succeed to write raft store file:/var/lib/taos//mnode/sync/raft_store.json, term:13
03/10 02:04:12.329363 00000090 SYN vgId:1, succeed to write raft store file:/var/lib/taos//mnode/sync/raft_store.json, term:13
03/10 02:04:13.545751 00000098 DND ERROR failed to send status req since Sync not leader, epSet:{tdengine-2.taosd.experimentb.svc.cluster.local:6030, tdengine-0.taosd.experimentb.svc.cluster.local:6030, tdengine-1.taosd.experimentb.svc.cluster.local:6030}, inUse:0

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the official Kubernetes deployment steps in the linked documentation, then reproduce the failure with three mnode replicas by deleting the leader pod. Inspect the mnode election and Raft log output, especially the term, quorum, membership, and leader-status messages. Done means the remaining replicas elect a leader and the cluster becomes usable after the pod restarts.

Written by the indexing model from the issue text.

Assessment

Tech stack
c, kubernetes
Domain
databases, devops, distributed-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.