moby / moby/swarmkit

Problems with raft and quorum after connecting new managers

Open
#2,670 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Go
Stars
3.7k
Forks
676
Avg merge
4d 9h
Merged PRs (30d)
6

Description

Docker version:
Managers: 18.04.0-ce
Workers: 17.12.0-ce

My clusters have been around for a few months. I was trying to rebuild the cluster to get them on consistent and updated versions. In the process of that I created a new autoscale group and connected new 18.04.0-ce managers to my existing cluster.

During the process every new node would fail with context deadline exceeded.

Normally we run with 5 manager nodes. During the process I attempted to scale down to 1 management node. I removed the previous nodes with docker node rm.

At this point adding any additional nodes would continue to fail:
Error: rpc error: code = DeadlineExceeded desc = context deadline exceeded

At some point during this I tried to force a new cluster with docker swarm init --force-new-cluster.
That worked until a manager attempted to join, then it would say there were not enough managers and drop the cluster back offline.

This process repeated 2 or 3 times.

docker swarm init --force-new-cluster

Swarm initialized: current node (sdgd5lpphdzmjrlnl5cpkhihm) is now a manager.

docker node ls

Error response from daemon: rpc error: code = Unknown desc = The swarm does not have a leader. It's possible that too few managers are online. Make sure more than half of the managers are online.

I finally found this docker-for-aws thread and ran the suggested adjustment for snapshot duration: https://github.com/docker/for-aws/issues/81

That appeared to settle things out, adding new nodes afterwards was fine and functional

At this point i'm doing our production environment in two weeks, I can update this further if the issue arrises again.

CC @johnharris85

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No source file or test is named in the report. Start by examining the manager-join and Raft quorum behavior around docker swarm init --force-new-cluster, then compare the context deadline exceeded and no-leader failures with the snapshot-duration adjustment linked from docker-for-aws#81. Done would require a reproducible fix or a regression test documenting the failure.

Written by the indexing model from the issue text.

Assessment

Tech stack
docker, go
Domain
devops, distributed-systems
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
15/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.