chef / chef/chef-server

Lost leader after restarting chef-backend

Open
#1,473 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Aspect: Stability Component: chef-backend Triage: Try Reproducing Type: Bug
Dominant language
Erlang
Stars
303
Forks
211
Avg merge
1d 8h
Merged PRs (30d)
5

Description

Using a chef-backend cluster with 3 backend nodes - backend1, backend2 and backend3. backend1 is the current leader. When I stop chef-backend (chef-backendctl stop) on all three machines, the start it up again the cluster has no leader and never recovers.

Expected Behavior

Leader to be re-elected after restarting chef-backend on all three backend machines.

Current Behavior

Leader is never re-elected after starting the services and leaderl logs show:

2018-02-15_17:03:14.91161 [I] <0.715.0> leader_elector:init => leader_elector is starting in state: initializing
2018-02-15_17:03:14.91226 [I] <0.716.0> status_updater:init => status_updater started
2018-02-15_17:03:14.97625 [I] <0.563.0> no_mod:no_fun => Application leaderl started on node 'leaderl@127.0.0.1'
2018-02-15_17:03:14.97720 [I] <0.563.0> no_mod:no_fun => Application eper started on node 'leaderl@127.0.0.1'
2018-02-15_17:03:14.97785 [I] <0.563.0> no_mod:no_fun => Application recon started on node 'leaderl@127.0.0.1'
2018-02-15_17:03:15.09875 [I] <0.730.0> key_watcher:handle_info => Initial start of watcher on behalf of leader_block_state for "cb/control/blocked_leaders"
2018-02-15_17:03:15.09967 [I] <0.729.0> leader_elector:do_connect => Connecting as node backend1 (10.40.10.56,da5d0b3758d410e434be81a1453b4c24)
2018-02-15_17:03:15.09997 [I] <0.729.0> leader_elector:do_connect => Leader no_leader, Boot no_bootstrap_node
2018-02-15_17:03:15.10002 [I] <0.729.0> leader_elector:do_connect => No leader in place.
2018-02-15_17:03:15.10048 [I] <0.729.0> leader_elector:do_connect => I am NOT a bootstrap node because bootstrap_key returned no_bootstrap_node so I'll wait for a leader.

cluster-status shows:

Name            IP           GUID                              Role                PG        ES
backend2  10.40.10.57  59540d34fa370a28ed3098cc78b2a245  waiting_for_leader  follower  not_master
backend1  10.40.10.56  da5d0b3758d410e434be81a1453b4c24  waiting_for_leader  leader    not_master
backend3  10.40.10.55  8c4ae52bf8a7b77ee65b638f7159199a  waiting_for_leader  follower  master

Steps to Reproduce (for bugs)

  1. Starting with backend1, then backend2 followed by backend3 "chef-backendctl stop"
  2. Bring back chef in the reverse order - staring with backend3, then backend2 followed by backend3 "chef-backendctl start"

Your Environment

  • Chef Server Version: chef-backend 2.0.1
  • Total/free RAM and disk space: Total RAM 8GB / free RAM 5.2 GB
  • Operating System and Version: ubuntu 14.04.5 LTS
  • If upgrading, previous Chef Server version: N/A
  • Running in a container? no

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the leader_elector log messages and the documented three-node restart sequence, using the chef-backend 2.0.1 environment details as context. Reproduce the stop/start order and inspect why all nodes remain waiting_for_leader; done means a leader is re-elected and cluster-status no longer reports waiting_for_leader.

Written by the indexing model from the issue text.

Assessment

Tech stack
erlang
Domain
backend, distributed-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
32/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.