Graylog2 / Graylog2/graylog2-server
Datanode - Datanode not recovering after previous cluster-manager isn't available/ready at start
- Dominant language
- Java
- Stars
- 8.1k
- Forks
- 1.1k
- Avg merge
- 1d 20h
- Merged PRs (30d)
- 217
Description
## Expected Behavior
After initially not finding the previous cluster-manager, the datanode should recover when the previous cluster-manager becomes available/ready.
## Current Behavior
Datanode stays in loop trying to discover/elect cluster-manager.
`
2025-04-04T06:38:48.212Z INFO [OpensearchProcessImpl] [2025-04-04T06:38:48,212][WARN ][o.o.c.c.ClusterFormationFailureHelper] [ansibleopensearch02] cluster-manager not discovered or elected yet, an election requires a node with id [9tukJ6SzSb6O5naDEQ0P6Q], have discovered [{ansibleopensearch02}{XtQ51mPFSvuHYAUOpoH_PQ}{2-WQWNYQTFaQGNq4Mj3y-Q}{ansibleopensearch02}{192.168.15.122:9300}{dimr}{shard_indexing_pressure_enabled=true}] which is not a quorum; discovery will continue using [127.0.0.1:9300, [::1]:9300, 127.0.1.1:9300] from hosts providers and [{ansibleopensearch02}{XtQ51mPFSvuHYAUOpoH_PQ}{2-WQWNYQTFaQGNq4Mj3y-Q}{ansibleopensearch02}{192.168.15.122:9300}{dimr}{shard_indexing_pressure_enabled=true}] from last-known cluster state; node term 9, last-accepted version 403 in term 9
2025-04-04T06:38:49.043Z INFO [OpensearchProcessImpl] [2025-04-04T06:38:49,043][INFO ][o.o.s.c.ConfigurationRepository] [ansibleopensearch02] Wait for cluster to be available ...
2025-04-04T06:38:50.044Z INFO [OpensearchProcessImpl] [2025-04-04T06:38:50,043][INFO ][o.o.s.c.ConfigurationRepository] [ansibleopensearch02] Wait for cluster to be available ...
2025-04-04T06:38:51.044Z INFO [OpensearchProcessImpl] [2025-04-04T06:38:51,044][INFO ][o.o.s.c.ConfigurationRepository] [ansibleopensearch02] Wait for cluster to be available ...
2025-04-04T06:38:52.044Z INFO [OpensearchProcessImpl] [2025-04-04T06:38:52,044][INFO ][o.o.s.c.ConfigurationRepository] [ansibleopensearch02] Wait for cluster to be available ...
2025-04-04T06:38:53.045Z INFO [OpensearchProcessImpl] [2025-04-04T06:38:53,044][INFO ][o.o.s.c.ConfigurationRepository] [ansibleopensearch02] Wait for cluster to be available ...
2025-04-04T06:38:54.045Z INFO [OpensearchProcessImpl] [2025-04-04T06:38:54,045][INFO ][o.o.s.c.ConfigurationRepository] [ansibleopensearch02] Wait for cluster to be available ...
2025-04-04T06:38:55.045Z INFO [OpensearchProcessImpl] [2025-04-04T06:38:55,045][INFO ][o.o.s.c.ConfigurationRepository] [ansibleopensearch02] Wait for cluster to be available ...
2025-04-04T06:38:56.046Z INFO [OpensearchProcessImpl] [2025-04-04T06:38:56,045][INFO ][o.o.s.c.ConfigurationRepository] [ansibleopensearch02] Wait for cluster to be available ...
2025-04-04T06:38:57.046Z INFO [OpensearchProcessImpl] [2025-04-04T06:38:57,046][INFO ][o.o.s.c.ConfigurationRepository] [ansibleopensearch02] Wait for cluster to be available ...
2025-04-04T06:38:58.046Z INFO [OpensearchProcessImpl] [2025-04-04T06:38:58,046][INFO ][o.o.s.c.ConfigurationRepository] [ansibleopensearch02] Wait for cluster to be available ...
2025-04-04T06:38:58.214Z INFO [OpensearchProcessImpl] [2025-04-04T06:38:58,214][WARN ][o.o.c.c.ClusterFormationFailureHelper] [ansibleopensearch02] cluster-manager not discovered or elected yet, an election requires a node with id [9tukJ6SzSb6O5naDEQ0P6Q], have discovered [{ansibleopensearch02}{XtQ51mPFSvuHYAUOpoH_PQ}{2-WQWNYQTFaQGNq4Mj3y-Q}{ansibleopensearch02}{192.168.15.122:9300}{dimr}{shard_indexing_pressure_enabled=true}] which is not a quorum; discovery will continue using [127.0.0.1:9300, [::1]:9300, 127.0.1.1:9300] from hosts providers and [{ansibleopensearch02}{XtQ51mPFSvuHYAUOpoH_PQ}{2-WQWNYQTFaQGNq4Mj3y-Q}{ansibleopensearch02}{192.168.15.122:9300}{dimr}{shard_indexing_pressure_enabled=true}] from last-known cluster state; node term 9, last-accepted version 403 in term 9
2025-04-04T06:38:59.047Z INFO [OpensearchProcessImpl] [2025-04-04T06:38:59,047][INFO ][o.o.s.c.ConfigurationRepository] [ansibleopensearch02] Wait for cluster to be available ...
2025-04-04T06:39:00.048Z INFO [OpensearchProcessImpl] [2025-04-04T06:39:00,047][INFO ][o.o.s.c.ConfigurationRepository] [ansibleopensearch02] Wait for cluster to be available ...
2025-04-04T06:39:01.048Z INFO [OpensearchProcessImpl] [2025-04-04T06:39:01,048][INFO ][o.o.s.c.ConfigurationRepository] [ansibleopensearch02] Wait for cluster to be available ...
2025-04-04T06:39:02.048Z INFO [OpensearchProcessImpl] [2025-04-04T06:39:02,048][INFO ][o.o.s.c.ConfigurationRepository] [ansibleopensearch02] Wait for cluster to be available ...
2025-04-04T06:39:03.049Z INFO [OpensearchProcessImpl] [2025-04-04T06:39:03,049][INFO ][o.o.s.c.ConfigurationRepository] [ansibleopensearch02] Wait for cluster to be available ...
2025-04-04T06:39:04.050Z INFO [OpensearchProcessImpl] [2025-04-04T06:39:04,050][INFO ][o.o.s.c.ConfigurationRepository] [ansibleopensearch02] Wait for cluster to be available ...
2025-04-04T06:39:05.050Z INFO [OpensearchProcessImpl] [2025-04-04T06:39:05,050][INFO ][o.o.s.c.ConfigurationRepository] [ansibleopensearch02] Wait for cluster to be available ...
2025-04-04T06:39:06.051Z INFO [OpensearchProcessImpl] [2025-04-04T06:39:06,051][INFO ][o.o.s.c.ConfigurationRepository] [ansibleopensearch02] Wait for cluster to be available ...
2025-04-04T06:39:07.051Z INFO [OpensearchProcessImpl] [2025-04-04T06:39:07,051][INFO ][o.o.s.c.ConfigurationRepository] [ansibleopensearch02] Wait for cluster to be available ...
2025-04-04T06:39:08.052Z INFO [OpensearchProcessImpl] [2025-04-04T06:39:08,051][INFO ][o.o.s.c.ConfigurationRepository] [ansibleopensearch02] Wait for cluster to be available ...
2025-04-04T06:39:08.215Z INFO [OpensearchProcessImpl] [2025-04-04T06:39:08,214][WARN ][o.o.c.c.ClusterFormationFailureHelper] [ansibleopensearch02] cluster-manager not discovered or elected yet, an election requires a node with id [9tukJ6SzSb6O5naDEQ0P6Q], have discovered [{ansibleopensearch02}{XtQ51mPFSvuHYAUOpoH_PQ}{2-WQWNYQTFaQGNq4Mj3y-Q}{ansibleopensearch02}{192.168.15.122:9300}{dimr}{shard_indexing_pressure_enabled=true}] which is not a quorum; discovery will continue using [127.0.0.1:9300, [::1]:9300, 127.0.1.1:9300] from hosts providers and [{ansibleopensearch02}{XtQ51mPFSvuHYAUOpoH_PQ}{2-WQWNYQTFaQGNq4Mj3y-Q}{ansibleopensearch02}{192.168.15.122:9300}{dimr}{shard_indexing_pressure_enabled=true}] from last-known cluster state; node term 9, last-accepted version 403 in term 9
`
## Steps to Reproduce (for bugs)
The behaviour occurs when the previous cluster-manager isn't available/ready during start-up:
- restart the service on the secondary datanode
- wait a few seconds
- restart the service on the current cluster-manager
## Context
This usually happens during automated updates where the servers reboot and/or restart their services at a similar time.
It can be fixed with another restart of the service on the secondary datanode.
## Your Environment
Happens in multiple environments (first entry and logs are from the test setup, second from production)
* Graylog Version: Graylog 6.1.10+a308be3, Graylog 6.1.5+e3ae3ce
* Java Version:
openjdk 17.0.14 2025-01-21
OpenJDK Runtime Environment (build 17.0.14+7-Ubuntu-120.04)
OpenJDK 64-Bit Server VM (build 17.0.14+7-Ubuntu-120.04, mixed mode, sharing)
openjdk 17.0.14 2025-01-21
OpenJDK Runtime Environment (build 17.0.14+7-Ubuntu-122.04.1)
OpenJDK 64-Bit Server VM (build 17.0.14+7-Ubuntu-122.04.1, mixed mode, sharing)
* OpenSearch Version: graylog-datanode 6.1.10-1, graylog-datanode 6.1.5-2
* MongoDB Version: 6.0.19, 7.0.16
* Operating System: 20.04.6 LTS (Focal Fossa), 22.04.5 LTS (Jammy Jellyfish)
* Browser version: n/a
General setup for both environments: Graylog Server on two servers graylog01 and graylog02, Graylog Datanode on two servers opensearch01 and opensearch02, MongoDB on graylog01, graylog02, and opensearch01.
Please let me know if you need any further information.
Thanks,
Anita
Contributor guide
Assessment
This issue has not been assessed yet.