canonical / canonical/microceph
Mon out of Quorum Recovery
- Dominant language
- Go
- Stars
- 396
- Forks
- 74
- Avg merge
- 2d 20h
- Merged PRs (30d)
- 7
Description
I am having a little bit trouble with a mon being out of quorum. I have a cluster of 3, 2 are still in quorum and i'm unsure how to recover.
I had a look at the ceph documentation and I was going to try removing and adding a mon, but it first requires this...
```bash
ceph-mon -i `hostname` --extract-monmap /tmp/monmap
```
I'm not sure what the equivalent would be for microceph? I tried ceph-mon without the dash in the middle, but the command appears to hang (following ceph documentation, you are required to systemctl stop ceph-mon.target on each node. I tried...
```bash
sudo systemctl stop snap.microceph.mon.service
```
But i'm not sure if this is equivalent to ceph documentation.
health status....
```bash
cluster:
id: a832124c-ee14-4a2d-b6d9-2a1b8326a560
health: HEALTH_WARN
1/3 mons down, quorum uby2,uby1
Degraded data redundancy: 8994/26982 objects degraded (33.333%), 16 pgs degraded, 16 pgs undersized
12 pgs not deep-scrubbed in time
16 pgs not scrubbed in time
services:
mon: 3 daemons, quorum uby2,uby1 (age 15s), out of quorum: uby3
mgr: uby1(active, since 25h), standbys: uby2
osd: 3 osds: 2 up (since 25h), 2 in (since 10d)
data:
pools: 1 pools, 16 pgs
objects: 8.99k objects, 27 GiB
usage: 46 GiB used, 431 GiB / 477 GiB avail
pgs: 8994/26982 objects degraded (33.333%)
16 active+undersized+degraded
io:
client: 0 B/s rd, 6.0 KiB/s wr, 0 op/s rd, 0 op/s wr
```
I'm not sure if this is an issue, or if my incompetence, but couldn't find documentation for this (i might be looking in the wrong place...)
It's been like this for 24hrs at least, but I've not been checking this cluster for a couple of weeks. When I check the logs, nothing stood out to me, except the last entry being 27th September, and i couldn't find bus errors, or corruption errors.
The last 20 entries...
```bash
2023-09-27T23:14:42.441+0100 7f24051ec640 4 rocksdb: [db/compaction/compaction_job.cc:1594] [default] [JOB 12248] Compacted 1@0 + 1@6 files to L6 => 28530906 bytes
2023-09-27T23:14:42.445+0100 7f24051ec640 4 rocksdb: (Original Log Time 2023/09/27-23:14:42.448932) [db/compaction/compaction_job.cc:812] [default] compacted to: base level 6 level multiplier 10.00 max bytes base 268435456 files[0 0 0 0 0 0 1] max score 0.00, MB/sec: 109.9 rd, 108.3 wr, level 6, files in(1, 1) out(1) MB in(0.7, 26.9) out(27.2), read-write-amplify(80.4) write-amplify(39.9) OK, records in: 24363, records dropped: 724 output_compression: NoCompression
2023-09-27T23:14:42.445+0100 7f24051ec640 4 rocksdb: (Original Log Time 2023/09/27-23:14:42.448999) EVENT_LOG_v1 {"time_micros": 1695852882448972, "job": 12248, "event": "compaction_finished", "compaction_time_micros": 263337, "compaction_time_cpu_micros": 147058, "output_level": 6, "num_output_files": 1, "total_output_size": 28530906, "num_input_records": 24363, "num_output_records": 23639, "num_subcompactions": 1, "output_compression": "NoCompression", "num_single_delete_mismatches": 0, "num_single_delete_fallthrough": 0, "lsm_state": [0, 0, 0, 0, 0, 0, 1]}
2023-09-27T23:14:42.445+0100 7f24051ec640 4 rocksdb: [file/delete_scheduler.cc:69] Deleted file /var/lib/ceph/mon/ceph-uby3/store.db/605896.sst immediately, rate_bytes_per_sec 0, total_trash_size 0 max_trash_db_ratio 0.250000
2023-09-27T23:14:42.445+0100 7f24051ec640 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1695852882449791, "job": 12248, "event": "table_file_deletion", "file_number": 605896}
2023-09-27T23:14:42.461+0100 7f24051ec640 4 rocksdb: [file/delete_scheduler.cc:69] Deleted file /var/lib/ceph/mon/ceph-uby3/store.db/605894.sst immediately, rate_bytes_per_sec 0, total_trash_size 0 max_trash_db_ratio 0.250000
2023-09-27T23:14:42.461+0100 7f24051ec640 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1695852882465134, "job": 12248, "event": "table_file_deletion", "file_number": 605894}
2023-09-27T23:14:42.461+0100 7f23fb9d9640 4 rocksdb: [db/db_impl/db_impl_compaction_flush.cc:1615] [default] Manual compaction starting
2023-09-27T23:14:42.461+0100 7f23fb9d9640 4 rocksdb: [db/db_impl/db_impl_compaction_flush.cc:1615] [default] Manual compaction starting
2023-09-27T23:14:42.461+0100 7f23fb9d9640 4 rocksdb: [db/db_impl/db_impl_compaction_flush.cc:1615] [default] Manual compaction starting
2023-09-27T23:14:42.461+0100 7f23fb9d9640 4 rocksdb: [db/db_impl/db_impl_compaction_flush.cc:1615] [default] Manual compaction starting
2023-09-27T23:14:42.461+0100 7f23fb9d9640 4 rocksdb: [db/db_impl/db_impl_compaction_flush.cc:1615] [default] Manual compaction starting
2023-09-27T23:15:33.170+0100 7f24059ed640 -1 received signal: Terminated from /sbin/init (PID: 1) UID: 0
2023-09-27T23:15:33.182+0100 7f24059ed640 -1 mon.uby3@2(peon) e4 *** Got Signal Terminated ***
2023-09-27T23:15:33.182+0100 7f24059ed640 1 mon.uby3@2(peon) e4 shutdown
2023-09-27T23:15:33.410+0100 7f2406c31d40 1 rocksdb: close waiting for compaction thread to stop
2023-09-27T23:15:33.418+0100 7f2406c31d40 1 rocksdb: close compaction thread to stopped
2023-09-27T23:15:33.434+0100 7f2406c31d40 4 rocksdb: [db/db_impl/db_impl.cc:446] Shutdown: canceling all background work
2023-09-27T23:15:33.486+0100 7f2406c31d40 4 rocksdb: [db/db_impl/db_impl.cc:625] Shutdown complete
```
i've grep'd the last lines to see if they are unusual entries and they don't seem to be, so i'm guessing i might have rebooted the server on the 27th before going to bed, or something like this and it fell over.
Systemctl status looks like this...
```bash
× snap.microceph.daemon.service - Service for snap application microceph.daemon
Loaded: loaded (/etc/systemd/system/snap.microceph.daemon.service; enabled; vendor preset: enabled)
Active: failed (Result: exit-code) since Sun 2023-10-08 19:29:40 BST; 2h 7min ago
Process: 3873 ExecStart=/usr/bin/snap run microceph.daemon (code=exited, status=1/FAILURE)
Main PID: 3873 (code=exited, status=1/FAILURE)
CPU: 87ms
Oct 08 19:29:40 uby3 systemd[1]: snap.microceph.daemon.service: Scheduled restart job, restart counter is at 6.
Oct 08 19:29:40 uby3 systemd[1]: Stopped Service for snap application microceph.daemon.
Oct 08 19:29:40 uby3 systemd[1]: snap.microceph.daemon.service: Start request repeated too quickly.
Oct 08 19:29:40 uby3 systemd[1]: snap.microceph.daemon.service: Failed with result 'exit-code'.
Oct 08 19:29:40 uby3 systemd[1]: Failed to start Service for snap application microceph.daemon.
```
This is the mon_status... I'm not sure if i should be concerned, but i am unable to run this command on the other ~~2 nodes~~ node without a socket warning, no file present...
```bash
admin_socket: exception getting command descriptions: [Errno 2] No such file or directory
```
mon_status....
```bash
{
"name": "uby1",
"rank": 1,
"state": "peon",
"election_epoch": 690,
"quorum": [
0,
1
],
"quorum_age": 787,
"features": {
"required_con": "2449958755906961412",
"required_mon": [
"kraken",
"luminous",
"mimic",
"osdmap-prune",
"nautilus",
"octopus",
"pacific",
"elector-pinging",
"quincy"
],
"quorum_con": "4540138320759226367",
"quorum_mon": [
"kraken",
"luminous",
"mimic",
"osdmap-prune",
"nautilus",
"octopus",
"pacific",
"elector-pinging",
"quincy"
]
},
"outside_quorum": [],
"extra_probe_peers": [],
"sync_provider": [],
"monmap": {
"epoch": 4,
"fsid": "a832124c-ee14-4a2d-b6d9-2a1b8326a560",
"modified": "2023-01-06T17:08:23.097085Z",
"created": "2023-01-06T17:06:46.508041Z",
"min_mon_release": 17,
"min_mon_release_name": "quincy",
"election_strategy": 1,
"disallowed_leaders: ": "",
"stretch_mode": false,
"tiebreaker_mon": "",
"removed_ranks: ": "",
"features": {
"persistent": [
"kraken",
"luminous",
"mimic",
"osdmap-prune",
"nautilus",
"octopus",
"pacific",
"elector-pinging",
"quincy"
],
"optional": []
},
"mons": [
{
"rank": 0,
"name": "uby2",
"public_addrs": {
"addrvec": [
{
"type": "v2",
"addr": "10.10.101.42:3300",
"nonce": 0
},
{
"type": "v1",
"addr": "10.10.101.42:6789",
"nonce": 0
}
]
},
"addr": "10.10.101.42:6789/0",
"public_addr": "10.10.101.42:6789/0",
"priority": 0,
"weight": 0,
"crush_location": "{}"
},
{
"rank": 1,
"name": "uby1",
"public_addrs": {
"addrvec": [
{
"type": "v2",
"addr": "10.10.101.41:3300",
"nonce": 0
},
{
"type": "v1",
"addr": "10.10.101.41:6789",
"nonce": 0
}
]
},
"addr": "10.10.101.41:6789/0",
"public_addr": "10.10.101.41:6789/0",
"priority": 0,
"weight": 0,
"crush_location": "{}"
},
{
"rank": 2,
"name": "uby3",
"public_addrs": {
"addrvec": [
{
"type": "v2",
"addr": "10.10.101.43:3300",
"nonce": 0
},
{
"type": "v1",
"addr": "10.10.101.43:6789",
"nonce": 0
}
]
},
"addr": "10.10.101.43:6789/0",
"public_addr": "10.10.101.43:6789/0",
"priority": 0,
"weight": 0,
"crush_location": "{}"
}
]
},
"feature_map": {
"mon": [
{
"features": "0x3f01cfbf7ffdffff",
"release": "luminous",
"num": 1
}
],
"osd": [
{
"features": "0x3f01cfbf7ffdffff",
"release": "luminous",
"num": 1
}
],
"client": [
{
"features": "0x3f01cfbf7ffdffff",
"release": "luminous",
"num": 3
}
],
"mgr": [
{
"features": "0x3f01cfbf7ffdffff",
"release": "luminous",
"num": 1
}
]
},
"stretch_mode": false
}
```
Update: just realised my second node is responding to mon_status, apologies...
```bash
{
"name": "uby2",
"rank": 0,
"state": "leader",
"election_epoch": 690,
"quorum": [
0,
1
],
"quorum_age": 1439,
"features": {
"required_con": "2449958755906961412",
"required_mon": [
"kraken",
"luminous",
"mimic",
"osdmap-prune",
"nautilus",
"octopus",
"pacific",
"elector-pinging",
"quincy"
],
"quorum_con": "4540138320759226367",
"quorum_mon": [
"kraken",
"luminous",
"mimic",
"osdmap-prune",
"nautilus",
"octopus",
"pacific",
"elector-pinging",
"quincy"
]
},
"outside_quorum": [],
"extra_probe_peers": [],
"sync_provider": [],
"monmap": {
"epoch": 4,
"fsid": "a832124c-ee14-4a2d-b6d9-2a1b8326a560",
"modified": "2023-01-06T17:08:23.097085Z",
"created": "2023-01-06T17:06:46.508041Z",
"min_mon_release": 17,
"min_mon_release_name": "quincy",
"election_strategy": 1,
"disallowed_leaders: ": "",
"stretch_mode": false,
"tiebreaker_mon": "",
"removed_ranks: ": "",
"features": {
"persistent": [
"kraken",
"luminous",
"mimic",
"osdmap-prune",
"nautilus",
"octopus",
"pacific",
"elector-pinging",
"quincy"
],
"optional": []
},
"mons": [
{
"rank": 0,
"name": "uby2",
"public_addrs": {
"addrvec": [
{
"type": "v2",
"addr": "10.10.101.42:3300",
"nonce": 0
},
{
"type": "v1",
"addr": "10.10.101.42:6789",
"nonce": 0
}
]
},
"addr": "10.10.101.42:6789/0",
"public_addr": "10.10.101.42:6789/0",
"priority": 0,
"weight": 0,
"crush_location": "{}"
},
{
"rank": 1,
"name": "uby1",
"public_addrs": {
"addrvec": [
{
"type": "v2",
"addr": "10.10.101.41:3300",
"nonce": 0
},
{
"type": "v1",
"addr": "10.10.101.41:6789",
"nonce": 0
}
]
},
"addr": "10.10.101.41:6789/0",
"public_addr": "10.10.101.41:6789/0",
"priority": 0,
"weight": 0,
"crush_location": "{}"
},
{
"rank": 2,
"name": "uby3",
"public_addrs": {
"addrvec": [
{
"type": "v2",
"addr": "10.10.101.43:3300",
"nonce": 0
},
{
"type": "v1",
"addr": "10.10.101.43:6789",
"nonce": 0
}
]
},
"addr": "10.10.101.43:6789/0",
"public_addr": "10.10.101.43:6789/0",
"priority": 0,
"weight": 0,
"crush_location": "{}"
}
]
},
"feature_map": {
"mon": [
{
"features": "0x3f01cfbf7ffdffff",
"release": "luminous",
"num": 1
}
],
"mds": [
{
"features": "0x3f01cfbf7ffdffff",
"release": "luminous",
"num": 2
}
],
"osd": [
{
"features": "0x3f01cfbf7ffdffff",
"release": "luminous",
"num": 1
}
],
"client": [
{
"features": "0x2f018fb87aa4aafe",
"release": "luminous",
"num": 2
},
{
"features": "0x3f01cfbf7ffdffff",
"release": "luminous",
"num": 1
}
],
"mgr": [
{
"features": "0x3f01cfbf7ffdffff",
"release": "luminous",
"num": 1
}
]
},
"stretch_mode": false
}
```
Contributor guide
Assessment
This issue has not been assessed yet.