cockroachdb / cockroachdb/cockroach

CockroachDB v24.3 cluster crashed

Open
#136,932 32 comments 0 reactions 1 assignee Claimed by @pav-kv View on GitHub
A-kv branch-release-24.3 C-bug O-community T-kv X-blathers-triaged
Dominant language
Go
Stars
32.5k
Forks
4.1k
PR merge metrics
PR metrics pending

Description

Hello.

My CockroachDB v24.3 has crashed with the following:

```
tk35c@mdzhlcrdb03:~> sudo tail -n1000 /var/log/cockroach/cockroach.log
I241206 22:22:14.983428 1 util/log/file_sync_buffer.go:237 ⋮ [config] file created at: 2024/12/06 22:22:14
I241206 22:22:14.983440 1 util/log/file_sync_buffer.go:237 ⋮ [config] running on machine: ‹mdzhlcrdb03›
I241206 22:22:14.983447 1 util/log/file_sync_buffer.go:237 ⋮ [config] binary: CockroachDB CCL v24.3.0 (x86_64-pc-linux-gnu, built 2024/11/21 17:04:13, go1.22.8 X:nocoverageredesign)
I241206 22:22:14.983455 1 util/log/file_sync_buffer.go:237 ⋮ [config] arguments: [‹/usr/local/bin/cockroach› ‹start› ‹--insecure› ‹--locality=region=ch,zone=zh› ‹--advertise-addr=mdzhlcrdb03› ‹--join=mdzhlcrdb01,mdzhlcrdb02,mdlplcrdb01› ‹--log-dir=/var/log/cockroach› ‹--cache=.35› ‹--max-sql-memory=.35›]
I241206 22:22:14.983472 1 util/log/file_sync_buffer.go:237 ⋮ [config] log format (utf8=✓): crdb-v2
I241206 22:22:14.983476 1 util/log/file_sync_buffer.go:237 ⋮ [config] line format: [IWEF]yymmdd hh:mm:ss.uuuuuu goid [chan@]file:line redactionmark \[tags\] [counter] msg
W241206 22:22:14.983097 1 1@cli/start.go:1439 ⋮ [n?] 1 ALL SECURITY CONTROLS HAVE BEEN DISABLED!
W241206 22:22:14.983097 1 1@cli/start.go:1439 ⋮ [n?] 1 +
W241206 22:22:14.983097 1 1@cli/start.go:1439 ⋮ [n?] 1 +This mode is intended for non-production testing only.
W241206 22:22:14.983097 1 1@cli/start.go:1439 ⋮ [n?] 1 +
W241206 22:22:14.983097 1 1@cli/start.go:1439 ⋮ [n?] 1 +In this mode:
W241206 22:22:14.983097 1 1@cli/start.go:1439 ⋮ [n?] 1 +- Your cluster is open to any client that can access ‹any of your IP addresses›.
W241206 22:22:14.983097 1 1@cli/start.go:1439 ⋮ [n?] 1 +- Intruders with access to your machine or network can observe client-server traffic.
W241206 22:22:14.983097 1 1@cli/start.go:1439 ⋮ [n?] 1 +- Intruders can log in without password and read or write any data in the cluster.
W241206 22:22:14.983097 1 1@cli/start.go:1439 ⋮ [n?] 1 +- Intruders can consume all your server's resources and cause unavailability.
I241206 22:22:14.983663 1 1@cli/start.go:1449 ⋮ [n?] 2 To start a secure server without mandating TLS for clients,
I241206 22:22:14.983663 1 1@cli/start.go:1449 ⋮ [n?] 2 +consider --accept-sql-without-tls instead. For other options, see:
I241206 22:22:14.983663 1 1@cli/start.go:1449 ⋮ [n?] 2 +
I241206 22:22:14.983663 1 1@cli/start.go:1449 ⋮ [n?] 2 +- https://go.crdb.dev/issue-v/53404/v24.3
I241206 22:22:14.983663 1 1@cli/start.go:1449 ⋮ [n?] 2 +- https://www.cockroachlabs.com/docs/v24.3/secure-a-cluster.html
I241206 22:22:14.984266 1 server/status/recorder.go:832 ⋮ [n?] 3 ‹available memory from cgroups (8.0 EiB) is unsupported, using system memory 62 GiB instead: ›
I241206 22:22:14.984303 1 1@cli/start.go:1478 ⋮ [n?] 4 ‹CockroachDB CCL v24.3.0 (x86_64-pc-linux-gnu, built 2024/11/21 17:04:13, go1.22.8 X:nocoverageredesign)›
I241206 22:22:14.985261 1 server/status/recorder.go:832 ⋮ [n?] 5 ‹available memory from cgroups (8.0 EiB) is unsupported, using system memory 62 GiB instead: ›
W241206 22:22:14.985304 1 1@cli/start.go:739 ⋮ [n?] 6 recommended default value of --max-go-memory (48 GiB) was truncated to 32 GiB, consider reducing --max-sql-memory (21 GiB) and / or --cache (21 GiB); total system/cgroup memory: 62 GiB.
I241206 22:22:14.985325 1 1@cli/start.go:556 ⋮ [n?] 7 soft memory limit of Go runtime is set to 32 GiB
I241206 22:22:14.985348 1 1@cli/start.go:593 ⋮ [n?] 8 GC target percentage of Go runtime is set to 300%
I241206 22:22:14.986962 1 server/status/recorder.go:832 ⋮ [T1,Vsystem,n?] 9 ‹available memory from cgroups (8.0 EiB) is unsupported, using system memory 62 GiB instead: ›
I241206 22:22:14.986984 1 server/config.go:669 ⋮ [T1,Vsystem,n?] 10 system total memory: 62 GiB
I241206 22:22:14.987009 1 server/config.go:671 ⋮ [T1,Vsystem,n?] 11 server configuration:
I241206 22:22:14.987009 1 server/config.go:671 ⋮ [T1,Vsystem,n?] 11 +‹max offset 500000000›
I241206 22:22:14.987009 1 server/config.go:671 ⋮ [T1,Vsystem,n?] 11 +‹cache size 21 GiB›
I241206 22:22:14.987009 1 server/config.go:671 ⋮ [T1,Vsystem,n?] 11 +‹SQL memory pool size 21 GiB›
I241206 22:22:14.987009 1 server/config.go:671 ⋮ [T1,Vsystem,n?] 11 +‹scan interval 10m0s›
I241206 22:22:14.987009 1 server/config.go:671 ⋮ [T1,Vsystem,n?] 11 +‹scan min idle time 10ms›
I241206 22:22:14.987009 1 server/config.go:671 ⋮ [T1,Vsystem,n?] 11 +‹scan max idle time 1s›
I241206 22:22:14.987009 1 server/config.go:671 ⋮ [T1,Vsystem,n?] 11 +‹event log enabled true›
I241206 22:22:14.987092 1 1@cli/start.go:1339 ⋮ [T1,Vsystem,n?] 12 using local environment variables:
I241206 22:22:14.987092 1 1@cli/start.go:1339 ⋮ [T1,Vsystem,n?] 12 +HTTPS_PROXY=‹inetproxy-nonprod.sn.six-group.net:8081›
I241206 22:22:14.987092 1 1@cli/start.go:1339 ⋮ [T1,Vsystem,n?] 12 +LANG=‹en_US.UTF-8›
I241206 22:22:14.987092 1 1@cli/start.go:1339 ⋮ [T1,Vsystem,n?] 12 +NO_PROXY=‹*.host.net›
I241206 22:22:14.987117 1 1@cli/start.go:1346 ⋮ [T1,Vsystem,n?] 13 process identity: ‹uid 18104 euid 18104 gid 18104 egid 18104›
I241206 22:22:14.994783 1 1@cli/start.go:1501 ⋮ [T1,Vsystem,n?] 14 GEOS loaded from directory ‹/usr/local/lib/cockroach›
I241206 22:22:14.994815 1 1@cli/start.go:785 ⋮ [T1,Vsystem,n?] 15 starting cockroach node
I241206 22:22:15.017026 96 server/config.go:881 ⋮ [T1,Vsystem,n?] 16 1 storage engine initialized
I241206 22:22:15.017056 96 server/config.go:884 ⋮ [T1,Vsystem,n?] 17 Pebble cache size: 21 GiB
I241206 22:22:15.017077 96 server/config.go:884 ⋮ [T1,Vsystem,n?] 18 store 0: max size 0 B, max open file limit 519288
I241206 22:22:15.017084 96 server/config.go:884 ⋮ [T1,Vsystem,n?] 19 store 0: /cockroach/cockroach-data: rw encrypted=false fs:{bdev=tmpfs fstype=tmpfs mountpoint=/run/user/1001 mountopts=rw,seclabel,nosuid,nodev,relatime,size=6541748k,nr_inodes=1635437,mode=700,uid=1001,gid=1001,inode64}
I241206 22:22:15.021281 162 util/cidr/cidr.go:299 ⋮ [T1,Vsystem,n?] 20 CIDR lookup updated with 0 destinations
I241206 22:22:15.093254 96 1@server/clock_monotonicity.go:61 ⋮ [T1,Vsystem,n?] 21 monitoring forward clock jumps based on server.clock.forward_jump_check.enabled
I241206 22:22:15.096903 96 1@server/clock_monotonicity.go:136 ⋮ [T1,Vsystem,n8] 22 Sleeping till wall time 1733523735096897549 to catches up to 1733523735593223350 to ensure monotonicity. Sub: 496.325801ms
I241206 22:22:15.096916 268 1@server/server.go:1773 ⋮ [T1,Vsystem,n8] 23 connecting to gossip network to verify cluster ID ‹"673392ac-8928-448a-b71e-c9272400a044"›
W241206 22:22:15.593320 96 1@cli/start.go:1300 ⋮ [T1,Vsystem,n8] 24 Running a server without --sql-addr, with a combined RPC/SQL listener, is deprecated.
W241206 22:22:15.593320 96 1@cli/start.go:1300 ⋮ [T1,Vsystem,n8] 24 +This feature will be removed in a later version of CockroachDB.
W241206 22:22:15.601784 96 2@gossip/gossip.go:1462 ⋮ [T1,Vsystem,n8] 25 no incoming or outgoing connections
I241206 22:22:15.602217 447 gossip/client.go:112 ⋮ [T1,Vsystem,n8] 26 unvalidated bootstrap gossip dial to ‹mdzhlcrdb01:26257›
I241206 22:22:15.606611 535 kv/kvserver/kvstorage/init.go:250 ⋮ [T1,Vsystem,n8,s8] 27 beginning range descriptor iteration
I241206 22:22:15.606845 447 gossip/client.go:136 ⋮ [T1,Vsystem,n8] 28 started gossip client to n0 (‹mdzhlcrdb01:26257›)
I241206 22:22:15.607516 268 1@server/server.go:1776 ⋮ [T1,Vsystem,n8] 29 node connected via gossip
I241206 22:22:15.607855 535 kv/kvserver/kvstorage/init.go:374 ⋮ [T1,Vsystem,n8,s8] 30 range descriptor iteration done: 3 range descriptors, 1 intents, 0 tombstones; stats: ‹seeked 16 times (10 internal); stepped 0 times (0 internal); blocks: 27KB cached, 25KB not cached (read time: 27.585µs); points: 8 (233B keys, 312B values); separated: 31 (1.1KB, 39B fetched)›
I241206 22:22:15.608744 535 kv/kvserver/kvstorage/init.go:498 ⋮ [T1,Vsystem,n8,s8] 31 loaded replica ID for 82/82 replicas
I241206 22:22:15.609366 535 kv/kvserver/kvstorage/init.go:515 ⋮ [T1,Vsystem,n8,s8] 32 loaded Raft state for 82/82 replicas
I241206 22:22:15.609396 535 kv/kvserver/kvstorage/init.go:542 ⋮ [T1,Vsystem,n8,s8] 33 loaded 82 replicas
I241206 22:22:15.609411 535 kv/kvserver/kvstorage/init.go:573 ⋮ [T1,Vsystem,n8,s8] 34 verified 82/82 replicas
I241206 22:22:15.610763 535 kv/kvserver/store.go:2320 ⋮ [T1,Vsystem,n8,s8] 35 initialized 82/82 replicas
I241206 22:22:15.611312 535 server/node.go:704 ⋮ [T1,Vsystem,n8] 36 initialized store s8 in 9ms (3 replicas)
I241206 22:22:15.611412 96 kv/kvserver/stores.go:255 ⋮ [T1,Vsystem,n8] 37 read 5 node addresses from persistent storage
I241206 22:22:15.612803 96 server/node.go:819 ⋮ [T1,Vsystem,n8] 38 started with engine type pebble
I241206 22:22:15.612833 96 server/node.go:821 ⋮ [T1,Vsystem,n8] 39 started with attributes []
I241206 22:22:15.612923 96 server/goroutinedumper/goroutinedumper.go:117 ⋮ [T1,Vsystem,n8] 40 writing goroutine dumps to ‹/var/log/cockroach/goroutine_dump›
I241206 22:22:15.612965 96 server/profiler/heapprofiler.go:72 ⋮ [T1,Vsystem,n8] 41 writing go heap profiles to ‹/var/log/cockroach/heap_profiler› at least every 1h0m0s
I241206 22:22:15.612983 96 server/profiler/memory_monitoring_profiler.go:66 ⋮ [T1,Vsystem,n8] 42 writing memory monitoring dumps to ‹/var/log/cockroach/heap_profiler› at least every 1h0m0s
I241206 22:22:15.612995 96 server/profiler/cgoprofiler.go:66 ⋮ [T1,Vsystem,n8] 43 to enable jmalloc profiling: "export MALLOC_CONF=prof:true" or "ln -s prof:true /etc/malloc.conf"
I241206 22:22:15.613009 96 server/profiler/statsprofiler.go:68 ⋮ [T1,Vsystem,n8] 44 writing memory stats to ‹/var/log/cockroach/heap_profiler› at last every 1h0m0s
I241206 22:22:15.613391 96 server/profiler/activequeryprofiler.go:88 ⋮ [T1,Vsystem,n8] 45 writing go query profiles to ‹/var/log/cockroach/heap_profiler›
I241206 22:22:15.613416 96 server/profiler/cpuprofiler.go:87 ⋮ [T1,Vsystem,n8] 46 writing cpu profile dumps to ‹/var/log/cockroach/pprof_dump›
I241206 22:22:15.613579 96 1@server/server.go:1970 ⋮ [T1,Vsystem,n8] 47 starting http server at ‹[::]:8080› (use: ‹mdzhlcrdb03:8080›)
I241206 22:22:15.613607 96 1@server/server.go:1978 ⋮ [T1,Vsystem,n8] 48 starting grpc/postgres server at ‹[::]:26257›
I241206 22:22:15.613627 96 1@server/server.go:1979 ⋮ [T1,Vsystem,n8] 49 advertising CockroachDB node at ‹mdzhlcrdb03:26257›
E241206 22:22:15.631984 563 2@rpc/peer.go:663 ⋮ [T1,Vsystem,n8,rnode=4,raddr=‹mdzhlcrdb03:26257›,class=default,rpc] 50 failed connection attempt (last connected 0s ago): grpc: ‹client requested node ID 4 doesn't match server node ID 8› [code 2/Unknown]
I241206 22:22:15.636605 1 1@cli/start.go:1042 ⋮ [T1,Vsystem,n8] 51 initiating hard shutdown of server
I241206 22:22:15.636941 1 1@cli/start.go:1120 ⋮ [T1,Vsystem,n8] 52 too early to drain; used hard shutdown instead
E241206 22:22:15.637508 1 1@cli/clierror/check.go:30 ⋮ [-] 53 ‹ERROR›: server startup failed: cockroach server exited with error: error recording initial status summaries: remote wall time is too far ahead (521.632924ms) to be trustworthy
```

It started with the node `mdzhlcrdb03` not being able to start anymore and, as an idiot, I decided to stop the service completely and to drop the `/cockroach` folder with the hope that the node would then restart freshly and sync the data from the other nodes.

Unfortunately, as I found out later, that's not the case but - instead - other zombie nodes with different IDs get added as shown in
> E241206 22:22:15.631984 563 2@rpc/peer.go:663 ⋮ [T1,Vsystem,n8,rnode=4,raddr=‹mdzhlcrdb03:26257›,class=default,rpc] 50 failed connection attempt (last connected 0s ago): grpc: ‹client requested node ID 4 doesn't match server node ID 8› [code 2/Unknown]

What's the procedure to get the cluster operational again when
> E241206 22:22:15.637508 1 1@cli/clierror/check.go:30 ⋮ [-] 53 ‹ERROR›: server startup failed: cockroach server exited with error: error recording initial status summaries: remote wall time is too far ahead (521.632924ms) to be trustworthy

occurs and how to remove and recreate a node if the service doesn't start at all because of the above?
If the service on the node isn't running, the commands to either drain or remove a node won't work since they can't connect to it, for what I've experienced.

Moreover, the UI becomes quite unusable as soon as even a single node in the cluster dies.

Kind regards and thanks.

Jira issue: CRDB-45306

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.