cockroachdb / cockroachdb/cockroach

kvserver: timeouts can starve snapshots

Open
#103,879 6 comments 0 reactions 0 assignees View on GitHub
C-bug P-2 T-kv
Dominant language
Go
Stars
32.5k
Forks
4.1k
PR merge metrics
PR metrics pending

Description

**Describe the problem**

See internal issue https://github.com/cockroachlabs/support/issues/2327#issuecomment-1562510835, observed on 22.2.9 but code hasn't changed materially on master.

The gist of it is that we saw an instance of a CC cluster where raft snaps for certain ranges never got through within the timeout. We compute the timeout based on the logical stats of the replica, however we do this even if the replica contains estimates, where we have no real error bounds on how far off the true size we are. We assume that we are able to send at ~10% of the configured max snapshot rate, which is certainly mostly true, but appears to not have been true here (for reasons still unknown).

We ended up allotting timeouts in the 1-2m range to snapshots that clearly needed more time.

We already have a maximum snapshot duration of 1h:

https://github.com/cockroachdb/cockroach/blob/80535e10ccde8effdb3787d474e5e0c962f428f4/pkg/kv/kvserver/replica_command.go#L62-L70

but it might be useful to also introduce a minimum that is by default higher than the guaranteed processing time of 1m:

https://github.com/cockroachdb/cockroach/blob/5c1909bdefc0b4b1238f44334ec66b6d6fc602c8/pkg/kv/kvserver/queue.go#L57-L63

Last but not least, context cancellation is perhaps too brute a mechanism, a keepalive type of approach where a snapshot is allowed to proceed just as long as it makes progress would avoid this class of problems altogether.

Jira issue: CRDB-28236

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.