cloudflare / cloudflare/cloudflared

💡 Feature Request: Add randomized delay to auto-update timer to prevent simultaneous fleet-wide HA tunnel restarts

Open
#1,645 0 comments 3 reactions 0 assignees View on GitHub
Priority: Normal Type: Feature Request
Dominant language
Go
Stars
15.6k
Forks
1.4k
PR merge metrics
No merged PRs in 30d

Description

**Describe the feature you'd like**
Add a randomized delay (e.g., `RandomizedDelaySec=12h`) to the default systemd `cloudflared-update.timer` file generated by the `cloudflared` service installer.

**What problem does it solve for you?**
Currently, the `cloudflared-update.timer` is configured with `OnCalendar=daily`, which triggers exactly at `00:00:00` local time (often `00:00:00 UTC` on standard cloud servers). For fleets running multiple `cloudflared` instances for High Availability (HA) failover, this causes every active replica to pull the update and restart at the exact same second.

This simultaneous fleet-wide restart defeats HA redundancy. All active tunnels drop their connections simultaneously to apply the binary update, resulting in a brief but total outage for routed traffic across the entire fleet.

**Describe alternatives you've considered**
The current workaround is using configuration management (Ansible, bash scripts, etc.) to manually inject a systemd drop-in override for `RandomizedDelaySec`, or writing custom scripts to assign a deterministically hashed offset per-node.

My experience with this is that it introduces unnecessary operational friction and configuration drift. Since HA tunnels are a core Cloudflare use-case, the default auto-update behavior should be safe for multi-node failover deployments out-of-the-box.

**Additional context**
Here is an example modification for the default generated timer template that would resolve the issue:

```ini
[Unit]
Description=Update cloudflared

[Timer]
OnCalendar=daily
RandomizedDelaySec=12h
Persistent=true

[Install]
WantedBy=timers.target

```

This issue strictly applies to multi-node setups handling mission-critical, continuous traffic (like DoH or heavy ingress), where even a few seconds of simultaneous downtime across all replicas is noticeable to clients.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.