Multiple DCs: Add new status information for monitoring multiple DC deployments
- Dominant language
- C++
- Stars
- 16.7k
- Forks
- 1.6k
- Avg merge
- 1d 20h
- Merged PRs (30d)
- 126
Description
The most basic information to monitor is the version lag between the primary DC and the remote DC.
Data distribution monitoring may also want to be separated by DC, although this may be tricky to implement.
Some basic machine fault tolerance calculations should be updated, and a new DC fault tolerance should be added.
Network latencies between DCs would be nice to have.
Since satellites may not be in active use in some configurations, their failure monitoring may need to be done differently.
Contributor guide
Research direction
No files, tests, or entry points are identified. First map the existing monitoring and status paths for multi-DC deployments, then clarify the desired version-lag, per-DC distribution, fault-tolerance, latency, and satellite-failure information before implementation can be scoped.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- databases, distributed-systems, observability-sre
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100