dapr / dapr/test-infra

[Chaos] Service health monitor

Open
#7 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
C#
Stars
17
Forks
26
PR merge metrics
No merged PRs in 30d

Description

Complete outages can be detected with other alarms. To detect partial failures, no service can have less than 3 healthy PODs for more than 50 minutes. This metric can be emitted by Failure Daemon.

Parents: dapr/dapr#1100 dapr/dapr#1155

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by locating the Failure Daemon in the test-infra repository and read parent issues dapr/dapr#1100 and dapr/dapr#1155 for the intended health metric. Define where the metric should be emitted and how the 3-healthy-POD threshold over 50 minutes will be detected; done means partial outages are reported without relying on complete-outage alarms.

Written by the indexing model from the issue text.

Assessment

Tech stack
kubernetes
Domain
observability
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.