Kuadrant / Kuadrant/dns-operator
Controller manager restarts more than N in Y duration
- Dominant language
- Go
- Stars
- 12
- Forks
- 23
- Avg merge
- 1d 4h
- Merged PRs (30d)
- 14
Description
come up with an alert query that will allow a user to specify that if the controller manager restarts more than N in Y duration an alert will fire.
This will need to be tested that it does fire when expected.
Also some details on potential causes, and how to investigate and troubleshoot.
recreate by killing the operator pod more than N times in Y duration.
Contributor guide
Research direction
Start by locating the controller-manager restart metric and existing alert-query tests in the repository; the issue names no files. Reproduce the condition by killing the operator pod more than N times in Y duration, then verify the configurable alert fires as expected and document potential causes and troubleshooting steps.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- go, kubernetes
- Domain
- devops, observability
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100