Recommendations for alerts in Production
- Dominant language
- Go
- Stars
- 4k
- Forks
- 818
- Avg merge
- 1h 14m
- Merged PRs (30d)
- 1
Description
So, this is great, we've just started deploying it to our development clusters and find that it is really useful. We currently use New Relic -> PagerDuty and I've create a simple Burrow agent to publish consumer group info to New Relic Insights.
I am curious if there are some best practices for establishing alerts based on status/lag.
Specifically I'm interested on what kinds of rules people create for successfully warning about potential issues and alerting for real issues that need attention.
Contributor guide
No contributing guide indexed for this repository
Research direction
The issue names no file, test, or entry point. Start by checking Burrow's existing documentation for consumer status and lag, along with the described New Relic-to-PagerDuty workflow; done would require agreed, documented warning and critical alert guidance.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- go, kafka
- Domain
- documentation, observability, stream-processing
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100