prometheus / prometheus/alertmanager
Feature request - Matcher continue on receiver failure
Nobody has claimed this yet.
- Dominant language
- Go
- Stars
- 8.6k
- Forks
- 2.5k
- Avg merge
- 2d 6h
- Merged PRs (30d)
- 61
Description
Setup
We have 2 Slack receivers: r1 and r2.
We have routes with matchers like so:
- matchers:
receiver: r2
routes:
- matchers:
- namespace!=""
receiver: r1
- matchers:
receiver: r2
(sorry for formatting, can't get it right, you'll get the idea no?)
So what this does is:
- Have a main fallback to
r1in case nothing matches. However in this example, it should never hit that due to the last matcher always matching - Have a matcher if
namespace is not emptywhich routes tor1 - Have a second matcher which matches everything else (because
continue=falseby default) and goes tor2
Use-case
Now imagine having a more complicated setup with dynamic settings. For example dynamic Slack channels as such:
channel: some-channel-{{ (index .Alerts 0).Labels.namespace }}
The use-case here is to dynamically route stuff based on alert labels.
It's easy to check if a namespace is present, then route to a receiver with such channel variable. However the problem is that there is no way to know if the channel exists and if the alert gets send.
This would result in an error like:
level=error ts=2022-06-16T18:51:31.451Z caller=dispatch.go:354 component=dispatcher msg="Notify for alerts failed" num_alerts=11 err="slack[0]: notify retry canceled due to unrecoverable error after 1 attempts: channel "****redacted****": unexpected status code 404: channel_not_found"
Feature request
- matchers:
receiver: r2
routes:
- matchers:
- namespace!=""
receiver: r1
continue_on_receiver_failure: true
- matchers:
receiver: r2
Having an option like continue_on_receiver_failure which would be the same as continue in its behaviour but only triggers when it hits an error while sending and then continues with other matchers.
The current log would not be a level=error anymore but a level=info and only goes into level=error if when no other matcher is going to catch/send it. For example keep the state and after the full evaluation of the routes, it should know if it did send it somewhere eventually.
Other info
I also think this is a very valuable thing to have in general. Let's say, some receiver endpoint like Slack is down. Then we can automatically fall back to an other receiver endpoint. Without having the need to constantly send it to both endpoints and introducing noise.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by tracing the receiver failure reported at dispatch.go:354 and the route evaluation behavior described in the issue. Review how matcher continuation and notification failures are currently handled, then determine how fallback delivery and logging should behave; done means the new option is documented, tested, and preserves existing behavior when unset.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- go
- Domain
- backend, observability
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100