hypothesis / hypothesis/product-backlog
Improve our alerting setup
Open
Epic
Operations
- Dominant language
- No language data
- Stars
- 122
- Forks
- 7
- PR merge metrics
- No merged PRs in 30d
Description
Identify and make any useful improvements to our alerting set up. For example are we getting alerted when our error rate or response time rises? What about when our application server CPUs are pegged or their disks are full etc?
In the past we've only been alerted when requests to healthcheck endpoints start failing.
See also:
* [STORY: Alert on persistent high load](https://github.com/hypothesis/product-backlog/issues/257)
* [STORY: Alert on root partition space/inode exhaustion](https://github.com/hypothesis/product-backlog/issues/254)
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.