hypothesis / hypothesis/product-backlog

Improve our alerting setup

Open
#545 0 comments 0 reactions 0 assignees View on GitHub
Epic Operations
Dominant language
No language data
Stars
122
Forks
7
PR merge metrics
No merged PRs in 30d

Description

Identify and make any useful improvements to our alerting set up. For example are we getting alerted when our error rate or response time rises? What about when our application server CPUs are pegged or their disks are full etc?

In the past we've only been alerted when requests to healthcheck endpoints start failing.

See also:

* [STORY: Alert on persistent high load](https://github.com/hypothesis/product-backlog/issues/257)
* [STORY: Alert on root partition space/inode exhaustion](https://github.com/hypothesis/product-backlog/issues/254)

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.