alphagov / alphagov/govuk-developer-docs
Document monitoring, metrics, tracing, observability and alerting
A pull request for this has already been merged.
- #4863 by @nicholsj — merged
- Dominant language
- Ruby
- Stars
- 143
- Forks
- 39
- Avg merge
- 1d 4h
- Merged PRs (30d)
- 12
Description
As an engineer on GOV.UK,
I want to know how to configure monitoring, metrics, tracing, observability and alerting for my applications,
so that we can enable proactive detection and resolution of issues, ensure optimal performance and enhance reliability by providing real time insight into applications health and behaviour, as well as to inform product decisions.
Current documentation
Logging
How logging works on GOV.UK
Request tracing
Monitoring
Debug underperforming search - I've asked Search team to review and probably remove this
How we handle errors
Pingdom
Sentry
Alerting
Pingdom Bouncer canary check
Router error ratio too high
Travel Advice or Drug and Medical Device email alerts not sent
Signon API user token expires soon
PagerDuty
Things that may contact on-call - I suggest the specifics get taken out of here and instead link to the relevant pages
[WIP] Missing documentation
- Grafana - looks like there used to be one at https://docs.publishing.service.gov.uk/manual/grafana.html. Access to Grafana is mentioned here . (example steps to create dashboards for Prometheus metrics are included in [this card])(https://trello.com/c/SQro2G8f/3476-monitor-how-long-content-datas-csv-exports-take-5)
- App metrics, Prometheus
- configuring Alertmanager alerts
[WIP] Documentation that could do with a refresh
- Pagerduty alerts section, AlertManager alerts section and the Monitoring section - consolidate all of this and more under a new section called
Monitoring and alerting.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the existing Logging, Monitoring and Alerting pages linked in the issue, including the Monitoring and alerting sections of the manual. Review the listed gaps for Grafana, app metrics and Prometheus, and Alertmanager. Done should be consolidated monitoring and alerting documentation covering the identified missing topics and refreshed links.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- grafana, prometheus
- Domain
- documentation, observability-sre
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100