[doc] Failure planning
- Dominant language
- Go
- Stars
- 28
- Forks
- 4
- PR merge metrics
- No merged PRs in 30d
Description
* identify potential failure scenarios
* External services
* EC2 / AZ
* Host disappear
* Network failure
* Account limits
* Dynamodb
* Connectivity
* Slow?
* Github (login, ssh keys, pages, groups)
* Connectivity
* Unsupported key
* API limits
* Missing group
* Route53 (cmd.io, dune)
* DNS failure
* Auth0
* Honeycomb
* Sentry
* Gliderlabs.io (Heroku)
* Ops notifications
* S3
* Connectivity
* Docker Hub
* Slow
* Connectivity
* Dune
* No Dockers
* Hosts gone
* Convox, ELB etc
* Uncaught panics
* Host resources
* Memory
* Disk space
* Bad TF deploy
* testing
* should we have an automated way of testing failure scenarios?
* how could they be run, and how often?
* Unittests ideally
* https://github.com/Shopify/toxiproxy
* Injecting responses?
* identify what happens when (x) service goes down
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.