coder / coder/internal

C210K: Scaletesting Grafana Dashboards

Aperta
#328 0 commenti 0 reazioni 1 assegnatario Rivendicata da @f0ssel Vedi su GitHub
Lingua principale
Nessun dato sulla lingua
Stelle
3
Fork
0
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

### Overview

We need to create new grafana dashboards for measuring monthly scaletest performance, defining pass/fail criteria, and identifying points of failure.

Key metrics to collect (from @johnstcn and @mafredri)
- Workspaces (coder prom)
- Basic resources: CPU, Disk, Mem (coder prom)
- Heap usage (coder prom)
- Number of goroutines (coder prom)
- HTTP Requests by path, status, duration (coder prom)
- Provisioners (coder prom)
- TCP performance [bytes in, bytes out] (kube)

### Requirements

Dashboards
- [ ] Coder Server
- [ ] Coder Proxy
- [ ] Coder Provisioners
- [ ] Workspace agents
- [ ] K8s

Nice to have tooling
- [ ] [Grafana annotations](https://grafana.com/docs/grafana/latest/dashboards/build-dashboards/annotate-visualizations/): to mark scaletest events
- [ ] [Template variables](https://grafana.com/docs/grafana/latest/dashboards/variables/): reusable resource queries

Metrics to summarize pass/fail
- [ ] API error rate over time -- HTTP response codes (coder prom)
[_key_ failure criteria]
- [ ] Pod restart counts (kube)
[_must_ be zero]
- [ ] Deviation of connected agents (coder prom)
- [ ] Acceptable CPU, memory thresholds (kube)

---

> _Don't sleep on heatmaps._
> _- Cian_

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.