Prometheus: expose cumulative worker utilisation instead of spot metric
Open
- Dominant language
- Python
- Stars
- 1.7k
- Forks
- 778
- Avg merge
- 2h 50m
- Merged PRs (30d)
- 3
Description
Tightly related to #7123.
In Prometheus, we currently expose cluster-wide utilisation over time using the spot measure of number of executing tasks on the workers. This means that with a sample time of e.g. 5s you may not notice imperfect cluster occupation if your tasks are shorter than 5s.
It should be changed to a monotonically-increasing number of tasks that are currently executing + tasks that exited computation (finished, erred, or rescheduled).
Contributor guide
Assessment
This issue has not been assessed yet.