Export count metric of tasks recomputed due to worker death
Open
- Dominant language
- Python
- Stars
- 1.7k
- Forks
- 778
- Avg merge
- 2h 50m
- Merged PRs (30d)
- 3
Description
For understanding how well graceful shutdown is working, it would be very useful if there was a metric, available on scheduler Prometheus endpoint, that said how many tasks had been recomputed due to worker death.
(If there are other reasons tasks can be recomputed, it might make sense to export `recomputed_task_count` metric that has labels for each reason a task can be recomputed.)
Contributor guide
Assessment
This issue has not been assessed yet.