Dashboards for training and evaluation loss
Open
enhancement
good first issue
- Dominant language
- Python
- Stars
- 18
- Forks
- 8
- PR merge metrics
- No merged PRs in 30d
Description
**Is your feature request related to a problem? Please describe.**
There is currently no end-2-end support for tensorboard or other training/evaluation loss dashboards.
**Describe the solution you'd like**
Solution should show real-time metrics for training jobs running:
- Basic metrics (training loss, validation loss)
- Health of job (maybe including email warning if job fails)
- GPU memory usage
Contributor guide
Assessment
This issue has not been assessed yet.