Optimizer state isn't preserved across runs
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 10.8k
- Forks
- 989
- Avg merge
- 6h 29m
- Merged PRs (30d)
- 85
Description
I often have to restart a run, either to fix something in my reward function, in response to an OOM or crash that broke training, etc. When I do, by restarting the training process the optimizer state is thrown away. I’m worried that this might lead to worse performance than just letting a run go all the way through. Is it easy to save the optimizer state along with the weights so we can truly resume as if nothing happened?
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
No files, tests, or entry points are identified in the issue. Start by locating the training checkpoint and resume flow, then determine how weights are saved and restored; done means a restarted run restores the optimizer state so training can continue as before.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100