RVC-Project / RVC-Project/Retrieval-based-Voice-Conversion-WebUI
Error in epoch rounds, recounting from 1
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 38.4k
- Forks
- 5.3k
- PR merge metrics
- No merged PRs in 30d
Description
During the training process of my model, because I did not reserve enough space on the hard drive, the model stopped with an error "OSError:No space left on device".
I cleaned it up and left enough space to start training again, but it should have trained for more than 400 rounds, but the epoch is counting from 1. Is this because the checkpoints are missing? Is there any negative impact on my training?
I really need to solve this problem urgently, thanks for your help!
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
No file, test, or entry point is identified. Start by tracing the training process and checkpoint handling to determine why training resumes at epoch 1 after the disk-space failure; done means documenting whether the restart is expected and whether checkpoint recovery is affected.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100