RVC-Project / RVC-Project/Retrieval-based-Voice-Conversion-WebUI
MacOS | MPS | torch metal work on Train -> step-1 then Step-2 but not on training the model (Step-3)
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 38.4k
- Forks
- 5.3k
- PR merge metrics
- No merged PRs in 30d
Description
Hi
is there a way to make the model train on torch metal, mps, even when passing the "mps" manually to the G and D models, it still uses the cpu, I also monitored the GPU performence, it peaks only on the first two steps, but on the model training it peaks the CPU
any idea?
and it stuck sometimes on "INFO:torch.nn.parallel.distributed:Reducer buckets have been rebuilt in this iteration." for a long time, I was trying to train on a 3 minutes speech, it took more than 5hrs for 20 epoch
Thanks is advance
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reproducing the Step-3 model-training path on macOS with the MPS device passed to both G and D, while monitoring CPU and GPU usage. Inspect the training entry point and device handling around the reported distributed reducer message; done means training actually uses MPS and completes without the reported prolonged stall.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning, performance
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100