RVC-Project / RVC-Project/Retrieval-based-Voice-Conversion-WebUI

MacOS | MPS | torch metal work on Train -> step-1 then Step-2 but not on training the model (Step-3)

Open
#991 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

help wanted
Dominant language
Python
Stars
38.4k
Forks
5.3k
PR merge metrics
No merged PRs in 30d

Description

Hi
is there a way to make the model train on torch metal, mps, even when passing the "mps" manually to the G and D models, it still uses the cpu, I also monitored the GPU performence, it peaks only on the first two steps, but on the model training it peaks the CPU
any idea?
and it stuck sometimes on "INFO:torch.nn.parallel.distributed:Reducer buckets have been rebuilt in this iteration." for a long time, I was trying to train on a 3 minutes speech, it took more than 5hrs for 20 epoch
Thanks is advance

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the Step-3 model-training path on macOS with the MPS device passed to both G and D, while monitoring CPU and GPU usage. Inspect the training entry point and device handling around the reported distributed reducer message; done means training actually uses MPS and completes without the reported prolonged stall.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.