How to get stable voice accent and style with HiFi cloning?
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 37.8k
- Forks
- 4.3k
- Avg merge
- 7m
- Merged PRs (30d)
- 1
Description
I have been testing HiFi cloning. I am getting slight variations in voice depending on input text or it seems to be off on random inputs with same prompt-wav file and same prompt-wav text with cfg value of 3.
What are the options to get a stable cloned voice ? I read about LoRA fine tuning in docs for a speaker but does that mean I will have to train a new model each time I need a new voice or can multiple voices exist simultaneously in the same model ? How does that work?
Is there any other way to get a stable voice output?
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
The report names no files, tests, or entry points. Begin with the HiFi cloning and LoRA fine-tuning documentation referenced in the issue, then define documentation or reproducible-test scope that explains voice stability and whether voices can coexist in one model.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- audio-video-rtc, machine-learning
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100