Target Accent issues with Ultimate Cloning Mode
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 37.8k
- Forks
- 4.3k
- Avg merge
- 7m
- Merged PRs (30d)
- 1
Description
Hi,
I just test your project on huggingface.com.
First of all I like to compliment to all the staff for very good quality result of cloning.
I have tested cloning from English reference voice to Italian voice and I find that result is very good.
After, I have select the some English voice, but I have insert the Transcript of Reference Audio also, because I expected that this way I would get better quality.
However, while without Transcript of Reference, the final result keep correcly the target accent like I need, when I choose Ultimate Cloning Mode I get the italian voice with English accent that is unuseful for my need,
I'm wondering if you plan to improve the preservation of the target accent in “Ultimate Cloning Mode”.
In my humble opinion, it might be helpful for some people to have the option of choosing whether or not to keep the target accent.
But in most cases, it would probably be more useful to preserve the target accent something it does quite well, but only if you don't specify “Ultimate Cloning Mode”
Thank you for your consideration
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
No files, tests, or entry points are identified. First reproduce the behavior on huggingface.com with an English reference voice, an Italian target, and the reference transcript enabled in Ultimate Cloning Mode; compare it with cloning without the transcript. Done would require preserving the target accent or clearly defining an option for choosing accent preservation.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- audio-video-rtc, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100