RVC-Project / RVC-Project/Retrieval-based-Voice-Conversion-WebUI
Incorrect Pronunciation of single word 'He' after Conversion
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 38.4k
- Forks
- 5.3k
- PR merge metrics
- No merged PRs in 30d
Description
Hello,
I encountered an issue with the RVC (1006NVIDIA) when converting a TTS-generated voice file. Specifically, the single word "He" is not pronounced correctly after conversion. Instead of the expected pronunciation, the output sounds more like "swee" or "sui."
Steps to Reproduce:
- I used a TTS engine to generate a voice file that includes the single word "He."
- I applied RVC to convert the voice file.
- The output consistently mispronounces "He" across different RVC models I tried.
What I've Tried:
- I tested the conversion with different TTS-generated voice files, but the issue persists.
- I used multiple RVC models to see if the problem was model-specific, but the result was the same.
- I adjusted the settings in RVC, but the issue remains unresolved.
Expected Result:
The RVC conversion should accurately pronounce "He" as it is in the original file.
Actual Result:
The word "He" is incorrectly converted to something like "swee" or "sui."
I have attached the original he.wav file for reference. Any help or suggestions to resolve this issue would be greatly appreciated.
he.zip
Thank you!
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reproducing the conversion with the attached he.wav file, the RVC (1006NVIDIA) setup, and the reported settings. Compare the original and converted audio across the mentioned TTS files and RVC models; done requires identifying why the single word “He” changes pronunciation and confirming that it is preserved after conversion.
Written by the indexing model from the issue text.
Assessment
- Domain
- audio-video-rtc, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100