RVC-Project / RVC-Project/Retrieval-based-Voice-Conversion-WebUI

Incorrect Pronunciation of single word 'He' after Conversion

Open
#2,278 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
38.4k
Forks
5.3k
PR merge metrics
No merged PRs in 30d

Description

Hello,

I encountered an issue with the RVC (1006NVIDIA) when converting a TTS-generated voice file. Specifically, the single word "He" is not pronounced correctly after conversion. Instead of the expected pronunciation, the output sounds more like "swee" or "sui."

Steps to Reproduce:
  1. I used a TTS engine to generate a voice file that includes the single word "He."
  2. I applied RVC to convert the voice file.
  3. The output consistently mispronounces "He" across different RVC models I tried.
What I've Tried:
  • I tested the conversion with different TTS-generated voice files, but the issue persists.
  • I used multiple RVC models to see if the problem was model-specific, but the result was the same.
  • I adjusted the settings in RVC, but the issue remains unresolved.
Expected Result:

The RVC conversion should accurately pronounce "He" as it is in the original file.

Actual Result:

The word "He" is incorrectly converted to something like "swee" or "sui."

I have attached the original he.wav file for reference. Any help or suggestions to resolve this issue would be greatly appreciated.
he.zip

Thank you!

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the conversion with the attached he.wav file, the RVC (1006NVIDIA) setup, and the reported settings. Compare the original and converted audio across the mentioned TTS files and RVC models; done requires identifying why the single word “He” changes pronunciation and confirming that it is preserved after conversion.

Written by the indexing model from the issue text.

Assessment

Domain
audio-video-rtc, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.