RVC-Project / RVC-Project/Retrieval-based-Voice-Conversion-WebUI
Use higher sample rate for inference
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 38.4k
- Forks
- 5.3k
- PR merge metrics
- No merged PRs in 30d
Description
I'm having issues with pronunciation and I believe that this is due to the input audio being downsampled to 16khz.
For example, I have an audio in 48khz that has clear pronunciation of some words, but when I downsample it to 16khz, it's hard to understand some words and I think the inference model has the same problem.
I think the solution would be to use a higher sample rate for inference.
So it this possible to use input audio with a higher sample rate?
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
The issue names no files, tests, or entry points; begin by locating where input audio is downsampled to 16 kHz and where inference accepts its sample rate. Determine whether higher-rate input is supported and validate pronunciation quality against the reported 48 kHz versus 16 kHz example.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- audio-video-rtc, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100