Issues with Hindi Language Voice Cloning in VoxCPM
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 37.8k
- Forks
- 4.3k
- Avg merge
- 7m
- Merged PRs (30d)
- 1
Description
When using the VoxCPM2 model for voice cloning on an Ubuntu GPU server, I encountered a few issues while testing Hindi voice generation.
1. Hindi generation using English reference audio
When I provide Hindi input text along with an English reference audio sample (e.g., a speech by Barack Obama), the model fails to generate proper Hindi audio output. In some cases, no meaningful audio is produced, or the output is unintelligible.
2. Extra word/artifact with Hindi reference audio
When I use a Hindi reference audio and generate Hindi speech, the output consistently contains an unexpected extra word or artifact at the beginning of the generated audio. This behavior is reproducible across multiple test samples.
3. Hi-Fi cloning observation
I also tested the Hi-Fi cloning setup (with prompt audio and transcript), and the above issues—especially the initial artifact in Hindi output—still persist.
Minimal reproducible example:
# -*- coding: utf-8 -*-
from voxcpm import VoxCPM
import soundfile as sf
# Load model
model = VoxCPM.from_pretrained(
"openbmb/VoxCPM2",
load_denoiser=False
)
# INPUTS
text = "हम बैकग्राउंड म्यूज़िक के साथ आवाज़ की स्पष्टता की जाँच कर रहे हैं।"
reference_audio = "Obama_reference.wav"
# Generate cloned speech
wav = model.generate(
text=text,
reference_wav_path=reference_audio,
cfg_value=2.0,
inference_timesteps=10,
)
# Save output
sf.write("output_hindi.wav", wav, model.tts_model.sample_rate)
print("Output saved as output_hindi.wav")
Environment details:
- OS: Ubuntu
- GPU: H100
- Python: 3.10
- Installation: Cloned from the official VoxCPM repository
Questions:
- Is cross-language voice cloning (e.g., English reference → Hindi output) currently supported or recommended?
- What could be causing the extra word or artifact at the beginning of Hindi outputs?
- Are there any best practices or parameter tuning recommendations for Hindi or multilingual voice cloning?
Any guidance would be greatly appreciated. Thank you!
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the provided Python reproducer at VoxCPM.from_pretrained(...).generate(...), comparing the English Obama_reference.wav case with the Hindi reference and the Hi-Fi cloning setup described. Done means determining whether cross-language cloning is supported and isolating or documenting the reproducible initial artifact, with guidance for Hindi parameters if applicable.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- audio-video-rtc, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100