Request for Guidance on Tamil Voice Cloning – Training Data, Metadata & Multi-Speaker Support
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 15.3k
- Forks
- 2.2k
- PR merge metrics
- No merged PRs in 30d
Description
Checks
- This template is only for research question, not usage problems, feature requests or bug reports.
- I have thoroughly reviewed the project documentation and read the related paper(s).
- I have searched for existing issues, including closed ones, no similar questions.
- I am using English to submit this issue to facilitate community communication.
Question details
Hello there,
I am currently working on a Tamil TTS project with the goal of achieving high-quality voice cloning using only training text. I have a few questions and would really appreciate your guidance and support:
Data Requirements: To achieve accurate and natural-sounding voice cloning in Tamil, how many hours of audio data would you recommend?
Emotion Metadata: Does the training process support or require emotional classification in the metadata (e.g., happy, sad), or is neutral speech sufficient?
Generalization Issue: In a previous experiment, I used 200 hours of audio. The model performed well when reproducing the trained voice, but it did not generalize effectively to other voices.
Configuration used: batch size = 700, 20 epochs, 1.3M training steps.
Given this setup, what would you suggest to enable voice cloning for new speakers?
I’d be grateful for any insights or suggestions. If needed, I’d also be happy to arrange a call or meeting to discuss this further.
@SWivid
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
The issue names no files, tests, or code entry points. Start with the project documentation and related papers referenced by the author, then examine the training configuration described in the issue. A concrete, scoped change and acceptance criteria would be needed before implementation can be considered done.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100