OpenMOSS / OpenMOSS/AnyGPT

Details on TTS evaluation?

Open
#42 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
882
Forks
76
PR merge metrics
No merged PRs in 30d

Description

Hello! Thanks for your wonderful work. Trying to reproduce your results on the TTS task, I'm wondering if you could provide more details about the evaluation of the TTS task, especially:

  • How many / Which samples are used in VCTK dataset
  • Which ASR model is used to convert the generated speech into text
  • How the WER is calculated; What kind of text normalization is applied before the calculation

Thanks!

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No files or tests are named. Start by locating the TTS evaluation entry point and VCTK sample selection, then identify the ASR model and WER text normalization used. Done means documenting these details well enough to reproduce the reported results.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
audio-video-rtc, documentation
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.