OpenMOSS / OpenMOSS/MOSS-TTSD

v0.7版本第一个人说话时间超过20s会出现音色突变

Open
#113 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
1.4k
Forks
138
PR merge metrics
No merged PRs in 30d

Description

在测试v0.7版本双人对话非流式输出的时候,发现如果第一个人说话时间超过20s会出现很严重的音色突变

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Reproduce the v0.7 two-speaker, non-streaming dialogue case with the first speaker talking for more than 20 seconds. Compare the generated audio around the 20-second point and determine what causes the severe timbre change; done when the output preserves a consistent voice throughout.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
audio-video-rtc, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.