OpenBMB / OpenBMB/VoxCPM

极致克隆模式,每次合成后音色有区别

Open
#262 7 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
37.8k
Forks
4.3k
Avg merge
7m
Merged PRs (30d)
1

Description

我用了极致克隆,每次生成的音色也还是有区别,能明显听出来。我是长文本合成,如果按段落进行合成,合成的质量会随时间逐渐降低,只有刚开始的两句话质量还挺好。我就改成了逐句合成,这样每句的质量是有保证了,但是每句话的音色又有区别,合成最后的音频时,每句话的音色和质量都不一样,能明显听出来。

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No file, test, or entry point is named. First reproduce the reported long-text paragraph synthesis and sentence-by-sentence synthesis, comparing voice consistency and quality across segments; done means the generated segments retain consistent timbre and quality.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
audio-video-rtc, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.