OpenBMB / OpenBMB/VoxCPM

1.5B模型在1min左右开始有几率发出啸叫声

Open
#103 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
37.8k
Forks
4.3k
Avg merge
7m
Merged PRs (30d)
1

Description

如题,第一次测试的时候默认设置,cfg2,步数10,字数2000左右,直出了6min无克隆的音频特别完美,当时很兴奋
可惜的是,之后的测试再也没有复现这个壮举,基本上在1min左右开始出现微弱啸叫声,3-4min左右变大无法接受
目前我的测试下,cfg1.5,是比较稳定的1min左右可以稳定没有啸叫,或者啸叫可以接受。步数的增加似乎没有能够有效控制啸叫,流失输出与非流失输出差别也不大

不过我除了cfg和步数外,并没有修改其他参数,是否有其他参数的调整可以增强稳定性?

其实对比0.5B的版本,已经是巨大的进步了,当时30s左右就开始啸叫,现在稳定在1min左右,特别棒!

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No source file or test is named. Start by reproducing the 1.5B model's long-generation behavior with the reported defaults, then compare cfg1.5 and cfg2, step counts, and dropout settings. Done means identifying a reproducible stability cause or documenting a parameter combination that prevents unacceptable howling.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
audio-video-rtc, machine-learning
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.