babysor / babysor/MockingBird

关于训练某个特定声音遇到的问题

Open
#541 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
36.9k
Forks
5.2k
PR merge metrics
No merged PRs in 30d

Description

基于社区版继续训练某个特定声音遇到的问题

本人基于社区版ceshi.pt,加入一些某个人的音频文件,以aidatatang_200zh数据集格式进行训练,到244k的时候测试发现如下问题
1.断句问题。比如“服务器”,会感觉服务跟器之间有明显停顿
2.多音字问题。比如“重装”,“重”会读成四声
3.语速快问题。用新模型进行合成之后发现,录出的语音比较快。

以上三个问题,不知道能否通过预处理或者修改代码进行优化。恳请路过的各位大神不吝赐教!

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reproducing the three reported symptoms with ceshi.pt after training on the added audio in the aidatatang_200zh format. Compare preprocessing and training behavior around 244k, and consider the work complete only when pauses, polyphonic pronunciation, and speaking speed are consistently improved and verified with test utterances.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
audio-video-rtc, machine-learning
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
15/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.