关于训练某个特定声音遇到的问题
Open
- Dominant language
- Python
- Stars
- 36.9k
- Forks
- 5.2k
- PR merge metrics
- No merged PRs in 30d
Description
基于社区版继续训练某个特定声音遇到的问题
本人基于社区版ceshi.pt,加入一些某个人的音频文件,以aidatatang_200zh数据集格式进行训练,到244k的时候测试发现如下问题
1.断句问题。比如“服务器”,会感觉服务跟器之间有明显停顿
2.多音字问题。比如“重装”,“重”会读成四声
3.语速快问题。用新模型进行合成之后发现,录出的语音比较快。
以上三个问题,不知道能否通过预处理或者修改代码进行优化。恳请路过的各位大神不吝赐教!
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reproducing the three reported symptoms with ceshi.pt after training on the added audio in the aidatatang_200zh format. Compare preprocessing and training behavior around 244k, and consider the work complete only when pauses, polyphonic pronunciation, and speaking speed are consistently improved and verified with test utterances.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- audio-video-rtc, machine-learning
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 15/100