babysor / babysor/MockingBird

分享日语95k训练模型

Open
#556 2 comments 8 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
36.9k
Forks
5.2k
PR merge metrics
No merged PRs in 30d

Description

链接:https://pan.baidu.com/s/1is8yodQ5QmgsIyLogKF24w?pwd=5yf9
提取码:5yf9
训练时采用的是java里的库来先对所有日语的汉字转换为片假名,然后将片假名转换为罗马音来进行输入...
使用时需要先用将所有日语转换为片假名...然后转为对应的罗马音输入就能合成了

效果:
![attention_step_92000_sample_1](https://user-images.githubusercontent.com/41981371/168515379-bc69acc3-5ffc-40ff-9f17-e3c9d2211c60.png)
![step-95000-mel-spectrogram_sample_1](https://user-images.githubusercontent.com/41981371/168515414-d2390d09-fff0-41d7-b1fe-d72121a142cb.png)
日语对应的罗马音:
![1c950a7b02087bf42f778c96e2d3572c11dfcf32](https://user-images.githubusercontent.com/41981371/168515454-e156ead7-6a05-4086-9899-58c2227f5320.png)

我使用的时候是前几个单词或最后几个单词可以正确识别并且效果还不错,但中间的单词不太好识别出来。而这种情况从一开始的18k完全收敛到95k也没有明显改善,不知道是因为数据集还不够大还是训练的时候哪里出错的原因。还请有相关训练经验的指教一下

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue provides a Baidu model link and training observations, but names no repository files or tests. Start by reviewing the Japanese preprocessing flow and the training configuration used for the 95k model; done would require determining whether the middle-word recognition problem is data-related or caused by training, with a reproducible conclusion.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, python, pytorch
Domain
audio-video-rtc, machine-learning
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.