分享日语95k训练模型
- Dominant language
- Python
- Stars
- 36.9k
- Forks
- 5.2k
- PR merge metrics
- No merged PRs in 30d
Description
链接:https://pan.baidu.com/s/1is8yodQ5QmgsIyLogKF24w?pwd=5yf9
提取码:5yf9
训练时采用的是java里的库来先对所有日语的汉字转换为片假名,然后将片假名转换为罗马音来进行输入...
使用时需要先用将所有日语转换为片假名...然后转为对应的罗马音输入就能合成了
效果:


日语对应的罗马音:

我使用的时候是前几个单词或最后几个单词可以正确识别并且效果还不错,但中间的单词不太好识别出来。而这种情况从一开始的18k完全收敛到95k也没有明显改善,不知道是因为数据集还不够大还是训练的时候哪里出错的原因。还请有相关训练经验的指教一下
Contributor guide
No contributing guide indexed for this repository
Research direction
The issue provides a Baidu model link and training observations, but names no repository files or tests. Start by reviewing the Japanese preprocessing flow and the training configuration used for the 95k model; done would require determining whether the middle-word recognition problem is data-related or caused by training, with a reproducible conclusion.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java, python, pytorch
- Domain
- audio-video-rtc, machine-learning
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100