babysor / babysor/MockingBird

aishell3中的语音合成效果较差

Open
#784 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
36.9k
Forks
5.2k
PR merge metrics
No merged PRs in 30d

Description

Preparing the encoder, the synthesizer and the vocoder...
Loaded encoder "pretrained.pt" trained to step 1594501
Synthesizer using device: cuda
Building hifigan
Loading 'vocoder/saved_models/pretrained/g_hifigan.pt'
Complete.
Removing weight norm...
Trainable Parameters: 0.000M
Loaded synthesizer "mandarin.pt" trained to step 75000
+----------+---+
| Tacotron | r |
+----------+---+
| 75k | 2 |
+----------+---+

Read ['江苏修鞋奶奶婉拒捐款一人养活患病老伴和儿子']
Synthesizing ['jiang1 su1 xiu1 xie2 nai3 nai3 wan3 ju4 juan1 kuan3 yi1 ren2 yang3 huo2 huan4 bing4 lao3 ban4 he2 er2 zi5']

| Generating 1/1

Done.
使用了百度网盘的合成器,请问是什么问题呢?

Contributor guide

No contributing guide indexed for this repository

Research direction

Start at the encoder, synthesizer, and vocoder initialization shown in the report, using pretrained.pt, mandarin.pt, and vocoder/saved_models/pretrained/g_hifigan.pt with the supplied AISHELL3 text. Compare the generated output with the expected synthesis quality and identify a reproducible cause and fix; no test or acceptance criteria are provided.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
audio-video-rtc, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.