babysor / babysor/MockingBird

Synthesizer loss increases/diverges under training with GPU

Open
#612 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
36.9k
Forks
5.2k
PR merge metrics
No merged PRs in 30d

Description

**Summary[问题简述(一句话)]**
If I use CPU to train the synthesizer, under the fine-tuning methodology, I get good results and the loss has been decreasing over time. However, when I moved the models over to an Ubuntu container, running ROCm for GPU acceleration using the AMD graphics cards, the loss actually diverges.

Has anyone else experienced this, and if so, how did you solve it?

**Env & To Reproduce[复现与环境]**
描述你用的环境、代码版本、模型
Ubuntu 20.04
ROCm 5.1, using RX580
Pytorch 1.11
aidatatang_200zh

**Screenshots[截图(如有)]**
If applicable, add screenshots to help

Contributor guide

No contributing guide indexed for this repository

Research direction

No source file or test is named. Start by reproducing synthesizer fine-tuning with the listed Ubuntu, ROCm 5.1, RX580, PyTorch 1.11, and aidatatang_200zh environment, then compare its loss behavior with CPU training. Done means documenting a confirmed cause and a reproducible resolution or workaround.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch, ubuntu
Domain
ai, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.