Synthesizer loss increases/diverges under training with GPU
- Dominant language
- Python
- Stars
- 36.9k
- Forks
- 5.2k
- PR merge metrics
- No merged PRs in 30d
Description
**Summary[问题简述(一句话)]**
If I use CPU to train the synthesizer, under the fine-tuning methodology, I get good results and the loss has been decreasing over time. However, when I moved the models over to an Ubuntu container, running ROCm for GPU acceleration using the AMD graphics cards, the loss actually diverges.
Has anyone else experienced this, and if so, how did you solve it?
**Env & To Reproduce[复现与环境]**
描述你用的环境、代码版本、模型
Ubuntu 20.04
ROCm 5.1, using RX580
Pytorch 1.11
aidatatang_200zh
**Screenshots[截图(如有)]**
If applicable, add screenshots to help
Contributor guide
No contributing guide indexed for this repository
Research direction
No source file or test is named. Start by reproducing synthesizer fine-tuning with the listed Ubuntu, ROCm 5.1, RX580, PyTorch 1.11, and aidatatang_200zh environment, then compare its loss behavior with CPU training. Done means documenting a confirmed cause and a reproducible resolution or workaround.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch, ubuntu
- Domain
- ai, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100