babysor / babysor/MockingBird

训练到150k,但loss仍在0.5左右,attention图时好时坏是为什么?loss为什么减不下去?

Open
#636 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
36.9k
Forks
5.2k
PR merge metrics
No merged PRs in 30d

Description

下载的aidatatang_200zh.tgz数据集,在pre.py预处理之后,进行生成器训练。下面是部分图:

![Z2RHQC821_BWWKP7}PX5D1](https://user-images.githubusercontent.com/105544339/177909320-c52432e5-9903-499e-871e-1b38ca2f833a.png)

![155000](https://user-images.githubusercontent.com/105544339/177909324-411c8d60-4d61-4354-b533-4c203eaad642.png)
step=155000⬆
![143000](https://user-images.githubusercontent.com/105544339/177909359-31efeb55-a3a4-4c3a-9839-63a1760f3a9c.png)
step=143000⬆

总觉得不对劲,这种情况是正常的吗?若不正常应该如何处理?

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with the preprocessing flow in pre.py and the generator-training entry point, then review the reported loss and attention images around steps 143000 and 155000. Determine whether the plateau and changing attention are expected for the aidatatang_200zh.tgz training setup; done requires a reproducible diagnosis and a clearly identified corrective action if they are not.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
ai, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.