RL nightly test failing train/reward < 0.9
Open
bug
- Dominant language
- Python
- Stars
- 2k
- Forks
- 561
- Avg merge
- 4d 5h
- Merged PRs (30d)
- 145
Description
**Describe the bug**
RL nightly test failing: vlm_grpo-qwen2.5-vl-3b-instruct-clevr-1n4g-dtensor2tp1.v1 with the train/reward value is less than 0.9.
**Expected behavior**
train/reward value should greater than 0.9
Contributor guide
Assessment
This issue has not been assessed yet.