llm_grpo_llama3_2_1b_instruct_1n8g_megatron_generation nightly test reward not stable
Open
accuracy
bug
community-request
waiting-on-maintainers
- Dominant language
- Python
- Stars
- 2k
- Forks
- 561
- Avg merge
- 4d 5h
- Merged PRs (30d)
- 145
Description
**Describe the bug**
Final step train reward for this test recipe is not always stable
**Steps/Code to reproduce bug**
Tests off the main branch.
Historical details
Another example was observed durign the v0.7.0 release testing:
```
Expected data["train/reward"]["500"] > 0.1
Value: 0.05
```
Please help make the nightly test stable.
Contributor guide
Assessment
This issue has not been assessed yet.