High logprob errors with deepscaler recipe + megatron backend
Open
bug
t-mcore
- Dominant language
- Python
- Stars
- 2k
- Forks
- 561
- Avg merge
- 4d 5h
- Merged PRs (30d)
- 145
Description
**Describe the bug**
Running the deepscaler recipe with Megatron yields very high logprob errors throughout training
**Steps/Code to reproduce bug**
Find the config used to reproduce [here](https://github.com/NVIDIA-NeMo/RL/blob/3d81caf0f5df13732ff97b51366762f0ecd6e8ab/examples/configs/recipes/llm/grpo-deepscaler-1.5b-8K-megatron.yaml)
Contributor guide
Assessment
This issue has not been assessed yet.