NVIDIA-NeMo / NVIDIA-NeMo/RL

Gym: Empty epochs if Gym agent fails

Open
#2,305 0 comments 0 reactions 0 assignees View on GitHub
bug
Dominant language
Python
Stars
2k
Forks
561
Avg merge
4d 5h
Merged PRs (30d)
145

Description

**Describe the bug**
When running Gym example `examples/nemo_gym/grpo_qwen3_30ba3b_instruct.yaml` I encountered strage bug. Config is outdated, using 4096 * 16 samples (instead of 64 * 16) which results in CPU OOM during rollout collection and killing Gym agent silently. This does not crash the training, but only results in "empty epochs" in logs.

**Steps/Code to reproduce bug**

Run `examples/nemo_gym/grpo_qwen3_30ba3b_instruct.yaml` with ` examples/nemo_gym/run_grpo_nemo_gym.py` script. Tested on 8 DGXH100 nodes

**Expected behavior**
More reasonable error, training crashes instead of producing empty runs

**Additional context**
I believe config should also be updated and number of rollouts decreased

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.