NVIDIA-NeMo / NVIDIA-NeMo/RL

Nano v3 dpo failed with cpu-offload after transformer v5 bump.

Open
#2,130 0 comments 0 reactions 1 assignee Claimed by @hemildesai View on GitHub
bug
Dominant language
Python
Stars
2k
Forks
561
Avg merge
4d 5h
Merged PRs (30d)
145

Description

For now we doesn't use cpu-offload in `examples/configs/recipes/llm/dpo-nanov3-30B3AB-1n8g-fsdp8ep8-automodel.yaml`.
After it fixed, need to add cpu-offload in this config.

Related PR: https://github.com/NVIDIA-NeMo/RL/pull/1962

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.