NVIDIA-NeMo / NVIDIA-NeMo/RL

Refactor using CP inside loss function

Open
#1,204 0 comments 0 reactions 1 assignee Claimed by @joyang-nv View on GitHub
enhancement t-pytdensor
Dominant language
Python
Stars
2k
Forks
561
Avg merge
4d 5h
Merged PRs (30d)
145

Description

Some nits for CP on dtensor. This could be confusing for developers. But it has to be the state right now. Just file an issue to track this.

1. Mcore explicitly use context parallel group.
2. Dtensor relies on the seq_index.

We will eventually switch to flex attention when it is ready for cp.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.