NVIDIA-NeMo / NVIDIA-NeMo/RL

PyT DTensor Path - Llama 70B with 4k seq gives OOM with sequence packing enabled

Open
#769 4 comments 0 reactions 1 assignee Assigned to @wangshangsam View on GitHub
bug memory issue t-pytdensor
Dominant language
Python
Stars
2k
Forks
561
Avg merge
4d 5h
Merged PRs (30d)
145

Description

This issue has no description.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.