NVIDIA-NeMo / NVIDIA-NeMo/RL

Performance issues for SFT on Qwen 30B-a3b

Open
#937 1 comment 1 reaction 1 assignee Claimed by @guyueh1 View on GitHub
Performance research t-mcore
Dominant language
Python
Stars
2k
Forks
561
Avg merge
4d 5h
Merged PRs (30d)
145

Description

Performance of SFT for Qwen 30B-a3b is ~30% lower than Qwen 14B probably because of it being an MoE model. Can we improve performance of SFT Qwen 30B? This is for megatron backend.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.