Megatron LoRA is slow with tulu3 dataset
Open
bug
- Dominant language
- Python
- Stars
- 2k
- Forks
- 561
- Avg merge
- 4d 5h
- Merged PRs (30d)
- 145
Description
Megatron LoRA's perf w/ squad dataset is fine, but is bad when using tulu3 dataset.
Repro:
`bash tests/test_suites/llm/sft-llama3.1-8b-1n8g-megatron-lora.sh` in https://github.com/NVIDIA-NeMo/RL/pull/1629.
Test Result:
https://wandb.ai/nvidia/lora-rl/panel/o7jy6zley?nw=nwuservadams
Contributor guide
Assessment
This issue has not been assessed yet.