gemma 27b slow on v2
Open
bug
t-pytdensor
- Dominant language
- Python
- Stars
- 2k
- Forks
- 561
- Avg merge
- 4d 5h
- Merged PRs (30d)
- 145
Description
I am noticing that gemma 27b is significantly slower on the v2 backend. I don't observe such a large slowdown with other models.
Here are three sets of runs:
https://wandb.ai/nvidia/nemo-rl?nw=drf8mhln88
* green: v2
* blue, red: v1
There is some in-run variance in perf (~8%), but the diff from v1 is much larger
Related automodel issue: https://github.com/NVIDIA-NeMo/Automodel/issues/432
Contributor guide
Assessment
This issue has not been assessed yet.