gpt-oss-20 OOMs on 8 B200s using default config
Open
bug
community-request
waiting-on-maintainers
- Dominant language
- Python
- Stars
- 2k
- Forks
- 562
- Avg merge
- 4d 1h
- Merged PRs (30d)
- 150
Description
Hi!
I'm using the provided config values for gpt-oss but despite TP+EP across 8 B200s (and using flex attention) I keep running into OOMs; I've also tried configuring megatron as well but that just seems to outright fail as it falls back to a naive attention implementation.
Contributor guide
Assessment
This issue has not been assessed yet.