deepseek-ai / deepseek-ai/DeepEP
Unable to saturate NVL BW w/ intranode on GB200
Open
- Dominant language
- Cuda
- Stars
- 10.1k
- Forks
- 1.4k
- Avg merge
- 4d 1h
- Merged PRs (30d)
- 2
Description
Are there any configuration knobs that need to be tuned to enable full bandwidth on GB200 nodes (800Gb/s)?
On a single host, w/ the intrahost kernel I get:
| Best config | BW (GB/s) | Latency (us)
-- | -- | -- | --
Dispatch | SMs 24, NVL chunk 16 | 169 | 690
Combine | SMs 24, NVL chunk 16 | 135 | 861
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.