deepseek-ai / deepseek-ai/DeepEP
Confues about the Low-latency kernels data issue
Open
- Dominant language
- Cuda
- Stars
- 10.1k
- Forks
- 1.4k
- Avg merge
- 4d 1h
- Merged PRs (30d)
- 2
Description
Hello DeepEP developers, appreciate the solid works!
I have two questions about the Low-latency kernels with pure RDMA displayed on homepage, thanks :
1, Is the low_latency data displayed the average, maximum or minimum value of all ranks? Which method makes more sense?
2, Why does the bandwidth decrease as the number of #EP increases?
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.