deepseek-ai / deepseek-ai/DeepEP
curious about the difference between latency and throughput
Open
- Dominant language
- Cuda
- Stars
- 10.1k
- Forks
- 1.4k
- Avg merge
- 4d 1h
- Merged PRs (30d)
- 2
Description
It seems that EP-DISPATCH and EP-COMBINE both work well in throughput tests while EP-DISPATCH works much better than EP-COMBINE in latency tests.
Why?
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.