deepseek-ai / deepseek-ai/DeepEP

`dispatch_wait_recv_cost_stats` is cumulating time from each warp, when instead computing a max across warps would be more helpful

Open
#473 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
Cuda
Stars
10.1k
Forks
1.4k
Avg merge
4d 1h
Merged PRs (30d)
2

Description

```
atomicAdd(reinterpret_cast(dispatch_wait_recv_cost_stats + src_rank), wait_recv_cost);
```
This op adds the wait time as seen by a given warp for a given src_rank. This time is being added by each warp.

I wonder what is the utility of this metric? Each warp is waiting in parallel. Are we trying to infer slow ranks through this metric?

More interesting would be the max across warps since that would indicate the actual wait time from a src_rank.

Then we could also get a max of this max across src_ranks to actually infer the pure network comms latency of the recv kernel.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.