mlcommons / mlcommons/chakra

Lack of GPU comm to compute dependency.

Open
#186 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

question
Dominant language
Python
Stars
196
Forks
84
PR merge metrics
No merged PRs in 30d

Description

In kineto traces we see GPU compute waits for communication to finish.

Image

However, after merging with trace_link there is no dependency between GPU compute and communication, even though HTA identifies the dependency, it is not added to the node.

The identified HTA dependencies are used to add GPU-to-CPU dependency using but not GPU comm to compute dependency. Why is that?

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by tracing how HTA dependencies are merged after trace_link, then compare the existing GPU-to-CPU dependency handling with GPU communication-to-compute dependencies. The work is complete when the identified GPU communication-to-compute dependency is added to the appropriate node and the resulting kineto trace reflects the dependency.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
distributed-systems, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.