NVIDIA / NVIDIA/TransformerEngine

did not get improvement from tp/sp overlap

Open
#1,587 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug
Dominant language
Python
Stars
3.5k
Forks
831
Avg merge
3d 11h
Merged PRs (30d)
65

Description

i run examples of te_comm_gemm_overlap.py, and remove backwards code, only forward code.

compared to tp allreduce, the tp/sp allgather + reduce scatter is slower

my gpu is 8 h20

is there any other args needs to change?

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with te_comm_gemm_overlap.py and run the forward-only comparison between TP allreduce and TP/SP allgather plus reduce-scatter on the reported 8 H20 GPUs. Inspect the example's available arguments and determine whether a configuration change explains the slower result; done means identifying the relevant setting or documenting why the overlap does not improve performance.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
distributed-systems, machine-learning, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.