deepspeedai / deepspeedai/DeepSpeed

[BUG] Pipeline parallel code breaks with sequence_parallel enabled

Open
#4,196 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug training
Dominant language
Python
Stars
43.1k
Forks
5k
Avg merge
4d 15h
Merged PRs (30d)
112

Description

Describe the bug
In Deepspeed pipeline parallel mode, activations are not send directly, instead it's split to chunks as PartitionedTensor, then allgathered into a full tensor.

It worked correctly, but if Megatron Deepspeed sequence parallel is enabled, the output activation tensor is already splited. It doesn't make sense to use PartitionTensor. Using PartitionTensor.full() will result in each tp rank worker recv same activation, and the tensor actually is noise.

To Reproduce
Combine Zero1 + TP + PP + SP.

Screenshots
image

image

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in the DeepSpeed pipeline-parallel path that exchanges activations as PartitionedTensor, then reproduce the Zero1 + TP + PP + SP combination described in the issue. Compare the sequence-parallel output with the PartitionedTensor handling and verify that each tensor-parallel rank receives the correct activation rather than a noisy allgathered tensor.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
distributed-systems, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
28/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.