deepspeedai / deepspeedai/DeepSpeed
[BUG] Pipeline parallel code breaks with sequence_parallel enabled
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 43.1k
- Forks
- 5k
- Avg merge
- 4d 15h
- Merged PRs (30d)
- 112
Description
Describe the bug
In Deepspeed pipeline parallel mode, activations are not send directly, instead it's split to chunks as PartitionedTensor, then allgathered into a full tensor.
It worked correctly, but if Megatron Deepspeed sequence parallel is enabled, the output activation tensor is already splited. It doesn't make sense to use PartitionTensor. Using PartitionTensor.full() will result in each tp rank worker recv same activation, and the tensor actually is noise.
To Reproduce
Combine Zero1 + TP + PP + SP.
Screenshots
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start in the DeepSpeed pipeline-parallel path that exchanges activations as PartitionedTensor, then reproduce the Zero1 + TP + PP + SP combination described in the issue. Compare the sequence-parallel output with the PartitionedTensor handling and verify that each tensor-parallel rank receives the correct activation rather than a noisy allgathered tensor.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- distributed-systems, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 28/100