NVIDIA / NVIDIA/Megatron-LM

[BUG] p2p communication order error and stuck when pp 2 and vpp 2 with remove pad

Open
#1,450 2 comments 0 reactions 0 assignees View on GitHub
bug community-request module: distributed waiting-on-maintainers
Dominant language
Python
Stars
17.9k
Forks
4.5k
Avg merge
4d 3h
Merged PRs (30d)
272

Description

**Describe the bug**
p2p communication order error and stuck when pp 2 and vpp 2 with remove pad

**To Reproduce**
When use `PP=2` and `VPP=2` with `config.variable_seq_lengths=True`, `config.batch_p2p_comm=True` and `config.overlap_p2p_comm=False`, current implementation of p2p_communication.py will cause incorrect behavior.

If we set `config.overlap_p2p_comm=True` and `config.batch_p2p_comm=False`, bug disappear.

You can use [verl](https://github.com/volcengine/verl/blob/main/examples/ppo_trainer/run_deepseek_megatron.sh) to reproduce this issues, you should set actor's/critic's `pipeline_model_parallel_size=2` and `virtual_pipeline_model_parallel_size=1`, and delete`config.overlap_p2p_comm` and `config.batch_p2p_comm` in `verl/utils/megatron_utils.py` to use original Megatron-LM configuration.

**Expected behavior**
Like this image below:

![vpp](https://github.com/user-attachments/assets/5150685a-a7e9-4632-a8ca-49ad82490fad)

After 2 devices finish at the dashed time, Device 1 should pass `output_tensor` and `input_tensor_grad` to Device 2, and because world size is 2, both devices have the same `next_rank` and `prev_rank`, the original ring communication becomes intercommunication, thus cause conflicts in p2p_communication. In detail, Device 1 passes `output_tensor` to `next_rank` and `input_tensor_grad` to `prev_rank`, and Device 2 receives `output_tensor_grad` from `next_rank` and `input_tensor` from `prev_rank`.

**Stack trace/logs**

Here is more log:

```txt
# Device 0
send_prev_shape_tensor: torch.Tensor([1673, 1, 3840], device='cuda:0'), send_next_shape_tensor: torch.Tensor([1702, 1, 3840], device='cuda:0')
recv_prev_shape_tensor: torch.Tensor([], device='cuda:0'), recv_next_shape_tensor: torch.Tensor([1664, 1, 3840], device='cuda:0')

# Device 1
send_prev_shape_tensor: torch.Tensor([1664, 1, 3840], device='cuda:0'), send_next_shape_tensor: torch.Tensor([1653, 1, 3840], device='cuda:0')
recv_prev_shape_tensor: torch.Tensor([1673, 1, 3840], device='cuda:0'), recv_next_shape_tensor: torch.Tensor([1702, 1, 3840], device='cuda:0') # Reverse Error
```

**Environment (please complete the following information):**
- Megatron-LM core_r0.11.0
- PyTorch 2.4.0
- CUDA 12.4
- NCCL 2.20.5

**Proposed fix**
PR see #1451

**Additional context**
Add any other context about the problem here.

Contributor guide

Open the contributing guide

Research direction

Start by reproducing the PP=2, VPP=2 case from the issue with variable sequence lengths, batched P2P communication enabled, and overlap disabled. Read p2p_communication.py and compare the send/receive ordering against the expected two-device intercommunication described here. Done means communication no longer reverses shapes or stalls, and the reported configuration completes successfully.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
distributed-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Clearly specified
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.