[Bug] qwen35 running with pp error
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 8.5k
- Forks
- 1.3k
- Avg merge
- 5h 36m
- Merged PRs (30d)
- 22
Description
Bug Description
File ".../slime/slime/backends/megatron_utils/actor.py", line 368, in train
return self.train_actor(rollout_id, rollout_data)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File ".../slime/slime/backends/megatron_utils/actor.py", line 440, in train_actor
self.compute_log_prob(
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File ".../slime/slime/backends/megatron_utils/actor.py", line 346, in compute_log_prob
return forward_only(
^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/dist-packages/torch/utils/_contextlib.py", line 120, in decorate_context
return func(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^
File ".../slime/slime/backends/megatron_utils/model.py", line 263, in forward_only
forward_data_store += forward_backward_func(
^^^^^^^^^^^^^^^^^^^^^^
File "/root/Megatron-LM/megatron/core/pipeline_parallel/schedules.py", line 2221, in forward_backward_pipelining_without_interleaving
output_tensor, num_tokens = forward_step(
^^^^^^^^^^^^^
File "/root/Megatron-LM/megatron/core/pipeline_parallel/schedules.py", line 417, in forward_step
output_tensor, loss_func = forward_step_func(data_iterator, model)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File ".../slime/slime/backends/megatron_utils/model.py", line 227, in forward_step
output_tensor = model(
^^^^^^
File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1775, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1786, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/Megatron-LM/megatron/core/distributed/data_parallel_base.py", line 22, in forward
return self.module(*inputs, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1775, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1786, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/Megatron-LM/megatron/core/transformer/module.py", line 456, in forward
outputs = self.module(*inputs, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1775, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1786, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File ".../Megatron-Bridge-slime/src/megatron/bridge/models/qwen_vl/modelling_qwen3_vl/model.py", line 484, in forward
output = self.language_model(
^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1775, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1786, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File ".../Megatron-Bridge-slime/src/megatron/bridge/models/qwen_vl/modelling_qwen3_vl/text_model.py", line 161, in forward
hidden_states = self.decoder(
^^^^^^^^^^^^^
File "/root/Megatron-LM/megatron/core/transformer/transformer_block.py", line 586, in __call__
return super().__call__(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/Megatron-LM/megatron/core/transformer/module.py", line 319, in __call__
return super().__call__(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1775, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1786, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File ".../Megatron-Bridge-slime/src/megatron/bridge/models/qwen_vl/modelling_qwen3_vl/transformer_block.py", line 721, in forward
hidden_states, context = layer(
^^^^^^
File "/root/Megatron-LM/megatron/core/transformer/transformer_layer.py", line 1044, in __call__
return super().__call__(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/Megatron-LM/megatron/core/transformer/module.py", line 319, in __call__
return super().__call__(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1775, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1786, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/Megatron-LM/megatron/core/transformer/transformer_layer.py", line 475, in forward
hidden_states, context = self._forward_attention(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/Megatron-LM/megatron/core/transformer/transformer_layer.py", line 549, in _forward_attention
attention_output_with_bias = self.self_attention(
^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1775, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1786, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File ".../Megatron-Bridge-slime/src/megatron/bridge/models/qwen_vl/modelling_qwen3_vl/attention.py", line 187, in forward
query = apply_rotary_pos_emb_absolute(
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File ".../Megatron-Bridge-slime/src/megatron/bridge/models/qwen_vl/modelling_qwen3_vl/rope.py", line 324, in apply_rotary_pos_emb_absolute
result = _apply_rotary_pos_emb_bshd(t, freqs, rotary_interleaved=config.rotary_interleaved)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/Megatron-LM/megatron/core/models/common/embeddings/rope_utils.py", line 125, in _apply_rotary_pos_emb_bshd
t = (t * cos_) + (_rotate_half(t, rotary_interleaved) * sin_)
~~^~~~~~
RuntimeError: The size of tensor a (4096) must match the size of tensor b (1280) at non-singleton dimension 0
Steps to Reproduce
examples/geo3k_vlm/run_geo3k_qwen35.sh
with image dataset
with setting:
--tensor-model-parallel-size 2
--sequence-parallel
--pipeline-model-parallel-size 2
--context-parallel-size 1
--expert-model-parallel-size 4
--expert-tensor-parallel-size 1
Expected Behavior
no
Actual Behavior
no
Environment
- slime version:
- Python version:
- PyTorch version:
- CUDA/ROCm version:
- GPU type and count:
- OS:
- SGLang version (if relevant):
- Megatron-LM version (if relevant):
Logs
Additional Context
No response
Pre-submission Checklist
- I have read the CONTRIBUTING.md and understand the collaboration scope.
- I have read the documentation and my issue is not addressed there.
- I have searched for existing issues and this is not a duplicate.
- I have provided a minimal, reproducible example.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reproducing examples/geo3k_vlm/run_geo3k_qwen35.sh with the image dataset and the listed parallelism settings. Trace the reported size mismatch through slime/backends/megatron_utils/model.py and the Qwen files under Megatron-Bridge-slime/src/megatron/bridge/models/qwen_vl/modelling_qwen3_vl/, especially rope.py and attention.py. Done means the run no longer fails in apply_rotary_pos_emb_absolute with those settings.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- distributed-systems, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100