modelscope / modelscope/ms-swift
Qwen3-VL-235B-A22B sft Error
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 15.7k
- Forks
- 1.7k
- Avg merge
- 1d 16h
- Merged PRs (30d)
- 136
Description
执行命令:
PYTORCH_CUDA_ALLOC_CONF='expandable_segments:True'
NPROC_PER_NODE=4
CUDA_VISIBLE_DEVICES=4,5,6,7
megatron sft
--load /data/workspacemodels_weight/Qwen3-VL-235B-A22B-Instruct-mcoreV2
--dataset /data/workspace/code/ms-swift/train_datasets/sft.jsonl
--split_dataset_ratio 0.01
--train_type lora
--lora_rank 8
--lora_alpha 32
--target_modules all-linear
--moe_permute_fusion true
--tensor_model_parallel_size 2
--expert_tensor_parallel_size 1
--expert_model_parallel_size 4
--moe_grouped_gemm true
--moe_shared_expert_overlap true
--moe_aux_loss_coeff 1e-6
--micro_batch_size 1
--global_batch_size 4
--recompute_granularity full
--recompute_method uniform
--recompute_num_layers 1
--max_epochs 1
--finetune true
--cross_entropy_loss_fusion true
--lr 1e-4
--lr_warmup_fraction 0.05
--min_lr 1e-5
--save megatron_output/Qwen3-VL-235B-A22B-Instruct
--eval_interval 200
--save_interval 200
--max_length 12288
--packing true
--num_workers 8
--dataset_num_proc 4
--no_save_optim true
--no_save_rng true
--sequence_parallel true
--attention_backend flash
日志如下:
00, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, -100, 34528, 65330, 264, 19082, 52142, 2337, 279, 61082, 11, 264, 1697, 12233, 264, 3691, 27215, 291, 27466, 48372, 323, 6303, 33289, 374, 3884, 1790, 311, 264, 3691, 29586, 10855, 448, 279, 4065, 23148, 6006, 1787, 13, 58556, 11, 279, 1697, 374, 11259, 389, 279, 10855, 594, 4303, 4479, 11, 1667, 264, 8061, 28202, 9099, 389, 279, 4910, 311, 4240, 279, 7310, 594, 14791, 13, 66487, 1283, 11, 807, 3019, 1495, 504, 279, 10855, 13, 1752, 279, 26313, 315, 279, 8500, 11, 279, 1697, 13352, 304, 279, 52142, 11, 62514, 279, 28202, 37500, 323, 1181, 33679, 11, 518, 825, 1459, 57118, 1495, 26753, 323, 2937, 40119, 264, 4478, 389, 279, 28202, 646, 1571, 13, 362, 6176, 7310, 374, 9434, 304, 279, 4004, 369, 949, 315, 279, 8500, 3918, 2147, 1784, 9217, 36721, 74349, 15012, 522, 9217, 29, 151645]
[INFO:swift] [LABELS] [-100 * 1431]In a rural driveway during the daytime, a person wearing a black hooded sweatshirt and blue jeans is seen next to a black pickup truck with the front passenger door open. Initially, the person is standing on the truck's running board, using a shop vacuum placed on the ground to clean the vehicle's interior. Shortly after, they step down from the truck. For the remainder of the sequence, the person stands in the driveway, manipulating the vacuum hose and its attachments, at one point bending down briefly and later resting a foot on the vacuum canister. A green vehicle is visible in the background for part of the sequence.B,O,a<|im_end|>
[INFO:swift] Dataset Token Length: 9496.770956±578.689826, min=4285.000000, max=10480.000000, size=847
[INFO:swift] Dataset Token Length: 9031.444444±720.746852, min=7240.000000, max=10216.000000, size=9
[INFO:swift] padding_to: 2
using world size: 4, data-parallel size: 2, context-parallel size: 1, hierarchical context-parallel sizes: None, tensor-model-parallel size: 2, encoder-tensor-model-parallel size: 0, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0
Number of virtual stages per pipeline stage: None
accumulate and all-reduce gradients in fp32 for bfloat16 data type.
using torch.bfloat16 for parameters ...
Loading: 100%|█████████████████████| 16617/16617 [00:18<00:00, 908.50it/s]
Loading: 100%|█████████████████████| 16617/16617 [00:18<00:00, 903.91it/s]
could not find arguments in the checkpoint ...
checkpoint version 3.0
WARNING:megatron.core.rerun_state_machine:RerunStateMachine disabled via CLI, ignoring machine state saved in checkpoint
successfully loaded checkpoint from /data/workspace/models_weight/Qwen3-VL-235B-A22B-Instruct-mcoreV2 [ t 1/2, p 1/1 ] at iteration 0
(min, max) time across ranks (ms):
load-checkpoint ................................: (199591.49, 199591.54)
[after model, optimizer, and learning rate scheduler are built] datetime: 2025-09-26 14:47:16
building train, validation, and test datasets ...
datasets target sizes (minimum size):
train: 844
validation: 16
test: 8
[after dataloaders are built] datetime: 2025-09-26 14:47:16
done with setup ...
(min, max) time across ranks (ms):
model-and-optimizer-setup ......................: (213294.60, 213295.14)
train/valid/test-data-iterators-setup ..........: (0.67, 0.73)
training ...
Setting rerun_state_machine.current_iteration to 0...
[before the start of training step] datetime: 2025-09-26 14:47:16
[INFO:swift] The training of Epoch 0 starts...
WARNING:megatron.core.utils:Caught IndexError in DataLoader worker process 0.
Original Traceback (most recent call last):
File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/worker.py", line 349, in _worker_loop
data = fetcher.fetch(index) # type: ignore[possibly-undefined]
^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/fetch.py", line 55, in fetch
return self.collate_fn(data)
^^^^^^^^^^^^^^^^^^^^^
File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1390, in data_collator
res = self._data_collator(batch, padding_to=padding_to)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 411, in _data_collator
res = super()._data_collator(batch, padding_to=padding_to)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1592, in _data_collator
batch[:] = [self.packing_row(batch)]
^^^^^^^^^^^^^^^^^^^^^^^
File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 384, in packing_row
position_ids.append(self._get_position_ids(r))
^^^^^^^^^^^^^^^^^^^^^^^^^
File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 402, in _get_position_ids
position_ids, _ = get_rope_index(
^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/site-packages/transformers/models/qwen3_vl_moe/modeling_qwen3_vl_moe.py", line 1092, in get_rope_index
input_ids = input_ids[attention_mask[i] == 1]
~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^
IndexError: The shape of the mask [1042] at index 0 does not match the shape of the indexed tensor [1518] at index 0
['Traceback (most recent call last):\n', ' File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/trainer.py", line 141, in forward_step\n data = get_batch(data_iterator)\n ^^^^^^^^^^^^^^^^^^^^^^^^\n', ' File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/utils.py", line 123, in get_batch\n batch = get_batch_on_this_tp_rank(data_iterator)\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n', ' File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/utils.py", line 33, in get_batch_on_this_tp_rank\n data = next(data_iterator)\n ^^^^^^^^^^^^^^^^^^^\n', ' File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/core/rerun_state_machine.py", line 1025, in next\n n: Any = next(self.iterable)\n ^^^^^^^^^^^^^^^^^^^\n', ' File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 104, in new_cyclic_iter\n x = [next(it) for _ in range(num_microbatches - n_batch % num_microbatches)]\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n', ' File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 104, in \n x = [next(it) for _ in range(num_microbatches - n_batch % num_microbatches)]\n ^^^^^^^^\n', ' File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 733, in next\n data = self._next_data()\n ^^^^^^^^^^^^^^^^^\n', ' File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 1515, in _next_data\n return self._process_data(data, worker_id)\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n', ' File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 1550, in _process_data\n data.reraise()\n', ' File "/usr/local/lib/python3.11/site-packages/torch/_utils.py", line 750, in reraise\n raise exception\n', 'IndexError: Caught IndexError in DataLoader worker process 0.\nOriginal Traceback (most recent call last):\n File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/worker.py", line 349, in _worker_loop\n data = fetcher.fetch(index) # type: ignore[possibly-undefined]\n ^^^^^^^^^^^^^^^^^^^^\n File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/fetch.py", line 55, in fetch\n return self.collate_fn(data)\n ^^^^^^^^^^^^^^^^^^^^^\n File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1390, in data_collator\n res = self._data_collator(batch, padding_to=padding_to)\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 411, in _data_collator\n res = super()._data_collator(batch, padding_to=padding_to)\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1592, in _data_collator\n batch[:] = [self.packing_row(batch)]\n ^^^^^^^^^^^^^^^^^^^^^^^\n File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 384, in packing_row\n position_ids.append(self._get_position_ids(r))\n ^^^^^^^^^^^^^^^^^^^^^^^^^\n File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 402, in _get_position_ids\n position_ids, _ = get_rope_index(\n ^^^^^^^^^^^^^^^\n File "/usr/local/lib/python3.11/site-packages/transformers/models/qwen3_vl_moe/modeling_qwen3_vl_moe.py", line 1092, in get_rope_index\n input_ids = input_ids[attention_mask[i] == 1]\n ~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^\nIndexError: The shape of the mask [1042] at index 0 does not match the shape of the indexed tensor [1518] at index 0\n\n']
[INFO:swift] images_dir: /data/workspace/models_weight/Qwen3-VL-235B-A22B-Instruct-mcoreV2/megatron_output/Qwen3-VL-235B-A22B-Instruct/v1-20250926-142332/images
WARNING:megatron.core.utils:Caught IndexError in DataLoader worker process 0.
Original Traceback (most recent call last):
File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/worker.py", line 349, in _worker_loop
data = fetcher.fetch(index) # type: ignore[possibly-undefined]
^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/fetch.py", line 55, in fetch
return self.collate_fn(data)
^^^^^^^^^^^^^^^^^^^^^
File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1390, in data_collator
res = self._data_collator(batch, padding_to=padding_to)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 411, in _data_collator
res = super()._data_collator(batch, padding_to=padding_to)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1592, in _data_collator
batch[:] = [self.packing_row(batch)]
^^^^^^^^^^^^^^^^^^^^^^^
File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 384, in packing_row
position_ids.append(self._get_position_ids(r))
^^^^^^^^^^^^^^^^^^^^^^^^^
File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 402, in _get_position_ids
position_ids, _ = get_rope_index(
^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/site-packages/transformers/models/qwen3_vl_moe/modeling_qwen3_vl_moe.py", line 1092, in get_rope_index
input_ids = input_ids[attention_mask[i] == 1]
~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^
IndexError: The shape of the mask [1042] at index 0 does not match the shape of the indexed tensor [1518] at index 0
['Traceback (most recent call last):\n', ' File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/trainer.py", line 141, in forward_step\n data = get_batch(data_iterator)\n ^^^^^^^^^^^^^^^^^^^^^^^^\n', ' File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/utils.py", line 123, in get_batch\n batch = get_batch_on_this_tp_rank(data_iterator)\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n', ' File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/utils.py", line 33, in get_batch_on_this_tp_rank\n data = next(data_iterator)\n ^^^^^^^^^^^^^^^^^^^\n', ' File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/core/rerun_state_machine.py", line 1025, in next\n n: Any = next(self.iterable)\n ^^^^^^^^^^^^^^^^^^^\n', ' File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 104, in new_cyclic_iter\n x = [next(it) for _ in range(num_microbatches - n_batch % num_microbatches)]\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n', ' File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 104, in \n x = [next(it) for _ in range(num_microbatches - n_batch % num_microbatches)]\n ^^^^^^^^\n', ' File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 733, in next\n data = self._next_data()\n ^^^^^^^^^^^^^^^^^\n', ' File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 1515, in _next_data\n return self._process_data(data, worker_id)\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n', ' File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 1550, in _process_data\n data.reraise()\n', ' File "/usr/local/lib/python3.11/site-packages/torch/_utils.py", line 750, in reraise\n raise exception\n', 'IndexError: Caught IndexError in DataLoader worker process 0.\nOriginal Traceback (most recent call last):\n File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/worker.py", line 349, in _worker_loop\n data = fetcher.fetch(index) # type: ignore[possibly-undefined]\n ^^^^^^^^^^^^^^^^^^^^\n File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/fetch.py", line 55, in fetch\n return self.collate_fn(data)\n ^^^^^^^^^^^^^^^^^^^^^\n File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1390, in data_collator\n res = self._data_collator(batch, padding_to=padding_to)\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 411, in _data_collator\n res = super()._data_collator(batch, padding_to=padding_to)\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1592, in _data_collator\n batch[:] = [self.packing_row(batch)]\n ^^^^^^^^^^^^^^^^^^^^^^^\n File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 384, in packing_row\n position_ids.append(self._get_position_ids(r))\n ^^^^^^^^^^^^^^^^^^^^^^^^^\n File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 402, in _get_position_ids\n position_ids, _ = get_rope_index(\n ^^^^^^^^^^^^^^^\n File "/usr/local/lib/python3.11/site-packages/transformers/models/qwen3_vl_moe/modeling_qwen3_vl_moe.py", line 1092, in get_rope_index\n input_ids = input_ids[attention_mask[i] == 1]\n ~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^\nIndexError: The shape of the mask [1042] at index 0 does not match the shape of the indexed tensor [1518] at index 0\n\n']
[rank1]: Traceback (most recent call last):
[rank1]: File "/data/workspace/vlm/ms-swift/swift/cli/_megatron/sft.py", line 5, in
[rank1]: megatron_sft_main()
[rank1]: File "/data/workspace/vlm/ms-swift/swift/megatron/train/sft.py", line 79, in megatron_sft_main
[rank1]: return MegatronSft(args).main()
[rank1]: ^^^^^^^^^^^^^^^^^^^^^^^^
[rank1]: File "/data/workspace/vlm/ms-swift/swift/llm/base.py", line 49, in main
[rank1]: result = self.run()
[rank1]: ^^^^^^^^^^
[rank1]: File "/data/workspace/vlm/ms-swift/swift/megatron/train/sft.py", line 69, in run
[rank1]: self.trainer.train(train_dataset, val_dataset, data_collator)
[rank1]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 758, in train
[rank1]: pretrain(
[rank1]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/training/training.py", line 864, in pretrain
[rank1]: iteration, num_floating_point_operations_so_far = train(
[rank1]: ^^^^^^
[rank1]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/training/training.py", line 2279, in train
[rank1]: ) = train_step(
[rank1]: ^^^^^^^^^^^
[rank1]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 311, in train_step
[rank1]: return self._origin_train_step(forward_step_func, new_data_iterator, model, optimizer, opt_param_scheduler,
[rank1]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank1]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/training/training.py", line 1395, in train_step
[rank1]: losses_reduced = forward_backward_func(
[rank1]: ^^^^^^^^^^^^^^^^^^^^^^
[rank1]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/core/pipeline_parallel/schedules.py", line 500, in forward_backward_no_pipelining
[rank1]: output_tensor, num_tokens = forward_step(
[rank1]: ^^^^^^^^^^^^^
[rank1]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/core/pipeline_parallel/schedules.py", line 289, in forward_step
[rank1]: output_tensor, loss_func = forward_step_func(data_iterator, model)
[rank1]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank1]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/trainer.py", line 141, in forward_step
[rank1]: data = get_batch(data_iterator)
[rank1]: ^^^^^^^^^^^^^^^^^^^^^^^^
[rank1]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/utils.py", line 123, in get_batch
[rank1]: batch = get_batch_on_this_tp_rank(data_iterator)
[rank1]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank1]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/utils.py", line 33, in get_batch_on_this_tp_rank
[rank1]: data = next(data_iterator)
[rank1]: ^^^^^^^^^^^^^^^^^^^
[rank1]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/core/rerun_state_machine.py", line 1025, in next
[rank1]: n: Any = next(self.iterable)
[rank1]: ^^^^^^^^^^^^^^^^^^^
[rank1]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 104, in new_cyclic_iter
[rank1]: x = [next(it) for _ in range(num_microbatches - n_batch % num_microbatches)]
[rank1]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank1]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 104, in
[rank1]: x = [next(it) for _ in range(num_microbatches - n_batch % num_microbatches)]
[rank1]: ^^^^^^^^
[rank1]: File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 733, in next
[rank1]: data = self._next_data()
[rank1]: ^^^^^^^^^^^^^^^^^
[rank1]: File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 1515, in _next_data
[rank1]: return self._process_data(data, worker_id)
[rank1]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank1]: File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 1550, in _process_data
[rank1]: data.reraise()
[rank1]: File "/usr/local/lib/python3.11/site-packages/torch/_utils.py", line 750, in reraise
[rank1]: raise exception
[rank1]: IndexError: Caught IndexError in DataLoader worker process 0.
[rank1]: Original Traceback (most recent call last):
[rank1]: File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/worker.py", line 349, in _worker_loop
[rank1]: data = fetcher.fetch(index) # type: ignore[possibly-undefined]
[rank1]: ^^^^^^^^^^^^^^^^^^^^
[rank1]: File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/fetch.py", line 55, in fetch
[rank1]: return self.collate_fn(data)
[rank1]: ^^^^^^^^^^^^^^^^^^^^^
[rank1]: File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1390, in data_collator
[rank1]: res = self._data_collator(batch, padding_to=padding_to)
[rank1]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank1]: File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 411, in _data_collator
[rank1]: res = super()._data_collator(batch, padding_to=padding_to)
[rank1]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank1]: File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1592, in _data_collator
[rank1]: batch[:] = [self.packing_row(batch)]
[rank1]: ^^^^^^^^^^^^^^^^^^^^^^^
[rank1]: File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 384, in packing_row
[rank1]: position_ids.append(self._get_position_ids(r))
[rank1]: ^^^^^^^^^^^^^^^^^^^^^^^^^
[rank1]: File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 402, in _get_position_ids
[rank1]: position_ids, _ = get_rope_index(
[rank1]: ^^^^^^^^^^^^^^^
[rank1]: File "/usr/local/lib/python3.11/site-packages/transformers/models/qwen3_vl_moe/modeling_qwen3_vl_moe.py", line 1092, in get_rope_index
[rank1]: input_ids = input_ids[attention_mask[i] == 1]
[rank1]: ~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^
[rank1]: IndexError: The shape of the mask [1042] at index 0 does not match the shape of the indexed tensor [1518] at index 0
[rank0]: Traceback (most recent call last):
[rank0]: File "/data/workspace/vlm/ms-swift/swift/cli/_megatron/sft.py", line 5, in
[rank0]: megatron_sft_main()
[rank0]: File "/data/workspace/vlm/ms-swift/swift/megatron/train/sft.py", line 79, in megatron_sft_main
[rank0]: return MegatronSft(args).main()
[rank0]: ^^^^^^^^^^^^^^^^^^^^^^^^
[rank0]: File "/data/workspace/vlm/ms-swift/swift/llm/base.py", line 49, in main
[rank0]: result = self.run()
[rank0]: ^^^^^^^^^^
[rank0]: File "/data/workspace/vlm/ms-swift/swift/megatron/train/sft.py", line 69, in run
[rank0]: self.trainer.train(train_dataset, val_dataset, data_collator)
[rank0]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 758, in train
[rank0]: pretrain(
[rank0]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/training/training.py", line 864, in pretrain
[rank0]: iteration, num_floating_point_operations_so_far = train(
[rank0]: ^^^^^^
[rank0]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/training/training.py", line 2279, in train
[rank0]: ) = train_step(
[rank0]: ^^^^^^^^^^^
[rank0]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 311, in train_step
[rank0]: return self._origin_train_step(forward_step_func, new_data_iterator, model, optimizer, opt_param_scheduler,
[rank0]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank0]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/training/training.py", line 1395, in train_step
[rank0]: losses_reduced = forward_backward_func(
[rank0]: ^^^^^^^^^^^^^^^^^^^^^^
[rank0]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/core/pipeline_parallel/schedules.py", line 500, in forward_backward_no_pipelining
[rank0]: output_tensor, num_tokens = forward_step(
[rank0]: ^^^^^^^^^^^^^
[rank0]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/core/pipeline_parallel/schedules.py", line 289, in forward_step
[rank0]: output_tensor, loss_func = forward_step_func(data_iterator, model)
[rank0]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank0]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/trainer.py", line 141, in forward_step
[rank0]: data = get_batch(data_iterator)
[rank0]: ^^^^^^^^^^^^^^^^^^^^^^^^
[rank0]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/utils.py", line 123, in get_batch
[rank0]: batch = get_batch_on_this_tp_rank(data_iterator)
[rank0]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank0]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/utils.py", line 33, in get_batch_on_this_tp_rank
[rank0]: data = next(data_iterator)
[rank0]: ^^^^^^^^^^^^^^^^^^^
[rank0]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/core/rerun_state_machine.py", line 1025, in next
[rank0]: n: Any = next(self.iterable)
[rank0]: ^^^^^^^^^^^^^^^^^^^
[rank0]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 104, in new_cyclic_iter
[rank0]: x = [next(it) for _ in range(num_microbatches - n_batch % num_microbatches)]
[rank0]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank0]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 104, in
[rank0]: x = [next(it) for _ in range(num_microbatches - n_batch % num_microbatches)]
[rank0]: ^^^^^^^^
[rank0]: File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 733, in next
[rank0]: data = self._next_data()
[rank0]: ^^^^^^^^^^^^^^^^^
[rank0]: File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 1515, in _next_data
[rank0]: return self._process_data(data, worker_id)
[rank0]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank0]: File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 1550, in _process_data
[rank0]: data.reraise()
[rank0]: File "/usr/local/lib/python3.11/site-packages/torch/_utils.py", line 750, in reraise
[rank0]: raise exception
[rank0]: IndexError: Caught IndexError in DataLoader worker process 0.
[rank0]: Original Traceback (most recent call last):
[rank0]: File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/worker.py", line 349, in _worker_loop
[rank0]: data = fetcher.fetch(index) # type: ignore[possibly-undefined]
[rank0]: ^^^^^^^^^^^^^^^^^^^^
[rank0]: File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/fetch.py", line 55, in fetch
[rank0]: return self.collate_fn(data)
[rank0]: ^^^^^^^^^^^^^^^^^^^^^
[rank0]: File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1390, in data_collator
[rank0]: res = self._data_collator(batch, padding_to=padding_to)
[rank0]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank0]: File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 411, in _data_collator
[rank0]: res = super()._data_collator(batch, padding_to=padding_to)
[rank0]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank0]: File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1592, in _data_collator
[rank0]: batch[:] = [self.packing_row(batch)]
[rank0]: ^^^^^^^^^^^^^^^^^^^^^^^
[rank0]: File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 384, in packing_row
[rank0]: position_ids.append(self._get_position_ids(r))
[rank0]: ^^^^^^^^^^^^^^^^^^^^^^^^^
[rank0]: File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 402, in _get_position_ids
[rank0]: position_ids, _ = get_rope_index(
[rank0]: ^^^^^^^^^^^^^^^
[rank0]: File "/usr/local/lib/python3.11/site-packages/transformers/models/qwen3_vl_moe/modeling_qwen3_vl_moe.py", line 1092, in get_rope_index
[rank0]: input_ids = input_ids[attention_mask[i] == 1]
[rank0]: ~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^
[rank0]: IndexError: The shape of the mask [1042] at index 0 does not match the shape of the indexed tensor [1518] at index 0
WARNING:megatron.core.utils:Caught IndexError in DataLoader worker process 0.
Original Traceback (most recent call last):
File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/worker.py", line 349, in _worker_loop
data = fetcher.fetch(index) # type: ignore[possibly-undefined]
^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/fetch.py", line 55, in fetch
return self.collate_fn(data)
^^^^^^^^^^^^^^^^^^^^^
File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1390, in data_collator
res = self._data_collator(batch, padding_to=padding_to)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 411, in _data_collator
res = super()._data_collator(batch, padding_to=padding_to)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1592, in _data_collator
batch[:] = [self.packing_row(batch)]
^^^^^^^^^^^^^^^^^^^^^^^
File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 384, in packing_row
position_ids.append(self._get_position_ids(r))
^^^^^^^^^^^^^^^^^^^^^^^^^
File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 402, in _get_position_ids
position_ids, _ = get_rope_index(
^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/site-packages/transformers/models/qwen3_vl_moe/modeling_qwen3_vl_moe.py", line 1092, in get_rope_index
input_ids = input_ids[attention_mask[i] == 1]
~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^
IndexError: The shape of the mask [1040] at index 0 does not match the shape of the indexed tensor [1521] at index 0
['Traceback (most recent call last):\n', ' File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/trainer.py", line 141, in forward_step\n data = get_batch(data_iterator)\n ^^^^^^^^^^^^^^^^^^^^^^^^\n', ' File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/utils.py", line 123, in get_batch\n batch = get_batch_on_this_tp_rank(data_iterator)\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n', ' File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/utils.py", line 33, in get_batch_on_this_tp_rank\n data = next(data_iterator)\n ^^^^^^^^^^^^^^^^^^^\n', ' File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/core/rerun_state_machine.py", line 1025, in next\n n: Any = next(self.iterable)\n ^^^^^^^^^^^^^^^^^^^\n', ' File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 104, in new_cyclic_iter\n x = [next(it) for _ in range(num_microbatches - n_batch % num_microbatches)]\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n', ' File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 104, in \n x = [next(it) for _ in range(num_microbatches - n_batch % num_microbatches)]\n ^^^^^^^^\n', ' File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 733, in next\n data = self._next_data()\n ^^^^^^^^^^^^^^^^^\n', ' File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 1515, in _next_data\n return self._process_data(data, worker_id)\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n', ' File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 1550, in _process_data\n data.reraise()\n', ' File "/usr/local/lib/python3.11/site-packages/torch/_utils.py", line 750, in reraise\n raise exception\n', 'IndexError: Caught IndexError in DataLoader worker process 0.\nOriginal Traceback (most recent call last):\n File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/worker.py", line 349, in _worker_loop\n data = fetcher.fetch(index) # type: ignore[possibly-undefined]\n ^^^^^^^^^^^^^^^^^^^^\n File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/fetch.py", line 55, in fetch\n return self.collate_fn(data)\n ^^^^^^^^^^^^^^^^^^^^^\n File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1390, in data_collator\n res = self._data_collator(batch, padding_to=padding_to)\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 411, in _data_collator\n res = super()._data_collator(batch, padding_to=padding_to)\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1592, in _data_collator\n batch[:] = [self.packing_row(batch)]\n ^^^^^^^^^^^^^^^^^^^^^^^\n File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 384, in packing_row\n position_ids.append(self._get_position_ids(r))\n ^^^^^^^^^^^^^^^^^^^^^^^^^\n File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 402, in _get_position_ids\n position_ids, _ = get_rope_index(\n ^^^^^^^^^^^^^^^\n File "/usr/local/lib/python3.11/site-packages/transformers/models/qwen3_vl_moe/modeling_qwen3_vl_moe.py", line 1092, in get_rope_index\n input_ids = input_ids[attention_mask[i] == 1]\n ~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^\nIndexError: The shape of the mask [1040] at index 0 does not match the shape of the indexed tensor [1521] at index 0\n\n']
[rank2]: Traceback (most recent call last):
[rank2]: File "/data/workspace/vlm/ms-swift/swift/cli/_megatron/sft.py", line 5, in
[rank2]: megatron_sft_main()
[rank2]: File "/data/workspace/vlm/ms-swift/swift/megatron/train/sft.py", line 79, in megatron_sft_main
[rank2]: return MegatronSft(args).main()
[rank2]: ^^^^^^^^^^^^^^^^^^^^^^^^
[rank2]: File "/data/workspace/vlm/ms-swift/swift/llm/base.py", line 49, in main
[rank2]: result = self.run()
[rank2]: ^^^^^^^^^^
[rank2]: File "/data/workspace/vlm/ms-swift/swift/megatron/train/sft.py", line 69, in run
[rank2]: self.trainer.train(train_dataset, val_dataset, data_collator)
[rank2]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 758, in train
[rank2]: pretrain(
[rank2]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/training/training.py", line 864, in pretrain
[rank2]: iteration, num_floating_point_operations_so_far = train(
[rank2]: ^^^^^^
[rank2]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/training/training.py", line 2279, in train
[rank2]: ) = train_step(
[rank2]: ^^^^^^^^^^^
[rank2]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 311, in train_step
[rank2]: return self._origin_train_step(forward_step_func, new_data_iterator, model, optimizer, opt_param_scheduler,
[rank2]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank2]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/training/training.py", line 1395, in train_step
[rank2]: losses_reduced = forward_backward_func(
[rank2]: ^^^^^^^^^^^^^^^^^^^^^^
[rank2]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/core/pipeline_parallel/schedules.py", line 500, in forward_backward_no_pipelining
[rank2]: output_tensor, num_tokens = forward_step(
[rank2]: ^^^^^^^^^^^^^
[rank2]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/core/pipeline_parallel/schedules.py", line 289, in forward_step
[rank2]: output_tensor, loss_func = forward_step_func(data_iterator, model)
[rank2]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank2]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/trainer.py", line 141, in forward_step
[rank2]: data = get_batch(data_iterator)
[rank2]: ^^^^^^^^^^^^^^^^^^^^^^^^
[rank2]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/utils.py", line 123, in get_batch
[rank2]: batch = get_batch_on_this_tp_rank(data_iterator)
[rank2]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank2]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/utils.py", line 33, in get_batch_on_this_tp_rank
[rank2]: data = next(data_iterator)
[rank2]: ^^^^^^^^^^^^^^^^^^^
[rank2]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/core/rerun_state_machine.py", line 1025, in next
[rank2]: n: Any = next(self.iterable)
[rank2]: ^^^^^^^^^^^^^^^^^^^
[rank2]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 104, in new_cyclic_iter
[rank2]: x = [next(it) for _ in range(num_microbatches - n_batch % num_microbatches)]
[rank2]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank2]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 104, in
[rank2]: x = [next(it) for _ in range(num_microbatches - n_batch % num_microbatches)]
[rank2]: ^^^^^^^^
[rank2]: File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 733, in next
[rank2]: data = self._next_data()
[rank2]: ^^^^^^^^^^^^^^^^^
[rank2]: File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 1515, in _next_data
[rank2]: return self._process_data(data, worker_id)
[rank2]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank2]: File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 1550, in _process_data
[rank2]: data.reraise()
[rank2]: File "/usr/local/lib/python3.11/site-packages/torch/_utils.py", line 750, in reraise
[rank2]: raise exception
[rank2]: IndexError: Caught IndexError in DataLoader worker process 0.
[rank2]: Original Traceback (most recent call last):
[rank2]: File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/worker.py", line 349, in _worker_loop
[rank2]: data = fetcher.fetch(index) # type: ignore[possibly-undefined]
[rank2]: ^^^^^^^^^^^^^^^^^^^^
[rank2]: File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/fetch.py", line 55, in fetch
[rank2]: return self.collate_fn(data)
[rank2]: ^^^^^^^^^^^^^^^^^^^^^
[rank2]: File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1390, in data_collator
[rank2]: res = self._data_collator(batch, padding_to=padding_to)
[rank2]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank2]: File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 411, in _data_collator
[rank2]: res = super()._data_collator(batch, padding_to=padding_to)
[rank2]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank2]: File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1592, in _data_collator
[rank2]: batch[:] = [self.packing_row(batch)]
[rank2]: ^^^^^^^^^^^^^^^^^^^^^^^
[rank2]: File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 384, in packing_row
[rank2]: position_ids.append(self._get_position_ids(r))
[rank2]: ^^^^^^^^^^^^^^^^^^^^^^^^^
[rank2]: File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 402, in _get_position_ids
[rank2]: position_ids, _ = get_rope_index(
[rank2]: ^^^^^^^^^^^^^^^
[rank2]: File "/usr/local/lib/python3.11/site-packages/transformers/models/qwen3_vl_moe/modeling_qwen3_vl_moe.py", line 1092, in get_rope_index
[rank2]: input_ids = input_ids[attention_mask[i] == 1]
[rank2]: ~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^
[rank2]: IndexError: The shape of the mask [1040] at index 0 does not match the shape of the indexed tensor [1521] at index 0
WARNING:megatron.core.utils:Caught IndexError in DataLoader worker process 0.
Original Traceback (most recent call last):
File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/worker.py", line 349, in _worker_loop
data = fetcher.fetch(index) # type: ignore[possibly-undefined]
^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/fetch.py", line 55, in fetch
return self.collate_fn(data)
^^^^^^^^^^^^^^^^^^^^^
File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1390, in data_collator
res = self._data_collator(batch, padding_to=padding_to)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 411, in _data_collator
res = super()._data_collator(batch, padding_to=padding_to)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1592, in _data_collator
batch[:] = [self.packing_row(batch)]
^^^^^^^^^^^^^^^^^^^^^^^
File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 384, in packing_row
position_ids.append(self._get_position_ids(r))
^^^^^^^^^^^^^^^^^^^^^^^^^
File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 402, in _get_position_ids
position_ids, _ = get_rope_index(
^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/site-packages/transformers/models/qwen3_vl_moe/modeling_qwen3_vl_moe.py", line 1092, in get_rope_index
input_ids = input_ids[attention_mask[i] == 1]
~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^
IndexError: The shape of the mask [1040] at index 0 does not match the shape of the indexed tensor [1521] at index 0
['Traceback (most recent call last):\n', ' File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/trainer.py", line 141, in forward_step\n data = get_batch(data_iterator)\n ^^^^^^^^^^^^^^^^^^^^^^^^\n', ' File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/utils.py", line 123, in get_batch\n batch = get_batch_on_this_tp_rank(data_iterator)\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n', ' File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/utils.py", line 33, in get_batch_on_this_tp_rank\n data = next(data_iterator)\n ^^^^^^^^^^^^^^^^^^^\n', ' File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/core/rerun_state_machine.py", line 1025, in next\n n: Any = next(self.iterable)\n ^^^^^^^^^^^^^^^^^^^\n', ' File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 104, in new_cyclic_iter\n x = [next(it) for _ in range(num_microbatches - n_batch % num_microbatches)]\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n', ' File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 104, in \n x = [next(it) for _ in range(num_microbatches - n_batch % num_microbatches)]\n ^^^^^^^^\n', ' File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 733, in next\n data = self._next_data()\n ^^^^^^^^^^^^^^^^^\n', ' File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 1515, in _next_data\n return self._process_data(data, worker_id)\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n', ' File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 1550, in _process_data\n data.reraise()\n', ' File "/usr/local/lib/python3.11/site-packages/torch/_utils.py", line 750, in reraise\n raise exception\n', 'IndexError: Caught IndexError in DataLoader worker process 0.\nOriginal Traceback (most recent call last):\n File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/worker.py", line 349, in _worker_loop\n data = fetcher.fetch(index) # type: ignore[possibly-undefined]\n ^^^^^^^^^^^^^^^^^^^^\n File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/fetch.py", line 55, in fetch\n return self.collate_fn(data)\n ^^^^^^^^^^^^^^^^^^^^^\n File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1390, in data_collator\n res = self._data_collator(batch, padding_to=padding_to)\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 411, in _data_collator\n res = super()._data_collator(batch, padding_to=padding_to)\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1592, in _data_collator\n batch[:] = [self.packing_row(batch)]\n ^^^^^^^^^^^^^^^^^^^^^^^\n File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 384, in packing_row\n position_ids.append(self._get_position_ids(r))\n ^^^^^^^^^^^^^^^^^^^^^^^^^\n File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 402, in _get_position_ids\n position_ids, _ = get_rope_index(\n ^^^^^^^^^^^^^^^\n File "/usr/local/lib/python3.11/site-packages/transformers/models/qwen3_vl_moe/modeling_qwen3_vl_moe.py", line 1092, in get_rope_index\n input_ids = input_ids[attention_mask[i] == 1]\n ~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^\nIndexError: The shape of the mask [1040] at index 0 does not match the shape of the indexed tensor [1521] at index 0\n\n']
[rank3]: Traceback (most recent call last):
[rank3]: File "/data/workspace/vlm/ms-swift/swift/cli/_megatron/sft.py", line 5, in
[rank3]: megatron_sft_main()
[rank3]: File "/data/workspace/vlm/ms-swift/swift/megatron/train/sft.py", line 79, in megatron_sft_main
[rank3]: return MegatronSft(args).main()
[rank3]: ^^^^^^^^^^^^^^^^^^^^^^^^
[rank3]: File "/data/workspace/vlm/ms-swift/swift/llm/base.py", line 49, in main
[rank3]: result = self.run()
[rank3]: ^^^^^^^^^^
[rank3]: File "/data/workspace/vlm/ms-swift/swift/megatron/train/sft.py", line 69, in run
[rank3]: self.trainer.train(train_dataset, val_dataset, data_collator)
[rank3]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 758, in train
[rank3]: pretrain(
[rank3]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/training/training.py", line 864, in pretrain
[rank3]: iteration, num_floating_point_operations_so_far = train(
[rank3]: ^^^^^^
[rank3]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/training/training.py", line 2279, in train
[rank3]: ) = train_step(
[rank3]: ^^^^^^^^^^^
[rank3]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 311, in train_step
[rank3]: return self._origin_train_step(forward_step_func, new_data_iterator, model, optimizer, opt_param_scheduler,
[rank3]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank3]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/training/training.py", line 1395, in train_step
[rank3]: losses_reduced = forward_backward_func(
[rank3]: ^^^^^^^^^^^^^^^^^^^^^^
[rank3]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/core/pipeline_parallel/schedules.py", line 500, in forward_backward_no_pipelining
[rank3]: output_tensor, num_tokens = forward_step(
[rank3]: ^^^^^^^^^^^^^
[rank3]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/core/pipeline_parallel/schedules.py", line 289, in forward_step
[rank3]: output_tensor, loss_func = forward_step_func(data_iterator, model)
[rank3]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank3]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/trainer.py", line 141, in forward_step
[rank3]: data = get_batch(data_iterator)
[rank3]: ^^^^^^^^^^^^^^^^^^^^^^^^
[rank3]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/utils.py", line 123, in get_batch
[rank3]: batch = get_batch_on_this_tp_rank(data_iterator)
[rank3]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank3]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/utils.py", line 33, in get_batch_on_this_tp_rank
[rank3]: data = next(data_iterator)
[rank3]: ^^^^^^^^^^^^^^^^^^^
[rank3]: File "/mnt/workspace/.cache/modelscope/hub/_github/Megatron-LM/megatron/core/rerun_state_machine.py", line 1025, in next
[rank3]: n: Any = next(self.iterable)
[rank3]: ^^^^^^^^^^^^^^^^^^^
[rank3]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 104, in new_cyclic_iter
[rank3]: x = [next(it) for _ in range(num_microbatches - n_batch % num_microbatches)]
[rank3]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank3]: File "/data/workspace/vlm/ms-swift/swift/megatron/trainers/base.py", line 104, in
[rank3]: x = [next(it) for _ in range(num_microbatches - n_batch % num_microbatches)]
[rank3]: ^^^^^^^^
[rank3]: File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 733, in next
[rank3]: data = self._next_data()
[rank3]: ^^^^^^^^^^^^^^^^^
[rank3]: File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 1515, in _next_data
[rank3]: return self._process_data(data, worker_id)
[rank3]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank3]: File "/usr/local/lib/python3.11/site-packages/torch/utils/data/dataloader.py", line 1550, in _process_data
[rank3]: data.reraise()
[rank3]: File "/usr/local/lib/python3.11/site-packages/torch/_utils.py", line 750, in reraise
[rank3]: raise exception
[rank3]: IndexError: Caught IndexError in DataLoader worker process 0.
[rank3]: Original Traceback (most recent call last):
[rank3]: File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/worker.py", line 349, in _worker_loop
[rank3]: data = fetcher.fetch(index) # type: ignore[possibly-undefined]
[rank3]: ^^^^^^^^^^^^^^^^^^^^
[rank3]: File "/usr/local/lib/python3.11/site-packages/torch/utils/data/_utils/fetch.py", line 55, in fetch
[rank3]: return self.collate_fn(data)
[rank3]: ^^^^^^^^^^^^^^^^^^^^^
[rank3]: File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1390, in data_collator
[rank3]: res = self._data_collator(batch, padding_to=padding_to)
[rank3]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank3]: File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 411, in _data_collator
[rank3]: res = super()._data_collator(batch, padding_to=padding_to)
[rank3]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank3]: File "/data/workspace/vlm/ms-swift/swift/llm/template/base.py", line 1592, in _data_collator
[rank3]: batch[:] = [self.packing_row(batch)]
[rank3]: ^^^^^^^^^^^^^^^^^^^^^^^
[rank3]: File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 384, in packing_row
[rank3]: position_ids.append(self._get_position_ids(r))
[rank3]: ^^^^^^^^^^^^^^^^^^^^^^^^^
[rank3]: File "/data/workspace/vlm/ms-swift/swift/llm/template/template/qwen.py", line 402, in _get_position_ids
[rank3]: position_ids, _ = get_rope_index(
[rank3]: ^^^^^^^^^^^^^^^
[rank3]: File "/usr/local/lib/python3.11/site-packages/transformers/models/qwen3_vl_moe/modeling_qwen3_vl_moe.py", line 1092, in get_rope_index
[rank3]: input_ids = input_ids[attention_mask[i] == 1]
[rank3]: ~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^
[rank3]: IndexError: The shape of the mask [1040] at index 0 does not match the shape of the indexed tensor [1521] at index 0
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with swift/llm/template/template/qwen.py, especially packing_row and _get_position_ids, then trace the collator in swift/llm/template/base.py. Reproduce the provided megatron sft command and inspect the inputs passed to transformers' get_rope_index; done means the mask and input_ids shapes agree and training proceeds without the DataLoader IndexError.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- ai, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100