THUDM / THUDM/slime

[Question] qwen3-vl转化

Open
#1,863 4 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

question
Dominant language
Python
Stars
8.5k
Forks
1.3k
Avg merge
5h 36m
Merged PRs (30d)
22

Description

Your Question

slime/scripts/models没有vlm相关的配置参数,如果我想训练qwen3-vl系列的模型,该怎么从hf转化为Megatron 格式

What I've Tried

我看了操作教程但没发现vlm适配

Environment (if relevant)
  • slime version:
  • Python version:
  • PyTorch version:
  • CUDA/ROCm version:
  • GPU type and count:
  • OS:
Additional Context

No response

Pre-submission Checklist

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reviewing the model-conversion guidance and the contents of slime/scripts/models, then compare the existing configuration coverage with Qwen3-VL requirements. The issue is complete when a documented, working path exists for converting Qwen3-VL models from Hugging Face to Megatron format, including the needed VLM configuration.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.