[Feature Request] Qwen3VL GRPO, SFT training
Open
enhancement
external
multimodal
new model
r0.6.0
x-ss
- Dominant language
- Python
- Stars
- 2k
- Forks
- 561
- Avg merge
- 4d 5h
- Merged PRs (30d)
- 145
Description
**Additional context**
Our customer would like to apply RL methods (GRPO, GSPO, and SPO) to VLM with MoE (such as Qwen3-VL).
Would it be possible to extend the current VLM support to Qwen3-VL?
(cc. @terrykong, @snowmanwwg )
Contributor guide
Assessment
This issue has not been assessed yet.