modelscope / modelscope/ms-swift
支持 Vision-centric 的多轮对话 Attention Masking
Open
Nobody has claimed this yet.
enhancement
stale
- Dominant language
- Python
- Stars
- 15.7k
- Forks
- 1.7k
- Avg merge
- 1d 16h
- Merged PRs (30d)
- 136
Description
Checklist / 检查清单
- I have searched existing issues, and this is a new feature request. / 我已经搜索过现有的 issues,确认这是一个新的 Feature Request。
Feature Request Description / Feature Request 描述
感谢 ms-swift 团队的框架支持,希望能够参考 Molmo2,支持同一图像/视频的多轮QA进行独立 Causal Mask 防止信息泄露,对图像和视频理解相关训练增加效率。
Pull Request / Pull Request 信息
No response
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
No files, tests, or entry points are identified. Start by reading the feature description and the linked Molmo2 paper, then locate the image/video multimodal training and attention-mask implementation. Done means multi-turn QA over the same image or video uses independent causal masks that prevent information leakage and improves training efficiency.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100