modelscope / modelscope/ms-swift

支持 Vision-centric 的多轮对话 Attention Masking

Open
#7,932 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement stale
Dominant language
Python
Stars
15.7k
Forks
1.7k
Avg merge
1d 16h
Merged PRs (30d)
136

Description

Checklist / 检查清单
  • I have searched existing issues, and this is a new feature request. / 我已经搜索过现有的 issues,确认这是一个新的 Feature Request。
Feature Request Description / Feature Request 描述

感谢 ms-swift 团队的框架支持,希望能够参考 Molmo2,支持同一图像/视频的多轮QA进行独立 Causal Mask 防止信息泄露,对图像和视频理解相关训练增加效率。

Image
Pull Request / Pull Request 信息

No response

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No files, tests, or entry points are identified. Start by reading the feature description and the linked Molmo2 paper, then locate the image/video multimodal training and attention-mask implementation. Done means multi-turn QA over the same image or video uses independent causal masks that prevent information leakage and improves training efficiency.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.