modelscope / modelscope/DiffSynth-Studio

wan2.1 是否支持基于多个 reference 帧的训练

Open
#785 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
13.1k
Forks
1.3k
Avg merge
13h 12m
Merged PRs (30d)
45

Description

目前的 wan2.1(包括其他版本的 wan)只基于一张 reference_image 进行训练,是否支持多帧的 reference 进行训练?这部分的代码实现似乎比较简单,但看起来没有实现这个功能,是因为多帧训练表现不好吗

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by locating the wan2.1 training entry point and the code path that consumes the single reference_image. Compare this behavior with other wan versions and inspect existing training tests or examples, if present. Done would require a clear decision on multi-frame reference support, its expected behavior, and evidence about training quality or limitations.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.