modelscope / modelscope/DiffSynth-Studio

是否支持图生720P视频的训练任务

Open
#565 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
13.1k
Forks
1.3k
Avg merge
13h 12m
Merged PRs (30d)
45

Description

您好我这边在数据处理阶段用的指令如下:
CUDA_VISIBLE_DEVICES="0" python ${project_dir}/examples/wanvideo/train_wan_t2v.py
--task data_process
--dataset_path $data_dir
--output_path ./models
--text_encoder_path "$bash_checkpoint/models_t5_umt5-xxl-enc-bf16.pth"
--vae_path "$bash_checkpoint/Wan2.1_VAE.pth"
--image_encoder_path "$bash_checkpoint/models_clip_open-clip-xlm-roberta-large-vit-huge-14.pth"
--num_frames 144
--height 720
--width 1280
--tiled
报错:
File "/mnt/workspace/guiji.hc/code/DiffSynth-Studio-main/examples/wanvideo/train_wan_t2v.py", line 599, in
data_process(args)
File "/mnt/workspace/guiji.hc/code/DiffSynth-Studio-main/examples/wanvideo/train_wan_t2v.py", line 541, in data_process
trainer.test(model, dataloader)
File "/mnt/workspace/guiji.hc/pythonEnv/diffstudio/lib/python3.10/site-packages/lightning/pytorch/trainer/trainer.py", line 775, in test
return call._call_and_handle_interrupt(
File "/mnt/workspace/guiji.hc/pythonEnv/diffstudio/lib/python3.10/site-packages/lightning/pytorch/trainer/call.py", line 48, in _call_and_handle_interrupt
return trainer_fn(*args, **kwargs)
File "/mnt/workspace/guiji.hc/pythonEnv/diffstudio/lib/python3.10/site-packages/lightning/pytorch/trainer/trainer.py", line 817, in _test_impl
results = self._run(model, ckpt_path=ckpt_path)
File "/mnt/workspace/guiji.hc/pythonEnv/diffstudio/lib/python3.10/site-packages/lightning/pytorch/trainer/trainer.py", line 1012, in _run
results = self._run_stage()
File "/mnt/workspace/guiji.hc/pythonEnv/diffstudio/lib/python3.10/site-packages/lightning/pytorch/trainer/trainer.py", line 1049, in _run_stage
return self._evaluation_loop.run()
File "/mnt/workspace/guiji.hc/pythonEnv/diffstudio/lib/python3.10/site-packages/lightning/pytorch/loops/utilities.py", line 179, in _decorator
return loop_run(self, *args, **kwargs)
File "/mnt/workspace/guiji.hc/pythonEnv/diffstudio/lib/python3.10/site-packages/lightning/pytorch/loops/evaluation_loop.py", line 145, in run
self._evaluation_step(batch, batch_idx, dataloader_idx, dataloader_iter)
File "/mnt/workspace/guiji.hc/pythonEnv/diffstudio/lib/python3.10/site-packages/lightning/pytorch/loops/evaluation_loop.py", line 437, in _evaluation_step
output = call._call_strategy_hook(trainer, hook_name, *step_args)
File "/mnt/workspace/guiji.hc/pythonEnv/diffstudio/lib/python3.10/site-packages/lightning/pytorch/trainer/call.py", line 328, in _call_strategy_hook
output = fn(*args, **kwargs)
File "/mnt/workspace/guiji.hc/pythonEnv/diffstudio/lib/python3.10/site-packages/lightning/pytorch/strategies/strategy.py", line 425, in test_step
return self.lightning_module.test_step(*args, **kwargs)
File "/mnt/workspace/guiji.hc/code/DiffSynth-Studio-main/examples/wanvideo/train_wan_t2v.py", line 152, in test_step
image_emb = self.pipe.encode_image(first_frame, None, num_frames, height, width)
File "/mnt/workspace/guiji.hc/code/DiffSynth-Studio-main/diffsynth/pipelines/wan_video.py", line 221, in encode_image
msk = msk.view(1, msk.shape[1] // 4, 4, height//8, width//8)
RuntimeError: shape '[1, 36, 4, 90, 160]' is invalid for input of size 2116800

请问是否支持720P&30帧率的图生视频任务,现在好像只能训练480p&16帧率,感谢~

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Reproduce the command in examples/wanvideo/train_wan_t2v.py with the reported 720×1280 and 144-frame settings. Read data_process and test_step there, then inspect diffsynth/pipelines/wan_video.py around encode_image and its mask reshape; done means the data-processing run completes for the requested resolution and frame count without the reported shape error.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.