modelscope / modelscope/DiffSynth-Studio
Wan 2.1 FP8 model weights causing color issue - BF16 has no issue
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 13.1k
- Forks
- 1.3k
- Avg merge
- 13h 12m
- Merged PRs (30d)
- 45
Description
This is how I load. By the way text to video doesnt have this. E.g. tested on WAN 2.1 14B Text-to-Video
I hope you can help me @Artiprocher
elif model_choice == "14B_image_480p":
clip_path = get_common_file(os.path.join("models", "models_clip_open-clip-xlm-roberta-large-vit-huge-14.pth"),
os.path.join("models", "Wan-AI", "Wan2.1-I2V-14B-480P", "models_clip_open-clip-xlm-roberta-large-vit-huge-14.pth"))
t5_path = get_common_file(os.path.join("models", "models_t5_umt5-xxl-enc-bf16.pth"),
os.path.join("models", "Wan-AI", "Wan2.1-I2V-14B-480P", "models_t5_umt5-xxl-enc-bf16.pth"))
vae_path = get_common_file(os.path.join("models", "Wan2.1_VAE.pth"),
os.path.join("models", "Wan-AI", "Wan2.1-I2V-14B-480P", "Wan2.1_VAE.pth"))
model_manager.load_models([clip_path], torch_dtype=torch.float32)
model_manager.load_models(
[
[
os.path.join("models", "Wan-AI", "Wan2.1-I2V-14B-480P", "diffusion_pytorch_model-00001-of-00007.safetensors"),
os.path.join("models", "Wan-AI", "Wan2.1-I2V-14B-480P", "diffusion_pytorch_model-00002-of-00007.safetensors"),
os.path.join("models", "Wan-AI", "Wan2.1-I2V-14B-480P", "diffusion_pytorch_model-00003-of-00007.safetensors"),
os.path.join("models", "Wan-AI", "Wan2.1-I2V-14B-480P", "diffusion_pytorch_model-00004-of-00007.safetensors"),
os.path.join("models", "Wan-AI", "Wan2.1-I2V-14B-480P", "diffusion_pytorch_model-00005-of-00007.safetensors"),
os.path.join("models", "Wan-AI", "Wan2.1-I2V-14B-480P", "diffusion_pytorch_model-00006-of-00007.safetensors"),
os.path.join("models", "Wan-AI", "Wan2.1-I2V-14B-480P", "diffusion_pytorch_model-00007-of-00007.safetensors"),
],
t5_path,
vae_path,
],
torch_dtype=torch.float8_e4m3fn,
)
pipe = WanVideoPipeline.from_model_manager(model_manager, torch_dtype=torch.bfloat16, device=device)
try:
num_persistent_val = int(num_persistent)
except:
print("[CMD] Warning: could not parse num_persistent value, defaulting to 6000000000")
num_persistent_val = 6000000000
print(f"num_persistent_val {num_persistent_val}")
pipe.enable_vram_management(num_persistent_param_in_dit=num_persistent_val)
print("[CMD] Model loaded successfully.")
return pipe
Now I will show BF16 vs FP8
https://github.com/user-attachments/assets/60aaea12-3fbc-480e-b9f4-ba7d25827e57
https://github.com/user-attachments/assets/814d1318-416d-453b-8977-509017030c51
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Reproduce the Wan 2.1 image-to-video result using the shown model-loading code, comparing FP8 with BF16 and the text-to-video case. Start by tracing WanVideoPipeline and model_manager precision handling, then identify where the color difference is introduced. Done means the FP8 image-to-video output no longer has the reported color issue without regressing BF16 or text-to-video behavior.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100