WAN 2.2 i2v: Second run is 4 5to 5 times slower on AMD GPU (ROCm)
- Dominant language
- Python
- Stars
- 133k
- Forks
- 15.7k
- Avg merge
- 1d 7h
- Merged PRs (30d)
- 158
Description
### Custom Node Testing
- [x] I have tried disabling custom nodes and the issue persists (see [how to disable custom nodes](https://docs.comfy.org/troubleshooting/custom-node-issues#step-1%3A-test-with-all-custom-nodes-disabled) if you need help)
### Expected Behavior
Second run to generate a video with Wan Image to Video should be equal fast as the first run.
### Actual Behavior
Second run to generate a video with Wan Image to Video is significant slowe than the first run.
First run video generation time of a resolution of 1024x576 is around 300 seconds on Linux, and 450 to 600 seconds on Windows on a dual boot system. Second run is then on both 45 minutes and more. That's four to five times slower. With a very small video size the difference is between 60 sconds and 190 seconds.
### Steps to Reproduce
Load the attached Wan Image to Video workflow on a PC with AMD card and ROCm 7.2. See attached assets.
Let it run through.
Change the image.
Let it run through again.
Watch the generation times.
### Debug Logs
```powershell
G:\comfyamd\ComfyUI_windows_portable>set PYTORCH_ALLOC_CONF=garbage_collection_threshold:0.9,max_split_size_mb:512
G:\comfyamd\ComfyUI_windows_portable>.\python_embeded\python.exe -s ComfyUI\main.py --windows-standalone-build --disable-pinned-memory --use-split-cross-attention --highvram
G:\comfyamd\ComfyUI_windows_portable\python_embeded\Lib\site-packages\requests\__init__.py:113: RequestsDependencyWarning: urllib3 (2.6.3) or chardet (6.0.0.post1)/charset_normalizer (3.4.4) doesn't match a supported version!
warnings.warn(
[START] Security scan
[DONE] Security scan
## ComfyUI-Manager: installing dependencies done.
** ComfyUI startup time: 2026-02-27 10:26:56.330
** Platform: Windows
** Python version: 3.12.10 (tags/v3.12.10:0cc8128, Apr 8 2025, 12:21:36) [MSC v.1943 64 bit (AMD64)]
** Python executable: G:\comfyamd\ComfyUI_windows_portable\python_embeded\python.exe
** ComfyUI Path: G:\comfyamd\ComfyUI_windows_portable\ComfyUI
** ComfyUI Base Folder Path: G:\comfyamd\ComfyUI_windows_portable\ComfyUI
** User directory: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\user
** ComfyUI-Manager config path: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\user\default\ComfyUI-Manager\config.ini
** Log path: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\user\comfyui.log
Prestartup times for custom nodes:
0.0 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\rgthree-comfy
7.0 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\ComfyUI-Manager
Found comfy_kitchen backend cuda: {'available': True, 'disabled': True, 'unavailable_reason': None, 'capabilities': ['apply_rope', 'apply_rope1', 'dequantize_nvfp4', 'dequantize_per_tensor_fp8', 'quantize_nvfp4', 'quantize_per_tensor_fp8', 'scaled_mm_nvfp4']}
Found comfy_kitchen backend eager: {'available': True, 'disabled': False, 'unavailable_reason': None, 'capabilities': ['apply_rope', 'apply_rope1', 'dequantize_nvfp4', 'dequantize_per_tensor_fp8', 'quantize_nvfp4', 'quantize_per_tensor_fp8', 'scaled_mm_nvfp4']}
Found comfy_kitchen backend triton: {'available': False, 'disabled': True, 'unavailable_reason': "ImportError: No module named 'triton'", 'capabilities': []}
Checkpoint files will always be loaded safely.
Total VRAM 32624 MB, total RAM 63140 MB
pytorch version: 2.9.1+rocmsdk20260116
Set: torch.backends.cudnn.enabled = False for better AMD performance.
AMD arch: gfx1201
ROCm version: (7, 2)
Set vram state to: HIGH_VRAM
Device: cuda:0 AMD Radeon AI PRO R9700 : native
Using async weight offloading with 2 streams
Using split optimization for attention
Python version: 3.12.10 (tags/v3.12.10:0cc8128, Apr 8 2025, 12:21:36) [MSC v.1943 64 bit (AMD64)]
ComfyUI version: 0.15.0
ComfyUI frontend version: 1.39.16
[Prompt Server] web root: G:\comfyamd\ComfyUI_windows_portable\python_embeded\Lib\site-packages\comfyui_frontend_package\static
Traceback (most recent call last):
File "G:\comfyamd\ComfyUI_windows_portable\ComfyUI\nodes.py", line 2224, in load_custom_node
module_spec.loader.exec_module(module)
File "", line 999, in exec_module
File "", line 488, in _call_with_frames_removed
File "G:\comfyamd\ComfyUI_windows_portable\ComfyUI\comfy_extras\nodes_glsl.py", line 64, in
_check_opengl_availability()
File "G:\comfyamd\ComfyUI_windows_portable\ComfyUI\comfy_extras\nodes_glsl.py", line 35, in _check_opengl_availability
raise RuntimeError(
RuntimeError: OpenGL dependencies not available.
Please install the updated requirements.txt file by running:
G:\comfyamd\ComfyUI_windows_portable\python_embeded\python.exe -s -m pip install -r G:\comfyamd\ComfyUI_windows_portable\ComfyUI\requirements.txt
If you are on the portable package you can run: update\update_comfyui.bat to solve this problem.
Cannot import G:\comfyamd\ComfyUI_windows_portable\ComfyUI\comfy_extras\nodes_glsl.py module for custom nodes: OpenGL dependencies not available.
Please install the updated requirements.txt file by running:
G:\comfyamd\ComfyUI_windows_portable\python_embeded\python.exe -s -m pip install -r G:\comfyamd\ComfyUI_windows_portable\ComfyUI\requirements.txt
If you are on the portable package you can run: update\update_comfyui.bat to solve this problem.
ComfyUI-FreeMemory nodes loaded: Image, Latent, Model, CLIP, and String versions available.
ComfyUI-GGUF: Allowing full torch compile
### Loading: ComfyUI-Impact-Pack (V8.28.2)
[Impact Pack] Wildcard total size (0.00 MB) is within cache limit (50.00 MB). Using full cache mode.
[Impact Pack] Wildcards loading done.
### Loading: ComfyUI-Manager (V3.37.1)
[ComfyUI-Manager] network_mode: public
### ComfyUI Revision: 150 [b874bd2b] *DETACHED | Released on '2026-02-24'
Error loading module AILab_QwenVL: cannot import name 'AutoModelForVision2Seq' from 'transformers' (G:\comfyamd\ComfyUI_windows_portable\python_embeded\Lib\site-packages\transformers\__init__.py)
Error loading module AILab_QwenVL_GGUF_PromptEnhancer: No module named 'llama_cpp'
Error loading module AILab_QwenVL_PromptEnhancer: cannot import name 'ATTENTION_MODES' from 'AILab_QwenVL' (G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\ComfyUI-QwenVL\AILab_QwenVL.py)
[ComfyUI-Manager] default cache updated: https://raw.githubusercontent.com/ltdrdata/ComfyUI-Manager/main/alter-list.json
Note: NumExpr detected 24 cores but "NUMEXPR_MAX_THREADS" not set, so enforcing safe limit of 16.
NumExpr defaulting to 16 threads.
[ComfyUI-Manager] default cache updated: https://raw.githubusercontent.com/ltdrdata/ComfyUI-Manager/main/github-stats.json
[ComfyUI-Manager] default cache updated: https://raw.githubusercontent.com/ltdrdata/ComfyUI-Manager/main/extension-node-map.json
[ComfyUI-Manager] default cache updated: https://raw.githubusercontent.com/ltdrdata/ComfyUI-Manager/main/model-list.json
[ComfyUI-Manager] default cache updated: https://raw.githubusercontent.com/ltdrdata/ComfyUI-Manager/main/custom-node-list.json
FizzleDorf Custom Nodes: Loaded
██████╗ ██╗ ██╗ █████╗ ███╗ ██╗ ██████╗ ███╗ ██╗
██╔══██╗╚██╗ ██╔╝██╔══██╗████╗ ██║ ██╔═══██╗████╗ ██║
██████╔╝ ╚████╔╝ ███████║██╔██╗ ██║ ██║ ██║██╔██╗ ██║
██╔══██╗ ╚██╔╝ ██╔══██║██║╚██╗██║ ██║ ██║██║╚██╗██║
██║ ██║ ██║ ██║ ██║██║ ╚████║ ╚██████╔╝██║ ╚████║
╚═╝ ╚═╝ ╚═╝ ╚═╝ ╚═╝╚═╝ ╚═══╝ ╚═════╝ ╚═╝ ╚═══╝
████████╗██╗ ██╗███████╗ ██╗███╗ ██╗███████╗██╗██████╗ ███████╗
╚══██╔══╝██║ ██║██╔════╝ ██║████╗ ██║██╔════╝██║██╔══██╗██╔════╝
██║ ███████║█████╗ ██║██╔██╗ ██║███████╗██║██║ ██║█████╗
██║ ██╔══██║██╔══╝ ██║██║╚██╗██║╚════██║██║██║ ██║██╔══╝
██║ ██║ ██║███████╗ ██║██║ ╚████║███████║██║██████╔╝███████╗
╚═╝ ╚═╝ ╚═╝╚══════╝ ╚═╝╚═╝ ╚═══╝╚══════╝╚═╝╚═════╝ ╚══════╝
⚡ R Y A N O N T H E I N S I D E ⚡
[RyanOnTheInside] Copying extension file: widget_hotkey.js
[RyanOnTheInside] Failed to load .nodes.masks.temporal_masks: No module named 'pymunk'
[RyanOnTheInside] Failed to load .nodes.masks.optical_flow_masks: No module named 'pymunk'
[RyanOnTheInside] Failed to load .nodes.masks.particle_system_masks: No module named 'pymunk'
[RyanOnTheInside] Failed to load .nodes.masks.flex_masks: No module named 'pymunk'
Checking for AdvancedLivePortrait at: None
Checking for Advanced-ControlNet at: None
Checking for AnimateDiff-Evolved at: None
ComfyUI-AdvancedLivePortrait not found. FlexExpressionEditor will not be available. Install ComfyUI-AdvancedLivePortrait and restart ComfyUI.
ComfyUI-Advanced-ControlNet not found. Advanced-ControlNet feature nodes will not be available. Install ComfyUI-Advanced-ControlNet and restart ComfyUI.
ComfyUI-AnimateDiff-Evolved not found. AnimateDiff feature nodes will not be available. Install ComfyUI-AnimateDiff-Evolved and restart ComfyUI.
[comfyui_ryanontheinside.acestep] [ACE-Step Patches] Patched resolve_areas_and_cond_masks_multidim for 1D latent support
[comfyui_ryanontheinside.acestep] [ACE-Step Patches] Patched comfy.utils.reshape_mask for 1D latent support
[Crystools ERROR] pynvml is not installed. No module named 'pynvml'
[Crystools WARNING] No GPU monitoring libraries available.
[rgthree-comfy] Loaded 48 magnificent nodes. 🎉
[rgthree-comfy] ComfyUI's new Node 2.0 rendering may be incompatible with some rgthree-comfy nodes and features, breaking some rendering as well as losing the ability to access a node's properties (a vital part of many nodes). It also appears to run MUCH more slowly spiking CPU usage and causing jankiness and unresponsiveness, especially with large workflows. Personally I am not planning to use the new Nodes 2.0 and, unfortunately, am not able to invest the time to investigate and overhaul rgthree-comfy where needed. If you have issues when Nodes 2.0 is enabled, I'd urge you to switch it off as well and join me in hoping ComfyUI is not planning to deprecate the existing, stable canvas rendering all together.
Error: modelscope not installed. Please run: pip install modelscope
Import times for custom nodes:
0.0 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\websocket_image_save.py
0.0 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\HeartMuLa_ComfyUI
0.0 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\comfyui-simple-prompt-batcher
0.0 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\comfyui_auto_prompt_schedule
0.0 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\z-image-turbo
0.0 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\ComfyUI-FreeMemory
0.0 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\ComfyUI-VFI
0.0 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\Crystools-MonitorOnly
0.0 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\ComfyMath
0.0 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\ComfyUI-Custom-Scripts
0.0 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\Derfuu_ComfyUI_ModdedNodes
0.0 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\comfyui-frame-interpolation
0.0 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\ComfyUI-QwenVL
0.0 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\rgthree-comfy
0.0 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\ComfyUI-GGUF
0.1 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\ComfyUI-VideoHelperSuite
0.1 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\ComfyUI-LTXVideo
0.1 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\comfyui_ryanonyheinside
0.3 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\comfyui-levelpixel
0.3 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\ComfyUI-Impact-Pack
0.5 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\ComfyUI-Manager
0.6 seconds: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\ComfyUI_FizzNodes
WARNING: some comfy_extras/ nodes did not import correctly. This may be because they are missing some dependencies.
IMPORT FAILED: nodes_glsl.py
This issue might be caused by new missing dependencies added the last time you updated ComfyUI.
Please run the update script: update/update_comfyui.bat
Context impl SQLiteImpl.
Will assume non-transactional DDL.
Assets scan(roots=['models']) completed in 0.015s (created=0, skipped_existing=65, orphans_pruned=0, total_seen=65)
Starting server
To see the GUI go to: http://127.0.0.1:8188
[DEPRECATION WARNING] Detected import of deprecated legacy API: /scripts/ui.js. This is likely caused by a custom node extension using outdated APIs. Please update your extensions or contact the extension author for an updated version.
[DEPRECATION WARNING] Detected import of deprecated legacy API: /extensions/core/clipspace.js. This is likely caused by a custom node extension using outdated APIs. Please update your extensions or contact the extension author for an updated version.
[DEPRECATION WARNING] Detected import of deprecated legacy API: /extensions/core/groupNode.js. This is likely caused by a custom node extension using outdated APIs. Please update your extensions or contact the extension author for an updated version.
[DEPRECATION WARNING] Detected import of deprecated legacy API: /extensions/core/widgetInputs.js. This is likely caused by a custom node extension using outdated APIs. Please update your extensions or contact the extension author for an updated version.
[DEPRECATION WARNING] Detected import of deprecated legacy API: /scripts/ui/components/button.js. This is likely caused by a custom node extension using outdated APIs. Please update your extensions or contact the extension author for an updated version.
[DEPRECATION WARNING] Detected import of deprecated legacy API: /scripts/ui/components/buttonGroup.js. This is likely caused by a custom node extension using outdated APIs. Please update your extensions or contact the extension author for an updated version.
got prompt
Using split attention in VAE
Using split attention in VAE
VAE load device: cuda:0, offload device: cpu, dtype: torch.bfloat16
Found quantization metadata version 1
Using MixedPrecisionOps for text encoder
Requested to load WanTEModel
loaded completely; 6419.48 MB loaded, full load: True
CLIP/text encoder model load device: cuda:0, offload device: cpu, current: cuda:0, dtype: torch.float16
Requested to load WanVAE
loaded completely; 242.03 MB loaded, full load: True
### vae.encode done
gguf qtypes: F16 (694), Q4_K (280), Q6_K (120), F32 (1)
model weight dtype torch.float16, manual cast: None
model_type FLOW
Requested to load WAN21
Unloaded partially: 1232.48 MB freed, 5187.00 MB remains loaded, 112.00 MB buffer reserved, lowvram patches: 0
loaded completely; 9337.19 MB loaded, full load: True
100%|████████████████████████████████████████████████████████████████████████████████████| 2/2 [00:12<00:00, 6.44s/it]
gguf qtypes: F16 (694), Q4_K (280), Q6_K (120), F32 (1)
model weight dtype torch.float16, manual cast: None
model_type FLOW
Requested to load WAN21
Unloaded partially: 4305.35 MB freed, 5034.20 MB remains loaded, 161.55 MB buffer reserved, lowvram patches: 278
loaded completely; 9337.19 MB loaded, full load: True
100%|████████████████████████████████████████████████████████████████████████████████████| 2/2 [00:12<00:00, 6.11s/it]
Requested to load WanVAE
loaded completely; 242.03 MB loaded, full load: True
Loading RIFE model from: G:\comfyamd\ComfyUI_windows_portable\ComfyUI\custom_nodes\ComfyUI-VFI\rife\train_log\flownet.pkl
Prompt executed in 64.19 seconds
FETCH ComfyRegistry Data [DONE]
[ComfyUI-Manager] default cache updated: https://api.comfy.org/nodes
FETCH DATA from: https://raw.githubusercontent.com/ltdrdata/ComfyUI-Manager/main/custom-node-list.json [DONE]
[ComfyUI-Manager] All startup tasks have been completed.
got prompt
Requested to load WanTEModel
loaded completely; 6419.48 MB loaded, full load: True
### vae.encode done
Unloaded partially: 8569.84 MB freed, 769.71 MB remains loaded, 276.88 MB buffer reserved, lowvram patches: 394
loaded completely; 9337.19 MB loaded, full load: True
100%|████████████████████████████████████████████████████████████████████████████████████| 2/2 [00:25<00:00, 12.77s/it]
loaded completely; 9337.19 MB loaded, full load: True
100%|████████████████████████████████████████████████████████████████████████████████████| 2/2 [00:13<00:00, 6.74s/it]
Requested to load WanVAE
loaded completely; 242.03 MB loaded, full load: True
Prompt executed in 192.19 seconds
```
### Other
**WAN 2.2 i2v: Second run is 4 5to 5 times slower on AMD GPUs (ROCm)**
**Environment:**
- AMD Radeon AI PRO R9700 (32GB VRAM, gfx1201)
- ROCm 7.2, Windows and Linux (Docker) both affected
- ComfyUI 0.15, WAN 2.2 i2v with dual GGUF loaders
- 1024x576 @ 81 frames
**Problem:**
First run completes in ~300 seconds. Second run (same workflow, same settings) takes 30-45 minutes or hangs indefinitely at the WanImageToVideo node during `vae.encode`.
**Root Cause identified:**
Through extensive instrumentation we traced the hang to `CausalConv3d.forward` in `comfy/ldm/wan/vae.py`.
The WAN VAE encoder processes video in chunks: chunk 0 is always 1 frame, subsequent chunks are 4 frames with a temporal cache (`feat_map`). On the first run, chunk 0 uses the fast path (`cache_x is None and x.shape[2] == 1`). On the second run, chunks arrive with `x.shape[2]=4` AND a populated cache, creating a new tensor shape that ROCm has never compiled a kernel for.
A minimal reproduction confirms the issue:
```python
import torch, time
device = "cuda:0"
x1 = torch.randn(1, 96, 1, 576, 1024, device=device, dtype=torch.float16)
conv = torch.nn.Conv3d(96, 96, 3, padding=1).to(device).half()
start = time.time(); conv(x1); print(f"1 frame: {time.time()-start:.3f}s")
x2 = torch.randn(1, 96, 6, 576, 1024, device=device, dtype=torch.float16)
start = time.time(); conv(x2); print(f"6 frames first call: {time.time()-start:.3f}s")
start = time.time(); conv(x2); print(f"6 frames second call: {time.time()-start:.3f}s")
```
Output:
```
1 frame: 7.640s
6 frames first call: 83.488s
6 frames second call: 0.001s
```
ROCm compiles kernels on first use of a new tensor shape. The kernel cache works within a session, but is invalidated when the VAE is unloaded and reloaded between runs (which ComfyUI does to free VRAM for the diffusion model).
Additional observation: Between Run 1 and Run 2, ComfyUI unloads significantly more VRAM before loading WAN21 (Unloaded partially: 8569 MB in Run 2 vs 1232 MB in Run 1). This suggests a memory management issue where ComfyUI's state differs between runs, causing more aggressive model offloading on subsequent runs. This may be a contributing factor or separate bug.
Additional observation2: Qwen image generation (also using WAN VAE) shows a similar pattern, but not at the vae, but the ksampler. – KSampler runs in ~4 seconds on first run, ~40 seconds on subsequent run. And then i have to reboot to get the 4 seconds back. It is not enough to simply close ComfyUI. Whether this is related to the same root cause or the increased `Unloaded partially` behavior between runs is unclear and warrants further investigation. Most probably this is another issue. I nevertheless wanted to mention it here already.
**What we tried:**
- `--disable-smart-memory`, `--disable-pinned-memory`, `--disable-async-offload` – no effect
- `MIOPEN_USER_DB_PATH` persistent cache – no effect
- `PYTORCH_TUNABLEOP_ENABLED=1` – made everything slower
- `AMD_COMGR_CACHE` persistent kernel cache – not supported on Windows ROCm
- `HIP_LAUNCH_BLOCKING=1` – made everything slower
- `HSA_ENABLE_SDMA=0` – no effect
- `--highvram` – prevents VAE unloading but causes OOM in KSampler
- Frame-by-frame processing in `CausalConv3d.forward` (always 1-frame tensors → always fast path) – VAE becomes fast but increases VRAM usage by ~1GB causing OOM in KSampler
- COMFYUI_ENABLE_MIOPEN=1 – no effect
- Attempted to rewrite `CausalConv3d.forward` inspired by LTX Video's approach (per-instance cache with thread ID). This caused excessive VRAM usage.
- Attempted frame-by-frame processing splitting the 4-frame tensor into 4 individual 1-frame calls. Runtime dropped from 45 minutes to ~14 minutes on second run, but broke Qwen image workflow due to incorrect tensor size. A more targeted version with spatial kernel size check (`kernel_size[1] == 3`) and pre-allocated output tensor reduced VRAM overhead but still caused OOM in KSampler.
And quite a few more things that i cannot count anymore.
**Two bugs found as a side effect of the investigation:**
1. `comfy/ldm/wan/vae.py`, `WanVAE.encode`: `feat_map` is sized using `count_conv3d(self.decoder)` instead of `count_conv3d(self.encoder)` – wrong cache size for the encoder.
2. `comfy/sd.py`: `VAE_KL_MEM_RATIO = 2.73` for AMD/ROCm massively overestimates VRAM requirements, causing unnecessary model offloading. Value should be ~1.0-1.3 for modern ROCm. 1.0 worked just fine here.
**Why we're stuck:**
I am simply running out of ideas. Frame-by-frame processing in CausalConv3d.forward (always 1-frame tensors → always fast path) – VAE becomes fast on second run but increases VRAM usage by ~1GB causing OOM in KSampler. Not tested to completion at full resolution. The kernel cache cannot be made persistent on Windows ROCm. Preventing VAE unloading requires too much VRAM. The root issue is that ROCm recompiles kernels when tensor shapes change and does not persist this cache across model load/unload cycles. A proper fix likely requires either a ROCm-level kernel cache solution or a redesign of how the WAN VAE temporal cache interacts with the chunked encoding.
Assets contains the test_rocm.py, the wan image to video workflow and two example images in png format.
[assets.zip](https://github.com/user-attachments/files/25599817/assets.zip)

Contributor guide
Assessment
This issue has not been assessed yet.