Comfy-Org / Comfy-Org/ComfyUI

[ROCm/Windows][RX 9070 XT] Severe GPU/UI stalling during VAE Decode with DynamicVRAM enabled

Open
#16,062 8 comments 0 reactions 0 assignees View on GitHub
Potential Bug
Dominant language
Python
Stars
133k
Forks
15.7k
Avg merge
1d 6h
Merged PRs (30d)
155

Description

### Custom Node Testing

- [x] I have tried disabling custom nodes and the issue persists (see [how to disable custom nodes](https://docs.comfy.org/troubleshooting/custom-node-issues#step-1%3A-test-with-all-custom-nodes-disabled) if you need help)

### Expected Behavior

Normal image generation

### Actual Behavior

[Title.md](https://github.com/user-attachments/files/31791990/Title.md)

### Steps to Reproduce

[cyberIllustrious.json](https://github.com/user-attachments/files/31792569/cyberIllustrious.json)

### Debug Logs

```powershell
[INFO] setup plugin alembic.autogenerate.schemas
[INFO] setup plugin alembic.autogenerate.tables
[INFO] setup plugin alembic.autogenerate.types
[INFO] setup plugin alembic.autogenerate.constraints
[INFO] setup plugin alembic.autogenerate.defaults
[INFO] setup plugin alembic.autogenerate.comments
[INFO] Adding extra search path checkpoints F:\ComfyUI-Shared\models\checkpoints
[INFO] Adding extra search path classifiers F:\ComfyUI-Shared\models\classifiers
[INFO] Adding extra search path clip_vision F:\ComfyUI-Shared\models\clip_vision
[INFO] Adding extra search path configs F:\ComfyUI-Shared\models\configs
[INFO] Adding extra search path controlnet F:\ComfyUI-Shared\models\controlnet
[INFO] Adding extra search path controlnet F:\ComfyUI-Shared\models\t2i_adapter
[INFO] Adding extra search path diffusers F:\ComfyUI-Shared\models\diffusers
[INFO] Adding extra search path diffusion_models F:\ComfyUI-Shared\models\diffusion_models
[INFO] Adding extra search path embeddings F:\ComfyUI-Shared\models\embeddings
[INFO] Adding extra search path gligen F:\ComfyUI-Shared\models\gligen
[INFO] Adding extra search path hypernetworks F:\ComfyUI-Shared\models\hypernetworks
[INFO] Adding extra search path latent_upscale_models F:\ComfyUI-Shared\models\latent_upscale_models
[INFO] Adding extra search path loras F:\ComfyUI-Shared\models\loras
[INFO] Adding extra search path model_patches F:\ComfyUI-Shared\models\model_patches
[INFO] Adding extra search path audio_encoders F:\ComfyUI-Shared\models\audio_encoders
[INFO] Adding extra search path photomaker F:\ComfyUI-Shared\models\photomaker
[INFO] Adding extra search path style_models F:\ComfyUI-Shared\models\style_models
[INFO] Adding extra search path text_encoders F:\ComfyUI-Shared\models\text_encoders
[INFO] Adding extra search path upscale_models F:\ComfyUI-Shared\models\upscale_models
[INFO] Adding extra search path background_removal F:\ComfyUI-Shared\models\background_removal
[INFO] Adding extra search path frame_interpolation F:\ComfyUI-Shared\models\frame_interpolation
[INFO] Adding extra search path geometry_estimation F:\ComfyUI-Shared\models\geometry_estimation
[INFO] Adding extra search path optical_flow F:\ComfyUI-Shared\models\optical_flow
[INFO] Adding extra search path detection F:\ComfyUI-Shared\models\detection
[INFO] Adding extra search path vae F:\ComfyUI-Shared\models\vae
[INFO] Adding extra search path vae_approx F:\ComfyUI-Shared\models\vae_approx
[INFO] Adding extra search path onnx F:\ComfyUI-Shared\models\onnx
[INFO] Adding extra search path sams F:\ComfyUI-Shared\models\sams
[INFO] Adding extra search path ultralytics F:\ComfyUI-Shared\models\ultralytics
[INFO] Adding extra search path clip F:\ComfyUI-Shared\models\clip
[INFO] Adding extra search path unet F:\ComfyUI-Shared\models\unet
[INFO] Setting output directory to: F:\ComfyUI-Shared\output
[INFO] Setting input directory to: F:\ComfyUI-Shared\input
[START] Security scan
[DONE] Security scan
** ComfyUI startup time: 2026-09-03 21:25:38.986
** Platform: Windows
** Python version: 3.12.12 (main, Feb 12 2026, 00:40:26) [MSC v.1944 64 bit (AMD64)]
** Python executable: C:\Users\brf-2\ComfyUI-Installs\ComfyUI\ComfyUI\.venv\Scripts\python.exe
** ComfyUI Path: C:\Users\brf-2\ComfyUI-Installs\ComfyUI\ComfyUI
** ComfyUI Base Folder Path: C:\Users\brf-2\ComfyUI-Installs\ComfyUI\ComfyUI
** User directory: C:\Users\brf-2\ComfyUI-Installs\ComfyUI\ComfyUI\user
** ComfyUI-Manager config path: C:\Users\brf-2\ComfyUI-Installs\ComfyUI\ComfyUI\user\__manager\config.ini
** Log path: C:\Users\brf-2\ComfyUI-Installs\ComfyUI\ComfyUI\user\comfyui.log
[INFO] [PRE] ComfyUI-Manager
[INFO] Found comfy_kitchen backend eager: {'available': True, 'disabled': False, 'unavailable_reason': None, 'capabilities': ['adaln', 'apply_rope', 'apply_rope1', 'apply_rope1_', 'apply_rope_', 'apply_rope_split_half', 'apply_rope_split_half1', 'apply_rope_split_half1_', 'apply_rope_split_half_', 'convrot_w4a4_linear', 'dequantize_convrot_w4a4_weight', 'dequantize_int8_convrot_weight', 'dequantize_int8_convrot_weight_dtype', 'dequantize_int8_embedding', 'dequantize_int8_simple', 'dequantize_int8_simple_dtype', 'dequantize_mxfp8', 'dequantize_nvfp4', 'dequantize_per_tensor_fp8', 'dequantize_w4a8_int8_weight', 'gemv_awq_w4a16', 'int8_linear', 'na3d', 'prepare_int4_weight_for_int8_linear', 'quantize_and_rotate_rowwise', 'quantize_convrot_w4a4_weight', 'quantize_int8_convrot_weight', 'quantize_int8_rowwise', 'quantize_int8_tensorwise', 'quantize_mxfp8', 'quantize_nvfp4', 'quantize_per_tensor_fp8', 'quantize_svdquant_w4a4', 'quantize_w4a8_int8_weight', 'rms_adaln', 'rms_rope', 'rms_rope1', 'rms_rope1_', 'rms_rope_', 'rms_rope_split_half', 'rms_rope_split_half1', 'rms_rope_split_half1_', 'rms_rope_split_half_', 'rotate_int8_convrot_weight', 'scaled_mm_mxfp8', 'scaled_mm_nvfp4', 'scaled_mm_svdquant_w4a4', 'stochastic_rounding_fp8', 'w4a8_int8_linear']}
[INFO] Found comfy_kitchen backend triton: {'available': False, 'disabled': True, 'unavailable_reason': "ImportError: No module named 'triton'", 'capabilities': []}
[INFO] Found comfy_kitchen backend hip: {'available': True, 'disabled': False, 'unavailable_reason': None, 'capabilities': ['adaln', 'apply_rope', 'apply_rope1', 'apply_rope1_', 'apply_rope_', 'apply_rope_split_half', 'apply_rope_split_half1', 'apply_rope_split_half1_', 'apply_rope_split_half_', 'convrot_w4a4_linear', 'dequantize_convrot_w4a4_weight', 'dequantize_int8_convrot_weight_dtype', 'dequantize_int8_simple_dtype', 'dequantize_per_tensor_fp8', 'dequantize_w4a8_int8_weight', 'gemv_awq_w4a16', 'int8_linear', 'na3d', 'quantize_and_rotate_rowwise', 'quantize_convrot_w4a4_weight', 'quantize_int8_convrot_weight', 'quantize_int8_rowwise', 'quantize_int8_tensorwise', 'quantize_per_tensor_fp8', 'quantize_svdquant_w4a4', 'quantize_w4a8_int8_weight', 'rms_adaln', 'rms_rope', 'rms_rope1', 'rms_rope1_', 'rms_rope_', 'rms_rope_split_half', 'rms_rope_split_half1', 'rms_rope_split_half1_', 'rms_rope_split_half_', 'scaled_mm_svdquant_w4a4', 'stochastic_rounding_fp8', 'w4a8_int8_linear']}
[INFO] Found comfy_kitchen backend cuda: {'available': True, 'disabled': True, 'unavailable_reason': None, 'capabilities': ['adaln', 'apply_rope', 'apply_rope1', 'apply_rope1_', 'apply_rope_', 'apply_rope_split_half', 'apply_rope_split_half1', 'apply_rope_split_half1_', 'apply_rope_split_half_', 'convrot_w4a4_linear', 'dequantize_convrot_w4a4_weight', 'dequantize_int8_convrot_weight', 'dequantize_int8_convrot_weight_dtype', 'dequantize_int8_simple', 'dequantize_int8_simple_dtype', 'dequantize_nvfp4', 'dequantize_per_tensor_fp8', 'dequantize_w4a8_int8_weight', 'gemv_awq_w4a16', 'na3d', 'prepare_int4_weight_for_int8_linear', 'quantize_and_rotate_rowwise', 'quantize_convrot_w4a4_weight', 'quantize_int8_convrot_weight', 'quantize_int8_rowwise', 'quantize_int8_tensorwise', 'quantize_mxfp8', 'quantize_nvfp4', 'quantize_per_tensor_fp8', 'quantize_svdquant_w4a4', 'quantize_w4a8_int8_weight', 'rms_adaln', 'rms_rope', 'rms_rope1', 'rms_rope1_', 'rms_rope_', 'rms_rope_split_half', 'rms_rope_split_half1', 'rms_rope_split_half1_', 'rms_rope_split_half_', 'rotate_int8_convrot_weight', 'scaled_mm_svdquant_w4a4', 'stochastic_rounding_fp8', 'w4a8_int8_linear']}
[INFO] Checkpoint files will always be loaded safely.
[INFO] Total VRAM 16304 MB, total RAM 32134 MB
[INFO] pytorch version: 2.12.0+rocm7.14.0
[INFO] Set: torch.backends.cudnn.enabled = False for better AMD performance.
[INFO] AMD arch: gfx1201
[INFO] ROCm version: (7, 14)
[INFO] Set vram state to: NORMAL_VRAM
[INFO] Device: cuda:0 AMD Radeon RX 9070 XT : native
[INFO] Using async weight offloading with 2 streams
[INFO] Enabled pinned memory 12853.0
[INFO] Using pytorch attention
[INFO] aimdo: src-win/cuda-detour.c:38:INFO:aimdo_setup_hooks: installing 6 hooks
[INFO] aimdo: src-win/shmem-detect.c:80:INFO:comfy-aimdo WDDM adapter match: AMD Radeon RX 9070 XT runtime_luid=00000000:0000ec8d dxgi_luid=00000000:0000ec8d
[INFO] aimdo: src/control.c:277:INFO:comfy-aimdo inited for GPU: AMD Radeon RX 9070 XT (VRAM: 16304 MB)
[INFO] DynamicVRAM support detected and enabled
[INFO] Python version: 3.12.12 (main, Feb 12 2026, 00:40:26) [MSC v.1944 64 bit (AMD64)]
[INFO] ComfyUI version: 0.34.3
[INFO] comfy-aimdo version: 0.4.15
[INFO] comfy-kitchen version: 0.2.31
[INFO] comfyui-frontend-package version: 1.49.6
[INFO] comfyui-workflow-templates version: 0.11.54
[INFO] comfyui-embedded-docs version: 0.5.10
[INFO] comfy-kitchen version: 0.2.31
[INFO] comfy-aimdo version: 0.4.15
[INFO] [Prompt Server] web root: C:\Users\brf-2\ComfyUI-Installs\ComfyUI\ComfyUI\.venv\Lib\site-packages\comfyui_frontend_package\static
[INFO] Asset seeder disabled
[INFO] [START] ComfyUI-Manager
[ComfyUI-Manager] Using GitPython backend
[INFO] [ComfyUI-Manager] network_mode: public
[WARNING] [ComfyUI-Manager] The matrix sharing feature has been disabled because the `matrix-nio` dependency is not installed.
To use this feature, please run the following command:
C:\Users\brf-2\ComfyUI-Installs\ComfyUI\ComfyUI\.venv\Scripts\python.exe -m pip install matrix-nio

[INFO] No OpenGL_accelerate module loaded: No module named 'OpenGL_accelerate'
[INFO]
Import times for custom nodes:
[INFO] 0.0 seconds: C:\Users\brf-2\ComfyUI-Installs\ComfyUI\ComfyUI\custom_nodes\websocket_image_save.py
[INFO]
[INFO] Context impl SQLiteImpl.
[INFO] Will assume non-transactional DDL.
[INFO] Using RAM pressure cache.
[INFO] Starting server

[INFO] To see the GUI go to: http://127.0.0.1:8188
[INFO] got prompt
[INFO] model weight dtype torch.float16, manual cast: None
[INFO] model_type EPS
[INFO] Using split attention in VAE
[INFO] Using split attention in VAE
[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.bfloat16
[INFO] CLIP/text encoder model load device: cuda:0, offload device: cpu, current: cpu, dtype: torch.float16
[INFO] Requested to load SDXLClipModel
[INFO] Model SDXLClipModel prepared for dynamic VRAM loading. 1560MB Staged. 0 patches attached. Force pre-loaded 180 weights: 400 KB.
[INFO] Model SDXLClipModel prepared for dynamic VRAM loading. 1560MB Staged. 0 patches attached. Force pre-loaded 180 weights: 400 KB.
[INFO] Requested to load SDXL
[INFO] Model SDXL prepared for dynamic VRAM loading. 4896MB Staged. 0 patches attached. Force pre-loaded 512 weights: 1197 KB.
0%| | 0/30 [00:00

Contributor guide

Open the contributing guide

Research direction

Start by loading the attached cyberIllustrious.json workflow on the reported RX 9070 XT setup and compare VAE Decode with DynamicVRAM enabled and disabled. Review Title.md and the supplied startup and execution logs, focusing on the DynamicVRAM and AutoencoderKL stages. Done means normal image generation without severe GPU or UI stalling.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.