rtx 5090 - Extreme slowdowns once the process reaches ksampler
- Dominant language
- Python
- Stars
- 133k
- Forks
- 15.7k
- Avg merge
- 1d 7h
- Merged PRs (30d)
- 158
Description
### Custom Node Testing
- [x] I have tried disabling custom nodes and the issue persists (see [how to disable custom nodes](https://docs.comfy.org/troubleshooting/custom-node-issues#step-1%3A-test-with-all-custom-nodes-disabled) if you need help)
### Your question
Hi,
I recently rebuilt my computer and installed the portable version of ComfyUI and Iβm experiencing unusually slow generation times on a high-end setup.
(originally I had the electron version installed by mistake, that had the exact same issues though)
I'm running a 5090 with 32GB VRAM, and even simple prompts using the WAN2.2 fp8_scaled models are taking nearly 2 minutes per image.
For now I'm just running the Wan2.2 text to video default workflow you can download from the templates section.
System Details:
GPU: RTX 5090 (32GB VRAM)
ComfyUI version: 0.3.48 (portable)
Python: 3.12.10 (Comfy portable Python)
PyTorch: 2.7.1+cu128
I tried modifying my commandline flags for run_nvidia_gpu.bat:
```.\python_embeded\python.exe -s ComfyUI\main.py --windows-standalone-build --fp16-vae --use-pytorch-cross-attention
pause
```
Here's a screenshot of the default workflow:
I'm getting ~2 minutes per image (KSamplerAdvanced reports 112s per frame)
Logs show:
```VAE load device: cuda:0, offload device: cpu
CLIP/text encoder model load device: cuda:0, offload device: cpu
Using scaled fp8: fp8 matrix mult: False, scale input: False
```
This happens with batch_size = 1 and resolution 1280x704
I have not installed Triton or Sage yet as I need to follow a tutorial on doing so, and I'd like to resolve my slow issues before I add more potential complications.
I also find that my regular image generation is abnormally slow for a 5090.
I tried using --gpu-only but that ended up giving me the error:
```KSamplerAdvanced
Allocation on device
This error means you ran out of memory on your GPU.```
My page file size for the drive where comfy is installed is:
Initial size 32768
Max size 65536
### Logs
```powershell
C:...\Documents\comfyui portable>.\python_embeded\python.exe -s ComfyUI\main.py --windows-standalone-build --fp16-vae --use-pytorch-cross-attention
[START] Security scan
[DONE] Security scan
## ComfyUI-Manager: installing dependencies done.
** ComfyUI startup time: 2025-08-04 16:33:03.870
** Platform: Windows
** Python version: 3.12.10 (tags/v3.12.10:0cc8128, Apr 8 2025, 12:21:36) [MSC v.1943 64 bit (AMD64)]
** Python executable: C:...\Documents\comfyui portable\python_embeded\python.exe
** ComfyUI Path: C:...\Documents\comfyui portable\ComfyUI
** ComfyUI Base Folder Path: C:...\Documents\comfyui portable\ComfyUI
** User directory: C:...\Documents\comfyui portable\ComfyUI\user
** ComfyUI-Manager config path: C:...\Documents\comfyui portable\ComfyUI\user\default\ComfyUI-Manager\config.ini
** Log path: C:...\Documents\comfyui portable\ComfyUI\user\comfyui.log
Prestartup times for custom nodes:
0.0 seconds: C:...\Documents\comfyui portable\ComfyUI\custom_nodes\rgthree-comfy
1.5 seconds: C:...\Documents\comfyui portable\ComfyUI\custom_nodes\comfyui-manager
Checkpoint files will always be loaded safely.
Total VRAM 32607 MB, total RAM 63033 MB
pytorch version: 2.7.1+cu128
Set vram state to: NORMAL_VRAM
Device: cuda:0 NVIDIA GeForce RTX 5090 : cudaMallocAsync
Using pytorch attention
Python version: 3.12.10 (tags/v3.12.10:0cc8128, Apr 8 2025, 12:21:36) [MSC v.1943 64 bit (AMD64)]
ComfyUI version: 0.3.48
ComfyUI frontend version: 1.23.4
[Prompt Server] web root: C:...\Documents\comfyui portable\python_embeded\Lib\site-packages\comfyui_frontend_package\static
### Loading: ComfyUI-Manager (V3.35)
[ComfyUI-Manager] network_mode: public
### ComfyUI Revision: 150 [bff60b5c] *DETACHED | Released on '2025-08-01'
[ComfyUI-Manager] default cache updated: https://raw.githubusercontent.com/ltdrdata/ComfyUI-Manager/main/alter-list.json
[ComfyUI-Manager] default cache updated: https://raw.githubusercontent.com/ltdrdata/ComfyUI-Manager/main/model-list.json
[ComfyUI-Manager] default cache updated: https://raw.githubusercontent.com/ltdrdata/ComfyUI-Manager/main/github-stats.json
[ComfyUI-Manager] default cache updated: https://raw.githubusercontent.com/ltdrdata/ComfyUI-Manager/main/extension-node-map.json
[ComfyUI-Manager] default cache updated: https://raw.githubusercontent.com/ltdrdata/ComfyUI-Manager/main/custom-node-list.json
[rgthree-comfy] Loaded 48 epic nodes. π
Import times for custom nodes:
0.0 seconds: C:...\Documents\comfyui portable\ComfyUI\custom_nodes\websocket_image_save.py
0.0 seconds: C:...\Documents\comfyui portable\ComfyUI\custom_nodes\comfyui-custom-scripts
0.0 seconds: C:...\Documents\comfyui portable\ComfyUI\custom_nodes\comfyui_essentials
0.0 seconds: C:...\Documents\comfyui portable\ComfyUI\custom_nodes\comfyui-kjnodes
0.0 seconds: C:...\Documents\comfyui portable\ComfyUI\custom_nodes\comfyui-videohelpersuite
0.1 seconds: C:...\Documents\comfyui portable\ComfyUI\custom_nodes\comfyui-manager
0.2 seconds: C:...\Documents\comfyui portable\ComfyUI\custom_nodes\rgthree-comfy
Context impl SQLiteImpl.
Will assume non-transactional DDL.
No target revision found.
Starting server
To see the GUI go to: http://127.0.0.1:8188
FETCH ComfyRegistry Data: 5/93
FETCH ComfyRegistry Data: 10/93
FETCH ComfyRegistry Data: 15/93
FETCH ComfyRegistry Data: 20/93
FETCH ComfyRegistry Data: 25/93
FETCH ComfyRegistry Data: 30/93
FETCH ComfyRegistry Data: 35/93
FETCH ComfyRegistry Data: 40/93
FETCH ComfyRegistry Data: 45/93
FETCH ComfyRegistry Data: 50/93
FETCH ComfyRegistry Data: 55/93
FETCH ComfyRegistry Data: 60/93
FETCH ComfyRegistry Data: 65/93
FETCH ComfyRegistry Data: 70/93
FETCH ComfyRegistry Data: 75/93
FETCH ComfyRegistry Data: 80/93
FETCH ComfyRegistry Data: 85/93
FETCH ComfyRegistry Data: 90/93
FETCH ComfyRegistry Data [DONE]
[ComfyUI-Manager] default cache updated: https://api.comfy.org/nodes
FETCH DATA from: https://raw.githubusercontent.com/ltdrdata/ComfyUI-Manager/main/custom-node-list.json [DONE]
[ComfyUI-Manager] All startup tasks have been completed.
got prompt
Using pytorch attention in VAE
Using pytorch attention in VAE
VAE load device: cuda:0, offload device: cpu, dtype: torch.float16
Using scaled fp8: fp8 matrix mult: False, scale input: False
Requested to load WanTEModel
loaded completely 9.5367431640625e+25 6419.477203369141 True
CLIP/text encoder model load device: cuda:0, offload device: cpu, current: cuda:0, dtype: torch.float16
Using scaled fp8: fp8 matrix mult: True, scale input: True
model weight dtype torch.float16, manual cast: None
model_type FLOW
Requested to load WAN21
loaded completely 14876.058919891357 13627.512924194336 True
80%|ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ | 8/10 [16:15<04:06, 123.29s/it]
```
### Other
_No response_
Contributor guide
Assessment
This issue has not been assessed yet.