lllyasviel / lllyasviel/stable-diffusion-webui-forge

[Bug]: Forge keeps using more and more VRAM with every generation

Open
#343 15 comments 3 reactions 0 assignees View on GitHub
AMD
Dominant language
Python
Stars
13k
Forks
1.7k
PR merge metrics
No merged PRs in 30d

Description

### Checklist

- [X] The issue exists after disabling all extensions
- [X] The issue exists on a clean installation of webui
- [ ] The issue is caused by an extension, but I believe it is caused by a bug in the webui
- [X] The issue exists in the current version of the webui
- [X] The issue has not been reported before recently
- [ ] The issue has been reported before but has not been fixed yet

### What happened?

With every subsequent generation the VRAM used by forge increases until it nears the max vram of my card. At this point the excessive vram usage will cause screen flickering or a blackscreen. Restarting the application frequently prevents the issue from arising.

### Steps to reproduce the problem

1. Generate multiple images (about 3 in 1024*1024 sdxl)
2. Drivers cause screenflickering/blackscreen because they don't have enough vram to work properly?

### What should have happened?

WebUI should use the same amount of VRAM on each generation

### What browsers do you use to access the UI ?

Mozilla Firefox

### Sysinfo

[sysinfo-2024-02-20-21-45.json](https://github.com/lllyasviel/stable-diffusion-webui-forge/files/14351454/sysinfo-2024-02-20-21-45.json)

### Console logs

```Shell
Python 3.10.12 (main, Nov 20 2023, 15:14:05) [GCC 11.4.0]
Version: f0.0.14v1.8.0rc-latest-184-g43c9e3b5
Commit hash: 43c9e3b5ce1642073c7a9684e36b45489eeb4a49
Legacy Preprocessor init warning: Unable to install insightface automatically. Please try run `pip install insightface` manually.
Launching Web UI with arguments: --listen --enable-insecure-extension-access --theme dark
Total VRAM 8176 MB, total RAM 15834 MB
Set vram state to: NORMAL_VRAM
Device: cuda:0 AMD Radeon RX 6600M : native
VAE dtype: torch.float32
2024-02-20 22:30:54.971808: I external/local_tsl/tsl/cuda/cudart_stub.cc:31] Could not find cuda drivers on your machine, GPU will not be used.
2024-02-20 22:30:55.079085: E external/local_xla/xla/stream_executor/cuda/cuda_dnn.cc:9261] Unable to register cuDNN factory: Attempting to register factory for plugin cuDNN when one has already been registered
2024-02-20 22:30:55.079167: E external/local_xla/xla/stream_executor/cuda/cuda_fft.cc:607] Unable to register cuFFT factory: Attempting to register factory for plugin cuFFT when one has already been registered
2024-02-20 22:30:55.096649: E external/local_xla/xla/stream_executor/cuda/cuda_blas.cc:1515] Unable to register cuBLAS factory: Attempting to register factory for plugin cuBLAS when one has already been registered
2024-02-20 22:30:55.138842: I external/local_tsl/tsl/cuda/cudart_stub.cc:31] Could not find cuda drivers on your machine, GPU will not be used.
2024-02-20 22:30:55.139490: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations.
To enable the following instructions: AVX2 FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags.
2024-02-20 22:30:55.990032: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT
Using sub quadratic optimization for cross attention, if you have memory or speed issues try using: --attention-split
ControlNet preprocessor location: /home/user/stable-diffusion-webui-forge/models/ControlNetPreprocessor
Tag Autocomplete: Could not locate model-keyword extension, Lora trigger word completion will be limited to those added through the extra networks menu.
[-] ADetailer initialized. version: 24.1.2, num models: 9
Loading weights [821aa5537f] from /home/user/stable-diffusion-webui-forge/models/Stable-diffusion/autismmixSDXL_autismmixPony.safetensors
2024-02-20 22:31:01,670 - ControlNet - INFO - ControlNet UI callback registered.
Running on local URL: http://0.0.0.0:7860

To create a public link, set `share=True` in `launch()`.
Startup time: 15.4s (prepare environment: 2.7s, import torch: 4.0s, import gradio: 0.9s, setup paths: 3.2s, other imports: 0.4s, load scripts: 3.2s, create ui: 0.8s, gradio launch: 0.2s).
model_type EPS
UNet ADM Dimension 2816
Using split attention in VAE
Working with z of shape (1, 4, 32, 32) = 4096 dimensions.
Using split attention in VAE
extra {'cond_stage_model.clip_l.text_projection', 'cond_stage_model.clip_g.transformer.text_model.embeddings.position_ids', 'cond_stage_model.clip_l.logit_scale'}
To load target model SDXLClipModel
Begin to load 1 model
Moving model(s) has taken 0.63 seconds
Model loaded in 8.5s (load weights from disk: 1.2s, forge load real models: 5.6s, calculate empty prompt: 1.7s).
WARNING: Invalid HTTP request received.
To load target model SDXLClipModel
Begin to load 1 model
unload clone 0
Moving model(s) has taken 0.71 seconds
To load target model SDXL
Begin to load 1 model
Moving model(s) has taken 9.94 seconds
100%|███████████████████████████████████████████| 20/20 [00:36<00:00, 1.80s/it]
To load target model AutoencoderKL██████████████| 20/20 [00:33<00:00, 1.75s/it]
Begin to load 1 model
Moving model(s) has taken 0.44 seconds
Warning: Ran out of memory when regular VAE decoding, retrying with tiled VAE decoding.
Total progress: 100%|███████████████████████████| 20/20 [00:40<00:00, 2.02s/it]
To load target model SDXLClipModel██████████████| 20/20 [00:40<00:00, 1.75s/it]
Begin to load 1 model
Moving model(s) has taken 0.30 seconds
To load target model SDXL
Begin to load 1 model
Moving model(s) has taken 1.92 seconds
100%|███████████████████████████████████████████| 20/20 [00:34<00:00, 1.74s/it]
To load target model AutoencoderKL██████████████| 20/20 [00:33<00:00, 1.74s/it]
Begin to load 1 model
Moving model(s) has taken 0.49 seconds
Warning: Ran out of memory when regular VAE decoding, retrying with tiled VAE decoding.
Total progress: 100%|███████████████████████████| 20/20 [00:41<00:00, 2.08s/it]
To load target model SDXL███████████████████████| 20/20 [00:41<00:00, 1.74s/it]
Begin to load 1 model
Moving model(s) has taken 1.35 seconds
100%|███████████████████████████████████████████| 20/20 [00:37<00:00, 1.87s/it]
To load target model AutoencoderKL██████████████| 20/20 [00:35<00:00, 1.84s/it]
Begin to load 1 model
Moving model(s) has taken 0.44 seconds
Warning: Ran out of memory when regular VAE decoding, retrying with tiled VAE decoding.
Total progress: 100%|███████████████████████████| 20/20 [00:42<00:00, 2.14s/it]
To load target model SDXL███████████████████████| 20/20 [00:42<00:00, 1.84s/it]
Begin to load 1 model
Moving model(s) has taken 1.36 seconds
100%|███████████████████████████████████████████| 20/20 [00:37<00:00, 1.87s/it]
To load target model AutoencoderKL██████████████| 20/20 [00:35<00:00, 1.89s/it]
Begin to load 1 model
Moving model(s) has taken 0.56 seconds
Warning: Ran out of memory when regular VAE decoding, retrying with tiled VAE decoding.
Total progress: 100%|███████████████████████████| 20/20 [00:42<00:00, 2.12s/it]
To load target model SDXL███████████████████████| 20/20 [00:42<00:00, 1.89s/it]
Begin to load 1 model
Moving model(s) has taken 1.41 seconds
100%|███████████████████████████████████████████| 20/20 [00:39<00:00, 2.00s/it]
To load target model AutoencoderKL██████████████| 20/20 [00:37<00:00, 1.96s/it]
Begin to load 1 model
Moving model(s) has taken 0.47 seconds
Warning: Ran out of memory when regular VAE decoding, retrying with tiled VAE decoding.
Total progress: 100%|███████████████████████████| 20/20 [00:44<00:00, 2.24s/it]
To load target model SDXL███████████████████████| 20/20 [00:44<00:00, 1.96s/it]
Begin to load 1 model
Moving model(s) has taken 1.44 seconds
100%|███████████████████████████████████████████| 20/20 [00:36<00:00, 1.83s/it]
To load target model AutoencoderKL██████████████| 20/20 [00:34<00:00, 1.83s/it]
Begin to load 1 model
Moving model(s) has taken 0.44 seconds
Warning: Ran out of memory when regular VAE decoding, retrying with tiled VAE decoding.
Total progress: 100%|███████████████████████████| 20/20 [00:41<00:00, 2.08s/it]
To load target model SDXLClipModel██████████████| 20/20 [00:41<00:00, 1.83s/it]
Begin to load 1 model
Moving model(s) has taken 0.31 seconds
To load target model SDXL
Begin to load 1 model
Moving model(s) has taken 1.77 seconds
100%|███████████████████████████████████████████| 20/20 [00:44<00:00, 2.21s/it]
To load target model AutoencoderKL██████████████| 20/20 [00:42<00:00, 2.29s/it]
Begin to load 1 model
Moving model(s) has taken 0.75 seconds
Warning: Ran out of memory when regular VAE decoding, retrying with tiled VAE decoding.
Total progress: 100%|███████████████████████████| 20/20 [00:49<00:00, 2.49s/it]
To load target model SDXLClipModel██████████████| 20/20 [00:49<00:00, 2.29s/it]
Begin to load 1 model
Moving model(s) has taken 0.48 seconds
To load target model SDXL
Begin to load 1 model
Moving model(s) has taken 1.69 seconds
100%|███████████████████████████████████████████| 20/20 [00:38<00:00, 1.90s/it]
To load target model AutoencoderKL██████████████| 20/20 [00:35<00:00, 1.82s/it]
Begin to load 1 model
Moving model(s) has taken 0.39 seconds
Warning: Ran out of memory when regular VAE decoding, retrying with tiled VAE decoding.
Total progress: 100%|███████████████████████████| 20/20 [00:42<00:00, 2.13s/it]
Total progress: 100%|███████████████████████████| 20/20 [00:42<00:00, 1.82s/it]
```

### Additional information

_No response_

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reviewing the linked sysinfo-2024-02-20-21-45.json and the console logs, then reproduce the reported sequence of repeated 1024×1024 SDXL generations. Compare VRAM usage across generations and investigate the repeated model/VAE loading and out-of-memory messages. Done means VRAM remains stable across generations without screen flickering or a black screen.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.