invoke-ai / invoke-ai/InvokeAI
[bug]: Crashing Krea2 non-standard checkpoints after recent update.
- Dominant language
- Python
- Stars
- 28.2k
- Forks
- 3k
- Avg merge
- 6d 5h
- Merged PRs (30d)
- 19
Description
### Is there an existing issue for this problem?
- [x] I have searched the existing issues
### Install method
Docker image on unRAID server
### Operating system
Linux (unRAID)
### GPU vendor
AMD (ROCm)
### GPU model
AI PRO R9700
### GPU VRAM
32 GiB
### Version number
v6.14.1-post1
### Browser
Edge Version 150.0.4078.48 (Official build) (64-bit)
### System Information
|
-- | --
version | "6.14.0"
dependencies |
absl-py | "2.4.0"
accelerate | "1.14.0"
annotated-doc | "0.0.4"
annotated-types | "0.7.0"
anyio | "4.14.1"
argon2-cffi | "25.1.0"
argon2-cffi-bindings | "25.1.0"
arrow | "1.4.0"
asttokens | "3.0.1"
async-lru | "2.3.0"
attrs | "26.1.0"
babel | "2.18.0"
bcrypt | "3.2.2"
beautifulsoup4 | "4.15.0"
bidict | "0.23.1"
bitsandbytes | "0.49.2"
blake3 | "1.0.9"
bleach | "6.4.0"
certifi | "2026.6.17"
cffi | "2.0.0"
charset-normalizer | "3.4.7"
click | "8.4.2"
coloredlogs | "15.0.1"
comm | "0.2.3"
compel | "2.4.0"
contourpy | "1.3.3"
cryptography | "49.0.0"
CUDA | "N/A"
cycler | "0.12.1"
debugpy | "1.8.21"
decorator | "5.3.1"
defusedxml | "0.7.1"
Deprecated | "1.3.1"
diffusers | "0.39.0"
dnspython | "2.8.0"
dynamicprompts | "0.31.0"
ecdsa | "0.19.2"
einops | "0.8.2"
email-validator | "2.3.0"
executing | "2.2.1"
fastapi | "0.141.1"
fastapi-events | "0.12.2"
fastjsonschema | "2.21.2"
filelock | "3.29.4"
flatbuffers | "25.12.19"
fonttools | "4.63.0"
fqdn | "1.5.1"
fsspec | "2026.6.0"
gguf | "0.19.0"
h11 | "0.16.0"
hf-xet | "1.5.1"
httpcore | "1.0.9"
httptools | "0.8.0"
httpx | "0.28.1"
huggingface_hub | "1.21.0"
humanfriendly | "10.0"
idna | "3.18"
ImageIO | "2.37.4"
imageio-ffmpeg | "0.6.0"
importlib_metadata | "9.0.0"
InvokeAI | "6.14.0"
ipykernel | "7.3.0"
ipython | "9.15.0"
ipython_pygments_lexers | "1.1.1"
isoduration | "20.11.0"
jax | "0.7.1"
jaxlib | "0.7.1"
jedi | "0.20.0"
Jinja2 | "3.1.6"
json5 | "0.15.0"
jsonpointer | "3.1.1"
jsonschema | "4.26.0"
jsonschema-specifications | "2025.9.1"
jupyter-events | "0.12.1"
jupyter-lsp | "2.3.1"
jupyter_builder | "1.0.2"
jupyter_client | "8.9.1"
jupyter_core | "5.9.1"
jupyter_server | "2.20.0"
jupyter_server_terminals | "0.5.4"
jupyterlab | "4.6.0"
jupyterlab_pygments | "0.3.0"
jupyterlab_server | "2.28.0"
kiwisolver | "1.5.0"
lark | "1.3.1"
markdown-it-py | "4.2.0"
MarkupSafe | "3.0.3"
matplotlib | "3.11.0"
matplotlib-inline | "0.2.2"
mdurl | "0.1.2"
mediapipe | "0.10.14"
mistral_common | "1.11.6"
mistune | "3.3.2"
ml_dtypes | "0.5.4"
mpmath | "1.3.0"
nbclient | "0.11.0"
nbconvert | "7.17.1"
nbformat | "5.10.4"
nest-asyncio2 | "1.7.2"
networkx | "3.6.1"
notebook | "7.6.0"
notebook_shim | "0.2.4"
numpy | "1.26.4"
onnx | "1.16.1"
onnxruntime | "1.19.2"
opencv-contrib-python | "4.11.0.86"
opt_einsum | "3.4.0"
packaging | "26.2"
pandocfilters | "1.5.1"
parso | "0.8.7"
passlib | "1.7.4"
pexpect | "4.9.0"
picklescan | "1.0.4"
pillow | "12.2.0"
platformdirs | "4.10.0"
prometheus_client | "0.25.0"
prompt_toolkit | "3.0.52"
protobuf | "4.25.9"
psutil | "7.2.2"
ptyprocess | "0.7.0"
pure_eval | "0.2.3"
pyasn1 | "0.6.3"
pycountry | "26.2.16"
pycparser | "3.0"
pydantic | "2.13.4"
pydantic-extra-types | "2.11.1"
pydantic-settings | "2.14.2"
pydantic_core | "2.46.4"
Pygments | "2.20.0"
pyparsing | "3.3.2"
PyPatchMatch | "1.0.2"
python-dateutil | "2.9.0.post0"
python-dotenv | "1.2.2"
python-engineio | "4.13.3"
python-jose | "3.5.0"
python-json-logger | "4.1.0"
python-multipart | "0.0.32"
python-socketio | "5.16.3"
PyWavelets | "1.9.0"
PyYAML | "6.0.3"
pyzmq | "27.1.0"
referencing | "0.37.0"
regex | "2026.5.9"
requests | "2.34.2"
rfc3339-validator | "0.1.4"
rfc3986-validator | "0.1.1"
rfc3987-syntax | "1.1.0"
rich | "15.0.0"
rpds-py | "2026.5.1"
rsa | "4.9.1"
safetensors | "0.8.0"
scipy | "1.17.1"
semver | "3.0.4"
Send2Trash | "2.1.0"
sentencepiece | "0.2.0"
setuptools | "82.0.1"
shellingham | "1.5.4"
simple-websocket | "1.1.0"
six | "1.17.0"
sounddevice | "0.5.5"
soupsieve | "2.8.4"
spandrel | "0.4.2"
stack-data | "0.6.3"
starlette | "0.48.0"
sympy | "1.14.0"
terminado | "0.18.1"
tiktoken | "0.13.0"
tinycss2 | "1.5.1"
tokenizers | "0.22.2"
torch | "2.10.0+rocm7.1"
torchsde | "0.2.6"
torchvision | "0.25.0+rocm7.1"
tornado | "6.5.7"
tqdm | "4.68.3"
traitlets | "5.15.1"
trampoline | "0.1.2"
transformers | "5.5.4"
triton-rocm | "3.6.0"
typer | "0.25.1"
typing-inspection | "0.4.2"
typing_extensions | "4.15.0"
tzdata | "2026.2"
uri-template | "1.3.0"
urllib3 | "2.7.0"
uvicorn | "0.49.0"
uvloop | "0.22.1"
watchfiles | "1.2.0"
wcwidth | "0.8.1"
webcolors | "25.10.0"
webencodings | "0.5.1"
websocket-client | "1.9.0"
websockets | "16.0"
wrapt | "2.2.2"
wsproto | "1.3.2"
zipp | "4.1.0"
config |
schema_version | "4.0.3"
legacy_models_yaml_path | null
host | "0.0.0.0"
port | 9090
allow_origins | []
allow_credentials | true
allow_methods |
0 | "*"
allow_headers |
0 | "*"
ssl_certfile | null
ssl_keyfile | null
base_url | null
forwarded_allow_ips | "127.0.0.1"
http_compression_level | 9
log_tokenization | false
patchmatch | true
models_dir | "models"
convert_cache_dir | "models/.convert_cache"
download_cache_dir | "models/.download_cache"
legacy_conf_dir | "configs"
db_dir | "databases"
outputs_dir | "outputs"
image_subfolder_strategy | "flat"
custom_nodes_dir | "nodes"
style_presets_dir | "style_presets"
workflow_thumbnails_dir | "workflow_thumbnails"
log_handlers |
0 | "console"
log_format | "color"
log_level | "info"
log_sql | false
log_level_network | "warning"
use_memory_db | false
dev_reload | false
profile_graphs | false
profile_prefix | null
profiles_dir | "profiles"
max_cache_ram_gb | 22
max_cache_vram_gb | null
log_memory_usage | false
model_cache_keep_alive_min | 0
device_working_mem_gb | 3
enable_partial_loading | false
keep_ram_copy_of_weights | false
ram | null
vram | null
lazy_offload | true
pytorch_cuda_alloc_conf | null
device | "auto"
generation_devices | "auto"
offload_text_encoders_to_idle_gpus | true
precision | "bfloat16"
sequential_guidance | false
wan_memory_optimization | false
pid_memory_optimization | false
attention_type | "torch-sdp"
attention_slice_size | "auto"
force_tiled_decode | true
pil_compress_level | 1
max_queue_size | 10000
session_queue_mode | "round_robin"
clear_queue_on_startup | true
max_queue_history | null
allow_nodes | null
deny_nodes | null
node_cache_size | 512
hashing_algorithm | "blake3_single"
remote_api_tokens | null
scan_models_on_startup | false
allow_private_download_urls | false
download_proxy | null
unsafe_disable_picklescan | false
allow_unknown_models | true
multiuser | false
strict_password_checking | false
external_alibabacloud_api_key | null
external_alibabacloud_base_url | null
external_gemini_api_key | null
external_openai_api_key | null
external_gemini_base_url | null
external_openai_base_url | null
external_seedream_api_key | null
external_seedream_base_url | null
set_config_fields |
0 | "max_cache_ram_gb"
1 | "precision"
2 | "attention_type"
3 | "host"
4 | "force_tiled_decode"
5 | "legacy_models_yaml_path"
6 | "port"
7 | "keep_ram_copy_of_weights"
8 | "enable_partial_loading"
9 | "clear_queue_on_startup"
### What happened
[bug]: FP8 and other quantized checkpoints are expanded to full precision on load — resident size is identical regardless of quantization level (ROCm)
Is there an existing issue for this problem?
- [x] I have searched the existing issues
Operating system
Linux (unRAID 7.x, Docker)
GPU vendor
AMD (ROCm)
GPU model
<!-- FILL IN: exact card, gfx1201 / RDNA 4 -->
GPU VRAM
32 GB
Version number
6.14.0 and 6.14.1-post1 — identical behaviour on both.
ghcr.io/invoke-ai/invokeai:6.14.0-rocm—sha256:032161d7873e3d3f13f699fa9a8a10ec3f5a3e86b662f07843f263fd1a245580ghcr.io/invoke-ai/invokeai:main-rocm(6.14.1-post1) —sha256:2bb1a68525a8c6e567ee499ac391a8289c469414fc9b9bea74eeb37caa7d70de
Browser
<!-- FILL IN -->
What happened
On ROCm, quantized checkpoints appear to be dequantized to full precision during
loading. The memory saving that quantization is chosen for is lost entirely, and the
resident footprint is the same no matter which quantization is used.
GGUF checkpoints are unaffected — they load at approximately their on-disk size.
Evidence 1: three different Z Image quantizations, one resident size
Three distinct Z Image checkpoints, three different file sizes, all reporting the
same loaded size:
Model ID | Size on disk | Total model size as loaded
-- | -- | --
d9006a28-cc54-49d3-93d9-3e1652d2cdf5 | 6.16 GiB | 11,739.56 MB
00ce5266-f09d-4287-a4ea-d8ae89f56263 | 6.54 GiB | 11,739.56 MB
f50e705c-df8a-496f-9acf-491fba3d885f | 12.57 GiB | 11,739.56 MB
Two independent FP8 Krea-2 checkpoints both land at exactly 24,452.35 MB. The GGUF
build of the same model family stays at its on-disk size.
precision: bfloat16 is set in invokeai.yaml, and ~2x is what FP8 → bf16 expansion
would produce.
Consequence
The model no longer fits in the cache, so it is evicted and re-staged on every
generation:
- GGUF Krea-2 (12.69 GiB resident, ~21 GiB working set with the encoder):
stays cached, repeatedin 0.00scache hits, 22 s per 960x1344 image. - FP8 Krea-2 (23.88 GiB resident, ~32 GiB working set): never a single cache
hit, full re-stage every generation, 82 s per image, and OOM-killed outright at
container limits below 26 GiB.
The staging phase is host-RAM bound with the GPU idle
Sampling the container cgroup's memory.current and amdgpu's mem_info_vram_used
every 2 seconds during an FP8 load:
13:58:57 sys= 8.2 GiB vram= 0.6 GiB <- load begins
13:59:07 sys= 20.3 GiB vram= 0.6 GiB
13:59:19 sys= 26.0 GiB vram= 0.6 GiB <- pinned at container limit
... 45 seconds pinned, GPU idle at 0.6 GiB ...
14:00:06 sys= 25.6 GiB vram= 1.2 GiB <- transfer to VRAM begins
14:00:12 sys= 6.2 GiB vram= 20.6 GiB
14:00:16 sys= 2.5 GiB vram= 26.0 GiB <- steady state
The weights are fully materialised in system RAM, then copied to VRAM in ~6 seconds.
Steady-state host RAM afterwards is 2.5 GiB. So host RAM must be at least
model-sized during loading even when VRAM is abundant, and the expansion above
doubles what that costs.
What you expected to happen
That an FP8 checkpoint would occupy roughly its on-disk size in memory, and that
selecting a smaller quantization would reduce the memory footprint.
How to reproduce the problem
- ROCm container image on a system with 32 GB VRAM.
- Install two quantizations of the same model — e.g. a 6 GiB and a 12 GiB Z Image
checkpoint, or an FP8 and a Q8_0 GGUF Krea-2. - Load each and compare the
Total model sizevalue in the[MODEL CACHE] Loaded modellog line against the file size on disk. - The quantized non-GGUF checkpoints report a resident size unrelated to their file
size.
Possible cause
The ROCm image does not contain rocminfo, so bitsandbytes cannot detect the GPU
architecture and falls back. Logged at every startup:
Could not detect ROCm GPU architecture: [Errno 2] No such file or directory: 'rocminfo'
ROCm GPU architecture detection failed despite ROCm being available.
Could not detect ROCm warp size: [Errno 2] No such file or directory: 'rocminfo'.
Defaulting to 64. (some 4-bit functions may not work!)
If quantized kernels are unavailable because the architecture could not be
identified, dequantizing the weights at load would be a plausible fallback — and
would explain why GGUF (which has its own dequantization path) behaves differently.
This is speculation on my part; I have not instrumented it. But rocminfo being
absent from the official ROCm image looks like a packaging defect regardless.
Environment
MALLOC_MMAP_THRESHOLD_=1048576
DISABLE_PINNED_MEMORY=1
PYTORCH_HIP_ALLOC_CONF=garbage_collection_threshold:0.8,max_split_size_mb:128
HSA_OVERRIDE_GFX_VERSION=12.0.1
HSA_ENABLE_SDMA=0
PYTORCH_TUNABLEOP_ENABLED=0
TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=0
MIOPEN_FIND_MODE=FAST
GPU_DRIVER=rocm
schema_version: 4.0.3
max_cache_ram_gb: 14.0
keep_ram_copy_of_weights: false
device_working_mem_gb: 3.0
model_cache_keep_alive_min: 0
enable_partial_loading: false
force_tiled_decode: true
device: auto
precision: bfloat16
attention_type: torch-sdp
Host: 32 GB system RAM (31 GiB usable), AMD Ryzen 7 3800X (8C/16T), container limited
to 26 GiB and 6 CPUs. Nothing else memory-heavy running. Components in the Krea-2
workflow: Qwen3-VL 4B encoder (8,464.46 MB), Qwen Image VAE (242.03 MB).
What this is not
Recorded so others don't repeat the dead ends I worked through:
- Not a leak. Host RAM returns to 2.5–7.8 GiB after every load; nothing
accumulates across generations. - Not a regression. 6.14.0 and 6.14.1-post1 are identical. Earlier versions
could not be tested — Krea-2 support does not exist in 6.13.x. - Not fixable via
max_cache_ram_gb. Tested at 8.0, 12.0 and 14.0. - Not glibc fragmentation. The FAQ's
MALLOC_MMAP_THRESHOLD_=1048576workaround
was in place throughout, withDISABLE_PINNED_MEMORY=1. - Not a GPU fault. No
amdgpuring timeouts, resets or segfaults indmesg.
All OOM kills areCONSTRAINT_MEMCG, contained to the container cgroup, with no
Python traceback.
Three smaller issues noticed along the way
Happy to split these out.
1. The RAM cache statistics block never updates. Byte-identical across
consecutive generations (hits: 2 / misses: 4 / cached: 1 / cleared: 1) despite each
run loading several models.
2. Model load time is attributed to the denoise node. krea2_denoise was billed
at 65.2 s while the 9 sampling steps took ~20 s (2.26 s/it) and the Loaded model
line reported 3.16 s. The ~45 s of staging is folded into denoise and invisible in
the graph stats — the timing table points at sampling when the cost is in loading.
This made the problem substantially harder to locate.
3. force_tiled_decode: true appeared to have no effect. VRAM still spiked to
31.6 GiB at VAE decode with it enabled. Possibly not wired to the Qwen Image VAE
path.
### What you expected to happen
I expected to be able to run the models.
### How to reproduce the problem
Start the app load any model which is not quantized.
### Additional context
_No response_
### Discord username
C_Dollerup
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reproducing the load comparison described for the FP8, quantized Z Image, and GGUF checkpoints, using the precision setting in invokeai.yaml and recording resident memory and cache behavior. Trace the model-loading path responsible for the FP8-to-bfloat16 expansion. Done means quantized checkpoints retain an appropriately reduced memory footprint, avoid repeated restaging, and GGUF behavior remains unaffected.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- docker, python, pytorch
- Domain
- machine-learning, performance
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 38/100