invoke-ai / invoke-ai/InvokeAI

[bug]: Crashing Krea2 non-standard checkpoints after recent update.

Open
#9,585 9 comments 0 reactions 0 assignees View on GitHub
bug
Dominant language
Python
Stars
28.2k
Forks
3k
Avg merge
6d 5h
Merged PRs (30d)
19

Description

### Is there an existing issue for this problem?

- [x] I have searched the existing issues

### Install method

Docker image on unRAID server

### Operating system

Linux (unRAID)

### GPU vendor

AMD (ROCm)

### GPU model

AI PRO R9700

### GPU VRAM

32 GiB

### Version number

v6.14.1-post1

### Browser

Edge Version 150.0.4078.48 (Official build) (64-bit)

### System Information


  |  
-- | --
version | "6.14.0"
dependencies |  
absl-py | "2.4.0"
accelerate | "1.14.0"
annotated-doc | "0.0.4"
annotated-types | "0.7.0"
anyio | "4.14.1"
argon2-cffi | "25.1.0"
argon2-cffi-bindings | "25.1.0"
arrow | "1.4.0"
asttokens | "3.0.1"
async-lru | "2.3.0"
attrs | "26.1.0"
babel | "2.18.0"
bcrypt | "3.2.2"
beautifulsoup4 | "4.15.0"
bidict | "0.23.1"
bitsandbytes | "0.49.2"
blake3 | "1.0.9"
bleach | "6.4.0"
certifi | "2026.6.17"
cffi | "2.0.0"
charset-normalizer | "3.4.7"
click | "8.4.2"
coloredlogs | "15.0.1"
comm | "0.2.3"
compel | "2.4.0"
contourpy | "1.3.3"
cryptography | "49.0.0"
CUDA | "N/A"
cycler | "0.12.1"
debugpy | "1.8.21"
decorator | "5.3.1"
defusedxml | "0.7.1"
Deprecated | "1.3.1"
diffusers | "0.39.0"
dnspython | "2.8.0"
dynamicprompts | "0.31.0"
ecdsa | "0.19.2"
einops | "0.8.2"
email-validator | "2.3.0"
executing | "2.2.1"
fastapi | "0.141.1"
fastapi-events | "0.12.2"
fastjsonschema | "2.21.2"
filelock | "3.29.4"
flatbuffers | "25.12.19"
fonttools | "4.63.0"
fqdn | "1.5.1"
fsspec | "2026.6.0"
gguf | "0.19.0"
h11 | "0.16.0"
hf-xet | "1.5.1"
httpcore | "1.0.9"
httptools | "0.8.0"
httpx | "0.28.1"
huggingface_hub | "1.21.0"
humanfriendly | "10.0"
idna | "3.18"
ImageIO | "2.37.4"
imageio-ffmpeg | "0.6.0"
importlib_metadata | "9.0.0"
InvokeAI | "6.14.0"
ipykernel | "7.3.0"
ipython | "9.15.0"
ipython_pygments_lexers | "1.1.1"
isoduration | "20.11.0"
jax | "0.7.1"
jaxlib | "0.7.1"
jedi | "0.20.0"
Jinja2 | "3.1.6"
json5 | "0.15.0"
jsonpointer | "3.1.1"
jsonschema | "4.26.0"
jsonschema-specifications | "2025.9.1"
jupyter-events | "0.12.1"
jupyter-lsp | "2.3.1"
jupyter_builder | "1.0.2"
jupyter_client | "8.9.1"
jupyter_core | "5.9.1"
jupyter_server | "2.20.0"
jupyter_server_terminals | "0.5.4"
jupyterlab | "4.6.0"
jupyterlab_pygments | "0.3.0"
jupyterlab_server | "2.28.0"
kiwisolver | "1.5.0"
lark | "1.3.1"
markdown-it-py | "4.2.0"
MarkupSafe | "3.0.3"
matplotlib | "3.11.0"
matplotlib-inline | "0.2.2"
mdurl | "0.1.2"
mediapipe | "0.10.14"
mistral_common | "1.11.6"
mistune | "3.3.2"
ml_dtypes | "0.5.4"
mpmath | "1.3.0"
nbclient | "0.11.0"
nbconvert | "7.17.1"
nbformat | "5.10.4"
nest-asyncio2 | "1.7.2"
networkx | "3.6.1"
notebook | "7.6.0"
notebook_shim | "0.2.4"
numpy | "1.26.4"
onnx | "1.16.1"
onnxruntime | "1.19.2"
opencv-contrib-python | "4.11.0.86"
opt_einsum | "3.4.0"
packaging | "26.2"
pandocfilters | "1.5.1"
parso | "0.8.7"
passlib | "1.7.4"
pexpect | "4.9.0"
picklescan | "1.0.4"
pillow | "12.2.0"
platformdirs | "4.10.0"
prometheus_client | "0.25.0"
prompt_toolkit | "3.0.52"
protobuf | "4.25.9"
psutil | "7.2.2"
ptyprocess | "0.7.0"
pure_eval | "0.2.3"
pyasn1 | "0.6.3"
pycountry | "26.2.16"
pycparser | "3.0"
pydantic | "2.13.4"
pydantic-extra-types | "2.11.1"
pydantic-settings | "2.14.2"
pydantic_core | "2.46.4"
Pygments | "2.20.0"
pyparsing | "3.3.2"
PyPatchMatch | "1.0.2"
python-dateutil | "2.9.0.post0"
python-dotenv | "1.2.2"
python-engineio | "4.13.3"
python-jose | "3.5.0"
python-json-logger | "4.1.0"
python-multipart | "0.0.32"
python-socketio | "5.16.3"
PyWavelets | "1.9.0"
PyYAML | "6.0.3"
pyzmq | "27.1.0"
referencing | "0.37.0"
regex | "2026.5.9"
requests | "2.34.2"
rfc3339-validator | "0.1.4"
rfc3986-validator | "0.1.1"
rfc3987-syntax | "1.1.0"
rich | "15.0.0"
rpds-py | "2026.5.1"
rsa | "4.9.1"
safetensors | "0.8.0"
scipy | "1.17.1"
semver | "3.0.4"
Send2Trash | "2.1.0"
sentencepiece | "0.2.0"
setuptools | "82.0.1"
shellingham | "1.5.4"
simple-websocket | "1.1.0"
six | "1.17.0"
sounddevice | "0.5.5"
soupsieve | "2.8.4"
spandrel | "0.4.2"
stack-data | "0.6.3"
starlette | "0.48.0"
sympy | "1.14.0"
terminado | "0.18.1"
tiktoken | "0.13.0"
tinycss2 | "1.5.1"
tokenizers | "0.22.2"
torch | "2.10.0+rocm7.1"
torchsde | "0.2.6"
torchvision | "0.25.0+rocm7.1"
tornado | "6.5.7"
tqdm | "4.68.3"
traitlets | "5.15.1"
trampoline | "0.1.2"
transformers | "5.5.4"
triton-rocm | "3.6.0"
typer | "0.25.1"
typing-inspection | "0.4.2"
typing_extensions | "4.15.0"
tzdata | "2026.2"
uri-template | "1.3.0"
urllib3 | "2.7.0"
uvicorn | "0.49.0"
uvloop | "0.22.1"
watchfiles | "1.2.0"
wcwidth | "0.8.1"
webcolors | "25.10.0"
webencodings | "0.5.1"
websocket-client | "1.9.0"
websockets | "16.0"
wrapt | "2.2.2"
wsproto | "1.3.2"
zipp | "4.1.0"
config |  
schema_version | "4.0.3"
legacy_models_yaml_path | null
host | "0.0.0.0"
port | 9090
allow_origins | []
allow_credentials | true
allow_methods |  
0 | "*"
allow_headers |  
0 | "*"
ssl_certfile | null
ssl_keyfile | null
base_url | null
forwarded_allow_ips | "127.0.0.1"
http_compression_level | 9
log_tokenization | false
patchmatch | true
models_dir | "models"
convert_cache_dir | "models/.convert_cache"
download_cache_dir | "models/.download_cache"
legacy_conf_dir | "configs"
db_dir | "databases"
outputs_dir | "outputs"
image_subfolder_strategy | "flat"
custom_nodes_dir | "nodes"
style_presets_dir | "style_presets"
workflow_thumbnails_dir | "workflow_thumbnails"
log_handlers |  
0 | "console"
log_format | "color"
log_level | "info"
log_sql | false
log_level_network | "warning"
use_memory_db | false
dev_reload | false
profile_graphs | false
profile_prefix | null
profiles_dir | "profiles"
max_cache_ram_gb | 22
max_cache_vram_gb | null
log_memory_usage | false
model_cache_keep_alive_min | 0
device_working_mem_gb | 3
enable_partial_loading | false
keep_ram_copy_of_weights | false
ram | null
vram | null
lazy_offload | true
pytorch_cuda_alloc_conf | null
device | "auto"
generation_devices | "auto"
offload_text_encoders_to_idle_gpus | true
precision | "bfloat16"
sequential_guidance | false
wan_memory_optimization | false
pid_memory_optimization | false
attention_type | "torch-sdp"
attention_slice_size | "auto"
force_tiled_decode | true
pil_compress_level | 1
max_queue_size | 10000
session_queue_mode | "round_robin"
clear_queue_on_startup | true
max_queue_history | null
allow_nodes | null
deny_nodes | null
node_cache_size | 512
hashing_algorithm | "blake3_single"
remote_api_tokens | null
scan_models_on_startup | false
allow_private_download_urls | false
download_proxy | null
unsafe_disable_picklescan | false
allow_unknown_models | true
multiuser | false
strict_password_checking | false
external_alibabacloud_api_key | null
external_alibabacloud_base_url | null
external_gemini_api_key | null
external_openai_api_key | null
external_gemini_base_url | null
external_openai_base_url | null
external_seedream_api_key | null
external_seedream_base_url | null
set_config_fields |  
0 | "max_cache_ram_gb"
1 | "precision"
2 | "attention_type"
3 | "host"
4 | "force_tiled_decode"
5 | "legacy_models_yaml_path"
6 | "port"
7 | "keep_ram_copy_of_weights"
8 | "enable_partial_loading"
9 | "clear_queue_on_startup"

### What happened

[bug]: FP8 and other quantized checkpoints are expanded to full precision on load — resident size is identical regardless of quantization level (ROCm)


Is there an existing issue for this problem?



  • [x] I have searched the existing issues


Operating system


Linux (unRAID 7.x, Docker)


GPU vendor


AMD (ROCm)


GPU model


<!-- FILL IN: exact card, gfx1201 / RDNA 4 -->

GPU VRAM


32 GB


Version number


6.14.0 and 6.14.1-post1 — identical behaviour on both.



  • ghcr.io/invoke-ai/invokeai:6.14.0-rocmsha256:032161d7873e3d3f13f699fa9a8a10ec3f5a3e86b662f07843f263fd1a245580

  • ghcr.io/invoke-ai/invokeai:main-rocm (6.14.1-post1) — sha256:2bb1a68525a8c6e567ee499ac391a8289c469414fc9b9bea74eeb37caa7d70de


Browser


<!-- FILL IN -->


What happened


On ROCm, quantized checkpoints appear to be dequantized to full precision during
loading. The memory saving that quantization is chosen for is lost entirely, and the
resident footprint is the same no matter which quantization is used.


GGUF checkpoints are unaffected — they load at approximately their on-disk size.


Evidence 1: three different Z Image quantizations, one resident size


Three distinct Z Image checkpoints, three different file sizes, all reporting the
same loaded size:

Model ID | Size on disk | Total model size as loaded
-- | -- | --
d9006a28-cc54-49d3-93d9-3e1652d2cdf5 | 6.16 GiB | 11,739.56 MB
00ce5266-f09d-4287-a4ea-d8ae89f56263 | 6.54 GiB | 11,739.56 MB
f50e705c-df8a-496f-9acf-491fba3d885f | 12.57 GiB | 11,739.56 MB

Two independent FP8 Krea-2 checkpoints both land at exactly 24,452.35 MB. The GGUF
build of the same model family stays at its on-disk size.


precision: bfloat16 is set in invokeai.yaml, and ~2x is what FP8 → bf16 expansion
would produce.


Consequence


The model no longer fits in the cache, so it is evicted and re-staged on every
generation:



  • GGUF Krea-2 (12.69 GiB resident, ~21 GiB working set with the encoder):
    stays cached, repeated in 0.00s cache hits, 22 s per 960x1344 image.

  • FP8 Krea-2 (23.88 GiB resident, ~32 GiB working set): never a single cache
    hit, full re-stage every generation, 82 s per image, and OOM-killed outright at
    container limits below 26 GiB.


The staging phase is host-RAM bound with the GPU idle


Sampling the container cgroup's memory.current and amdgpu's mem_info_vram_used
every 2 seconds during an FP8 load:


13:58:57  sys=  8.2 GiB  vram=  0.6 GiB    <- load begins

13:59:07 sys= 20.3 GiB vram= 0.6 GiB
13:59:19 sys= 26.0 GiB vram= 0.6 GiB <- pinned at container limit
... 45 seconds pinned, GPU idle at 0.6 GiB ...
14:00:06 sys= 25.6 GiB vram= 1.2 GiB <- transfer to VRAM begins
14:00:12 sys= 6.2 GiB vram= 20.6 GiB
14:00:16 sys= 2.5 GiB vram= 26.0 GiB <- steady state

The weights are fully materialised in system RAM, then copied to VRAM in ~6 seconds.
Steady-state host RAM afterwards is 2.5 GiB. So host RAM must be at least
model-sized during loading even when VRAM is abundant, and the expansion above
doubles what that costs.


What you expected to happen


That an FP8 checkpoint would occupy roughly its on-disk size in memory, and that
selecting a smaller quantization would reduce the memory footprint.


How to reproduce the problem



  1. ROCm container image on a system with 32 GB VRAM.

  2. Install two quantizations of the same model — e.g. a 6 GiB and a 12 GiB Z Image
    checkpoint, or an FP8 and a Q8_0 GGUF Krea-2.

  3. Load each and compare the Total model size value in the
    [MODEL CACHE] Loaded model log line against the file size on disk.

  4. The quantized non-GGUF checkpoints report a resident size unrelated to their file
    size.


Possible cause


The ROCm image does not contain rocminfo, so bitsandbytes cannot detect the GPU
architecture and falls back. Logged at every startup:


Could not detect ROCm GPU architecture: [Errno 2] No such file or directory: 'rocminfo'

ROCm GPU architecture detection failed despite ROCm being available.
Could not detect ROCm warp size: [Errno 2] No such file or directory: 'rocminfo'.
Defaulting to 64. (some 4-bit functions may not work!)

If quantized kernels are unavailable because the architecture could not be
identified, dequantizing the weights at load would be a plausible fallback — and
would explain why GGUF (which has its own dequantization path) behaves differently.


This is speculation on my part; I have not instrumented it. But rocminfo being
absent from the official ROCm image looks like a packaging defect regardless.




Environment


MALLOC_MMAP_THRESHOLD_=1048576

DISABLE_PINNED_MEMORY=1
PYTORCH_HIP_ALLOC_CONF=garbage_collection_threshold:0.8,max_split_size_mb:128
HSA_OVERRIDE_GFX_VERSION=12.0.1
HSA_ENABLE_SDMA=0
PYTORCH_TUNABLEOP_ENABLED=0
TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=0
MIOPEN_FIND_MODE=FAST
GPU_DRIVER=rocm

schema_version: 4.0.3

max_cache_ram_gb: 14.0
keep_ram_copy_of_weights: false
device_working_mem_gb: 3.0
model_cache_keep_alive_min: 0
enable_partial_loading: false
force_tiled_decode: true
device: auto
precision: bfloat16
attention_type: torch-sdp

Host: 32 GB system RAM (31 GiB usable), AMD Ryzen 7 3800X (8C/16T), container limited
to 26 GiB and 6 CPUs. Nothing else memory-heavy running. Components in the Krea-2
workflow: Qwen3-VL 4B encoder (8,464.46 MB), Qwen Image VAE (242.03 MB).


What this is not


Recorded so others don't repeat the dead ends I worked through:



  • Not a leak. Host RAM returns to 2.5–7.8 GiB after every load; nothing
    accumulates across generations.

  • Not a regression. 6.14.0 and 6.14.1-post1 are identical. Earlier versions
    could not be tested — Krea-2 support does not exist in 6.13.x.

  • Not fixable via max_cache_ram_gb. Tested at 8.0, 12.0 and 14.0.

  • Not glibc fragmentation. The FAQ's MALLOC_MMAP_THRESHOLD_=1048576 workaround
    was in place throughout, with DISABLE_PINNED_MEMORY=1.

  • Not a GPU fault. No amdgpu ring timeouts, resets or segfaults in dmesg.
    All OOM kills are CONSTRAINT_MEMCG, contained to the container cgroup, with no
    Python traceback.




Three smaller issues noticed along the way


Happy to split these out.


1. The RAM cache statistics block never updates. Byte-identical across
consecutive generations (hits: 2 / misses: 4 / cached: 1 / cleared: 1) despite each
run loading several models.


2. Model load time is attributed to the denoise node. krea2_denoise was billed
at 65.2 s while the 9 sampling steps took ~20 s (2.26 s/it) and the Loaded model
line reported 3.16 s. The ~45 s of staging is folded into denoise and invisible in
the graph stats — the timing table points at sampling when the cost is in loading.
This made the problem substantially harder to locate.


3. force_tiled_decode: true appeared to have no effect. VRAM still spiked to
31.6 GiB at VAE decode with it enabled. Possibly not wired to the Qwen Image VAE
path.

### What you expected to happen

I expected to be able to run the models.

### How to reproduce the problem

Start the app load any model which is not quantized.

### Additional context

_No response_

### Discord username

C_Dollerup

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reproducing the load comparison described for the FP8, quantized Z Image, and GGUF checkpoints, using the precision setting in invokeai.yaml and recording resident memory and cache behavior. Trace the model-loading path responsible for the FP8-to-bfloat16 expansion. Done means quantized checkpoints retain an appropriately reduced memory footprint, avoid repeated restaging, and GGUF behavior remains unaffected.

Written by the indexing model from the issue text.

Assessment

Tech stack
docker, python, pytorch
Domain
machine-learning, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.