HIP error: invalid device function
- Dominant language
- Python
- Stars
- 133k
- Forks
- 15.7k
- Avg merge
- 1d 7h
- Merged PRs (30d)
- 158
Description
### Expected Behavior
Render something
### Actual Behavior
Fails to queue anything including default workflow.
### Steps to Reproduce
Step 1. Load default workflow
Step 2. Queue
Step 3. Failure
### Debug Logs
7800 XT, tried using both stable and unstable versions of ROCm (6.2 & 6.3)
Very confusing because this error is related to unsupported GPU features, but my GPU is RDNA3 which as far as I know is fully supported by current ROCm.
Here is the output of `rocminfo`:
```ROCk module is loaded
=====================
HSA System Attributes
=====================
Runtime Version: 1.1
Runtime Ext Version: 1.6
System Timestamp Freq.: 1000.000000MHz
Sig. Max Wait Duration: 18446744073709551615 (0xFFFFFFFFFFFFFFFF) (timestamp count)
Machine Model: LARGE
System Endianness: LITTLE
Mwaitx: DISABLED
DMAbuf Support: YES
==========
HSA Agents
==========
*******
Agent 1
*******
Name: AMD Ryzen 7 7800X3D 8-Core Processor
Uuid: CPU-XX
Marketing Name: AMD Ryzen 7 7800X3D 8-Core Processor
Vendor Name: CPU
Feature: None specified
Profile: FULL_PROFILE
Float Round Mode: NEAR
Max Queue Number: 0(0x0)
Queue Min Size: 0(0x0)
Queue Max Size: 0(0x0)
Queue Type: MULTI
Node: 0
Device Type: CPU
Cache Info:
L1: 32768(0x8000) KB
Chip ID: 0(0x0)
ASIC Revision: 0(0x0)
Cacheline Size: 64(0x40)
Max Clock Freq. (MHz): 5050
BDFID: 0
Internal Node ID: 0
Compute Unit: 16
SIMDs per CU: 0
Shader Engines: 0
Shader Arrs. per Eng.: 0
WatchPts on Addr. Ranges:1
Memory Properties:
Features: None
Pool Info:
Pool 1
Segment: GLOBAL; FLAGS: FINE GRAINED
Size: 31948540(0x1e77efc) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:4KB
Alloc Alignment: 4KB
Accessible by all: TRUE
Pool 2
Segment: GLOBAL; FLAGS: KERNARG, FINE GRAINED
Size: 31948540(0x1e77efc) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:4KB
Alloc Alignment: 4KB
Accessible by all: TRUE
Pool 3
Segment: GLOBAL; FLAGS: COARSE GRAINED
Size: 31948540(0x1e77efc) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:4KB
Alloc Alignment: 4KB
Accessible by all: TRUE
ISA Info:
*******
Agent 2
*******
Name: gfx1101
Uuid: GPU-aacb34feaea8812a
Marketing Name: AMD Radeon RX 7800 XT
Vendor Name: AMD
Feature: KERNEL_DISPATCH
Profile: BASE_PROFILE
Float Round Mode: NEAR
Max Queue Number: 128(0x80)
Queue Min Size: 64(0x40)
Queue Max Size: 131072(0x20000)
Queue Type: MULTI
Node: 1
Device Type: GPU
Cache Info:
L1: 32(0x20) KB
L2: 4096(0x1000) KB
L3: 65536(0x10000) KB
Chip ID: 29822(0x747e)
ASIC Revision: 0(0x0)
Cacheline Size: 128(0x80)
Max Clock Freq. (MHz): 2124
BDFID: 768
Internal Node ID: 1
Compute Unit: 60
SIMDs per CU: 2
Shader Engines: 3
Shader Arrs. per Eng.: 2
WatchPts on Addr. Ranges:4
Coherent Host Access: FALSE
Memory Properties:
Features: KERNEL_DISPATCH
Fast F16 Operation: TRUE
Wavefront Size: 32(0x20)
Workgroup Max Size: 1024(0x400)
Workgroup Max Size per Dimension:
x 1024(0x400)
y 1024(0x400)
z 1024(0x400)
Max Waves Per CU: 32(0x20)
Max Work-item Per CU: 1024(0x400)
Grid Max Size: 4294967295(0xffffffff)
Grid Max Size per Dimension:
x 4294967295(0xffffffff)
y 4294967295(0xffffffff)
z 4294967295(0xffffffff)
Max fbarriers/Workgrp: 32
Packet Processor uCode:: 462
SDMA engine uCode:: 27
IOMMU Support:: None
Pool Info:
Pool 1
Segment: GLOBAL; FLAGS: COARSE GRAINED
Size: 16760832(0xffc000) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:2048KB
Alloc Alignment: 4KB
Accessible by all: FALSE
Pool 2
Segment: GLOBAL; FLAGS: EXTENDED FINE GRAINED
Size: 16760832(0xffc000) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:2048KB
Alloc Alignment: 4KB
Accessible by all: FALSE
Pool 3
Segment: GROUP
Size: 64(0x40) KB
Allocatable: FALSE
Alloc Granule: 0KB
Alloc Recommended Granule:0KB
Alloc Alignment: 0KB
Accessible by all: FALSE
ISA Info:
ISA 1
Name: amdgcn-amd-amdhsa--gfx1101
Machine Models: HSA_MACHINE_MODEL_LARGE
Profiles: HSA_PROFILE_BASE
Default Rounding Mode: NEAR
Default Rounding Mode: NEAR
Fast f16: TRUE
Workgroup Max Size: 1024(0x400)
Workgroup Max Size per Dimension:
x 1024(0x400)
y 1024(0x400)
z 1024(0x400)
Grid Max Size: 4294967295(0xffffffff)
Grid Max Size per Dimension:
x 4294967295(0xffffffff)
y 4294967295(0xffffffff)
z 4294967295(0xffffffff)
FBarrier Max Size: 32
*******
Agent 3
*******
Name: gfx1036
Uuid: GPU-XX
Marketing Name: AMD Radeon Graphics
Vendor Name: AMD
Feature: KERNEL_DISPATCH
Profile: BASE_PROFILE
Float Round Mode: NEAR
Max Queue Number: 128(0x80)
Queue Min Size: 64(0x40)
Queue Max Size: 131072(0x20000)
Queue Type: MULTI
Node: 2
Device Type: GPU
Cache Info:
L1: 16(0x10) KB
L2: 256(0x100) KB
Chip ID: 5710(0x164e)
ASIC Revision: 1(0x1)
Cacheline Size: 128(0x80)
Max Clock Freq. (MHz): 2200
BDFID: 4608
Internal Node ID: 2
Compute Unit: 2
SIMDs per CU: 2
Shader Engines: 1
Shader Arrs. per Eng.: 1
WatchPts on Addr. Ranges:4
Coherent Host Access: FALSE
Memory Properties: APU
Features: KERNEL_DISPATCH
Fast F16 Operation: TRUE
Wavefront Size: 32(0x20)
Workgroup Max Size: 1024(0x400)
Workgroup Max Size per Dimension:
x 1024(0x400)
y 1024(0x400)
z 1024(0x400)
Max Waves Per CU: 32(0x20)
Max Work-item Per CU: 1024(0x400)
Grid Max Size: 4294967295(0xffffffff)
Grid Max Size per Dimension:
x 4294967295(0xffffffff)
y 4294967295(0xffffffff)
z 4294967295(0xffffffff)
Max fbarriers/Workgrp: 32
Packet Processor uCode:: 21
SDMA engine uCode:: 9
IOMMU Support:: None
Pool Info:
Pool 1
Segment: GLOBAL; FLAGS: COARSE GRAINED
Size: 15974268(0xf3bf7c) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:2048KB
Alloc Alignment: 4KB
Accessible by all: FALSE
Pool 2
Segment: GLOBAL; FLAGS: EXTENDED FINE GRAINED
Size: 15974268(0xf3bf7c) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:2048KB
Alloc Alignment: 4KB
Accessible by all: FALSE
Pool 3
Segment: GROUP
Size: 64(0x40) KB
Allocatable: FALSE
Alloc Granule: 0KB
Alloc Recommended Granule:0KB
Alloc Alignment: 0KB
Accessible by all: FALSE
ISA Info:
ISA 1
Name: amdgcn-amd-amdhsa--gfx1036
Machine Models: HSA_MACHINE_MODEL_LARGE
Profiles: HSA_PROFILE_BASE
Default Rounding Mode: NEAR
Default Rounding Mode: NEAR
Fast f16: TRUE
Workgroup Max Size: 1024(0x400)
Workgroup Max Size per Dimension:
x 1024(0x400)
y 1024(0x400)
z 1024(0x400)
Grid Max Size: 4294967295(0xffffffff)
Grid Max Size per Dimension:
x 4294967295(0xffffffff)
y 4294967295(0xffffffff)
z 4294967295(0xffffffff)
FBarrier Max Size: 32
*** Done ***
```
```powershell
This error with all models:
{
"prompt_id": "a902c515-6f4b-4f69-bd65-01cb4ec73f04",
"node_id": "NegativeCLIP_Base",
"node_type": "CLIPTextEncode",
"executed": [],
"exception_message": "HIP error: invalid device function\nHIP kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.\nFor debugging consider passing AMD_SERIALIZE_KERNEL=3\nCompile with \u0060TORCH_USE_HIP_DSA\u0060 to enable device-side assertions.\n",
"exception_type": "RuntimeError",
"traceback": [
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/execution.py\u0022, line 327, in execute\n output_data, output_ui, has_subgraph = get_output_data(obj, input_data_all, execution_block_cb=execution_block_cb, pre_execute_cb=pre_execute_cb)\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/execution.py\u0022, line 202, in get_output_data\n return_values = _map_node_over_list(obj, input_data_all, obj.FUNCTION, allow_interrupt=True, execution_block_cb=execution_block_cb, pre_execute_cb=pre_execute_cb)\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/execution.py\u0022, line 174, in _map_node_over_list\n process_inputs(input_dict, i)\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/execution.py\u0022, line 163, in process_inputs\n results.append(getattr(obj, func)(**inputs))\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/nodes.py\u0022, line 69, in encode\n return (clip.encode_from_tokens_scheduled(tokens), )\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/comfy/sd.py\u0022, line 148, in encode_from_tokens_scheduled\n pooled_dict = self.encode_from_tokens(tokens, return_pooled=return_pooled, return_dict=True)\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/comfy/sd.py\u0022, line 210, in encode_from_tokens\n o = self.cond_stage_model.encode_token_weights(tokens)\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/comfy/sdxl_clip.py\u0022, line 60, in encode_token_weights\n g_out, g_pooled = self.clip_g.encode_token_weights(token_weight_pairs_g)\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/comfy/sd1_clip.py\u0022, line 45, in encode_token_weights\n o = self.encode(to_encode)\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/comfy/sd1_clip.py\u0022, line 252, in encode\n return self(tokens)\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/venv/lib/python3.10/site-packages/torch/nn/modules/module.py\u0022, line 1736, in _wrapped_call_impl\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/venv/lib/python3.10/site-packages/torch/nn/modules/module.py\u0022, line 1747, in _call_impl\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/comfy/sd1_clip.py\u0022, line 224, in forward\n outputs = self.transformer(tokens, attention_mask_model, intermediate_output=self.layer_idx, final_layer_norm_intermediate=self.layer_norm_hidden_state, dtype=torch.float32)\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/venv/lib/python3.10/site-packages/torch/nn/modules/module.py\u0022, line 1736, in _wrapped_call_impl\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/venv/lib/python3.10/site-packages/torch/nn/modules/module.py\u0022, line 1747, in _call_impl\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/comfy/clip_model.py\u0022, line 137, in forward\n x = self.text_model(*args, **kwargs)\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/venv/lib/python3.10/site-packages/torch/nn/modules/module.py\u0022, line 1736, in _wrapped_call_impl\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/venv/lib/python3.10/site-packages/torch/nn/modules/module.py\u0022, line 1747, in _call_impl\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/comfy/clip_model.py\u0022, line 101, in forward\n x = self.embeddings(input_tokens, dtype=dtype)\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/venv/lib/python3.10/site-packages/torch/nn/modules/module.py\u0022, line 1736, in _wrapped_call_impl\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/venv/lib/python3.10/site-packages/torch/nn/modules/module.py\u0022, line 1747, in _call_impl\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/comfy/clip_model.py\u0022, line 82, in forward\n return self.token_embedding(input_tokens, out_dtype=dtype) \u002B comfy.ops.cast_to(self.position_embedding.weight, dtype=dtype, device=input_tokens.device)\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/venv/lib/python3.10/site-packages/torch/nn/modules/module.py\u0022, line 1736, in _wrapped_call_impl\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/venv/lib/python3.10/site-packages/torch/nn/modules/module.py\u0022, line 1747, in _call_impl\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/comfy/ops.py\u0022, line 203, in forward\n return self.forward_comfy_cast_weights(*args, **kwargs)\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/comfy/ops.py\u0022, line 199, in forward_comfy_cast_weights\n return torch.nn.functional.embedding(input, weight, self.padding_idx, self.max_norm, self.norm_type, self.scale_grad_by_freq, self.sparse).to(dtype=output_dtype)\n",
" File \u0022/opt/stabilitymatrix/Data/Packages/ComfyUI/venv/lib/python3.10/site-packages/torch/nn/functional.py\u0022, line 2551, in embedding\n"
],
"current_inputs": {
"clip": [
"\u003Ccomfy.sd.CLIP object at 0x766f88883e50\u003E"
],
"text": [
""
]
},
"current_outputs": [
"SaveImage",
"EmptyLatentImage",
"NegativeCLIP_Base",
"VAEDecode_1",
"Sampler",
"CheckpointLoader_Base",
"PositiveCLIP_Base"
],
"timestamp": 1738306666190
}
```
### Other
_No response_
Contributor guide
Assessment
This issue has not been assessed yet.