Multiple instances of the same model (MiniMax H3) cause "Cannot set version_counter for inference tensor"
- Dominant language
- Python
- Stars
- 133k
- Forks
- 15.7k
- Avg merge
- 1d 7h
- Merged PRs (30d)
- 158
Description
### Custom Node Testing
- [x] I have tried disabling custom nodes and the issue persists (see [how to disable custom nodes](https://docs.comfy.org/troubleshooting/custom-node-issues#step-1%3A-test-with-all-custom-nodes-disabled) if you need help)
### Expected Behavior
I would expect a workflow to fully evaluate instead of failing with a runtime error.
### Actual Behavior
The workflow failed with `RuntimeError: Cannot set version_counter for inference tensor`
### Steps to Reproduce
I created a workflow with two MiniMax H3 sub-graphs, using the included "MiniMax H3: Text to Video" template. The first H3 sub-graph's video output was chained to Save Video -> Get Any Video Frame (frame_index = -1) -> first_frame of second H3 sub-graph -> Save Video.
Running the first individual sub-graph works, the second sometimes works and sometimes fails, and unloading the models consistently allows the second sub-graph to succeed.
### Debug Logs
```powershell
[INFO] setup plugin alembic.autogenerate.schemas
[INFO] setup plugin alembic.autogenerate.tables
[INFO] setup plugin alembic.autogenerate.types
[INFO] setup plugin alembic.autogenerate.constraints
[INFO] setup plugin alembic.autogenerate.defaults
[INFO] setup plugin alembic.autogenerate.comments
[INFO] setup plugin alembic.autogenerate.checkconstraint_byname
[INFO] Setting base directory to: /data
(null): No such file or directory
(null): No such file or directory
[INFO] Found comfy_kitchen backend cuda: {'available': True, 'disabled': True, 'unavailable_reason': None, 'capabilities': ['adaln', 'apply_rope', 'apply_rope1', 'apply_rope1_', 'apply_rope_', 'apply_rope_split_half', 'apply_rope_split_half1', 'apply_rope_split_half1_', 'apply_rope_split_half_', 'convrot_w4a4_linear', 'dequantize_convrot_w4a4_weight', 'dequantize_int8_convrot_weight', 'dequantize_int8_convrot_weight_dtype', 'dequantize_int8_simple', 'dequantize_int8_simple_dtype', 'dequantize_nvfp4', 'dequantize_per_tensor_fp8', 'dequantize_w4a8_int8_weight', 'gemv_awq_w4a16', 'na3d', 'prepare_int4_weight_for_int8_linear', 'quantize_and_rotate_rowwise', 'quantize_convrot_w4a4_weight', 'quantize_int8_convrot_weight', 'quantize_int8_rowwise', 'quantize_int8_tensorwise', 'quantize_mxfp8', 'quantize_nvfp4', 'quantize_per_tensor_fp8', 'quantize_svdquant_w4a4', 'quantize_w4a8_int8_weight', 'rms_adaln', 'rms_rope', 'rms_rope1', 'rms_rope1_', 'rms_rope_', 'rms_rope_split_half', 'rms_rope_split_half1', 'rms_rope_split_half1_', 'rms_rope_split_half_', 'rotate_int8_convrot_weight', 'scaled_mm_svdquant_w4a4', 'stochastic_rounding_fp8', 'w4a8_int8_linear']}
[INFO] Found comfy_kitchen backend triton: {'available': True, 'disabled': True, 'unavailable_reason': None, 'capabilities': ['adaln', 'apply_rope', 'apply_rope1', 'apply_rope1_', 'apply_rope_', 'apply_rope_split_half', 'apply_rope_split_half1', 'apply_rope_split_half1_', 'apply_rope_split_half_', 'dequantize_nvfp4', 'dequantize_per_tensor_fp8', 'int8_linear', 'na3d', 'quantize_and_rotate_rowwise', 'quantize_int8_rowwise', 'quantize_mxfp8', 'quantize_nvfp4', 'quantize_per_tensor_fp8', 'rms_adaln', 'rms_rope', 'rms_rope1', 'rms_rope1_', 'rms_rope_', 'rms_rope_split_half', 'rms_rope_split_half1', 'rms_rope_split_half1_', 'rms_rope_split_half_', 'w4a8_int8_linear']}
[INFO] Found comfy_kitchen backend eager: {'available': True, 'disabled': False, 'unavailable_reason': None, 'capabilities': ['adaln', 'apply_rope', 'apply_rope1', 'apply_rope1_', 'apply_rope_', 'apply_rope_split_half', 'apply_rope_split_half1', 'apply_rope_split_half1_', 'apply_rope_split_half_', 'convrot_w4a4_linear', 'dequantize_convrot_w4a4_weight', 'dequantize_int8_convrot_weight', 'dequantize_int8_convrot_weight_dtype', 'dequantize_int8_embedding', 'dequantize_int8_simple', 'dequantize_int8_simple_dtype', 'dequantize_mxfp8', 'dequantize_nvfp4', 'dequantize_per_tensor_fp8', 'dequantize_w4a8_int8_weight', 'gemv_awq_w4a16', 'int8_linear', 'na3d', 'prepare_int4_weight_for_int8_linear', 'quantize_and_rotate_rowwise', 'quantize_convrot_w4a4_weight', 'quantize_int8_convrot_weight', 'quantize_int8_rowwise', 'quantize_int8_tensorwise', 'quantize_mxfp8', 'quantize_nvfp4', 'quantize_per_tensor_fp8', 'quantize_svdquant_w4a4', 'quantize_w4a8_int8_weight', 'rms_adaln', 'rms_rope', 'rms_rope1', 'rms_rope1_', 'rms_rope_', 'rms_rope_split_half', 'rms_rope_split_half1', 'rms_rope_split_half1_', 'rms_rope_split_half_', 'rotate_int8_convrot_weight', 'scaled_mm_mxfp8', 'scaled_mm_nvfp4', 'scaled_mm_svdquant_w4a4', 'stochastic_rounding_fp8', 'w4a8_int8_linear']}
[INFO] Found comfy_kitchen backend hip: {'available': True, 'disabled': False, 'unavailable_reason': None, 'capabilities': ['adaln', 'apply_rope', 'apply_rope1', 'apply_rope1_', 'apply_rope_', 'apply_rope_split_half', 'apply_rope_split_half1', 'apply_rope_split_half1_', 'apply_rope_split_half_', 'convrot_w4a4_linear', 'dequantize_convrot_w4a4_weight', 'dequantize_int8_convrot_weight_dtype', 'dequantize_int8_simple_dtype', 'dequantize_per_tensor_fp8', 'dequantize_w4a8_int8_weight', 'gemv_awq_w4a16', 'int8_linear', 'na3d', 'quantize_and_rotate_rowwise', 'quantize_convrot_w4a4_weight', 'quantize_int8_convrot_weight', 'quantize_int8_rowwise', 'quantize_int8_tensorwise', 'quantize_per_tensor_fp8', 'quantize_svdquant_w4a4', 'quantize_w4a8_int8_weight', 'rms_adaln', 'rms_rope', 'rms_rope1', 'rms_rope1_', 'rms_rope_', 'rms_rope_split_half', 'rms_rope_split_half1', 'rms_rope_split_half1_', 'rms_rope_split_half_', 'scaled_mm_svdquant_w4a4', 'stochastic_rounding_fp8', 'w4a8_int8_linear']}
[INFO] Checkpoint files will always be loaded safely.
[INFO] Total VRAM 20464 MB, total RAM 128507 MB
[INFO] pytorch version: 2.13.0+rocm7.2
[INFO] Set: torch.backends.cudnn.enabled = False for better AMD performance.
[INFO] AMD arch: gfx1100
[INFO] ROCm version: (7, 2)
[INFO] Set vram state to: NORMAL_VRAM
[INFO] Device: cuda:0 AMD Radeon Graphics : native
[INFO] Using async weight offloading with 2 streams
[INFO] Enabled pinned memory 112123
[INFO] Using pytorch attention
[INFO] Python version: 3.13.15 (main, Aug 10 2026, 21:10:14) [GCC 14.2.0]
[INFO] ComfyUI version: 0.33.1
[INFO] comfy-aimdo version: 0.4.13
[INFO] comfy-kitchen version: 0.2.31
[INFO] comfyui-frontend-package version: 1.48.7
[INFO] comfyui-workflow-templates version: 0.11.41
[INFO] comfyui-embedded-docs version: 0.5.9
[INFO] comfy-kitchen version: 0.2.31
[INFO] comfy-aimdo version: 0.4.13
[INFO] [Prompt Server] web root: /opt/venv/lib/python3.13/site-packages/comfyui_frontend_package/static
[INFO] Asset seeder disabled
[INFO] No OpenGL_accelerate module loaded: No module named 'OpenGL_accelerate'
[INFO] Context impl SQLiteImpl.
[INFO] Will assume non-transactional DDL.
[INFO] Using RAM pressure cache.
[INFO] Starting server
[INFO] To see the GUI go to: http://127.0.0.1:8188
[...]
[INFO] Requested to load MiniMaxH3TEModel_
[ERROR] !!! Exception during processing !!! Cannot set version_counter for inference tensor
[ERROR] Traceback (most recent call last):
File "/opt/ComfyUI/execution.py", line 545, in execute
output_data, output_ui, has_subgraph, has_pending_tasks = await get_output_data(prompt_id, unique_id, obj, input_data_all, execution_block_cb=execution_block_cb, pre_execute_cb=pre_execute_cb, v3_data=v3_data)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/ComfyUI/execution.py", line 344, in get_output_data
return_values = await _async_map_node_over_list(prompt_id, unique_id, obj, input_data_all, obj.FUNCTION, allow_interrupt=True, execution_block_cb=execution_block_cb, pre_execute_cb=pre_execute_cb, v3_data=v3_data)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/ComfyUI/execution.py", line 318, in _async_map_node_over_list
await process_inputs(input_dict, i)
File "/opt/ComfyUI/execution.py", line 306, in process_inputs
result = f(**inputs)
File "/opt/ComfyUI/comfy_api/internal/__init__.py", line 149, in wrapped_func
return method(locked_class, **inputs)
File "/opt/ComfyUI/comfy_api/latest/_io.py", line 1990, in EXECUTE_NORMALIZED
to_return = cls.execute(*args, **kwargs)
File "/opt/ComfyUI/comfy_extras/nodes_minimax_h3.py", line 142, in execute
cond = clip.encode_from_tokens_scheduled(tokens)
File "/opt/ComfyUI/comfy/sd.py", line 340, in encode_from_tokens_scheduled
pooled_dict = self.encode_from_tokens(tokens, return_pooled=return_pooled, return_dict=True)
File "/opt/ComfyUI/comfy/sd.py", line 404, in encode_from_tokens
self.load_model(tokens)
~~~~~~~~~~~~~~~^^^^^^^^
File "/opt/ComfyUI/comfy/sd.py", line 462, in load_model
model_management.load_models_gpu([self.patcher], memory_required=memory_used)
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/ComfyUI/comfy/model_management.py", line 971, in load_models_gpu
free_memory(total_memory_required[device] * 1.1 + extra_mem,
~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
device,
^^^^^^^
for_dynamic=free_for_dynamic,
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
pins_required=total_pins_required.get(device, 0))
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/ComfyUI/comfy/model_management.py", line 889, in free_memory
if memory_to_free > 0 and current_loaded_models[i].model_unload(memory_to_free):
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^
File "/opt/ComfyUI/comfy/model_management.py", line 811, in model_unload
self.model.detach(unpatch_weights)
~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^
File "/opt/ComfyUI/comfy/model_patcher.py", line 1299, in detach
self.unpatch_model(self.offload_device, unpatch_weights=unpatch_all)
~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/ComfyUI/comfy/model_patcher.py", line 1156, in unpatch_model
self.model.to(device_to)
~~~~~~~~~~~~~^^^^^^^^^^^
File "/opt/venv/lib/python3.13/site-packages/torch/nn/modules/module.py", line 1383, in to
return self._apply(convert)
~~~~~~~~~~~^^^^^^^^^
File "/opt/venv/lib/python3.13/site-packages/torch/nn/modules/module.py", line 933, in _apply
module._apply(fn)
~~~~~~~~~~~~~^^^^
File "/opt/venv/lib/python3.13/site-packages/torch/nn/modules/module.py", line 933, in _apply
module._apply(fn)
~~~~~~~~~~~~~^^^^
File "/opt/venv/lib/python3.13/site-packages/torch/nn/modules/module.py", line 933, in _apply
module._apply(fn)
~~~~~~~~~~~~~^^^^
[Previous line repeated 2 more times]
File "/opt/ComfyUI/comfy/ops.py", line 1445, in _apply
return _quantized_apply(self, fn, recurse)
File "/opt/ComfyUI/comfy/ops.py", line 1104, in _quantized_apply
module.register_parameter(key, torch.nn.Parameter(p, requires_grad=False))
~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/venv/lib/python3.13/site-packages/torch/nn/parameter.py", line 60, in __new__
t = data.detach().requires_grad_(requires_grad)
~~~~~~~~~~~^^
RuntimeError: Cannot set version_counter for inference tensor
```
### Other
I pointed an LLM at the error message and the ComfyUI repository. It produced the following bug writeup and patch. I've applied the patch to my local version and verified it does fix the error for me, but I have no idea at all whether the code here is the level of quality that ComfyUI maintainers would want.
LLM-generated bug writeup
### Environment
- ComfyUI v0.33.1 (also applies to current `master`, the relevant code is unchanged)
- torch 2.13.0+rocm7.2, comfy-kitchen 0.2.31, Python 3.13
- AMD GPU, so **non-dynamic VRAM** path (`ModelPatcher`, not `ModelPatcherDynamic`), lowvram partial loading in use
- Workflow: MiniMax-H3 with fp8-quantized text encoder / diffusion model
### Root cause
It's an interaction between `torch.inference_mode()`, wrapper tensor subclasses (`comfy_kitchen.tensor.QuantizedTensor`) and the way `comfy/ops.py:_quantized_apply` re-wraps parameters.
1. `torch.nn.Parameter(x, requires_grad=False)` takes the "custom tensor" path for a `QuantizedTensor` and calls `x.detach()`.
2. `detach()` on a wrapper subclass goes through `__torch_dispatch__`, which constructs a **brand new wrapper** (`_make_wrapper_subclass`). Whether that wrapper is an *inference tensor* depends only on whether `inference_mode` is enabled **at that moment**, not on `x`.
3. `detach` is a view op, so the autograd layer then tries to make the result share `x`'s version counter. If `x` is a normal (non-inference) tensor and the freshly built wrapper is an inference tensor, torch rejects this with `Cannot set version_counter for inference tensor`.
So **`torch.nn.Parameter(qt)` fails whenever `qt` is a non-inference `QuantizedTensor` and inference mode is on.** Plain tensors are unaffected (`_make_subclass` path). Verified directly against the installed torch/comfy-kitchen:
| `QuantizedTensor` created | `Parameter()` called | result |
|---|---|---|
| outside inference mode | outside | OK |
| inside inference mode | inside | OK |
| inside inference mode | outside | OK (result becomes a normal tensor) |
| outside inference mode | **inside** | **FAIL** |
| `.clone()` of a normal one | inside | OK (clone rebuilds the wrapper under the current mode) |
| `.to(other_device)` of a normal one | inside | OK (same reason) |
| `.to(same_device)` of a normal one (returns `self`) | inside | **FAIL** |
How ComfyUI hits it:
- Node execution runs inside `with torch.inference_mode():` (`execution.py`), so weights are normally inference tensors and everything is fine.
- Any `.to()` of a quantized model that happens **outside** inference mode turns every `QuantizedTensor` weight into a normal tensor: `_quantized_apply` calls `torch.nn.Parameter(p)` (or `p.clone()` via its existing guard), and the new wrapper is built in normal mode. The obvious case is `unload_all_models()` being called from `main.py`'s prompt worker loop (Unload Models button / `free_memory` flag), which runs after the `inference_mode` block has exited; anything else that moves a model between prompts would do the same.
- The next time that model is touched *inside* inference mode, `.to()` onto a device the weight **already lives on** returns the very same tensor (`_handle_to` returns `qt` unchanged), and `_quantized_apply` then does `torch.nn.Parameter(p)` on a normal `QuantizedTensor` under inference mode → crash. With lowvram partial loading this is guaranteed to happen in `unpatch_model()`'s `self.model.to(offload_device)`: the offloaded weights are already on the CPU. That is exactly the traceback above.
- `comfy.utils.set_attr_param` (used when restoring weight backups in `unpatch_model`) has the same `torch.nn.Parameter(value)` call and the same problem.
I reproduced it outside the UI with a small fp8 `mixed_precision_ops` model and the real `ModelPatcher` / `load_models_gpu` / `free_memory` / `unload_all_models` calls: load → `unload_all_models()` outside inference mode → (inside inference mode) move half the layers to GPU → `model.to("cpu")` → `Cannot set version_counter for inference tensor`. With the patch below the same sequence succeeds and the model still runs forward.
### Fix
Two small changes:
**`comfy/utils.py`** – new helper `make_param(value)` that `set_attr_param` now uses:
- Inside inference mode, if `value` is a wrapper subclass (`type(value) is not torch.Tensor` and it has `__tensor_flatten__`) that is **not** an inference tensor, rebuild the wrapper around the same inner tensors with `__tensor_flatten__` / `__tensor_unflatten__` before calling `torch.nn.Parameter`. The rebuilt wrapper is created under the current mode, so it is an inference tensor and the `detach()` inside `Parameter()` succeeds. No data is copied (storage is shared).
- Outside inference mode the previous behaviour (clone inference tensors) is kept unchanged.
**`comfy/ops.py`** – `_quantized_apply`:
- If `fn(param) is param` (the move was a no-op, e.g. `.to()` onto the device the weight already lives on — the common case for offloaded lowvram weights), leave the existing `Parameter` in place instead of re-wrapping it. This is both cheaper and avoids the failing `Parameter()` call entirely. The one case where re-wrapping a no-op result still matters — outside inference mode with an inference tensor, which the old code cloned — is preserved.
- Otherwise wrap with `comfy.utils.make_param(p)`, so the inference/normal mismatch is handled in the one remaining place too.
### Notes for reviewers
- The patch makes the `Parameter()` call sites robust to *however* a quantized weight ended up as a non-inference tensor; it doesn't try to stop models from being moved outside inference mode (e.g. `unload_all_models()` in `main.py`'s worker loop after the `inference_mode` block has closed). If you'd rather also make that path run under `inference_mode`, that would be a reasonable additional change, but this fix is sufficient on its own.
- Training (`nodes_train.py`, `torch.inference_mode(False)`) is unaffected: outside inference mode the code path is the same as before (inference tensors are still cloned into normal ones).
- Verification was done against torch 2.13 / comfy-kitchen 0.2.31 with a standalone script; I couldn't run the repo's test suite on this machine (release container only, no pytest/ruff), so please run `ruff` + `tests-unit` on your side.
LLM-generated patch
```patch
diff --git a/comfy/ops.py b/comfy/ops.py
index 73ae4667..626eef81 100644
--- a/comfy/ops.py
+++ b/comfy/ops.py
@@ -1099,9 +1099,9 @@ def _quantized_apply(module, fn, recurse=True):
if param is None:
continue
p = fn(param)
- if (not torch.is_inference_mode_enabled()) and p.is_inference():
- p = p.clone()
- module.register_parameter(key, torch.nn.Parameter(p, requires_grad=False))
+ if p is param and (torch.is_inference_mode_enabled() or not p.is_inference()):
+ continue
+ module.register_parameter(key, comfy.utils.make_param(p))
for key, buf in module._buffers.items():
if buf is not None:
module._buffers[key] = fn(buf)
diff --git a/comfy/utils.py b/comfy/utils.py
index 61c2a22d..552edcc8 100644
--- a/comfy/utils.py
+++ b/comfy/utils.py
@@ -935,11 +935,19 @@ def set_attr(obj, attr, value):
return prev
def set_attr_param(obj, attr, value):
+ return set_attr(obj, attr, make_param(value))
+
+def make_param(value):
# Clone inference tensors (created under torch.inference_mode) since
# their version counter is frozen and nn.Parameter() cannot wrap them.
- if (not torch.is_inference_mode_enabled()) and value.is_inference():
+ if torch.is_inference_mode_enabled():
+ if (not value.is_inference()) and type(value) is not torch.Tensor and hasattr(value, "__tensor_flatten__"):
+ attrs, ctx = value.__tensor_flatten__()
+ inner = {a: getattr(value, a) for a in attrs}
+ value = type(value).__tensor_unflatten__(inner, ctx, value.shape, value.stride())
+ elif value.is_inference():
value = value.clone()
- return set_attr(obj, attr, torch.nn.Parameter(value, requires_grad=False))
+ return torch.nn.Parameter(value, requires_grad=False)
def set_attr_buffer(obj, attr, value):
obj, name = resolve_attr(obj, attr)
```
Contributor guide
Research direction
Reproduce the two MiniMax H3 sub-graph workflow and start at comfy/ops.py:_quantized_apply, following the unload path through comfy/model_patcher.py and comfy/model_management.py. Compare the failing model transition with the traceback; done means both sub-graphs complete without the inference-tensor RuntimeError, including when the models are unloaded.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- ai, backend
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 52/100