Comfy-Org / Comfy-Org/ComfyUI

Generate LTX2 Prompt w/ gemma4 generates nonsense on 7900XTX

Open
#15,680 2 comments 0 reactions 0 assignees View on GitHub
Potential Bug
Dominant language
Python
Stars
133k
Forks
15.7k
Avg merge
1d 7h
Merged PRs (30d)
158

Description

### Custom Node Testing

- [x] I have tried disabling custom nodes and the issue persists (see [how to disable custom nodes](https://docs.comfy.org/troubleshooting/custom-node-issues#step-1%3A-test-with-all-custom-nodes-disabled) if you need help)

### Expected Behavior

“Generate LTX2 Prompt” node generates a coherent prompt given the gemma4-e2b model. Attached screenshot loads the gemma4-e2b model onto the CPU, which works (albeit slowly).

Image

### Actual Behavior

Node generates incoherent nonsense, seemingly a mix of random Markdown and LaTeX. Screenshot is same workflow with the model loaded on GPU (a 7900XTX).

Image

### Steps to Reproduce

Workflow attached.

[generate-ltx25-prompt.json](https://github.com/user-attachments/files/31123835/generate-ltx25-prompt.json)

### Debug Logs

```powershell
32m[INFO]0m setup plugin alembic.autogenerate.schemas
32m[INFO]0m setup plugin alembic.autogenerate.tables
32m[INFO]0m setup plugin alembic.autogenerate.types
32m[INFO]0m setup plugin alembic.autogenerate.constraints
32m[INFO]0m setup plugin alembic.autogenerate.defaults
32m[INFO]0m setup plugin alembic.autogenerate.comments
32m[INFO]0m setup plugin alembic.autogenerate.checkconstraint_byname
32m[INFO]0m Found comfy_kitchen backend cuda: {'available': True, 'disabled': True, 'unavailable_reason': None, 'capabilities': ['adaln', 'apply_rope', 'apply_rope1', 'apply_rope1_', 'apply_rope_', 'apply_rope_split_half', 'apply_rope_split_half1', 'apply_rope_split_half1_', 'apply_rope_split_half_', 'convrot_w4a4_linear', 'dequantize_convrot_w4a4_weight', 'dequantize_int8_convrot_weight', 'dequantize_int8_convrot_weight_dtype', 'dequantize_int8_simple', 'dequantize_int8_simple_dtype', 'dequantize_nvfp4', 'dequantize_per_tensor_fp8', 'dequantize_w4a8_int8_weight', 'gemv_awq_w4a16', 'na3d', 'prepare_int4_weight_for_int8_linear', 'quantize_and_rotate_rowwise', 'quantize_convrot_w4a4_weight', 'quantize_int8_convrot_weight', 'quantize_int8_rowwise', 'quantize_int8_tensorwise', 'quantize_mxfp8', 'quantize_nvfp4', 'quantize_per_tensor_fp8', 'quantize_svdquant_w4a4', 'quantize_w4a8_int8_weight', 'rms_adaln', 'rms_rope', 'rms_rope1', 'rms_rope1_', 'rms_rope_', 'rms_rope_split_half', 'rms_rope_split_half1', 'rms_rope_split_half1_', 'rms_rope_split_half_', 'rotate_int8_convrot_weight', 'scaled_mm_svdquant_w4a4', 'stochastic_rounding_fp8', 'w4a8_int8_linear']}
32m[INFO]0m Found comfy_kitchen backend triton: {'available': True, 'disabled': True, 'unavailable_reason': None, 'capabilities': ['adaln', 'apply_rope', 'apply_rope1', 'apply_rope1_', 'apply_rope_', 'apply_rope_split_half', 'apply_rope_split_half1', 'apply_rope_split_half1_', 'apply_rope_split_half_', 'dequantize_nvfp4', 'dequantize_per_tensor_fp8', 'int8_linear', 'na3d', 'quantize_and_rotate_rowwise', 'quantize_int8_rowwise', 'quantize_mxfp8', 'quantize_nvfp4', 'quantize_per_tensor_fp8', 'rms_adaln', 'rms_rope', 'rms_rope1', 'rms_rope1_', 'rms_rope_', 'rms_rope_split_half', 'rms_rope_split_half1', 'rms_rope_split_half1_', 'rms_rope_split_half_', 'w4a8_int8_linear']}
32m[INFO]0m Found comfy_kitchen backend eager: {'available': True, 'disabled': False, 'unavailable_reason': None, 'capabilities': ['adaln', 'apply_rope', 'apply_rope1', 'apply_rope1_', 'apply_rope_', 'apply_rope_split_half', 'apply_rope_split_half1', 'apply_rope_split_half1_', 'apply_rope_split_half_', 'convrot_w4a4_linear', 'dequantize_convrot_w4a4_weight', 'dequantize_int8_convrot_weight', 'dequantize_int8_convrot_weight_dtype', 'dequantize_int8_embedding', 'dequantize_int8_simple', 'dequantize_int8_simple_dtype', 'dequantize_mxfp8', 'dequantize_nvfp4', 'dequantize_per_tensor_fp8', 'dequantize_w4a8_int8_weight', 'gemv_awq_w4a16', 'int8_linear', 'na3d', 'prepare_int4_weight_for_int8_linear', 'quantize_and_rotate_rowwise', 'quantize_convrot_w4a4_weight', 'quantize_int8_convrot_weight', 'quantize_int8_rowwise', 'quantize_int8_tensorwise', 'quantize_mxfp8', 'quantize_nvfp4', 'quantize_per_tensor_fp8', 'quantize_svdquant_w4a4', 'quantize_w4a8_int8_weight', 'rms_adaln', 'rms_rope', 'rms_rope1', 'rms_rope1_', 'rms_rope_', 'rms_rope_split_half', 'rms_rope_split_half1', 'rms_rope_split_half1_', 'rms_rope_split_half_', 'rotate_int8_convrot_weight', 'scaled_mm_mxfp8', 'scaled_mm_nvfp4', 'scaled_mm_svdquant_w4a4', 'stochastic_rounding_fp8', 'w4a8_int8_linear']}
32m[INFO]0m Found comfy_kitchen backend hip: {'available': True, 'disabled': False, 'unavailable_reason': None, 'capabilities': ['adaln', 'apply_rope', 'apply_rope1', 'apply_rope1_', 'apply_rope_', 'apply_rope_split_half', 'apply_rope_split_half1', 'apply_rope_split_half1_', 'apply_rope_split_half_', 'convrot_w4a4_linear', 'dequantize_convrot_w4a4_weight', 'dequantize_int8_convrot_weight_dtype', 'dequantize_int8_simple_dtype', 'dequantize_per_tensor_fp8', 'dequantize_w4a8_int8_weight', 'gemv_awq_w4a16', 'int8_linear', 'na3d', 'quantize_and_rotate_rowwise', 'quantize_convrot_w4a4_weight', 'quantize_int8_convrot_weight', 'quantize_int8_rowwise', 'quantize_int8_tensorwise', 'quantize_per_tensor_fp8', 'quantize_svdquant_w4a4', 'quantize_w4a8_int8_weight', 'rms_adaln', 'rms_rope', 'rms_rope1', 'rms_rope1_', 'rms_rope_', 'rms_rope_split_half', 'rms_rope_split_half1', 'rms_rope_split_half1_', 'rms_rope_split_half_', 'scaled_mm_svdquant_w4a4', 'stochastic_rounding_fp8', 'w4a8_int8_linear']}
32m[INFO]0m Checkpoint files will always be loaded safely.
32m[INFO]0m Total VRAM 24560 MB, total RAM 64015 MB
32m[INFO]0m pytorch version: 2.13.0+rocm7.2
32m[INFO]0m Set: torch.backends.cudnn.enabled = False for better AMD performance.
32m[INFO]0m AMD arch: gfx1100
32m[INFO]0m ROCm version: (7, 2)
32m[INFO]0m Set vram state to: NORMAL_VRAM
32m[INFO]0m Device: cuda:0 Radeon RX 7900 XTX : native
32m[INFO]0m Using async weight offloading with 2 streams
32m[INFO]0m Enabled pinned memory 57613.0
32m[INFO]0m Using pytorch attention
32m[INFO]0m Python version: 3.14.7 (main, Aug 16 2026, 17:31:03) [GCC 12.2.0]
32m[INFO]0m ComfyUI version: 0.33.0
32m[INFO]0m comfy-aimdo version: 0.4.13
32m[INFO]0m comfy-kitchen version: 0.2.31
32m[INFO]0m comfyui-frontend-package version: 1.49.6
32m[INFO]0m comfyui-workflow-templates version: 0.11.41
32m[INFO]0m comfyui-embedded-docs version: 0.5.10
32m[INFO]0m comfy-kitchen version: 0.2.31
32m[INFO]0m comfy-aimdo version: 0.4.13
32m[INFO]0m [Prompt Server] web root: /home/yahweasel/ComfyUI/venv/lib/python3.14/site-packages/comfyui_frontend_package/static
32m[INFO]0m Asset seeder disabled
32m[INFO]0m No OpenGL_accelerate module loaded: No module named 'OpenGL_accelerate'
32m[INFO]0m Skipping loading of custom nodes
32m[INFO]0m Context impl SQLiteImpl.
32m[INFO]0m Will assume non-transactional DDL.
32m[INFO]0m Using RAM pressure cache.
32m[INFO]0m Starting server

32m[INFO]0m To see the GUI go to: http://127.0.0.1:7821
32m[INFO]0m got prompt
32m[INFO]0m Requested to load Gemma4TEModel_
32m[INFO]0m loaded completely; 9741.68 MB loaded, full load: True
32m[INFO]0m CLIP/text encoder model load device: cpu, offload device: cpu, current: cpu, dtype: torch.float16
Generating tokens: 100%|██████████| 60/60[[00:38<00:00,s11.55it/s]
32m[INFO]0m 32mPrompt executed in 40.68 seconds0m
32m[INFO]0m got prompt
32m[INFO]0m 32mPrompt executed in 0.00 seconds0m
32m[INFO]0m got prompt
32m[INFO]0m 32mPrompt executed in 0.00 seconds0m
32m[INFO]0m got prompt
32m[INFO]0m Requested to load Gemma4TEModel_
32m[INFO]0m loaded completely; 9741.68 MB loaded, full load: True
32m[INFO]0m CLIP/text encoder model load device: cuda:0, offload device: cpu, current: cuda:0, dtype: torch.float16
Generating tokens: 100%|██████████| 60/60[[00:02<00:00,s25.13it/s]
32m[INFO]0m 32mPrompt executed in 12.13 seconds0m
```

### Other

Here's my packages (I just updated everything before reporting to make sure nothing was weird on my end):

```
$ pip3 list
Package Version
------------------------------------------ --------------
aiohappyeyeballs 2.7.1
aiohttp 3.14.3
aiosignal 1.4.0
alembic 1.19.1
annotated-doc 0.0.5
annotated-types 0.8.0
anyio 4.14.2
attrs 26.1.0
av 18.1.0
blake3 1.0.9
certifi 2026.7.22
charset-normalizer 3.5.1
click 8.4.2
comfy-aimdo 0.4.13
comfy-angle 0.1.0
comfy-kitchen 0.2.31
comfyui-embedded-docs 0.5.10
comfyui_frontend_package 1.49.6
comfyui_workflow_templates 0.11.41
comfyui-workflow-templates-core 0.3.312
comfyui-workflow-templates-json 0.1.47
comfyui-workflow-templates-media-api 0.3.84
comfyui-workflow-templates-media-assets-01 0.1.28
comfyui-workflow-templates-media-image 0.3.160
comfyui-workflow-templates-media-other 0.3.229
comfyui-workflow-templates-media-video 0.3.101
distlib 0.4.3
einops 0.8.2
filelock 3.32.3
frozenlist 1.8.0
fsspec 2026.7.0
greenlet 3.5.5
h11 0.16.0
hf-xet 1.6.0
httpcore 1.0.9
httpx 0.28.1
huggingface_hub 1.27.0
idna 3.18
Jinja2 3.1.6
kornia 0.8.3
kornia_rs 0.1.14
Mako 1.4.1
markdown-it-py 4.2.0
MarkupSafe 3.0.3
mdurl 0.1.2
mpmath 1.3.0
multidict 6.7.1
networkx 3.6.1
numpy 2.5.2
packaging 26.3
pillow 12.3.0
pip 26.2.1
platformdirs 4.11.3
propcache 0.5.2
psutil 7.2.2
pydantic 2.13.4
pydantic_core 2.46.4
pydantic-settings 2.15.0
Pygments 2.20.0
PyOpenGL 3.1.10
python-discovery 1.5.2
python-dotenv 1.2.3
PyYAML 6.0.3
regex 2026.7.19
requests 2.34.2
rich 15.0.0
safetensors 0.8.0
scipy 1.18.0
sentencepiece 0.2.2
setuptools 78.1.0
shellingham 1.5.4
simpleeval 1.0.7
spandrel 0.4.2
SQLAlchemy 2.0.52
sympy 1.14.0
tokenizers 0.22.2
torch 2.13.0+rocm7.2
torchaudio 2.11.0+rocm7.2
torchcodec 0.16.0
torchsde 0.2.6
torchvision 0.28.0+rocm7.2
tqdm 4.70.0
trampoline 0.1.2
transformers 5.15.0
triton 3.7.1
triton-rocm 3.7.1
typer 0.27.1
typing_extensions 4.16.0
typing-inspection 0.4.4
urllib3 2.7.0
virtualenv 21.7.4
yarl 1.24.5
```

The only command-line arguments I'm running with are `--port 7821 --disable-all-custom-nodes`. I've tried a lot of the --disable/--enable arguments, but none of them have helped.

I don't know if this is specific to gemma4. I've never used this node before, and don't know what other models it's designed to work with, if any (this is a part from the official LTX2.5 workflow, of course).

Contributor guide

Open the contributing guide

Research direction

Start with the attached generate-ltx25-prompt.json workflow and reproduce the “Generate LTX2 Prompt” result with gemma4-e2b on CPU and the Radeon RX 7900 XTX. Compare the CPU and GPU execution paths shown in the logs, including Gemma4TEModel_ device placement. Done means the GPU workflow produces coherent output comparable to the CPU result.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
ai, backend, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
52/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.