Comfy-Org / Comfy-Org/ComfyUI

MiniMax-Music-3: fp16 DiT silently produces all-NaN audio on gfx1151, --fp32-unet fixes it

Open
#16,249 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
133k
Forks
15.7k
Avg merge
1d 6h
Merged PRs (30d)
155

Description

### Custom Node Testing

- [x] I have tried disabling custom nodes and the issue persists

Every run below used `--disable-all-custom-nodes`, and all three model files are the official Comfy-Org ones.

### Expected Behavior

MiniMax-Music-3 should produce audio on `gfx1151` the same way it does elsewhere. The DiT runs in fp16 by default, and nothing warns that this is unsafe.

### Actual Behavior

The take comes out silently corrupt. Every sample of the 82 s render is `-32768`, one distinct value in the whole file, which is what NaN in the latent becomes on int16 conversion. There is no error, no warning, and no crash. The prompt completes normally in 512.93 s.

`--fp32-unet` fixes it completely. Same seed, same prompt, same server, stock nodes:

| Flags (on top of `--disable-all-custom-nodes`) | Noise floor | Verdict |
| --- | --- | --- |
| none (fp16 DiT, the default) | `0.000265 dB`, 1 distinct sample value | all-NaN |
| `--disable-cuda-graphs` | `0.000265 dB` | all-NaN |
| `--disable-comfy-compiler` | `0.000265 dB` | all-NaN |
| `--fp32-unet` | `-52.6 / -52.4 dB` | clean |

**This is not #16222.** That one is fixed by `--disable-cuda-graphs`; this one is not, and it is the DiT rather than the AR text encoder. #16222 notes "DiT-side graphs not yet isolated", so the two may still share a cause, but the flag that fixes theirs does not fix this.

It is also not a duration limit, which is how it first looked. At 8 steps, 2050 audio frames fails every time while the much longer 3745 passes:

| Audio frames | Latent frames | Steps | fp16 result |
| --- | --- | --- | --- |
| 2000 | 6890 | 8 | clean 4/4 |
| 2050 | 7062 | 8 | NaN 0/4 |
| 3745 | 12902 | 8 | clean 2/2 |
| 3750 | 12919 | 30 | corrupt, `-31.7 / -33.4 dB` |

Latent length, conditioning content and step count all change which runs break.

### Steps to Reproduce

1. Download the three official files from [Comfy-Org/MiniMax-Music-3](https://huggingface.co/Comfy-Org/MiniMax-Music-3): `diffusion_models/minimax_music3_dit_fp16.safetensors`, `text_encoders/minimax_music3_text_encoder_pruned_bf16.safetensors`, `vae/minimax_music3_dav.safetensors`.
2. Start ComfyUI with `--disable-all-custom-nodes`.
3. Queue the workflow below (stock nodes only, `UNETLoader` rather than any GGUF loader). The caption and lyrics come from [these two files](https://github.com/felladrin/text-to-music-lab/tree/main/prompts/vocal); the exact text matters, since other prompts of the same length do not always break.
4. Check the output: `ffmpeg -i out.flac -af astats -f null -`. A clean take reads about `-53 dB`; this one reads `0.000265 dB`.
5. Restart with `--fp32-unet` and queue the same workflow. It comes out clean.

Workflow JSON (API format, no custom nodes)

```json
{
"unet": {
"class_type": "UNETLoader",
"inputs": {
"unet_name": "minimax_music3_dit_fp16.safetensors",
"weight_dtype": "default"
}
},
"clip": {
"class_type": "CLIPLoader",
"inputs": {
"clip_name": "minimax_music3_text_encoder_pruned_bf16.safetensors",
"type": "minimax",
"device": "default"
}
},
"vae": {
"class_type": "VAELoader",
"inputs": {
"vae_name": "minimax_music3_dav.safetensors"
}
},
"cond": {
"class_type": "MiniMaxMusic3TextEncode",
"inputs": {
"clip": [
"clip",
0
],
"caption": "Global Metadata\nBasic Attributes: bpm is 104. key is B, and scale is minor. Indie Synth-Pop / Dream Pop with a live rhythm section. Song with lead vocals.\nGlobal Emotional Progression: Starts intimate and hesitant, almost confessional, over sparse electric piano. Warms and widens through the first verse as the band arrives, opens fully in the chorus with a bright hopeful lift, pulls back to near-silence for the bridge, then returns for a final chorus that is louder and more resolved than the first.\nApplication Scenarios & Imagery: Driving home along a harbour road at night, end-credits of a coming-of-age film, late-summer playlists, walking through a city that is falling asleep.\nSonics & Production Profile: Warm modern indie production. Analog tape saturation on the drum bus, wide plate reverb on the vocal, gentle bus compression that keeps the verses breathing and lets the choruses bloom. Bass is round and forward. Highs are soft and slightly rolled off, never brittle. Clear separation between the intimate verses and the full choruses.\nVocal Details\nVocal Gender & Timbre: Female lead, warm alto with a slight natural rasp on held notes and an airy breathy top. Close-miked and intimate in the verses, fuller and more open in the choruses.\nVocal Style: Half-spoken and conversational in the verses, sitting just behind the beat. Opens into sustained melodic phrasing in the choruses with clear diction and a small upward lift at the end of each line. Never belted or strained.\nHarmony/Backing Vocals: Double-tracked lead in the choruses with a third above, plus soft stacked wordless \"ooh\" pads under the final chorus. No harmonies at all in the first verse.\nVocal FX: Warm plate reverb with a medium tail, a short slapback delay on the ends of chorus lines, and light tape doubling. No autotune artefacts, no vocoder, no heavy pitch correction.\nArrangement\nInstrument Lifecycle Description (Primary/Secondary Layering):\nPrimary: A warm electric piano opens alone and carries the harmony throughout. A round fingered electric bass enters at the first verse and anchors every section afterward. A chiming reverb-soaked electric guitar states the main instrumental hook in the intro and answers the vocal between chorus lines.\nSecondary: Live drums with brushed snare in the verses, switching to full stick in the choruses. Wide analog synth pads swell underneath from the first chorus onward. A subtle string section joins only in the final chorus. Shaker and tambourine add motion through the second verse.\nGroove & Foundation Progression: Drumless intro on electric piano and guitar. Verse one is brushed drums, bass and piano only. The chorus opens up with a full backbeat and an eighth-note tambourine. Bridge strips everything to electric piano and one vocal line. The final chorus arrives with the whole band, strings and stacked harmonies at once. The outro drops the drums and lets piano and guitar ring out.\nEmbellishments, Textures & Spatial FX: Reverse guitar swells into each chorus. Vinyl-crackle texture under the intro. A tape-stop just before the bridge. Long guitar delay trails across the outro fading into room noise.",
"lyrics": "[intro]\n\n[verse]\nStreetlights counting down the empty road\nEvery window closing one by one\nI keep the radio low\nSo I can hear the harbour hum\n\n[chorus]\nAnd the harbour lights are burning\nSomewhere out there past the rain\nI don't know what I am learning\nBut I'd learn it all again\n\n[verse]\nSalt on the windshield, salt in my throat\nCoffee going cold inside my hands\nThere's a ferry pulling out\nFull of people making plans\n\n[chorus]\nAnd the harbour lights are burning\nSomewhere out there past the rain\nI don't know what I am learning\nBut I'd learn it all again\n\n[bridge]\nMaybe morning is a place\nAnd not a time\nMaybe I have been driving\nToward it all my life\n\n[chorus]\nAnd the harbour lights are burning\nBrighter than they were before\nI am done with all my turning\nI am pulling into shore\n\n[outro]",
"seed": 8899,
"max_duration": 82.0,
"cfg_scale": 1.7,
"top_k": 50
}
},
"neg": {
"class_type": "ConditioningZeroOut",
"inputs": {
"conditioning": [
"cond",
0
]
}
},
"latent": {
"class_type": "EmptyMiniMaxMusic3LatentAudio",
"inputs": {
"seconds": [
"cond",
1
],
"batch_size": 1
}
},
"sampler": {
"class_type": "KSampler",
"inputs": {
"model": [
"unet",
0
],
"positive": [
"cond",
0
],
"negative": [
"neg",
0
],
"latent_image": [
"latent",
0
],
"seed": 8899,
"steps": 8,
"cfg": 1.7,
"sampler_name": "euler",
"scheduler": "simple",
"denoise": 1.0
}
},
"decode": {
"class_type": "VAEDecodeAudio",
"inputs": {
"samples": [
"sampler",
0
],
"vae": [
"vae",
0
]
}
},
"save": {
"class_type": "SaveAudio",
"inputs": {
"audio": [
"decode",
0
],
"filename_prefix": "bugreport/minimax82"
}
}
}
```

### Debug Logs

Full log of the failing run, startup to finish. No errors anywhere.

```powershell
[INFO] setup plugin alembic.autogenerate.schemas
[INFO] setup plugin alembic.autogenerate.tables
[INFO] setup plugin alembic.autogenerate.types
[INFO] setup plugin alembic.autogenerate.constraints
[INFO] setup plugin alembic.autogenerate.defaults
[INFO] setup plugin alembic.autogenerate.comments
[INFO] setup plugin alembic.ext.checkconstraint_byname
[INFO] Found comfy_kitchen backend hip: {'available': True, 'disabled': False, 'unavailable_reason': None, 'capabilities': ['adaln', 'apply_rope', 'apply_rope1', 'apply_rope1_', 'apply_rope_', 'apply_rope_split_half', 'apply_rope_split_half1', 'apply_rope_split_half1_', 'apply_rope_split_half_', 'convrot_w4a4_linear', 'dequantize_convrot_w4a4_weight', 'dequantize_int8_convrot_weight_dtype', 'dequantize_int8_simple_dtype', 'dequantize_per_tensor_fp8', 'dequantize_w4a8_int8_weight', 'gemv_awq_w4a16', 'int8_linear', 'na3d', 'quantize_and_rotate_rowwise', 'quantize_convrot_w4a4_weight', 'quantize_int8_convrot_weight', 'quantize_int8_rowwise', 'quantize_int8_tensorwise', 'quantize_per_tensor_fp8', 'quantize_svdquant_w4a4', 'quantize_w4a8_int8_weight', 'rms_adaln', 'rms_rope', 'rms_rope1', 'rms_rope1_', 'rms_rope_', 'rms_rope_split_half', 'rms_rope_split_half1', 'rms_rope_split_half1_', 'rms_rope_split_half_', 'scaled_mm_svdquant_w4a4', 'sol_attn', 'stochastic_rounding_fp8', 'w4a8_int8_linear']}
[INFO] Found comfy_kitchen backend triton: {'available': True, 'disabled': True, 'unavailable_reason': None, 'capabilities': ['adaln', 'apply_rope', 'apply_rope1', 'apply_rope1_', 'apply_rope_', 'apply_rope_split_half', 'apply_rope_split_half1', 'apply_rope_split_half1_', 'apply_rope_split_half_', 'dequantize_nvfp4', 'dequantize_per_tensor_fp8', 'int8_linear', 'na3d', 'quantize_and_rotate_rowwise', 'quantize_int8_rowwise', 'quantize_mxfp8', 'quantize_nvfp4', 'quantize_per_tensor_fp8', 'rms_adaln', 'rms_rope', 'rms_rope1', 'rms_rope1_', 'rms_rope_', 'rms_rope_split_half', 'rms_rope_split_half1', 'rms_rope_split_half1_', 'rms_rope_split_half_', 'w4a8_int8_linear']}
[INFO] Found comfy_kitchen backend cuda: {'available': True, 'disabled': True, 'unavailable_reason': None, 'capabilities': ['adaln', 'apply_rope', 'apply_rope1', 'apply_rope1_', 'apply_rope_', 'apply_rope_split_half', 'apply_rope_split_half1', 'apply_rope_split_half1_', 'apply_rope_split_half_', 'convrot_w4a4_linear', 'dequantize_convrot_w4a4_weight', 'dequantize_int8_convrot_weight', 'dequantize_int8_convrot_weight_dtype', 'dequantize_int8_simple', 'dequantize_int8_simple_dtype', 'dequantize_nvfp4', 'dequantize_per_tensor_fp8', 'dequantize_w4a8_int8_weight', 'gemv_awq_w4a16', 'na3d', 'prepare_int4_weight_for_int8_linear', 'quantize_and_rotate_rowwise', 'quantize_convrot_w4a4_weight', 'quantize_int8_convrot_weight', 'quantize_int8_rowwise', 'quantize_int8_tensorwise', 'quantize_mxfp8', 'quantize_nvfp4', 'quantize_per_tensor_fp8', 'quantize_svdquant_w4a4', 'quantize_w4a8_int8_weight', 'rms_adaln', 'rms_rope', 'rms_rope1', 'rms_rope1_', 'rms_rope_', 'rms_rope_split_half', 'rms_rope_split_half1', 'rms_rope_split_half1_', 'rms_rope_split_half_', 'rotate_int8_convrot_weight', 'scaled_mm_svdquant_w4a4', 'sol_attn', 'stochastic_rounding_fp8', 'w4a8_int8_linear']}
[INFO] Found comfy_kitchen backend eager: {'available': True, 'disabled': False, 'unavailable_reason': None, 'capabilities': ['adaln', 'apply_rope', 'apply_rope1', 'apply_rope1_', 'apply_rope_', 'apply_rope_split_half', 'apply_rope_split_half1', 'apply_rope_split_half1_', 'apply_rope_split_half_', 'convrot_w4a4_linear', 'dequantize_convrot_w4a4_weight', 'dequantize_int8_convrot_weight', 'dequantize_int8_convrot_weight_dtype', 'dequantize_int8_embedding', 'dequantize_int8_simple', 'dequantize_int8_simple_dtype', 'dequantize_mxfp8', 'dequantize_nvfp4', 'dequantize_per_tensor_fp8', 'dequantize_w4a8_int8_weight', 'gemv_awq_w4a16', 'int8_linear', 'na3d', 'prepare_int4_weight_for_int8_linear', 'quantize_and_rotate_rowwise', 'quantize_convrot_w4a4_weight', 'quantize_int8_convrot_weight', 'quantize_int8_rowwise', 'quantize_int8_tensorwise', 'quantize_mxfp8', 'quantize_nvfp4', 'quantize_per_tensor_fp8', 'quantize_svdquant_w4a4', 'quantize_w4a8_int8_weight', 'rms_adaln', 'rms_rope', 'rms_rope1', 'rms_rope1_', 'rms_rope_', 'rms_rope_split_half', 'rms_rope_split_half1', 'rms_rope_split_half1_', 'rms_rope_split_half_', 'rotate_int8_convrot_weight', 'scaled_mm_mxfp8', 'scaled_mm_nvfp4', 'scaled_mm_svdquant_w4a4', 'sol_attn', 'stochastic_rounding_fp8', 'w4a8_int8_linear']}
[INFO] Checkpoint files will always be loaded safely.
[INFO] Total VRAM 131072 MB, total RAM 127437 MB
[INFO] pytorch version: 2.13.0+rocm7.1
[INFO] Set: torch.backends.cudnn.enabled = False for better AMD performance.
[INFO] AMD arch: gfx1151
[INFO] ROCm version: (7, 1)
[INFO] Set vram state to: NORMAL_VRAM
[INFO] Device: cuda:0 AMD Radeon 8060S Graphics : native
[INFO] Using async weight offloading with 2 streams
[INFO] Enabled pinned memory 114692.0
[INFO] Using pytorch attention
[INFO] Python version: 3.12.3 (main, Aug 31 2026, 10:18:26) [GCC 13.3.0]
[INFO] ComfyUI version: 0.35.0
[INFO] comfy-aimdo version: 0.5.3
[INFO] comfy-kitchen version: 0.2.33
[INFO] comfyui-frontend-package version: 1.51.10
[INFO] comfyui-workflow-templates version: 0.11.57
[INFO] comfyui-embedded-docs version: 0.5.11
[INFO] comfy-kitchen version: 0.2.33
[INFO] comfy-aimdo version: 0.5.3
[INFO] [Prompt Server] web root: /home/victor/Repositories/ComfyUI/.venv/lib/python3.12/site-packages/comfyui_frontend_package/static
[INFO] Asset seeder disabled
[INFO] No OpenGL_accelerate module loaded: No module named 'OpenGL_accelerate'
[INFO] Skipping loading of custom nodes
[INFO] Context impl SQLiteImpl.
[INFO] Will assume non-transactional DDL.
[INFO] Using RAM pressure cache.
[INFO] Starting server
[INFO] To see the GUI go to: http://127.0.0.1:8188
[INFO] got prompt
[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.float32
[INFO] Requested to load MiniMaxMusic3TEModel
[INFO] loaded completely; 15921.75 MB loaded, full load: True
[INFO] CLIP/text encoder model load device: cuda:0, offload device: cpu, current: cuda:0, dtype: torch.float16
[INFO] model weight dtype torch.float16, manual cast: None
[INFO] model_type FLOW
[INFO] Requested to load MiniMaxMusic3
[INFO] loaded completely; 106038.68 MB usable, 4686.50 MB loaded, full load: True
[INFO] Requested to load MiniMaxMusic3DAV
[INFO] loaded completely; 206.60 MB loaded, full load: True
[INFO] Prompt executed in 512.93 seconds
```

### Other

Hardware is a Ryzen AI MAX+ 395 (Radeon 8060S, `gfx1151`), Ubuntu 24.04, ROCm 7.1.1, torch 2.13.0+rocm7.1, ComfyUI 0.35.0.

What the fault looks like, from a few days of bisecting:

- The whole latent goes non-finite in a single sampler step, all 903,936 elements at once from a clean previous step. An fp16 overflow does not grow that way.
- `PYTORCH_NO_CUDA_MEMORY_CACHING=1` makes it clean 2/2, while `HIP_LAUNCH_BLOCKING=1` (0/2) and async offload off (`NUM_STREAMS = 0`, 0/3) change nothing. Fresh `hipMalloc` pages come back zeroed and recycled allocator blocks do not, so this looks like a read of memory that was never written rather than numerics or a race.
- The first run in a fresh process fails later (step 3) than the runs after it (step 1), which fits the same reading.
- It hides from almost any perturbation: sampling a different length first in the same process, wrapping each DiT window in an `isfinite` check, or filling every `torch.empty` with NaN all make it disappear. That is why I could not name the buffer.

Stages I ruled out, so nobody repeats them:

- The AR text encoder is clean: 2100 frames all finite, per-frame norm flat from first frame to last.
- The DAV VAE is clean: decode is bit-identical GPU versus CPU at 30, 80, 82, 90 and 150 s.
- Attention is not it: forcing `SDPBackend.MATH` still gives NaN 0/3.
- `comfy_kitchen.apply_rope_split_half` is not it: swapping in `dit.py`'s own `_apply_rope` torch fallback gives the identical failure.
- Quantisation is not it: the unquantised fp16 safetensors above fails the same way a Q8_0 GGUF does.
- Upcasting only `F.linear`, or only `F.conv1d`, does not help, though full fp32 does.

I could not narrow it past "an fp16 path inside the DiT". Happy to run more probes on this machine if that would help.

Contributor guide

Open the contributing guide

Research direction

Start by reproducing the supplied API workflow with UNETLoader, KSampler, and the official MiniMax-Music-3 files on gfx1151, then compare the default run with --fp32-unet, --disable-cuda-graphs, and --disable-comfy-compiler. Check the resulting audio with ffmpeg astats and trace where the fp16 DiT output first becomes NaN. Done means fp16 produces valid audio or a clear warning identifies the unsafe configuration.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
ai, backend, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
55/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.