MiniMax H3-audio always has static and popping sounds.
- Dominant language
- Python
- Stars
- 133k
- Forks
- 15.7k
- Avg merge
- 1d 7h
- Merged PRs (30d)
- 158
Description
### Custom Node Testing
- [ ] I have tried disabling custom nodes and the issue persists (see [how to disable custom nodes](https://docs.comfy.org/troubleshooting/custom-node-issues#step-1%3A-test-with-all-custom-nodes-disabled) if you need help)
### Your question
I've tried ComfyUI version 0.31, and also version 0.33.
I've tried using the acceleration LoRA, and also without it.
I've tried the official workflow, and also custom workflows with additional nodes.
I've tried 8-step sampling, and also 20-step sampling.
I've tried the res_multistep sampler, and also the Euler sampler.
I've tried the pruned version of the main model, and also the 32GB INT8 version.
I've tried CUDA 13.3 + torch 2.13, and also CUDA 12.8 + torch 2.11.
I've tried enabling SageAttention, and also disabling it.
None of it works—the generated videos still have audio popping issues. Can anyone help me out?
My GPU is an RTX 4080 Super with 32GB of VRAM."
https://github.com/user-attachments/assets/a9a9b5e0-09f4-4cb7-a452-de3f62bbe60d
https://github.com/user-attachments/assets/717c3336-3226-4cfe-b858-b1cad6e276f6
### Logs
```powershell
```
### Other
_No response_
Contributor guide
Research direction
Start by reproducing the issue with the official workflow and the MiniMax H3-audio model, then compare results across the listed sampler, LoRA, model, CUDA/PyTorch, and SageAttention configurations. The audio output should be checked for static and popping; progress will require capturing useful logs because the issue currently provides none.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- audio-video-rtc, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 38/100