Comfy-Org / Comfy-Org/ComfyUI

Feature Request: Support for GAIR/daVinci-MagiHuman (text-image-audio → video)

Open
#13,130 0 comments 15 reactions 0 assignees View on GitHub
Feature
Dominant language
Python
Stars
133k
Forks
15.7k
Avg merge
1d 7h
Merged PRs (30d)
158

Description

### Feature Idea

Hi, I’d like to request support for the daVinci-MagiHuman model:

https://huggingface.co/GAIR/daVinci-MagiHuman

This model is quite unique:
- unified multimodal generation (text + image + audio → video)
- single-stream Transformer architecture
- supports audio-driven video generation

It would be a great addition to ComfyUI’s ecosystem, especially for:
- talking head generation
- audio-driven animation
- multimodal video workflows

### Existing Solutions

_No response_

### Other

_No response_

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.