Feature Request: Support for GAIR/daVinci-MagiHuman (text-image-audio → video)
Open
Feature
- Dominant language
- Python
- Stars
- 133k
- Forks
- 15.7k
- Avg merge
- 1d 7h
- Merged PRs (30d)
- 158
Description
### Feature Idea
Hi, I’d like to request support for the daVinci-MagiHuman model:
https://huggingface.co/GAIR/daVinci-MagiHuman
This model is quite unique:
- unified multimodal generation (text + image + audio → video)
- single-stream Transformer architecture
- supports audio-driven video generation
It would be a great addition to ComfyUI’s ecosystem, especially for:
- talking head generation
- audio-driven animation
- multimodal video workflows
### Existing Solutions
_No response_
### Other
_No response_
Contributor guide
Assessment
This issue has not been assessed yet.