huggingface / huggingface/diffusers

[New Project] Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment

Open
#11,480 3 comments 3 reactions 0 assignees View on GitHub
stale
Dominant language
Python
Stars
34.5k
Forks
7.3k
Avg merge
3d 3h
Merged PRs (30d)
91

Description

**Is your feature request related to a problem? Please describe.**
With Runway and Midjourney, just now releasing updates tackling visual reference input, to be able to do character, location, and style consistency in narrative images and video, maybe this project could very well be the open-source answer to this challenge?
https://github.com/Phantom-video/Phantom

![Image](https://github.com/user-attachments/assets/fd5da955-fe1c-48e9-b806-f8acd0464d1d)

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.