huggingface / huggingface/diffusers
[New Project] Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment
- Vorherrschende Sprache
- Python
- Sterne
- 34.5k
- Forks
- 7.3k
- Ø Merge
- 3 T. 3 Std.
- Gemergte PRs (30 T.)
- 91
Beschreibung
**Is your feature request related to a problem? Please describe.**
With Runway and Midjourney, just now releasing updates tackling visual reference input, to be able to do character, location, and style consistency in narrative images and video, maybe this project could very well be the open-source answer to this challenge?
https://github.com/Phantom-video/Phantom

Beitragsleitfaden
Rechercherichtung
The issue links the Phantom project but names no Diffusers files, tests, or entry points. Start by reviewing Phantom's approach and the repository's model-integration guidance, then determine the required scope and validation before proposing an implementation.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- python, pytorch
- Bereich
- machine-learning
- Issue-Typ
- Feature
- Schwierigkeit
- 5/5
- Geschätzter Aufwand
- Über eine Woche
- Aktivitätsstatus
- Veraltet
- Klarheit
- Muss geklärt werden
- Anfängerfreundlichkeit
- 20/100