huggingface / huggingface/diffusers
[New Project] Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment
- Lingua principale
- Python
- Stelle
- 34.5k
- Fork
- 7.3k
- Merge medio
- 3g 3h
- PR unite (30g)
- 91
Descrizione
**Is your feature request related to a problem? Please describe.**
With Runway and Midjourney, just now releasing updates tackling visual reference input, to be able to do character, location, and style consistency in narrative images and video, maybe this project could very well be the open-source answer to this challenge?
https://github.com/Phantom-video/Phantom

Guida per i contributori
Apri la guida per i contributori
Direzione di ricerca
The issue links the Phantom project but names no Diffusers files, tests, or entry points. Start by reviewing Phantom's approach and the repository's model-integration guidance, then determine the required scope and validation before proposing an implementation.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- python, pytorch
- Ambito
- machine-learning
- Tipo di issue
- Funzionalità
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Stato di attività
- Ferma
- Chiarezza
- Da chiarire
- Idoneità per principianti
- 20/100