huggingface / huggingface/diffusers

[New Project] Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment

Aperta
#11,480 3 commenti 3 reazioni 0 assegnatari Vedi su GitHub
stale
Lingua principale
Python
Stelle
34.5k
Fork
7.3k
Merge medio
3g 3h
PR unite (30g)
91

Descrizione

**Is your feature request related to a problem? Please describe.**
With Runway and Midjourney, just now releasing updates tackling visual reference input, to be able to do character, location, and style consistency in narrative images and video, maybe this project could very well be the open-source answer to this challenge?
https://github.com/Phantom-video/Phantom

![Image](https://github.com/user-attachments/assets/fd5da955-fe1c-48e9-b806-f8acd0464d1d)

Guida per i contributori

Apri la guida per i contributori

Direzione di ricerca

The issue links the Phantom project but names no Diffusers files, tests, or entry points. Start by reviewing Phantom's approach and the repository's model-integration guidance, then determine the required scope and validation before proposing an implementation.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
python, pytorch
Ambito
machine-learning
Tipo di issue
Funzionalità
Difficoltà
5/5
Tempo stimato
Più di una settimana
Stato di attività
Ferma
Chiarezza
Da chiarire
Idoneità per principianti
20/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.