huggingface / huggingface/diffusers

[New Project] Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment

Offen
#11,480 3 Kommentare 3 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
stale
Vorherrschende Sprache
Python
Sterne
34.5k
Forks
7.3k
Ø Merge
3 T. 3 Std.
Gemergte PRs (30 T.)
91

Beschreibung

**Is your feature request related to a problem? Please describe.**
With Runway and Midjourney, just now releasing updates tackling visual reference input, to be able to do character, location, and style consistency in narrative images and video, maybe this project could very well be the open-source answer to this challenge?
https://github.com/Phantom-video/Phantom

![Image](https://github.com/user-attachments/assets/fd5da955-fe1c-48e9-b806-f8acd0464d1d)

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

The issue links the Phantom project but names no Diffusers files, tests, or entry points. Start by reviewing Phantom's approach and the repository's model-integration guidance, then determine the required scope and validation before proposing an implementation.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
python, pytorch
Bereich
machine-learning
Issue-Typ
Feature
Schwierigkeit
5/5
Geschätzter Aufwand
Über eine Woche
Aktivitätsstatus
Veraltet
Klarheit
Muss geklärt werden
Anfängerfreundlichkeit
20/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.