huggingface / huggingface/diffusers

Integrate LaVie (IJCV 2024) and Cinemo (CVPR 2025) to diffusers

Offen
#10,967 1 Kommentar 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
stale
Vorherrschende Sprache
Python
Sterne
34.5k
Forks
7.3k
Ø Merge
3 T. 3 Std.
Gemergte PRs (30 T.)
91

Beschreibung

### Model/Pipeline/Scheduler description

I am one of the authors of LaVie (IJCV 2024) and Cinemo (CVPR 2025). We are considering submitting pull requests (PRs) to the diffusers repository for these two video models, both of which are based on U-Net architectures. We would like to know if `diffusers` is still open to supporting PRs for U-Net-based models. If so, we will proceed with the submissions. 

### Open source status

- [x] The model implementation is available.
- [x] The model weights are available (Only relevant if addition is not a scheduler).

### Provide useful links for the implementation

LaVie:  [Vchitect/LaVie: [IJCV 2024] LaVie: High-Quality Video Generation with Cascaded Latent Diffusion Models](https://github.com/Vchitect/LaVie)
Cinemo: [maxin-cn/Cinemo: [CVPR 2025] Cinemo: Consistent and Controllable Image Animation with Motion Diffusion Models](https://github.com/maxin-cn/Cinemo)

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

Start by reading the linked LaVie and Cinemo implementations and comparing their model structure with the diffusers contribution expectations. The issue first needs confirmation that U-Net-based models are supported; done would be an agreed integration plan and subsequent pull requests for both models.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
python, pytorch
Bereich
computer-vision, machine-learning
Issue-Typ
Feature
Schwierigkeit
5/5
Geschätzter Aufwand
Über eine Woche
Aktivitätsstatus
Veraltet
Klarheit
Muss geklärt werden
Anfängerfreundlichkeit
25/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.