huggingface / huggingface/diffusers

Integrate LaVie (IJCV 2024) and Cinemo (CVPR 2025) to diffusers

Aperta
#10,967 1 commento 0 reazioni 0 assegnatari Vedi su GitHub
stale
Lingua principale
Python
Stelle
34.5k
Fork
7.3k
Merge medio
3g 3h
PR unite (30g)
91

Descrizione

### Model/Pipeline/Scheduler description

I am one of the authors of LaVie (IJCV 2024) and Cinemo (CVPR 2025). We are considering submitting pull requests (PRs) to the diffusers repository for these two video models, both of which are based on U-Net architectures. We would like to know if `diffusers` is still open to supporting PRs for U-Net-based models. If so, we will proceed with the submissions. 

### Open source status

- [x] The model implementation is available.
- [x] The model weights are available (Only relevant if addition is not a scheduler).

### Provide useful links for the implementation

LaVie:  [Vchitect/LaVie: [IJCV 2024] LaVie: High-Quality Video Generation with Cascaded Latent Diffusion Models](https://github.com/Vchitect/LaVie)
Cinemo: [maxin-cn/Cinemo: [CVPR 2025] Cinemo: Consistent and Controllable Image Animation with Motion Diffusion Models](https://github.com/maxin-cn/Cinemo)

Guida per i contributori

Apri la guida per i contributori

Direzione di ricerca

Start by reading the linked LaVie and Cinemo implementations and comparing their model structure with the diffusers contribution expectations. The issue first needs confirmation that U-Net-based models are supported; done would be an agreed integration plan and subsequent pull requests for both models.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
python, pytorch
Ambito
computer-vision, machine-learning
Tipo di issue
Funzionalità
Difficoltà
5/5
Tempo stimato
Più di una settimana
Stato di attività
Ferma
Chiarezza
Da chiarire
Idoneità per principianti
25/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.