huggingface / huggingface/diffusers
Integrate LaVie (IJCV 2024) and Cinemo (CVPR 2025) to diffusers
- Lingua principale
- Python
- Stelle
- 34.5k
- Fork
- 7.3k
- Merge medio
- 3g 3h
- PR unite (30g)
- 91
Descrizione
### Model/Pipeline/Scheduler description
I am one of the authors of LaVie (IJCV 2024) and Cinemo (CVPR 2025). We are considering submitting pull requests (PRs) to the diffusers repository for these two video models, both of which are based on U-Net architectures. We would like to know if `diffusers` is still open to supporting PRs for U-Net-based models. If so, we will proceed with the submissions.
### Open source status
- [x] The model implementation is available.
- [x] The model weights are available (Only relevant if addition is not a scheduler).
### Provide useful links for the implementation
LaVie: [Vchitect/LaVie: [IJCV 2024] LaVie: High-Quality Video Generation with Cascaded Latent Diffusion Models](https://github.com/Vchitect/LaVie)
Cinemo: [maxin-cn/Cinemo: [CVPR 2025] Cinemo: Consistent and Controllable Image Animation with Motion Diffusion Models](https://github.com/maxin-cn/Cinemo)
Guida per i contributori
Apri la guida per i contributori
Direzione di ricerca
Start by reading the linked LaVie and Cinemo implementations and comparing their model structure with the diffusers contribution expectations. The issue first needs confirmation that U-Net-based models are supported; done would be an agreed integration plan and subsequent pull requests for both models.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- python, pytorch
- Ambito
- computer-vision, machine-learning
- Tipo di issue
- Funzionalità
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Stato di attività
- Ferma
- Chiarezza
- Da chiarire
- Idoneità per principianti
- 25/100