huggingface / huggingface/diffusers
Integrate LaVie (IJCV 2024) and Cinemo (CVPR 2025) to diffusers
- Vorherrschende Sprache
- Python
- Sterne
- 34.5k
- Forks
- 7.3k
- Ø Merge
- 3 T. 3 Std.
- Gemergte PRs (30 T.)
- 91
Beschreibung
### Model/Pipeline/Scheduler description
I am one of the authors of LaVie (IJCV 2024) and Cinemo (CVPR 2025). We are considering submitting pull requests (PRs) to the diffusers repository for these two video models, both of which are based on U-Net architectures. We would like to know if `diffusers` is still open to supporting PRs for U-Net-based models. If so, we will proceed with the submissions.
### Open source status
- [x] The model implementation is available.
- [x] The model weights are available (Only relevant if addition is not a scheduler).
### Provide useful links for the implementation
LaVie: [Vchitect/LaVie: [IJCV 2024] LaVie: High-Quality Video Generation with Cascaded Latent Diffusion Models](https://github.com/Vchitect/LaVie)
Cinemo: [maxin-cn/Cinemo: [CVPR 2025] Cinemo: Consistent and Controllable Image Animation with Motion Diffusion Models](https://github.com/maxin-cn/Cinemo)
Beitragsleitfaden
Rechercherichtung
Start by reading the linked LaVie and Cinemo implementations and comparing their model structure with the diffusers contribution expectations. The issue first needs confirmation that U-Net-based models are supported; done would be an agreed integration plan and subsequent pull requests for both models.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- python, pytorch
- Bereich
- computer-vision, machine-learning
- Issue-Typ
- Feature
- Schwierigkeit
- 5/5
- Geschätzter Aufwand
- Über eine Woche
- Aktivitätsstatus
- Veraltet
- Klarheit
- Muss geklärt werden
- Anfängerfreundlichkeit
- 25/100