huggingface / huggingface/diffusers
Make the input `UNet2DConditionModel` and the returned `UNetMotionModel` instances share weights
- Langage dominant
- Python
- Étoiles
- 34.5k
- Forks
- 7.3k
- Merge moyen
- 3 j 3 h
- PR mergées (30 j)
- 91
Description
**What API design would you like to have changed or added to the library? Why?**
[diffusers.UNetMotionModel](https://huggingface.co/docs/diffusers/main/en/api/models/unet-motion#diffusers.UNetMotionModel).from_unet2d
https://github.com/huggingface/diffusers/blob/6be43bd855772f5ffb064bc7c6049f578b3856f8/src/diffusers/models/unets/unet_motion_model.py#L431-L437
**New feature**
Add an option to make the input `UNet2DConditionModel` and the returned `UNetMotionModel` instances share weights (Res, SA, CA).

> Image from [PIA project page](https://pi-animator.github.io)
**What use case would this enable or better enable? Can you give us a code example?**
In the [example](https://huggingface.co/docs/diffusers/en/api/pipelines/pia#usage-example) in the current doc of PIA, users can provide an existing image to generate a video (img2vid). However, one can also perform text2vid by using an image pipeline to generate the input image first. In such case, it would be best if we only make **one copy** of the weights (Res, SA, CA) in both unets of image and video pipeline.
| Modules | Res | SA | CA | TA | `conv_in` (first 4 ch) | `conv_in` (last 5 ch) |
|-|-|-|-|-|-|-|
|`UNet2DConditionModel`|✅|✅|✅|❌|✅|❌|
|`UNetMotionModel`|✅|✅|✅|✅|✅|✅|
**Note**
The additional channels in `conv_in` takes neglectable amount of vram, and I don't know if it is possible to share only half of this module.
Guide de contribution
Ouvrir le guide de contribution
Piste de recherche
Commencez par src/diffusers/models/unets/unet_motion_model.py, en particulier au niveau du point d’entrée from_unet2d autour des lignes liées, et examinez comment l’exemple d’utilisation de PIA construit les pipelines d’image et de vidéo. Définissez le comportement de l’option concernant le partage des poids Res, SA et CA tout en préservant les modules spécifiques au mouvement et les différents canaux conv_in ; le travail est terminé lorsque le UNet2DConditionModel d’entrée et le UNetMotionModel renvoyé partagent ces poids.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Évaluation
- Stack technique
- python, pytorch
- Domaine
- machine-learning
- Type d'issue
- Fonctionnalité
- Difficulté
- 5/5
- Temps estimé
- Plus d'une semaine
- Activité
- À l'abandon
- Clarté
- Plutôt claire
- Accessibilité débutants
- 35/100