huggingface / huggingface/diffusers

Implementing training-free RB-Modulation pipeline for most used models

Aperta
#9,283 4 commenti 2 reazioni 0 assegnatari Vedi su GitHub
consider-for-modular-diffusers wip
Lingua principale
Python
Stelle
34.5k
Fork
7.3k
Merge medio
3g 3h
PR unite (30g)
91

Descrizione

### Model/Pipeline/Scheduler description

The [RB-Modulation algorithm](https://rb-modulation.github.io/) is **training-free** technique to produce image 2 image style and content transfer in diffusion model. It has two components:
1. Stochastic Optimization Control (SOC): This component requires an [evaluator](https://github.com/learn2phoenix/CSD) for the style at each timestep. Therefore, an evaluator model and control function pipeline has to be built.
2. AttentionFeatureAggregation (AFA): This needs a clip image encoder to concat the K,V features of the image and caption. A slight tweak has to be done in the forward pass of the existing models.

This will be an interesting implementation for edits as the paper shows promising results.

### Open source status

- [X] The model implementation is available.
- [X] The model weights are available (Only relevant if addition is not a scheduler).

### Provide useful links for the implementation

### RB-Modulation:
Title: RB-Modulation: Training-Free Personalization of Diffusion Models using Stochastic Optimal Control
Code Link: https://github.com/google/RB-Modulation
Authors: Litu Rout and Yujia Chen and Nataniel Ruiz and Abhishek Kumar and Constantine Caramanis and Sanjay Shakkottai and Wen-Sheng Chu
Authors GH Username: @LituRout, @IssacCyj

### Style Evaluator:
Title: Measuring Style Similarity in Diffusion Models
Code Link: https://github.com/learn2phoenix/CSD
Authors: Somepalli, Gowthami and Gupta, Anubhav and Gupta, Kamal and Palta, Shramay and Goldblum, Micah and Geiping, Jonas and Shrivastava, Abhinav and Goldstein, Tom
Authors Username: @somepago, @learn2phoenix

Guida per i contributori

Apri la guida per i contributori

Direzione di ricerca

Start by reviewing the RB-Modulation implementation and paper, then compare the repository's existing diffusion model and scheduler entry points with the SOC and AttentionFeatureAggregation components described here. Investigate the CSD evaluator and the linked Google/RB-Modulation code before defining the supported models and integration boundaries. Done means a working training-free image-to-image style and content transfer pipeline for the selected models.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
python, pytorch
Ambito
machine-learning
Tipo di issue
Funzionalità
Difficoltà
5/5
Tempo stimato
Più di una settimana
Stato di attività
Ferma
Chiarezza
Da chiarire
Idoneità per principianti
25/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.