huggingface / huggingface/diffusers

Cache text encoder embeds in pipelines

Offen
#10,078 3 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
community-examples stale
Vorherrschende Sprache
Python
Sterne
34.5k
Forks
7.3k
Ø Merge
3 T. 3 Std.
Gemergte PRs (30 T.)
91

Beschreibung

**Is your feature request related to a problem? Please describe.**

When reusing a prompt text encoder embeds are recomputed, this can be time consuming for something like T5-XXL with offloading or on CPU.

Text encoder embeds are relatively small, so keeping them in memory is feasible.
```python
import torch

clip_l = torch.randn([1, 77, 768])
t5_xxl = torch.randn([1, 512, 4096])
>>> clip_l.numel() * clip_l.dtype.itemsize
236544
>>> t5_xxl.numel() * t5_xxl.dtype.itemsize
8388608
```

**Describe the solution you'd like.**

MVP would be reusing the last text encoder embeds if the prompt hasn't changed, this behaviour is supported in community UIs. Ideally, supports multiple prompts, potentially serializable.

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

Start by tracing the pipeline text-encoding path and how prompts are converted into text encoder embeds. Define the cache behavior for unchanged prompts before considering multiple prompts or serialization. Done means repeated pipeline use avoids recomputing unchanged embeds while preserving correct results when prompts change.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
python, pytorch
Bereich
machine-learning, performance
Issue-Typ
Feature
Schwierigkeit
5/5
Geschätzter Aufwand
Über eine Woche
Aktivitätsstatus
Veraltet
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
35/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.