huggingface / huggingface/diffusers

Add support for JoyAI-Video-Edit

Offen
#14,524 1 Kommentar 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Vorherrschende Sprache
Python
Sterne
34.5k
Forks
7.3k
Ø Merge
3 T. 3 Std.
Gemergte PRs (30 T.)
91

Beschreibung

### Model/Pipeline/Scheduler description

[JoyAI-Video-Edit](https://github.com/jd-opensource/JoyAI-Video-Edit) is an instruction-guided video editing model based on the JoyAI streaming architecture.

It performs causal, chunk-wise video denoising with a per-layer KV cache. Each chunk attends to previously denoised chunks and may additionally attend to an optional reference image.

Prompt embeddings are produced by the external [MiMo-VL](https://huggingface.co/XiaomiMiMo/MiMo-VL-7B-RL-2508) model.

### Open source status

- [x] The model implementation is available.
- [x] The model weights are available (Only relevant if addition is not a scheduler).

### Provide useful links for the implementation

- Reference implementation: https://github.com/jd-opensource/JoyAI-Video-Edit
- Original checkpoint: https://huggingface.co/jdopensource/JoyAI-Video-Edit
- MiMo-VL: https://huggingface.co/XiaomiMiMo/MiMo-VL-7B-RL-2508

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

No repository files, tests, or entry points are named. Start by comparing the linked reference implementation and checkpoint with existing video model integrations in diffusers, then determine how MiMo-VL prompt embeddings, causal KV caching, chunk-wise denoising, and optional reference images fit the project. Done means JoyAI-Video-Edit is supported using the linked model resources.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
python, pytorch
Bereich
computer-vision, machine-learning
Issue-Typ
Feature
Schwierigkeit
5/5
Geschätzter Aufwand
Über eine Woche
Aktivitätsstatus
Ruhig
Klarheit
Muss geklärt werden
Anfängerfreundlichkeit
35/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.