huggingface / huggingface/diffusers

Add VideoInpaintPipeline for temporally-consistent diffusion-based video inpainting

Offen
#12,477 10 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
stale
Vorherrschende Sprache
Python
Sterne
34.5k
Forks
7.3k
Ø Merge
3 T. 3 Std.
Gemergte PRs (30 T.)
91

Beschreibung

## Feature request

Introduce a new pipeline to extend existing image inpainting capabilities (StableDiffusionInpaintPipeline) to videos. The goal is to provide a native, GPU-optimized API within Diffusers that performs temporally coherent video inpainting instead of independent per-frame processing.

## Motivation
Current video inpainting approaches in the community simply loop over frames and call the image inpainting pipeline repeatedly.

This leads to:

- Temporal flicker and inconsistent textures between frames.
- Poor GPU utilization and high memory overhead.
- Lack of tools to maintain motion coherence or reuse diffusion latents across time.
- A built-in VideoInpaintPipeline would make it possible to remove objects, restore scenes, or creatively edit videos using diffusion models while keeping motion and lighting consistent across frames.

## Your contribution

I plan to:

Implement VideoInpaintPipeline as a subclass of DiffusionPipeline, leveraging StableDiffusionInpaintPipeline under the hood.
Add temporal consistency mechanisms, such as latent reuse between frames and optional optical-flow–guided warping (RAFT / GMFlow).
Optimize performance through batched FP16 inference, scheduler noise reuse, and optional torch.compile acceleration.
Provide a clean user API compatible with existing pipelines:

```
from diffusers import VideoInpaintPipeline

pipe = VideoInpaintPipeline.from_pretrained(
"runwayml/stable-diffusion-inpainting",
use_optical_flow=True,
compile=True,
)

result = pipe(
video_path="input.mp4",
mask_path="mask.mp4",
prompt="replace background with a snowy mountain",
num_inference_steps=10,
)
result.video.save("output.mp4")
```
Contribute documentation and tests demonstrating temporal coherence, performance benchmarks, and example notebooks for real-world use.

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

Start by reading DiffusionPipeline and StableDiffusionInpaintPipeline, which the request identifies as the proposed base and underlying pipeline. Clarify the temporal-consistency, optical-flow, batching, and compilation requirements before implementation. Done should include the requested video and mask API, temporal-coherence tests, performance benchmarks, documentation, and example notebooks.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
python, pytorch
Bereich
computer-vision, machine-learning, performance
Issue-Typ
Feature
Schwierigkeit
5/5
Geschätzter Aufwand
Über eine Woche
Aktivitätsstatus
Veraltet
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
25/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.