huggingface / huggingface/diffusers

Add VideoInpaintPipeline for temporally-consistent diffusion-based video inpainting

Abierto
#12,477 10 comentarios 0 reacciones 0 asignados Ver en GitHub

Nadie ha tomado este issue todavía.

stale
Lenguaje dominante
Python
Estrellas
34.5k
Forks
7.3k
Merge medio
3 d 3 h
PR fusionados (30 d)
91

Descripción

Feature request

Introduce a new pipeline to extend existing image inpainting capabilities (StableDiffusionInpaintPipeline) to videos. The goal is to provide a native, GPU-optimized API within Diffusers that performs temporally coherent video inpainting instead of independent per-frame processing.

Motivation

Current video inpainting approaches in the community simply loop over frames and call the image inpainting pipeline repeatedly.

This leads to:

  • Temporal flicker and inconsistent textures between frames.
  • Poor GPU utilization and high memory overhead.
  • Lack of tools to maintain motion coherence or reuse diffusion latents across time.
  • A built-in VideoInpaintPipeline would make it possible to remove objects, restore scenes, or creatively edit videos using diffusion models while keeping motion and lighting consistent across frames.

Your contribution

I plan to:

Implement VideoInpaintPipeline as a subclass of DiffusionPipeline, leveraging StableDiffusionInpaintPipeline under the hood.
Add temporal consistency mechanisms, such as latent reuse between frames and optional optical-flow–guided warping (RAFT / GMFlow).
Optimize performance through batched FP16 inference, scheduler noise reuse, and optional torch.compile acceleration.
Provide a clean user API compatible with existing pipelines:

from diffusers import VideoInpaintPipeline

pipe = VideoInpaintPipeline.from_pretrained(
    "runwayml/stable-diffusion-inpainting",
    use_optical_flow=True,
    compile=True,
)

result = pipe(
    video_path="input.mp4",
    mask_path="mask.mp4",
    prompt="replace background with a snowy mountain",
    num_inference_steps=10,
)
result.video.save("output.mp4")

Contribute documentation and tests demonstrating temporal coherence, performance benchmarks, and example notebooks for real-world use.

Guía de contribución

Abrir la guía de contribución

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Línea de trabajo

Empieza leyendo DiffusionPipeline y StableDiffusionInpaintPipeline, que la solicitud identifica como la clase base propuesta y la pipeline subyacente. Aclara los requisitos de consistencia temporal, flujo óptico, batching y compilación antes de la implementación. Done debe incluir la API solicitada para vídeo y máscaras, pruebas de coherencia temporal, benchmarks de rendimiento, documentación y notebooks de ejemplo.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
python, pytorch
Área
computer-vision, machine-learning, performance
Tipo de issue
Nueva funcionalidad
Dificultad
5/5
Tiempo estimado
Más de una semana
Estado de actividad
Estancado
Claridad
Bastante claro
Aptitud para principiantes
25/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.