huggingface / huggingface/diffusers

cogvideo model/pipeline review

Abierto
#13,622 0 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
Python
Estrellas
34.5k
Forks
7.3k
Merge medio
3 d 3 h
PR fusionados (30 d)
91

Descripción

# `cogvideo` model/pipeline review

Commit tested: `0f1abc4ae8b0eb2a3b40e82a310507281144c423`

Review performed against the repository review rules.

Reviewed target pipelines, model files, lazy exports, docs/tests/examples, dtype/device paths, offload-facing behavior, attention processors, and coverage. Public imports/lazy loading looked consistent. I did not find separate actionable issues in `pipeline_output.py` or `autoencoder_kl_cogvideox.py`.

Execution: standalone repros were run with `.venv/Scripts/python.exe`; no full pytest suite was run.

Duplicate search: searched GitHub Issues and PRs for `cogvideo`, affected class/file names, and the specific failure modes. Exact duplicate found only for Issue 5: https://github.com/huggingface/diffusers/issues/9641. Related but not exact: #11133, #9972, #13586, PR #11368, PR #9333.

## Issue 1: `num_videos_per_prompt` is accepted but ignored

Affected code:
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox.py#L518
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox.py#L618
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox_fun_control.py#L564
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox_fun_control.py#L669
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox_image2video.py#L612
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox_image2video.py#L714
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox_video2video.py#L589
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox_video2video.py#L688

Problem:
All four pipelines expose `num_videos_per_prompt`, but each `__call__` resets it to `1` before prompt encoding and latent preparation. The text-only pipeline can already use the parameter correctly if that reset is removed. The conditioned pipelines also need image/video/control latents expanded to the effective batch, or they should reject values above `1`.

Impact:
Users requesting multiple videos per prompt silently get one video per prompt. Batch behavior and callback tensor shapes are also misleading.

Reproduction:
```python
import torch
from diffusers import AutoencoderKLCogVideoX, CogVideoXDDIMScheduler, CogVideoXPipeline, CogVideoXTransformer3DModel

def tiny_pipe():
transformer = CogVideoXTransformer3DModel(
num_attention_heads=4, attention_head_dim=8, in_channels=4, out_channels=4,
time_embed_dim=2, text_embed_dim=32, num_layers=1,
sample_width=2, sample_height=2, sample_frames=9, patch_size=2,
temporal_compression_ratio=4, max_text_seq_length=16,
)
vae = AutoencoderKLCogVideoX(
in_channels=3, out_channels=3,
down_block_types=("CogVideoXDownBlock3D",) * 4,
up_block_types=("CogVideoXUpBlock3D",) * 4,
block_out_channels=(8, 8, 8, 8), latent_channels=4,
layers_per_block=1, norm_num_groups=2, temporal_compression_ratio=4,
)
return CogVideoXPipeline(None, None, transformer, vae, CogVideoXDDIMScheduler())

pipe = tiny_pipe()
frames = pipe(
prompt_embeds=torch.zeros(1, 16, 32),
height=16, width=16, num_frames=5,
num_inference_steps=1, guidance_scale=1,
num_videos_per_prompt=2, output_type="latent",
).frames
print(frames.shape)
assert frames.shape[0] == 2, f"expected 2 videos, got {frames.shape[0]}"
```

Relevant precedent:
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/wan/pipeline_wan.py#L533-L560
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/mochi/pipeline_mochi.py#L635-L650

Suggested fix:
```python
# Text-to-video: remove the forced reset.
# num_videos_per_prompt = 1

# Conditioned pipelines should either expand conditioning latents:
def _repeat_to_effective_batch(tensor, batch_size, num_videos_per_prompt):
if tensor.shape[0] == 1:
return tensor.repeat_interleave(batch_size * num_videos_per_prompt, dim=0)
if tensor.shape[0] == batch_size:
return tensor.repeat_interleave(num_videos_per_prompt, dim=0)
return tensor

# Or reject unsupported requests until expansion is implemented:
if num_videos_per_prompt != 1:
raise ValueError("`num_videos_per_prompt > 1` is not currently supported by this conditioned CogVideoX pipeline.")
```

## Issue 2: `CogVideoXFunControlPipeline` crashes when `control_video_latents` is supplied

Affected code:
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox_fun_control.py#L557-L568
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox_fun_control.py#L751-L756

Problem:
The API documents `control_video_latents`, and `check_inputs` only rejects passing both raw control video and latents. But `__call__` unconditionally runs `prepare_control_latents(None, control_video)` after the preprocessing branch. When only `control_video_latents` is supplied, `control_video` is `None`, so the user tensor is discarded and the next `.permute()` crashes.

Impact:
The precomputed control-latent path is unusable.

Reproduction:
```python
import torch
from diffusers import AutoencoderKLCogVideoX, CogVideoXDDIMScheduler, CogVideoXFunControlPipeline, CogVideoXTransformer3DModel

transformer = CogVideoXTransformer3DModel(
num_attention_heads=4, attention_head_dim=8, in_channels=8, out_channels=4,
time_embed_dim=2, text_embed_dim=32, num_layers=1,
sample_width=2, sample_height=2, sample_frames=9, patch_size=2,
temporal_compression_ratio=4, max_text_seq_length=16,
)
vae = AutoencoderKLCogVideoX(
in_channels=3, out_channels=3,
down_block_types=("CogVideoXDownBlock3D",) * 4,
up_block_types=("CogVideoXUpBlock3D",) * 4,
block_out_channels=(8, 8, 8, 8), latent_channels=4,
layers_per_block=1, norm_num_groups=2, temporal_compression_ratio=4,
)
pipe = CogVideoXFunControlPipeline(None, None, transformer, vae, CogVideoXDDIMScheduler())

try:
pipe(
prompt_embeds=torch.zeros(1, 16, 32),
control_video_latents=torch.zeros(1, 4, 2, 2, 2),
height=16, width=16, num_inference_steps=1,
guidance_scale=1, output_type="latent",
)
except Exception as e:
print(type(e).__name__, e)
```

Relevant precedent:
The raw-image/video latent paths in the other CogVideoX pipelines keep the precomputed `latents` branch separate from preprocessing.

Suggested fix:
```python
if control_video_latents is None:
if control_video is None:
raise ValueError("Provide either `control_video` or `control_video_latents`.")
control_video = self.video_processor.preprocess_video(control_video, height=height, width=width)
control_video = control_video.to(device=device, dtype=prompt_embeds.dtype)
_, control_video_latents = self.prepare_control_latents(None, control_video)
else:
control_video_latents = control_video_latents.to(device=device, dtype=prompt_embeds.dtype)

control_video_latents = control_video_latents.permute(0, 2, 1, 3, 4)
```

## Issue 3: Supplied `prompt_embeds` and `latents` are not cast to execution dtype

Affected code:
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox.py#L282-L323
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox.py#L342-L348
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox_fun_control.py#L253-L332
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox_fun_control.py#L352-L358
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox_image2video.py#L263-L342
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox_image2video.py#L410-L416
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox_video2video.py#L269-L348
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox_video2video.py#L390-L396

Problem:
Generated prompt embeddings are cast, but user-supplied `prompt_embeds` and `negative_prompt_embeds` are returned unchanged. User-supplied `latents` are moved to device but not dtype. With a bf16/fp16 pipeline and fp32 tensors, transformer projections hit dtype mismatches.

Impact:
Documented advanced inputs break mixed precision inference.

Reproduction:
```python
import torch
from diffusers import AutoencoderKLCogVideoX, CogVideoXDDIMScheduler, CogVideoXPipeline, CogVideoXTransformer3DModel

transformer = CogVideoXTransformer3DModel(
num_attention_heads=4, attention_head_dim=8, in_channels=4, out_channels=4,
time_embed_dim=2, text_embed_dim=32, num_layers=1,
sample_width=2, sample_height=2, sample_frames=9, patch_size=2,
temporal_compression_ratio=4, max_text_seq_length=16,
)
vae = AutoencoderKLCogVideoX(
in_channels=3, out_channels=3,
down_block_types=("CogVideoXDownBlock3D",) * 4,
up_block_types=("CogVideoXUpBlock3D",) * 4,
block_out_channels=(8, 8, 8, 8), latent_channels=4,
layers_per_block=1, norm_num_groups=2, temporal_compression_ratio=4,
)
pipe = CogVideoXPipeline(None, None, transformer, vae, CogVideoXDDIMScheduler()).to(dtype=torch.bfloat16)

try:
pipe(
prompt_embeds=torch.zeros(1, 16, 32, dtype=torch.float32),
height=16, width=16, num_frames=5,
num_inference_steps=1, guidance_scale=1, output_type="latent",
)
except RuntimeError as e:
print(e)
```

Relevant precedent:
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/wan/pipeline_wan.py#L544-L547
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/mochi/pipeline_mochi.py#L461-L462

Suggested fix:
```python
# After prompt embedding selection in encode_prompt:
prompt_embeds = prompt_embeds.to(device=device, dtype=dtype)
if negative_prompt_embeds is not None:
negative_prompt_embeds = negative_prompt_embeds.to(device=device, dtype=dtype)

# In prepare_latents branches:
latents = latents.to(device=device, dtype=dtype)
```

## Issue 4: Spatial validation accepts sizes that later fail patchification

Affected code:
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox.py#L387-L388
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox_fun_control.py#L427-L428
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox_image2video.py#L477-L478
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox_video2video.py#L448-L449
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/models/transformers/cogvideox_transformer_3d.py#L431-L441

Problem:
Pipelines only require `height` and `width` to be divisible by `8`, the VAE scale factor. The transformer then patchifies latents with `patch_size=2`, so the original size must usually be divisible by `8 * 2 = 16`. For example, `24x24` passes validation but produces latent `3x3`, which fails later.

Impact:
Users get a late low-level tensor shape error instead of an actionable validation error.

Reproduction:
```python
import torch
from diffusers import AutoencoderKLCogVideoX, CogVideoXDDIMScheduler, CogVideoXPipeline, CogVideoXTransformer3DModel

transformer = CogVideoXTransformer3DModel(
num_attention_heads=4, attention_head_dim=8, in_channels=4, out_channels=4,
time_embed_dim=2, text_embed_dim=32, num_layers=1,
sample_width=2, sample_height=2, sample_frames=9, patch_size=2,
temporal_compression_ratio=4, max_text_seq_length=16,
)
vae = AutoencoderKLCogVideoX(
in_channels=3, out_channels=3,
down_block_types=("CogVideoXDownBlock3D",) * 4,
up_block_types=("CogVideoXUpBlock3D",) * 4,
block_out_channels=(8, 8, 8, 8), latent_channels=4,
layers_per_block=1, norm_num_groups=2, temporal_compression_ratio=4,
)
pipe = CogVideoXPipeline(None, None, transformer, vae, CogVideoXDDIMScheduler())

try:
pipe(
prompt_embeds=torch.zeros(1, 16, 32),
height=24, width=24, num_frames=5,
num_inference_steps=1, guidance_scale=1, output_type="latent",
)
except Exception as e:
print(type(e).__name__, e)
```

Relevant precedent:
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/wan/pipeline_wan.py#L497-L511
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/hunyuan_video/pipeline_hunyuan_video.py#L351

Suggested fix:
```python
spatial_multiple = self.vae_scale_factor_spatial * self.transformer.config.patch_size
if height % spatial_multiple != 0 or width % spatial_multiple != 0:
raise ValueError(
f"`height` and `width` have to be divisible by {spatial_multiple} "
f"because CogVideoX patchifies VAE latents with patch_size={self.transformer.config.patch_size}; "
f"got {height} and {width}."
)
```

## Issue 5: Attention backend selection cannot affect CogVideoX attention

Affected code:
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/models/transformers/cogvideox_transformer_3d.py#L26
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/models/transformers/cogvideox_transformer_3d.py#L95-L104
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/models/transformers/cogvideox_transformer_3d.py#L334-L354
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/models/attention_processor.py#L2277-L2330
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/models/attention_processor.py#L2346-L2401

Problem:
`CogVideoXAttnProcessor2_0` and `FusedCogVideoXAttnProcessor2_0` are shared legacy processors without `_attention_backend` / `_parallel_config`, and they call `F.scaled_dot_product_attention` directly. `CogVideoXTransformer3DModel.set_attention_backend(...)` therefore leaves them unchanged and cannot route through the dispatcher.

Impact:
CogVideoX cannot use the model-level attention backend infrastructure consistently, including alternate kernels and parallel attention integrations covered by the review rules.

Reproduction:
```python
from diffusers import CogVideoXTransformer3DModel

model = CogVideoXTransformer3DModel(
num_attention_heads=2, attention_head_dim=8, in_channels=4, out_channels=4,
time_embed_dim=2, text_embed_dim=8, num_layers=1,
sample_width=8, sample_height=8, sample_frames=8,
patch_size=2, temporal_compression_ratio=4, max_text_seq_length=8,
)
model.set_attention_backend("native")
processors = list(model.attn_processors.values())
print([type(p).__name__ for p in processors])
print([hasattr(p, "_attention_backend") for p in processors])
assert all(hasattr(p, "_attention_backend") for p in processors)
```

Relevant precedent:
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/models/transformers/transformer_wan.py#L69-L143
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/models/transformers/transformer_flux.py#L29-L123

Suggested fix:
Move CogVideoX attention processors into `cogvideox_transformer_3d.py` as model-local processors, add `_attention_backend` and `_parallel_config`, and replace direct SDPA calls with dispatcher calls while preserving text/video split and RoPE behavior:
```python
hidden_states = dispatch_attention_fn(
query,
key,
value,
attn_mask=attention_mask,
dropout_p=0.0,
is_causal=False,
backend=self._attention_backend,
parallel_config=self._parallel_config,
)
```

## Issue 6: Dynamic CFG uses scheduler timestep values as denoising progress

Affected code:
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox.py#L737-L740
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox_fun_control.py#L803-L806
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox_image2video.py#L847-L850
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/src/diffusers/pipelines/cogvideo/pipeline_cogvideox_video2video.py#L818-L821

Problem:
This is already tracked by https://github.com/huggingface/diffusers/issues/9641, so I am not presenting it as new. The dynamic CFG formula uses `t.item()` from scheduler timesteps, not the denoising loop index/progress. With standard timesteps like `980, 960, ...`, the expression does not represent normalized progress and can exceed the requested guidance scale because it computes `1 + guidance_scale * ...`.

Impact:
`use_dynamic_cfg=True` applies an unintuitive and scheduler-dependent guidance schedule across all CogVideoX pipelines.

Reproduction:
```python
import math
from diffusers import CogVideoXDDIMScheduler

guidance_scale = 6
num_inference_steps = 50
scheduler = CogVideoXDDIMScheduler()
scheduler.set_timesteps(num_inference_steps)

scales = [
1 + guidance_scale * (
(1 - math.cos(math.pi * ((num_inference_steps - t.item()) / num_inference_steps) ** 5.0)) / 2
)
for t in scheduler.timesteps
]
print(scheduler.timesteps[:5].tolist())
print([round(x, 3) for x in scales[:12]], max(scales))
assert max(scales) <= guidance_scale
```

Relevant precedent:
The intended behavior should be resolved in the existing issue: https://github.com/huggingface/diffusers/issues/9641

Suggested fix:
Use loop progress instead of scheduler timestep value, and confirm the intended max scale against the CogVideoX implementation:
```python
progress = (num_inference_steps - i) / num_inference_steps
self._guidance_scale = 1 + (guidance_scale - 1) * ((1 - math.cos(math.pi * progress**5.0)) / 2)
```

## Issue 7: Slow tests are missing for FunControl and Video-to-Video

Affected code:
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/tests/pipelines/cogvideo/test_cogvideox_fun_control.py#L1-L330
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/tests/pipelines/cogvideo/test_cogvideox_video2video.py#L1-L326

Problem:
Fast tests exist for these pipelines, but there are no `@slow` integration tests for `CogVideoXFunControlPipeline` or `CogVideoXVideoToVideoPipeline`. Text-to-video and image-to-video do have slow coverage.

Impact:
Checkpoint compatibility, preprocessing with real media, and end-to-end output regressions are not covered for two public CogVideoX pipelines.

Reproduction:
```python
from pathlib import Path

for path in [
"tests/pipelines/cogvideo/test_cogvideox_fun_control.py",
"tests/pipelines/cogvideo/test_cogvideox_video2video.py",
]:
text = Path(path).read_text()
print(path, "@slow" in text, "IntegrationTests" in text)
assert "@slow" in text and "IntegrationTests" in text
```

Relevant precedent:
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/tests/pipelines/cogvideo/test_cogvideox.py#L339-L357
https://github.com/huggingface/diffusers/blob/0f1abc4ae8b0eb2a3b40e82a310507281144c423/tests/pipelines/cogvideo/test_cogvideox_image2video.py#L351-L369

Suggested fix:
Add slow integration classes for `CogVideoXVideoToVideoPipeline` and `CogVideoXFunControlPipeline` with published checkpoints, fixed seeds, small media fixtures, and output slice assertions. Also add a fast regression test for the `control_video_latents` path from Issue 2.

Guía de contribución

Abrir la guía de contribución

Línea de trabajo

Start with the affected CogVideoX pipeline files listed in each issue section, especially pipeline_cogvideox.py and pipeline_cogvideox_fun_control.py, then run the provided standalone reproductions with .venv/Scripts/python.exe. Review the related WAN and Mochi pipeline precedents. Done means the reported batch, control-latent, mixed-dtype, and validation cases behave correctly with coverage for the affected paths.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
python, pytorch
Área
machine-learning
Tipo de issue
Error
Dificultad
4/5
Tiempo estimado
3-5 días
Estado de actividad
Tranquilo
Claridad
Bastante claro
Aptitud para principiantes
45/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.