huggingface / huggingface/diffusers

Attention masks are missing in SD3 to mask out text padding tokens

Abierto
#8,673 9 comentarios 0 reacciones 0 asignados Ver en GitHub
contributions-welcome wip
Lenguaje dominante
Python
Estrellas
34.5k
Forks
7.3k
Merge medio
3 d 3 h
PR fusionados (30 d)
91

Descripción

### Describe the bug

In the attention implementation of SD3, attention masks currently are not used. This will result in inconsistent outputs for the different values `max_seq_length` where padding exists in text tokens as the attention scores of padding tokens are non-zero. This issue has been discussed in https://github.com/huggingface/diffusers/discussions/8628, and is created to track the progress of fixing this problem.

Thanks @sayakpaul for the discussion.

### Reproduction

n/a

### Logs

_No response_

### System Info

n/a

### Who can help?

_No response_

Guía de contribución

Abrir la guía de contribución

Línea de trabajo

Start with the SD3 attention implementation and read the linked discussion for the existing context on padding tokens. Trace how attention scores are computed when text padding is present, then verify that outputs no longer vary with max_seq_length solely because of padding.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
python, pytorch
Área
machine-learning
Tipo de issue
Error
Dificultad
4/5
Tiempo estimado
3-5 días
Estado de actividad
Estancado
Claridad
Bastante claro
Aptitud para principiantes
35/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.