huggingface / huggingface/diffusers
Attention masks are missing in SD3 to mask out text padding tokens
- Lenguaje dominante
- Python
- Estrellas
- 34.5k
- Forks
- 7.3k
- Merge medio
- 3 d 3 h
- PR fusionados (30 d)
- 91
Descripción
### Describe the bug
In the attention implementation of SD3, attention masks currently are not used. This will result in inconsistent outputs for the different values `max_seq_length` where padding exists in text tokens as the attention scores of padding tokens are non-zero. This issue has been discussed in https://github.com/huggingface/diffusers/discussions/8628, and is created to track the progress of fixing this problem.
Thanks @sayakpaul for the discussion.
### Reproduction
n/a
### Logs
_No response_
### System Info
n/a
### Who can help?
_No response_
Guía de contribución
Línea de trabajo
Start with the SD3 attention implementation and read the linked discussion for the existing context on padding tokens. Trace how attention scores are computed when text padding is present, then verify that outputs no longer vary with max_seq_length solely because of padding.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- python, pytorch
- Área
- machine-learning
- Tipo de issue
- Error
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Estado de actividad
- Estancado
- Claridad
- Bastante claro
- Aptitud para principiantes
- 35/100