huggingface / huggingface/diffusers

Native NPU attention incorrectly converts additive attention masks to boolean masks

Aperta
#14,397 2 commenti 0 reazioni 0 assegnatari Vedi su GitHub
bug models
Lingua principale
Python
Stelle
34.5k
Fork
7.3k
Merge medio
3g 3h
PR unite (30g)
91

Descrizione

### Describe the bug

I found this issue while testing `dg845/LTX-2.3-Diffusers` with the [verl-omni LTX-2.3 NPU training script](https://github.com/verl-project/verl-omni/blob/main/examples/flowgrpo_trainer/ltx2/run_ltx2_3_t2av_lora_npu.sh).

This issue has only been observed and validated with LTX-2.3. During testing, the train-inference consistency error was around `1e-3`, while the usual empirical value in verl-omni is around `1e-5`. Further investigation traced the discrepancy to the handling of floating-point attention masks in `_native_npu_attention`.

Currently, every non-`None` attention mask is passed to `_maybe_modify_attn_mask_npu`. This helper assumes that the input is a boolean keep mask, converts it to `torch.bool`, and then inverts it because PyTorch SDPA and `npu_fusion_attention` use opposite boolean-mask polarities:

- PyTorch SDPA: `True` means attend.
- `npu_fusion_attention`: `True` means discard.

However, floating-point masks in PyTorch SDPA are additive biases:

- `0.0` means attend without changing the attention score.
- `-10000.0` or `-inf` means discard.

Casting an additive mask to boolean and then inverting it reverses its semantics:

```text
Additive mask: [0.0, 0.0, -10000.0, -10000.0]
Current NPU result: [True, True, False, False]
Expected block mask: [False, False, True, True]
```

As a result, valid positions may be discarded while masked positions may be attended, causing a silent train-inference consistency error instead of a runtime failure.

Affected code:

https://github.com/huggingface/diffusers/blob/main/src/diffusers/models/attention_dispatch.py

### Proposed fix

- Route floating-point additive masks through the existing PyTorch SDPA path.
- Keep boolean masks on the `npu_fusion_attention` path.
- Reject unsupported mask dtypes instead of silently converting them to boolean.

### Reproduction

The following isolated example demonstrates the incorrect conversion in the current helper. After the proposed routing change, floating-point masks will not be passed to this helper.

```python
import torch

from diffusers.models.attention_dispatch import _maybe_modify_attn_mask_npu

query = torch.zeros(1, 2, 1, 8)
key = torch.zeros(1, 4, 1, 8)

# PyTorch SDPA additive mask:
# 0.0 = attend, -10000.0 = discard
additive_mask = torch.tensor([[[[0.0, 0.0, -10000.0, -10000.0]]]])

actual = _maybe_modify_attn_mask_npu(query, key, additive_mask)
expected = torch.tensor(
[[[[False, False, True, True], [False, False, True, True]]]]
)

print("Actual:")
print(actual)

print("Expected:")
print(expected)

torch.testing.assert_close(actual, expected)
```

The assertion fails because the current result is:

```text
[[[[ True, True, False, False],
[ True, True, False, False]]]]
```

### System Info

- OS: Linux aarch64
- Hardware: Ascend NPU
- Python: 3.11
- PyTorch: 2.10.0
- torch-npu: 2.10.0
- CANN: 9.0.0
- diffusers: `main`

### Who can help?

@DN6 @yiyixuxu

Guida per i contributori

Apri la guida per i contributori

Direzione di ricerca

Start in src/diffusers/models/attention_dispatch.py by reading _native_npu_attention and _maybe_modify_attn_mask_npu, then run the isolated additive-mask reproduction from the issue. Done means floating-point masks use the PyTorch SDPA path, boolean masks retain the NPU path, and unsupported mask dtypes are rejected rather than silently converted.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
python, pytorch
Ambito
machine-learning
Tipo di issue
Bug
Difficoltà
3/5
Tempo stimato
1-2 giorni
Stato di attività
Tranquilla
Chiarezza
Specificata chiaramente
Idoneità per principianti
68/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.