intel / intel/torch-xpu-ops

Sporadic test_mem_eff_attention_large_seq_len_uniform_attention_xpu failures

Open
#3,326 3 comments 0 reactions 0 assignees View on GitHub
random skipped
Dominant language
Python
Stars
113
Forks
128
Avg merge
5d 9h
Merged PRs (30d)
112

Description

### 🐛 Describe the bug

Cases:
op_ut,third_party.torch-xpu-ops.test.xpu.test_transformers_xpu.TestSDPAFailureModesXPU,test_mem_eff_attention_large_seq_len_uniform_attention_xpu
op_ut,third_party.torch-xpu-ops.test.xpu.test_transformers_xpu.TestSDPAFailureModesXPU,test_mem_eff_attention_large_seq_len_uniform_attention_xpu
op_ut,third_party.torch-xpu-ops.test.xpu.test_transformers_xpu.TestSDPAFailureModesXPU,test_mem_eff_attention_large_seq_len_uniform_attention_xpu

This test fails randomly. It sometimes pass in nightly, then ticket is reopen because it reproduces again. It might be problem with the driver or specific CI machine.

Bug previously seen in:

- https://github.com/intel/torch-xpu-ops/issues/2010
- https://github.com/intel/torch-xpu-ops/issues/2110
- https://github.com/intel/torch-xpu-ops/issues/2024
- https://github.com/intel/torch-xpu-ops/issues/3259

### Versions

2.12

Contributor guide

Open the contributing guide

Research direction

Start with the failing test test_mem_eff_attention_large_seq_len_uniform_attention_xpu in third_party.torch-xpu-ops.test.xpu.test_transformers_xpu, and review the related issues 2010, 2110, 2024, and 3259. Reproduce it across nightly runs and compare the affected CI machines or driver versions; done means the failure cause is identified and the test is reliably passing or appropriately isolated.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning, testing-qa
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.