Sporadic test_mem_eff_attention_large_seq_len_uniform_attention_xpu failures
- Dominant language
- Python
- Stars
- 113
- Forks
- 128
- Avg merge
- 5d 9h
- Merged PRs (30d)
- 112
Description
### 🐛 Describe the bug
Cases:
op_ut,third_party.torch-xpu-ops.test.xpu.test_transformers_xpu.TestSDPAFailureModesXPU,test_mem_eff_attention_large_seq_len_uniform_attention_xpu
op_ut,third_party.torch-xpu-ops.test.xpu.test_transformers_xpu.TestSDPAFailureModesXPU,test_mem_eff_attention_large_seq_len_uniform_attention_xpu
op_ut,third_party.torch-xpu-ops.test.xpu.test_transformers_xpu.TestSDPAFailureModesXPU,test_mem_eff_attention_large_seq_len_uniform_attention_xpu
This test fails randomly. It sometimes pass in nightly, then ticket is reopen because it reproduces again. It might be problem with the driver or specific CI machine.
Bug previously seen in:
- https://github.com/intel/torch-xpu-ops/issues/2010
- https://github.com/intel/torch-xpu-ops/issues/2110
- https://github.com/intel/torch-xpu-ops/issues/2024
- https://github.com/intel/torch-xpu-ops/issues/3259
### Versions
2.12
Contributor guide
Research direction
Start with the failing test test_mem_eff_attention_large_seq_len_uniform_attention_xpu in third_party.torch-xpu-ops.test.xpu.test_transformers_xpu, and review the related issues 2010, 2110, 2024, and 3259. Reproduce it across nightly runs and compare the affected CI machines or driver versions; done means the failure cause is identified and the test is reliably passing or appropriately isolated.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- machine-learning, testing-qa
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100