[release/2.13] [LNL] test_torch_xpu.py::TestTorchDeviceTypeXPU::test_grad_scaling_accumulation_xpu AssertionError: Tensor-likes are not close!
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 113
- Forks
- 129
- Avg merge
- 5d 9h
- Merged PRs (30d)
- 112
Description
🐛 Describe the bug
test_torch_xpu.py::TestTorchDeviceTypeXPU::test_grad_scaling_accumulation_xpu fail on LNL (Windows). Tests passed on BMG PT2.13 and LNL PT2.12
Error Message
__________ TestTorchDeviceTypeXPU.test_grad_scaling_accumulation_xpu __________
[gw0] win32 -- Python 3.12.13 C:\Users\gta\miniforge3\envs\pytorch_2.13\python.exe
Traceback (most recent call last):
File "C:\Users\gta\repositories\pytorch\pytorch\third_party\torch-xpu-ops\test\xpu\test_torch_xpu.py", line 7097, in test_grad_scaling_accumulation
self._run_scaling_case(device.type, run, unskipped=2, skipped=0)
File "C:\Users\gta\repositories\pytorch\pytorch\third_party\torch-xpu-ops\test\xpu\test_torch_xpu.py", line 6810, in _run_scaling_case
self.assertEqual(c, s, atol=atol, rtol=1e-05)
File "C:\Users\gta\miniforge3\envs\pytorch_2.13\Lib\site-packages\torch\testing\_internal\common_utils.py", line 4571, in assertEqual
raise error_metas.pop()[0].to_error( # type: ignore[index]
AssertionError: Tensor-likes are not close!
Mismatched elements: 64 / 64 (100.0%)
Greatest absolute difference: 0.3699711859226227 at index (4, 3) (up to 1e-07 allowed)
Greatest relative difference: 50.752376556396484 at index (3, 6) (up to 1e-05 allowed)
To execute this test, run the following from the base repo dir:
PYTORCH_TEST_WITH_SLOW=1 python test\xpu\test_torch_xpu.py TestTorchDeviceTypeXPU.test_grad_scaling_accumulation_xpu
Versions
PyTorch version: 2.13.0+xpu
Is debug build: False
CUDA used to build PyTorch: None
ROCM used to build PyTorch: N/A
OS: Microsoft Windows 11 Pro (10.0.26100 64-bit)
GCC version: Could not collect
Clang version: Could not collect
CMake version: version 3.31.6
Libc version: N/A
Python version: 3.12.13 | packaged by conda-forge | (main, Mar 5 2026, 16:36:12) [MSC v.1944 64 bit (AMD64)] (64-bit runtime)
Python platform: Windows-11-10.0.26100-SP0
Is CUDA available: False
CUDA runtime version: No CUDA
CUDA_MODULE_LOADING set to: N/A
GPU models and configuration: No CUDA
Nvidia driver version: No CUDA
cuDNN version: No CUDA
Is XPU available: True
XPU used to build PyTorch: 20260000
Intel GPU driver version:
- 32.0.101.8826 (20260529000000.***+)
Intel GPU models onboard: - Intel(R) Arc(TM) 140V GPU (16GB)
Intel GPU models detected: - [0] _XpuDeviceProperties(name='Intel(R) Arc(TM) 140V GPU (16GB)', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero V2', type='gpu', device_id=0x64A0, uuid=8680a064-0400-0000-0002-000000000000, driver_version='1.15.37858', total_memory=16870MB, local_mem_size=128KB, last_level_cache_size=8192KB, max_compute_units=64, memory_clock_rate=0MHz, memory_bus_width=64-bit, gpu_eu_count=64, gpu_subslice_count=8, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1, is_integrated_gpu=1)
HIP runtime version: N/A
MIOpen runtime version: N/A
Is XNNPACK available: False
Caching allocator config: N/A
CPU:
Name: Intel(R) Core(TM) Ultra 7 258V
Manufacturer: GenuineIntel
Family: 774
Architecture: 9
ProcessorType: 3
DeviceID: CPU0
CurrentClockSpeed: 2200
MaxClockSpeed: 2200
L2CacheSize: 14336
L2CacheSpeed: None
Revision: None
Versions of relevant libraries:
[pip3] dpcpp-cpp-rt==2026.0.0
[pip3] intel-cmplr-lib-rt==2026.0.0
[pip3] intel-cmplr-lib-ur==2026.0.0
[pip3] intel-cmplr-lic-rt==2026.0.0
[pip3] intel-opencl-rt==2026.0.0
[pip3] intel-openmp==2026.0.0
[pip3] intel-pti==0.17.0
[pip3] intel-sycl-rt==2026.0.0
[pip3] mkl==2026.0.0
[pip3] mkl-include==2024.2.0
[pip3] mkl-static==2024.2.0
[pip3] mypy_extensions==1.1.0
[pip3] numpy==1.26.2
[pip3] onemkl-license==2026.0.0
[pip3] onemkl-sycl-blas==2026.0.0
[pip3] onemkl-sycl-dft==2026.0.0
[pip3] onemkl-sycl-lapack==2026.0.0
[pip3] onemkl-sycl-rng==2026.0.0
[pip3] onemkl-sycl-sparse==2026.0.0
[pip3] onnx==1.21.0
[pip3] onnx-ir==0.1.16
[pip3] onnxscript==0.6.2
[pip3] optree==0.13.0
[pip3] tbb==2023.0.0
[pip3] tbb-devel==2021.13.1
[pip3] tcmlib==1.5.0
[pip3] torch==2.13.0+xpu
[pip3] torchaudio==2.11.0+xpu
[pip3] torchvision==0.28.0+xpu
[pip3] triton-xpu==3.7.2
[pip3] umf==1.1.0
[conda] dpcpp-cpp-rt 2026.0.0 pypi_0 pypi
[conda] intel-cmplr-lib-rt 2026.0.0 pypi_0 pypi
[conda] intel-cmplr-lib-ur 2026.0.0 pypi_0 pypi
[conda] intel-cmplr-lic-rt 2026.0.0 pypi_0 pypi
[conda] intel-opencl-rt 2026.0.0 pypi_0 pypi
[conda] intel-openmp 2026.0.0 pypi_0 pypi
[conda] intel-pti 0.17.0 pypi_0 pypi
[conda] intel-sycl-rt 2026.0.0 pypi_0 pypi
[conda] mkl 2026.0.0 pypi_0 pypi
[conda] mkl-include 2024.2.0 pypi_0 pypi
[conda] mkl-static 2024.2.0 pypi_0 pypi
[conda] numpy 1.26.2 pypi_0 pypi
[conda] onemkl-license 2026.0.0 pypi_0 pypi
[conda] onemkl-sycl-blas 2026.0.0 pypi_0 pypi
[conda] onemkl-sycl-dft 2026.0.0 pypi_0 pypi
[conda] onemkl-sycl-lapack 2026.0.0 pypi_0 pypi
[conda] onemkl-sycl-rng 2026.0.0 pypi_0 pypi
[conda] onemkl-sycl-sparse 2026.0.0 pypi_0 pypi
[conda] optree 0.13.0 pypi_0 pypi
[conda] tbb 2023.0.0 pypi_0 pypi
[conda] tbb-devel 2021.13.1 pypi_0 pypi
[conda] tcmlib 1.5.0 pypi_0 pypi
[conda] torch 2.13.0+xpu pypi_0 pypi
[conda] torchaudio 2.11.0+xpu pypi_0 pypi
[conda] torchvision 0.28.0+xpu pypi_0 pypi
[conda] triton-xpu 3.7.2 pypi_0 pypi
[conda] umf 1.1.0 pypi_0 pypi
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with third_party/torch-xpu-ops/test/xpu/test_torch_xpu.py, especially TestTorchDeviceTypeXPU.test_grad_scaling_accumulation_xpu and _run_scaling_case around lines 7097 and 6810. Run PYTORCH_TEST_WITH_SLOW=1 python test\xpu\test_torch_xpu.py TestTorchDeviceTypeXPU.test_grad_scaling_accumulation_xpu on LNL Windows and compare the failure with the reported BMG PT2.13 and LNL PT2.12 results. Done means the test passes on the affected LNL setup without weakening its assertions.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- testing-qa
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 45/100