intel / intel/torch-xpu-ops

[release/2.13] [LNL] test_torch_xpu.py::TestTorchDeviceTypeXPU::test_grad_scaling_accumulation_xpu AssertionError: Tensor-likes are not close!

Open
#4,137 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

hw: LNL os: Linux os: Windows test: ut
Dominant language
Python
Stars
113
Forks
129
Avg merge
5d 9h
Merged PRs (30d)
112

Description

🐛 Describe the bug

test_torch_xpu.py::TestTorchDeviceTypeXPU::test_grad_scaling_accumulation_xpu fail on LNL (Windows). Tests passed on BMG PT2.13 and LNL PT2.12

Error Message

__________ TestTorchDeviceTypeXPU.test_grad_scaling_accumulation_xpu __________
[gw0] win32 -- Python 3.12.13 C:\Users\gta\miniforge3\envs\pytorch_2.13\python.exe
Traceback (most recent call last):
  File "C:\Users\gta\repositories\pytorch\pytorch\third_party\torch-xpu-ops\test\xpu\test_torch_xpu.py", line 7097, in test_grad_scaling_accumulation
    self._run_scaling_case(device.type, run, unskipped=2, skipped=0)
  File "C:\Users\gta\repositories\pytorch\pytorch\third_party\torch-xpu-ops\test\xpu\test_torch_xpu.py", line 6810, in _run_scaling_case
    self.assertEqual(c, s, atol=atol, rtol=1e-05)
  File "C:\Users\gta\miniforge3\envs\pytorch_2.13\Lib\site-packages\torch\testing\_internal\common_utils.py", line 4571, in assertEqual
    raise error_metas.pop()[0].to_error(  # type: ignore[index]
AssertionError: Tensor-likes are not close!

Mismatched elements: 64 / 64 (100.0%)
Greatest absolute difference: 0.3699711859226227 at index (4, 3) (up to 1e-07 allowed)
Greatest relative difference: 50.752376556396484 at index (3, 6) (up to 1e-05 allowed)

To execute this test, run the following from the base repo dir:
    PYTORCH_TEST_WITH_SLOW=1 python test\xpu\test_torch_xpu.py TestTorchDeviceTypeXPU.test_grad_scaling_accumulation_xpu
Versions

PyTorch version: 2.13.0+xpu
Is debug build: False
CUDA used to build PyTorch: None
ROCM used to build PyTorch: N/A

OS: Microsoft Windows 11 Pro (10.0.26100 64-bit)
GCC version: Could not collect
Clang version: Could not collect
CMake version: version 3.31.6
Libc version: N/A

Python version: 3.12.13 | packaged by conda-forge | (main, Mar 5 2026, 16:36:12) [MSC v.1944 64 bit (AMD64)] (64-bit runtime)
Python platform: Windows-11-10.0.26100-SP0
Is CUDA available: False
CUDA runtime version: No CUDA
CUDA_MODULE_LOADING set to: N/A
GPU models and configuration: No CUDA
Nvidia driver version: No CUDA
cuDNN version: No CUDA
Is XPU available: True
XPU used to build PyTorch: 20260000
Intel GPU driver version:

  • 32.0.101.8826 (20260529000000.***+)
    Intel GPU models onboard:
  • Intel(R) Arc(TM) 140V GPU (16GB)
    Intel GPU models detected:
  • [0] _XpuDeviceProperties(name='Intel(R) Arc(TM) 140V GPU (16GB)', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero V2', type='gpu', device_id=0x64A0, uuid=8680a064-0400-0000-0002-000000000000, driver_version='1.15.37858', total_memory=16870MB, local_mem_size=128KB, last_level_cache_size=8192KB, max_compute_units=64, memory_clock_rate=0MHz, memory_bus_width=64-bit, gpu_eu_count=64, gpu_subslice_count=8, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1, is_integrated_gpu=1)
    HIP runtime version: N/A
    MIOpen runtime version: N/A
    Is XNNPACK available: False
    Caching allocator config: N/A

CPU:
Name: Intel(R) Core(TM) Ultra 7 258V
Manufacturer: GenuineIntel
Family: 774
Architecture: 9
ProcessorType: 3
DeviceID: CPU0
CurrentClockSpeed: 2200
MaxClockSpeed: 2200
L2CacheSize: 14336
L2CacheSpeed: None
Revision: None

Versions of relevant libraries:
[pip3] dpcpp-cpp-rt==2026.0.0
[pip3] intel-cmplr-lib-rt==2026.0.0
[pip3] intel-cmplr-lib-ur==2026.0.0
[pip3] intel-cmplr-lic-rt==2026.0.0
[pip3] intel-opencl-rt==2026.0.0
[pip3] intel-openmp==2026.0.0
[pip3] intel-pti==0.17.0
[pip3] intel-sycl-rt==2026.0.0
[pip3] mkl==2026.0.0
[pip3] mkl-include==2024.2.0
[pip3] mkl-static==2024.2.0
[pip3] mypy_extensions==1.1.0
[pip3] numpy==1.26.2
[pip3] onemkl-license==2026.0.0
[pip3] onemkl-sycl-blas==2026.0.0
[pip3] onemkl-sycl-dft==2026.0.0
[pip3] onemkl-sycl-lapack==2026.0.0
[pip3] onemkl-sycl-rng==2026.0.0
[pip3] onemkl-sycl-sparse==2026.0.0
[pip3] onnx==1.21.0
[pip3] onnx-ir==0.1.16
[pip3] onnxscript==0.6.2
[pip3] optree==0.13.0
[pip3] tbb==2023.0.0
[pip3] tbb-devel==2021.13.1
[pip3] tcmlib==1.5.0
[pip3] torch==2.13.0+xpu
[pip3] torchaudio==2.11.0+xpu
[pip3] torchvision==0.28.0+xpu
[pip3] triton-xpu==3.7.2
[pip3] umf==1.1.0
[conda] dpcpp-cpp-rt 2026.0.0 pypi_0 pypi
[conda] intel-cmplr-lib-rt 2026.0.0 pypi_0 pypi
[conda] intel-cmplr-lib-ur 2026.0.0 pypi_0 pypi
[conda] intel-cmplr-lic-rt 2026.0.0 pypi_0 pypi
[conda] intel-opencl-rt 2026.0.0 pypi_0 pypi
[conda] intel-openmp 2026.0.0 pypi_0 pypi
[conda] intel-pti 0.17.0 pypi_0 pypi
[conda] intel-sycl-rt 2026.0.0 pypi_0 pypi
[conda] mkl 2026.0.0 pypi_0 pypi
[conda] mkl-include 2024.2.0 pypi_0 pypi
[conda] mkl-static 2024.2.0 pypi_0 pypi
[conda] numpy 1.26.2 pypi_0 pypi
[conda] onemkl-license 2026.0.0 pypi_0 pypi
[conda] onemkl-sycl-blas 2026.0.0 pypi_0 pypi
[conda] onemkl-sycl-dft 2026.0.0 pypi_0 pypi
[conda] onemkl-sycl-lapack 2026.0.0 pypi_0 pypi
[conda] onemkl-sycl-rng 2026.0.0 pypi_0 pypi
[conda] onemkl-sycl-sparse 2026.0.0 pypi_0 pypi
[conda] optree 0.13.0 pypi_0 pypi
[conda] tbb 2023.0.0 pypi_0 pypi
[conda] tbb-devel 2021.13.1 pypi_0 pypi
[conda] tcmlib 1.5.0 pypi_0 pypi
[conda] torch 2.13.0+xpu pypi_0 pypi
[conda] torchaudio 2.11.0+xpu pypi_0 pypi
[conda] torchvision 0.28.0+xpu pypi_0 pypi
[conda] triton-xpu 3.7.2 pypi_0 pypi
[conda] umf 1.1.0 pypi_0 pypi

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with third_party/torch-xpu-ops/test/xpu/test_torch_xpu.py, especially TestTorchDeviceTypeXPU.test_grad_scaling_accumulation_xpu and _run_scaling_case around lines 7097 and 6810. Run PYTORCH_TEST_WITH_SLOW=1 python test\xpu\test_torch_xpu.py TestTorchDeviceTypeXPU.test_grad_scaling_accumulation_xpu on LNL Windows and compare the failure with the reported BMG PT2.13 and LNL PT2.12 results. Done means the test passes on the affected LNL setup without weakening its assertions.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
testing-qa
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.