[inductor] generated kernel calls `empty_strided_cpu` with a negative stride → `Expected result >= 0`
- Dominant language
- Python
- Stars
- 103k
- Forks
- 29.5k
- PR merge metrics
- PR metrics pending
Description
### 🐛 Describe the bug
The Inductor-generated kernel emits `empty_strided_cpu((0, 0), (-1, 0), ...)` with a negative stride, which fails at runtime. Eager mode produces a valid tensor; only the `torch.compile(backend="inductor")` path fails. Ops involved: `float_power`, `empty_strided`, `gather`.
## Repro
```python
# Minimal repro for crash 31947ae5 (compiled-only: eager OK, inductor crashes)
import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(0)
t31856 = torch.randn([3, 3], dtype=torch.complex64)
def model_259():
t31857 = torch.float_power(t31856, 0.024490423949617878)
t31851 = torch.empty_strided((0, 0), (-1, 0), dtype=torch.bool, pin_memory=False, requires_grad=False)
t31858 = torch.gather(t31857, -2, t31851)
return t31858
# eager baseline -- succeeds
eager_out = model_259()
print("eager OK:", type(eager_out))
# inductor compile -- crashes
compiled = torch.compile(model_259, backend="inductor")
with torch.no_grad():
compiled_out = compiled()
print("compiled OK:", type(compiled_out))
```
Eager runs successfully (`eager OK` prints); the crash happens in the compiled call.
### Error logs
```
Traceback (most recent call last):
File "/home/crash_analysis/torch/programs//_bug_drafts/repro_31947ae5.py", line 20, in
compiled_out = compiled()
File "/home/.venv/lib/python3.10/site-packages/torch/_dynamo/eval_frame.py", line 1024, in compile_wrapper
return fn(*args, **kwargs)
File "/home/crash_analysis/torch/programs//_bug_drafts/repro_31947ae5.py", line 7, in model_259
def model_259():
File "/home/.venv/lib/python3.10/site-packages/torch/_dynamo/eval_frame.py", line 1263, in _fn
return fn(*args, **kwargs)
File "/home/.venv/lib/python3.10/site-packages/torch/_functorch/aot_autograd.py", line 1200, in forward
return compiled_fn(full_args)
File "/home/.venv/lib/python3.10/site-packages/torch/_functorch/_aot_autograd/runtime_wrappers.py", line 580, in runtime_wrapper
all_outs = call_func_at_runtime_with_args(
File "/home/.venv/lib/python3.10/site-packages/torch/_functorch/_aot_autograd/utils.py", line 138, in call_func_at_runtime_with_args
out = normalize_as_list(f(args))
File "/home/.venv/lib/python3.10/site-packages/torch/_functorch/_aot_autograd/runtime_wrappers.py", line 2298, in __call__
return self.compiled_fn(*args, **kwargs)
File "/home/.venv/lib/python3.10/site-packages/torch/_functorch/_aot_autograd/runtime_wrappers.py", line 783, in wrapper
return compiled_fn(runtime_args)
File "/home/.venv/lib/python3.10/site-packages/torch/_functorch/_aot_autograd/runtime_wrappers.py", line 1011, in inner_fn
outs = compiled_fn(args)
File "/home/.venv/lib/python3.10/site-packages/torch/_inductor/output_code.py", line 656, in __call__
return self.current_callable(inputs)
File "/tmp/torchinductor_lliu39/jh/cjhztpj5odswxjbccwnqoden5szrrtetszy5t3dzhw5nejt5kgrq.py", line 64, in call
buf4 = empty_strided_cpu((0, 0), (-1, 0), torch.bool)
RuntimeError: Expected result >= 0 to be true, but got false. (Could this error message be improved? If so, please report an enhancement request to PyTorch.)
```
### Versions
Collecting environment information...
PyTorch version: 2.11.0+cu130
Is debug build: False
CUDA used to build PyTorch: 13.0
ROCM used to build PyTorch: N/A
OS: Ubuntu 22.04.5 LTS (x86_64)
GCC version: (Ubuntu 11.4.0-1ubuntu1~22.04.3) 11.4.0
Clang version: 15.0.0 ([git@github.com](mailto:git@github.com):llvm/llvm-project.git 4ba6a9c9f65bbc8bd06e3652cb20fd4dfc846137)
CMake version: version 3.22.1
Libc version: glibc-2.35
Python version: 3.10.12 (main, Mar 3 2026, 11:56:32) [GCC 11.4.0] (64-bit runtime)
Python platform: Linux-6.8.0-94-generic-x86_64-with-glibc2.35
Is CUDA available: False
CUDA runtime version: No CUDA
CUDA_MODULE_LOADING set to: N/A
GPU models and configuration: No CUDA
Nvidia driver version: No CUDA
cuDNN version: No CUDA
Is XPU available: False
HIP runtime version: N/A
MIOpen runtime version: N/A
Is XNNPACK available: True
Caching allocator config: N/A
CPU:
Architecture: x86_64
CPU op-mode(s): 32-bit, 64-bit
Address sizes: 52 bits physical, 57 bits virtual
Byte Order: Little Endian
CPU(s): 384
On-line CPU(s) list: 0-383
Vendor ID: AuthenticAMD
Model name: AMD EPYC 9684X 96-Core Processor
CPU family: 25
Model: 17
Thread(s) per core: 2
Core(s) per socket: 96
Socket(s): 2
Stepping: 2
BogoMIPS: 5099.98
Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ht syscall nx mmxext fxsr_opt pdpe1gb rdtscp lm constant_tsc rep_good amd_lbr_v2 nopl nonstop_tsc cpuid extd_apicid aperfmperf rapl pni pclmulqdq monitor ssse3 fma cx16 pcid sse4_1 sse4_2 x2apic movbe popcnt aes xsave avx f16c rdrand lahf_lm cmp_legacy svm extapic cr8_legacy abm sse4a misalignsse 3dnowprefetch osvw ibs skinit wdt tce topoext perfctr_core perfctr_nb bpext perfctr_llc mwaitx cpb cat_l3 cdp_l3 hw_pstate ssbd mba perfmon_v2 ibrs ibpb stibp ibrs_enhanced vmmcall fsgsbase bmi1 avx2 smep bmi2 erms invpcid cqm rdt_a avx512f avx512dq rdseed adx smap avx512ifma clflushopt clwb avx512cd sha_ni avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local user_shstk avx512_bf16 clzero irperf xsaveerptr rdpru wbnoinvd amd_ppin cppc amd_ibpb_ret arat npt lbrv svm_lock nrip_save tsc_scale vmcb_clean flushbyasid decodeassists pausefilter pfthreshold avic v_vmsave_vmload vgif x2avic v_spec_ctrl vnmi avx512vbmi umip pku ospke avx512_vbmi2 gfni vaes vpclmulqdq avx512_vnni avx512_bitalg avx512_vpopcntdq la57 rdpid overflow_recov succor smca fsrm flush_l1d debug_swap ibpb_exit_to_user
Virtualization: AMD-V
L1d cache: 6 MiB (192 instances)
L1i cache: 6 MiB (192 instances)
L2 cache: 192 MiB (192 instances)
L3 cache: 2.3 GiB (24 instances)
NUMA node(s): 2
NUMA node0 CPU(s): 0-95,192-287
NUMA node1 CPU(s): 96-191,288-383
Vulnerability Gather data sampling: Not affected
Vulnerability Itlb multihit: Not affected
Vulnerability L1tf: Not affected
Vulnerability Mds: Not affected
Vulnerability Meltdown: Not affected
Vulnerability Mmio stale data: Not affected
Vulnerability Reg file data sampling: Not affected
Vulnerability Retbleed: Not affected
Vulnerability Spec rstack overflow: Mitigation; Safe RET
Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization
Vulnerability Spectre v2: Mitigation; Enhanced / Automatic IBRS; IBPB conditional; STIBP always-on; RSB filling; PBRSB-eIBRS Not affected; BHI Not affected
Vulnerability Srbds: Not affected
Vulnerability Tsx async abort: Not affected
Vulnerability Vmscape: Mitigation; IBPB before exit to userspace
Versions of relevant libraries:
[pip3] numpy==2.2.6
[pip3] nvidia-cublas==13.1.0.3
[pip3] nvidia-cuda-cupti==13.0.85
[pip3] nvidia-cuda-nvrtc==13.0.88
[pip3] nvidia-cuda-runtime==13.0.96
[pip3] nvidia-cudnn-cu13==9.19.0.56
[pip3] nvidia-cufft==12.0.0.61
[pip3] nvidia-curand==10.4.0.35
[pip3] nvidia-cusolver==12.0.4.66
[pip3] nvidia-cusparse==12.6.3.3
[pip3] nvidia-cusparselt-cu13==0.8.0
[pip3] nvidia-nccl-cu13==2.28.9
[pip3] nvidia-nvjitlink==13.0.88
[pip3] nvidia-nvtx==13.0.85
[pip3] optree==0.19.0
[pip3] torch==2.11.0
[pip3] triton==3.6.0
[conda] Could not collect
cc @chauhang @penguinwu
Contributor guide
Research direction
Run the minimal Python repro with torch.compile(..., backend="inductor") and inspect the generated output_code.py at the call that emits empty_strided_cpu((0, 0), (-1, 0), ...). Trace how the float_power, empty_strided, and gather operations produce that allocation. Done means the compiled call no longer fails and its result matches eager mode.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- compilers
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Clearly specified
- Newbie friendliness
- 48/100