pytorch / pytorch/pytorch

`torch.prod`: the gradient is `nan` where the true gradient is exactly `0` for `torch.func`

Open
#194,699 0 comments 0 reactions 0 assignees View on GitHub
bot-triaged module: derivatives module: functorch triaged
Dominant language
Python
Stars
103k
Forks
29.6k
PR merge metrics
PR metrics pending

Description

### 🐛 Describe the bug

The correct gradient is all-zero, as computed using the `backward()` function below. However, when using `torch.func.grad`, one gradient entry is `NaN`, which is not correct.

### Minimal reproducer

```python
import torch

x = torch.tensor([[1e13, 0., 0.],
[1e13, 1e13, 0.],
[1e13, 1e13, 1e13]], dtype=torch.float32)

print(torch.prod(x))
print(torch.func.grad(torch.prod)(x))

xe = x.clone().requires_grad_()
torch.prod(xe).backward()
print(xe.grad)
```

Output:
```
tensor(0.)
tensor([[0., 0., 0.],
[0., 0., nan], -> nan entry is incorrect
[0., 0., 0.]])
tensor([[0., 0., 0.],
[0., 0., 0.],
[0., 0., 0.]])
```

### Versions

Collecting environment information...
PyTorch version: 2.14.0.dev20260809+cpu
Is debug build: False
CUDA used to build PyTorch: Could not collect
ROCM used to build PyTorch: N/A

OS: Ubuntu 24.04.4 LTS (x86_64)
GCC version: (Ubuntu 13.3.0-6ubuntu2~24.04.1) 13.3.0
Clang version: 16.0.6 (https://github.com/llvm/llvm-project.git 7cbf1a2591520c2491aa35339f227775f4d3adf6)
CMake version: version 4.1.0
Libc version: glibc-2.39

Python version: 3.12.13 | packaged by conda-forge | (main, Mar 5 2026, 16:50:00) [GCC 14.3.0] (64-bit runtime)
Python platform: Linux-6.8.0-134-generic-x86_64-with-glibc2.39
Is CUDA available: False
CUDA runtime version: 12.8.93
CUDA_MODULE_LOADING set to: N/A
GPU models and configuration: No devices found.
Nvidia driver version: Could not collect
cuDNN version: Probably one of the following:
/usr/local/cuda-12.0/targets/x86_64-linux/lib/libcudnn.so.9.8.0
/usr/local/cuda-12.0/targets/x86_64-linux/lib/libcudnn_adv.so.9.8.0
/usr/local/cuda-12.0/targets/x86_64-linux/lib/libcudnn_cnn.so.9.8.0
/usr/local/cuda-12.0/targets/x86_64-linux/lib/libcudnn_engines_precompiled.so.9.8.0
/usr/local/cuda-12.0/targets/x86_64-linux/lib/libcudnn_engines_runtime_compiled.so.9.8.0
/usr/local/cuda-12.0/targets/x86_64-linux/lib/libcudnn_graph.so.9.8.0
/usr/local/cuda-12.0/targets/x86_64-linux/lib/libcudnn_heuristic.so.9.8.0
/usr/local/cuda-12.0/targets/x86_64-linux/lib/libcudnn_ops.so.9.8.0
/usr/local/cuda-12.8/targets/x86_64-linux/lib/libcudnn.so.9.8.0
/usr/local/cuda-12.8/targets/x86_64-linux/lib/libcudnn_adv.so.9.8.0
/usr/local/cuda-12.8/targets/x86_64-linux/lib/libcudnn_cnn.so.9.8.0
/usr/local/cuda-12.8/targets/x86_64-linux/lib/libcudnn_engines_precompiled.so.9.8.0
/usr/local/cuda-12.8/targets/x86_64-linux/lib/libcudnn_engines_runtime_compiled.so.9.8.0
/usr/local/cuda-12.8/targets/x86_64-linux/lib/libcudnn_graph.so.9.8.0
/usr/local/cuda-12.8/targets/x86_64-linux/lib/libcudnn_heuristic.so.9.8.0
/usr/local/cuda-12.8/targets/x86_64-linux/lib/libcudnn_ops.so.9.8.0
Is XPU available: False
HIP runtime version: N/A
MIOpen runtime version: N/A
Is XNNPACK available: False
Caching allocator config: N/A

CPU:
Architecture: x86_64
CPU op-mode(s): 32-bit, 64-bit
Address sizes: 52 bits physical, 57 bits virtual
Byte Order: Little Endian
CPU(s): 384
On-line CPU(s) list: 0-383
Vendor ID: AuthenticAMD
Model name: AMD EPYC 9654 96-Core Processor
CPU family: 25
Model: 17
Thread(s) per core: 2
Core(s) per socket: 96
Socket(s): 2
Stepping: 1
Frequency boost: enabled
CPU(s) scaling MHz: 41%
CPU max MHz: 3709.3569
CPU min MHz: 1500.0000
BogoMIPS: 4799.77
Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ht syscall nx mmxext fxsr_opt pdpe1gb rdtscp lm constant_tsc rep_good amd_lbr_v2 nopl nonstop_tsc cpuid extd_apicid aperfmperf rapl pni pclmulqdq monitor ssse3 fma cx16 pcid sse4_1 sse4_2 x2apic movbe popcnt aes xsave avx f16c rdrand lahf_lm cmp_legacy svm extapic cr8_legacy abm sse4a misalignsse 3dnowprefetch osvw ibs skinit wdt tce topoext perfctr_core perfctr_nb bpext perfctr_llc mwaitx cpb cat_l3 cdp_l3 hw_pstate ssbd mba perfmon_v2 ibrs ibpb stibp ibrs_enhanced vmmcall fsgsbase bmi1 avx2 smep bmi2 erms invpcid cqm rdt_a avx512f avx512dq rdseed adx smap avx512ifma clflushopt clwb avx512cd sha_ni avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local user_shstk avx512_bf16 clzero irperf xsaveerptr rdpru wbnoinvd amd_ppin cppc amd_ibpb_ret arat npt lbrv svm_lock nrip_save tsc_scale vmcb_clean flushbyasid decodeassists pausefilter pfthreshold avic v_vmsave_vmload vgif x2avic v_spec_ctrl vnmi avx512vbmi umip pku ospke avx512_vbmi2 gfni vaes vpclmulqdq avx512_vnni avx512_bitalg avx512_vpopcntdq la57 rdpid overflow_recov succor smca fsrm flush_l1d debug_swap ibpb_exit_to_user
Virtualization: AMD-V
L1d cache: 6 MiB (192 instances)
L1i cache: 6 MiB (192 instances)
L2 cache: 192 MiB (192 instances)
L3 cache: 768 MiB (24 instances)
NUMA node(s): 2
NUMA node0 CPU(s): 0-95,192-287
NUMA node1 CPU(s): 96-191,288-383
Vulnerability Gather data sampling: Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit: Not affected
Vulnerability L1tf: Not affected
Vulnerability Mds: Not affected
Vulnerability Meltdown: Not affected
Vulnerability Mmio stale data: Not affected
Vulnerability Reg file data sampling: Not affected
Vulnerability Retbleed: Not affected
Vulnerability Spec rstack overflow: Mitigation; Safe RET
Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization
Vulnerability Spectre v2: Mitigation; Enhanced / Automatic IBRS; IBPB conditional; STIBP always-on; PBRSB-eIBRS Not affected; BHI Not affected
Vulnerability Srbds: Not affected
Vulnerability Tsa: Vulnerable: Clear CPU buffers attempted, no microcode
Vulnerability Tsx async abort: Not affected
Vulnerability Vmscape: Mitigation; IBPB before exit to userspace

Versions of relevant libraries:
[pip3] numpy==2.2.6
[pip3] torch==2.14.0.dev20260809+cpu
[pip3] torchvision==0.29.0.dev20260810+cpu
[conda] numpy 2.2.6 pypi_0 pypi
[conda] torch 2.14.0.dev20260809+cpu pypi_0 pypi
[conda] torchvision 0.29.0.dev20260810+cpu pypi_0 pypi

cc @albanD @Chillee @samdow @kshitij12345

Contributor guide

Open the contributing guide

Research direction

Start with the minimal reproducer using torch.prod, torch.func.grad, and backward, then compare the gradient results for the zero-containing tensor. Trace the torch.func.grad path for torch.prod and add coverage for this case; done means the previously NaN entry is zero and matches backward().

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
52/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.