Lightning-AI / Lightning-AI/lightning-thunder

`Tensor.copy_` tries to copy onto an intermediate tensor in a canonicalized trace

Open
#1,192 2 comments 0 reactions 1 assignee View on GitHub

@crcrpar is already working on this.

Since Sep 30, 2024.

triage review
Dominant language
Python
Stars
1.5k
Forks
121
PR merge metrics
No merged PRs in 30d

Description

import torch
import thunder

@thunder.jit
def f(x):
    x.add_(1)
    return x.copy_(x.sin())

f(torch.tensor(0.0, device='cuda'))

The above results in the following error from nvFuser.

Traceback (most recent call last):
  File "/opt/pytorch/nvfuser/nvfuser/__init__.py", line 62, in __exit__
    self._finalize_definition()
RuntimeError: input_expr->isA<UnaryOp>() INTERNAL ASSERT FAILED at "/opt/pytorch/nvfuser/csrc/fusion.cpp":799, please report a bug with repro script to NVFuser at https://github.com/NVIDIA/Fuser/issues. expected unary op for aliased input
Exception raised from aliasOutputToInput at /opt/pytorch/nvfuser/csrc/fusion.cpp:799 (most recent call first):
  ...
Fusion definition
# CUDA devices:
#  0: NVIDIA RTX 6000 Ada Generation
# torch version: 2.5.0a0+gitda32021
# cuda version: 12.6
# nvfuser version: 0.2.10+gitc3f8037
import torch
from nvfuser import FusionDefinition, DataType

def nvfuser_fusion_id0(fd : FusionDefinition) -> None :
    T0 = fd.define_tensor(shape=[], contiguity=[], dtype=DataType.Float, is_cpu=False)
    S1 = fd.define_scalar(1.00000, dtype=DataType.Double)
    T2 = fd.ops.add(T0, S1)
    T3 = fd.ops.sin(T2)
    T4 = fd.ops.set(T2)
    fd.add_output(T4, T0)
    T5 = fd.ops.set(T3)
    fd.add_output(T5, T2)
    fd.add_output(T0)
    fd.add_output(T2)

with FusionDefinition() as fd:
    nvfuser_fusion_id0(fd)

I am not sure about how to interpret nvFuser's error message, but the problem would be trying to write the output of fd.ops.sin onto the output of fd.ops.add.

def computation(x):
  # x: "cuda:0 f32[]"
  [t1, t3] = nvFusion0(x)
    # t0 = prims.add(x, 1.0)  # t0: "cuda:0 f32[]"
    # t2 = prims.sin(t0)  # t2: "cuda:0 f32[]"
    # t1 = prims.copy_(t0, x)  # t1: "cuda:0 f32[]"
    # t3 = prims.copy_(t2, t0)  # t3: "cuda:0 f32[]"
  del x
  return t3

When we use x.neg() instead of x.sin(), the nvFuser executor somehow orders the copy onto t0 before the one from t0, and gets flagged as unsafe by _inplace_copy_sanity_check.

NotImplementedError: t1 = prims.copy_(t0, x)  # t1: "cuda:0 f32[]" trying to use <TensorProxy(name="t0", dtype=thunder.dtypes.float32, shape=())> (the 'copy_to' argument of 'prims.copy_') as input, which is not safe. There is a risk of accessing the wrong memory. If you are sure you don't want to use this check, it can be disabled by setting `disable_inplace_copy_check=True` in `thunder.jit`.
Python source, Execution trace
@thunder.jit
def f(x):
    x.add_(1)
    return x.copy_(x.neg())
def computation(x):
  # x: "cuda:0 f32[]"
  [t3, t1] = nvFusion0(x)
    # t0 = prims.add(x, 1.0)  # t0: "cuda:0 f32[]"
    # t2 = prims.neg(t0)  # t2: "cuda:0 f32[]"
    # t3 = prims.copy_(t2, t0)  # t3: "cuda:0 f32[]"
    # t1 = prims.copy_(t0, x)  # t1: "cuda:0 f32[]"
  del x
  return {'output': t3, 'flat_args': [t1]}

I presume this will be fixed by functionalizing Tensor.copy_ like other in-place ops, but doing so appropriately would involve somewhat big changes in thunder/core/functionalization.py

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.