NVIDIA / NVIDIA/cuda-python

[BUG]: cudaLaunchHostFunc doesn't work with CUDA Graph

Abierto
#790 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Nadie ha tomado este issue todavía.

bug cuda.bindings triage
Lenguaje dominante
Cython
Estrellas
3.4k
Forks
329
Merge medio
1 d 23 h
PR fusionados (30 d)
116

Descripción

Is this a duplicate?
Type of Bug

Runtime Error

Component

cuda.bindings

Describe the bug

Hi team,

I tested cuda.bindings.runtime.cudaLaunchHostFunc, it seemed not working well with CUDA Graph. However, the C++ cudaLaunchHostFunc can work with CUDA Graph.

Similar to https://github.com/cupy/cupy/issues/9274, the host function can be captured by CUDA graph, but when replying CUDA graph, only the first reply can call the host function, while more replays cannot. Is this expected?

Please see the repro steps below. Thanks!

How to Reproduce
import ctypes

import cuda.bindings.runtime as cudart
import torch


class Struct(ctypes.Structure):
    _fields_ = [
        ("a", ctypes.c_int),
    ]


def hostfunc(userData):
    data = Struct.from_address(userData)
    print(f"Hello, World! {data.a}")
    return 0


HostFn_t = ctypes.PYFUNCTYPE(ctypes.c_int, ctypes.c_void_p)


def main():
    data = Struct(a=1)

    # ctypes is managing the pointer value for us
    c_hostfunc = HostFn_t(hostfunc)
    cuda_hostfunc = cudart.cudaHostFn_t(_ptr=ctypes.addressof(c_hostfunc))

    # Run
    stream = torch.cuda.Stream()
    cudart_stream = cudart.cudaStream_t(stream.cuda_stream)

    g = torch.cuda.CUDAGraph()
    with torch.cuda.graph(g, stream=stream):
        (err, ) = cudart.cudaLaunchHostFunc(cudart_stream, cuda_hostfunc,
                                            ctypes.addressof(data))
        assert err == cudart.cudaError_t.cudaSuccess
    torch.cuda.synchronize()
    print("Graph captured", flush=True)

    with torch.cuda.stream(stream):
        for i in range(10):
            g.replay()

    torch.cuda.synchronize()


if __name__ == "__main__":
    main()
Expected behavior

The hostfunc should be called for 10 times. However, only the first time is successful:

Graph captured
Hello, World! 1
Caught signal 11 (Segmentation fault: address not mapped to object at address 0x2e906)
Segmentation fault (core dumped)
Operating System

Ubuntu 24.04.2 LTS

nvidia-smi output
Mon Aug  4 10:10:28 2025       
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 575.57.08              Driver Version: 575.57.08      CUDA Version: 12.9     |
|-----------------------------------------+------------------------+----------------------+
| GPU  Name                 Persistence-M | Bus-Id          Disp.A | Volatile Uncorr. ECC |
| Fan  Temp   Perf          Pwr:Usage/Cap |           Memory-Usage | GPU-Util  Compute M. |
|                                         |                        |               MIG M. |
|=========================================+========================+======================|
|   0  NVIDIA H20                     On  |   00000000:D1:00.0 Off |                    0 |
| N/A   37C    P0             73W /  500W |       0MiB /  97871MiB |      0%      Default |
|                                         |                        |             Disabled |
+-----------------------------------------+------------------------+----------------------+
|   1  NVIDIA H20                     On  |   00000000:DF:00.0 Off |                    0 |
| N/A   39C    P0             76W /  500W |       0MiB /  97871MiB |      0%      Default |
|                                         |                        |             Disabled |
+-----------------------------------------+------------------------+----------------------+
                                                                                         
+-----------------------------------------------------------------------------------------+
| Processes:                                                                              |
|  GPU   GI   CI              PID   Type   Process name                        GPU Memory |
|        ID   ID                                                               Usage      |
|=========================================================================================|
|  No running processes found                                                             |
+-----------------------------------------------------------------------------------------+

Guía de contribución

Abrir la guía de contribución

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Línea de trabajo

Comienza ejecutando la reproducción en Python para cuda.bindings.runtime.cudaLaunchHostFunc con CUDA Graph capture y replays repetidos de torch graph. Inspecciona el binding de runtime de cuda.bindings para este punto de entrada y compara su comportamiento con el comportamiento de C++ cudaLaunchHostFunc descrito en el issue. Se considera completado cuando la función host se ejecuta correctamente en los 10 replays del grafo sin producir un segmentation fault.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
python
Área
api
Tipo de issue
Error
Dificultad
4/5
Tiempo estimado
3-5 días
Estado de actividad
Estancado
Claridad
Bastante claro
Aptitud para principiantes
35/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.