[BUG]: cudaLaunchHostFunc doesn't work with CUDA Graph
Nadie ha tomado este issue todavía.
- Lenguaje dominante
- Cython
- Estrellas
- 3.4k
- Forks
- 329
- Merge medio
- 1 d 23 h
- PR fusionados (30 d)
- 116
Descripción
Is this a duplicate?
- I confirmed there appear to be no duplicate issues for this bug and that I agree to the Code of Conduct
Type of Bug
Runtime Error
Component
cuda.bindings
Describe the bug
Hi team,
I tested cuda.bindings.runtime.cudaLaunchHostFunc, it seemed not working well with CUDA Graph. However, the C++ cudaLaunchHostFunc can work with CUDA Graph.
Similar to https://github.com/cupy/cupy/issues/9274, the host function can be captured by CUDA graph, but when replying CUDA graph, only the first reply can call the host function, while more replays cannot. Is this expected?
Please see the repro steps below. Thanks!
How to Reproduce
import ctypes
import cuda.bindings.runtime as cudart
import torch
class Struct(ctypes.Structure):
_fields_ = [
("a", ctypes.c_int),
]
def hostfunc(userData):
data = Struct.from_address(userData)
print(f"Hello, World! {data.a}")
return 0
HostFn_t = ctypes.PYFUNCTYPE(ctypes.c_int, ctypes.c_void_p)
def main():
data = Struct(a=1)
# ctypes is managing the pointer value for us
c_hostfunc = HostFn_t(hostfunc)
cuda_hostfunc = cudart.cudaHostFn_t(_ptr=ctypes.addressof(c_hostfunc))
# Run
stream = torch.cuda.Stream()
cudart_stream = cudart.cudaStream_t(stream.cuda_stream)
g = torch.cuda.CUDAGraph()
with torch.cuda.graph(g, stream=stream):
(err, ) = cudart.cudaLaunchHostFunc(cudart_stream, cuda_hostfunc,
ctypes.addressof(data))
assert err == cudart.cudaError_t.cudaSuccess
torch.cuda.synchronize()
print("Graph captured", flush=True)
with torch.cuda.stream(stream):
for i in range(10):
g.replay()
torch.cuda.synchronize()
if __name__ == "__main__":
main()
Expected behavior
The hostfunc should be called for 10 times. However, only the first time is successful:
Graph captured
Hello, World! 1
Caught signal 11 (Segmentation fault: address not mapped to object at address 0x2e906)
Segmentation fault (core dumped)
Operating System
Ubuntu 24.04.2 LTS
nvidia-smi output
Mon Aug 4 10:10:28 2025
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 575.57.08 Driver Version: 575.57.08 CUDA Version: 12.9 |
|-----------------------------------------+------------------------+----------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|=========================================+========================+======================|
| 0 NVIDIA H20 On | 00000000:D1:00.0 Off | 0 |
| N/A 37C P0 73W / 500W | 0MiB / 97871MiB | 0% Default |
| | | Disabled |
+-----------------------------------------+------------------------+----------------------+
| 1 NVIDIA H20 On | 00000000:DF:00.0 Off | 0 |
| N/A 39C P0 76W / 500W | 0MiB / 97871MiB | 0% Default |
| | | Disabled |
+-----------------------------------------+------------------------+----------------------+
+-----------------------------------------------------------------------------------------+
| Processes: |
| GPU GI CI PID Type Process name GPU Memory |
| ID ID Usage |
|=========================================================================================|
| No running processes found |
+-----------------------------------------------------------------------------------------+
Guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Línea de trabajo
Comienza ejecutando la reproducción en Python para cuda.bindings.runtime.cudaLaunchHostFunc con CUDA Graph capture y replays repetidos de torch graph. Inspecciona el binding de runtime de cuda.bindings para este punto de entrada y compara su comportamiento con el comportamiento de C++ cudaLaunchHostFunc descrito en el issue. Se considera completado cuando la función host se ejecuta correctamente en los 10 replays del grafo sin producir un segmentation fault.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- python
- Área
- api
- Tipo de issue
- Error
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Estado de actividad
- Estancado
- Claridad
- Bastante claro
- Aptitud para principiantes
- 35/100