NVIDIA / NVIDIA/cuda-python

[BUG]: MemcpyNode.update() rejects stream-captured memcpy nodes

Open
#2,649 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

triage
Dominant language
Cython
Stars
3.4k
Forks
329
Avg merge
1d 23h
Merged PRs (30d)
116

Description

Is this a duplicate?
  • I confirmed there appear to be no duplicate issues for this bug and that I agree to the Code of Conduct
Type of Bug

Runtime Error

Component

cuda.core

Describe the bug

MemcpyNode.update() raises NotImplementedError on every memcpy node produced by stream capture. Buffer.copy_from and Buffer.copy_to lower to cuMemcpyAsync, which the driver records with CU_MEMORYTYPE_UNIFIED on both operands, and _is_supported_memcpy_descriptor in cuda_core/cuda/core/graph/_subclasses.pyx admits only CU_MEMORYTYPE_HOST and CU_MEMORYTYPE_DEVICE. Capturing a GraphBuilder is the primary way to build a graph in cuda.core, so update() is unavailable on most memcpy nodes a user ends up holding.

How to Reproduce
from cuda.core import Device
from cuda.core.graph import MemcpyNode

dev = Device()
dev.set_current()
stream = dev.create_stream()
src = dev.memory_resource.allocate(64, stream=stream)
dst = dev.memory_resource.allocate(64, stream=stream)
stream.sync()

builder = dev.create_graph_builder().begin_building()
dst.copy_from(src, stream=builder)
builder.end_building()

node = next(n for n in builder.graph_definition.nodes() if isinstance(n, MemcpyNode))
node.update(size=32)

Output:

Traceback (most recent call last):
  File "<stdin>", line 16, in <module>
  File "cuda/core/graph/_subclasses.pyx", line 874, in cuda.core.graph._subclasses.MemcpyNode.update
NotImplementedError: updating multidimensional, pitched, offset, or array-backed memcpy nodes is not supported
Expected behavior

node.update(size=32) should replace the copy size. The descriptor the driver recorded for this node is one-dimensional, unpitched and unoffset, so none of the reasons given in the error apply to it.

Operating System

Ubuntu 26.04 LTS

nvidia-smi output
Sun Aug 16 21:24:13 2026       
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 595.84                 Driver Version: 595.84         CUDA Version: 13.2     |
+-----------------------------------------+------------------------+----------------------+
| GPU  Name                 Persistence-M | Bus-Id          Disp.A | Volatile Uncorr. ECC |
| Fan  Temp   Perf          Pwr:Usage/Cap |           Memory-Usage | GPU-Util  Compute M. |
|                                         |                        |               MIG M. |
|=========================================+========================+======================|
|   0  NVIDIA GeForce RTX 3050 ...    Off |   00000000:01:00.0 Off |                  N/A |
| N/A   62C    P8              4W /   35W |      66MiB /   4096MiB |      0%      Default |
|                                         |                        |                  N/A |
+-----------------------------------------+------------------------+----------------------+

+-----------------------------------------------------------------------------------------+
| Processes:                                                                              |
|  GPU   GI   CI              PID   Type   Process name                        GPU Memory |
|        ID   ID                                                               Usage      |
|=========================================================================================|
|    0   N/A  N/A            6822      G   /usr/bin/gnome-shell                      1MiB |
|    0   N/A  N/A         1071386      G   /app/libexec/stremio/stremio              1MiB |
+-----------------------------------------------------------------------------------------+

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with _is_supported_memcpy_descriptor in cuda_core/cuda/core/graph/_subclasses.pyx and reproduce the GraphBuilder stream-capture example. Trace how CU_MEMORYTYPE_UNIFIED descriptors are handled by MemcpyNode.update(). Done means the captured one-dimensional memcpy accepts node.update(size=32) without raising NotImplementedError, with coverage for this reproduction.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
tooling
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Active
Clarity
Clearly specified
Newbie friendliness
72/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.