[FEA] cuda.core: support stream re-capture into an existing graph (CUDA 13.3)
@Andy-Jost is already working on this.
Since Jul 23, 2026.
- Dominant language
- Cython
- Stars
- 3.4k
- Forks
- 329
- Avg merge
- 1d 23h
- Merged PRs (30d)
- 116
Description
Summary
CUDA 13.3 adds
cuStreamBeginRecaptureToGraph(CUstream hStream, CUstreamCaptureMode mode, CUgraph hGraph, CUgraphRecaptureCallback callbackFunc, void* userData):
begin a stream capture that re-captures into an existing graph instead of creating a new
one. The registered callback receives, per recaptured node,
(void* data, CUgraphNode node, const CUgraphNodeParams* originalParams, const CUgraphNodeParams* recaptureParams, CUgraphRecaptureStatus status)
with status CU_GRAPH_RECAPTURE_{ELIGIBLE_FOR_UPDATE,INELIGIBLE_FOR_UPDATE,ERROR} (exact
invocation semantics to be confirmed against the 13.3 driver docs during design). Failures
surface as the new CUDA_ERROR_GRAPH_RECAPTURE_FAILURE.
This is update-by-recapture: rerun the capture-producing code and let the driver map it onto
the existing graph — complementing the setter-based node updates of #2352 / #2354 and the
general graph-updates task #1330. cuda.core's graph support is capture-based
(GraphBuilder), so this slots in naturally.
Underlying C APIs to cover
| Symbol | Purpose |
|---|---|
cuStreamBeginRecaptureToGraph(stream, mode, graph, callback, userData) |
begin re-capture into an existing CUgraph |
CUgraphRecaptureCallback |
per-node hook comparing original vs. recaptured CUgraphNodeParams |
CUgraphRecaptureStatus |
ELIGIBLE_FOR_UPDATE / INELIGIBLE_FOR_UPDATE / ERROR |
CUDA_ERROR_GRAPH_RECAPTURE_FAILURE |
new error code to map in cuda.core error handling |
Design sketch (draft — needs design-meeting review)
[!IMPORTANT]
Starting point only, not a settled design — review in the cuda.core design meeting.
gb = graph.recapture(stream, mode="global") # names TBD; wraps cuStreamBeginRecaptureToGraph
launch(stream, config, kernel, new_args) # replay the capture region
gb.end() # graph updated in place
Graph construction is the one flow where cuda.core deliberately uses a builder, so returning a
(re)capture-scoped GraphBuilder mirrors the existing API shape.
Open questions for the meeting:
- The C callback fires from within capture — exposing an optional Python hook means calling
back into Python mid-capture (GIL/perf story). Verify whether a NULL callback is accepted
and make the Python hook opt-in. - Should cuda.core collect and report per-node recapture statuses (updated vs. ineligible)
afterend()? - Interaction with
GraphBuilder's existing capture-state guards (is_…capturing,
conditional handles). - Version gating: 13.3+ only; actionable error otherwise.
References
- Driver docs (Graph Management): https://docs.nvidia.com/cuda/cuda-driver-api/group__CUDA__GRAPH.html
- Found during the CUDA 12.8 → 13.3 bindings vs. cuda.core gap sweep (2026-07-14)
-- Leo's bot
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.