NVIDIA / NVIDIA/cuda-python

[FEA] cuda.core: support stream batch memory operations (incl. CUDA 13.1 atomic reductions)

Offen
#2,364 0 Kommentare 0 Reaktionen 1 zugewiesene Person Auf GitHub ansehen

@juenglin arbeitet bereits daran.

Seit 24.7.2026.

cuda.core feature
Vorherrschende Sprache
Cython
Sterne
3.4k
Forks
329
Ø Merge
1 T. 23 Std.
Gemergte PRs (30 T.)
116

Beschreibung

Summary

CUDA 13.1 extends the stream batch memory operations with atomic reductions: a new op type
CU_STREAM_MEM_OP_ATOMIC_REDUCTION with params struct CUstreamMemOpAtomicReductionParams
(operation, flags, reductionOp, dataType, address, value, alias), plus enums
CUstreamAtomicReductionOpType (ADD/AND/OR) and CUstreamAtomicReductionDataType
(UNSIGNED_32/UNSIGNED_64), joined into CUstreamBatchMemOpParams_union.atomicReduction.
Device support is advertised by CU_DEVICE_ATTRIBUTE_ATOMIC_REDUCTION_SUPPORTED (13.1, see
also #2361).

cuda.core has no batch mem-op surface at all today — the only mentions are a
graph debug-dot print flag
(https://github.com/NVIDIA/cuda-python/blob/c000331de6c37aa4565af74b001271ffcf6d5c99/cuda_core/cuda/core/graph/_graph_builder.pyx#L123)
and a DeviceProperties support query
(https://github.com/NVIDIA/cuda-python/blob/c000331de6c37aa4565af74b001271ffcf6d5c99/cuda_core/cuda/core/_device.pyx#L803).
So this issue covers designing the base surface (the pre-12.8 wait/write-value ops) and
including the 13.1 atomic-reduction op in it — adding the new op alone is not possible.

Underlying C APIs to cover

Symbol Purpose
cuStreamBatchMemOp(CUstream, unsigned int count, CUstreamBatchMemOpParams* paramArray, unsigned int flags) execute a batch of memory operations on a stream (pre-12.8, unexposed)
CU_STREAM_MEM_OP_ATOMIC_REDUCTION + CUstreamMemOpAtomicReductionParams (13.1) atomic-reduction op entry
CUstreamAtomicReductionOpType, CUstreamAtomicReductionDataType (13.1) reduction op / operand type
CU_DEVICE_ATTRIBUTE_ATOMIC_REDUCTION_SUPPORTED (13.1) support gating

Design sketch (draft — needs design-meeting review)

[!IMPORTANT]
Starting point only, not a settled design — review in the cuda.core design meeting.

stream.batch_mem_op([                          # names TBD
    WaitValue(addr, value, condition=">="),
    WriteValue(addr, value),
    AtomicReduction(address=addr, value=1, op="add", dtype="u64"),  # 13.1+
])

Open questions for the meeting:

  1. Scope: stream-level only, or also the graph batch-mem-op node type?
  2. Typed per-op dataclasses (above) vs. lower-level pass-through of binding structs.
  3. Semantics/purpose of the alias field (clarify against driver docs during design).
  4. Version/device gating (op availability differs across 12.x/13.x and devices).

References

-- Leo's bot

Beitragsleitfaden

Beitragsleitfaden öffnen

Erste Schritte

  1. Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
  3. Forke das Repository und arbeite in einem Branch.
  4. Öffne einen Pull Request, der die Issue-Nummer nennt.

Bewertung

Dieses Issue wurde noch nicht bewertet.

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.