NVIDIA / NVIDIA/cuda-python

[FEA] cuda.core: support stream batch memory operations (incl. CUDA 13.1 atomic reductions)

Ouverte
#2,364 0 commentaires 0 réactions 1 personne assignée Voir sur GitHub

@juenglin y travaille déjà.

Depuis le 24/7/2026.

cuda.core feature
Langage dominant
Cython
Étoiles
3.4k
Forks
329
Merge moyen
1 j 21 h
PR mergées (30 j)
113

Description

Summary

CUDA 13.1 extends the stream batch memory operations with atomic reductions: a new op type
CU_STREAM_MEM_OP_ATOMIC_REDUCTION with params struct CUstreamMemOpAtomicReductionParams
(operation, flags, reductionOp, dataType, address, value, alias), plus enums
CUstreamAtomicReductionOpType (ADD/AND/OR) and CUstreamAtomicReductionDataType
(UNSIGNED_32/UNSIGNED_64), joined into CUstreamBatchMemOpParams_union.atomicReduction.
Device support is advertised by CU_DEVICE_ATTRIBUTE_ATOMIC_REDUCTION_SUPPORTED (13.1, see
also #2361).

cuda.core has no batch mem-op surface at all today — the only mentions are a
graph debug-dot print flag
(https://github.com/NVIDIA/cuda-python/blob/c000331de6c37aa4565af74b001271ffcf6d5c99/cuda_core/cuda/core/graph/_graph_builder.pyx#L123)
and a DeviceProperties support query
(https://github.com/NVIDIA/cuda-python/blob/c000331de6c37aa4565af74b001271ffcf6d5c99/cuda_core/cuda/core/_device.pyx#L803).
So this issue covers designing the base surface (the pre-12.8 wait/write-value ops) and
including the 13.1 atomic-reduction op in it — adding the new op alone is not possible.

Underlying C APIs to cover

Symbol Purpose
cuStreamBatchMemOp(CUstream, unsigned int count, CUstreamBatchMemOpParams* paramArray, unsigned int flags) execute a batch of memory operations on a stream (pre-12.8, unexposed)
CU_STREAM_MEM_OP_ATOMIC_REDUCTION + CUstreamMemOpAtomicReductionParams (13.1) atomic-reduction op entry
CUstreamAtomicReductionOpType, CUstreamAtomicReductionDataType (13.1) reduction op / operand type
CU_DEVICE_ATTRIBUTE_ATOMIC_REDUCTION_SUPPORTED (13.1) support gating

Design sketch (draft — needs design-meeting review)

[!IMPORTANT]
Starting point only, not a settled design — review in the cuda.core design meeting.

stream.batch_mem_op([                          # names TBD
    WaitValue(addr, value, condition=">="),
    WriteValue(addr, value),
    AtomicReduction(address=addr, value=1, op="add", dtype="u64"),  # 13.1+
])

Open questions for the meeting:

  1. Scope: stream-level only, or also the graph batch-mem-op node type?
  2. Typed per-op dataclasses (above) vs. lower-level pass-through of binding structs.
  3. Semantics/purpose of the alias field (clarify against driver docs during design).
  4. Version/device gating (op availability differs across 12.x/13.x and devices).

References

-- Leo's bot

Guide de contribution

Ouvrir le guide de contribution

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Évaluation

Cette issue n'a pas encore été évaluée.

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.