cuda.core: support "3D copies with attributes" (follow-on to issue #2365)
Nessuno ha ancora preso questa issue.
- Lingua principale
- Cython
- Stelle
- 3.4k
- Fork
- 329
- Merge medio
- 1g 21h
- PR unite (30g)
- 113
Descrizione
Summary
#2365 tracked two CUDA 13.2 attribute-carrying copy entry points together. #2636 implemented the flat 1D case (cuMemcpyWithAttributesAsync, exposed via Buffer.copy_to/Buffer.copy_from's new options keyword). This issue splits off the remaining, unimplemented half:
cuMemcpy3DWithAttributesAsync(CUDA_MEMCPY3D_BATCH_OP* op, unsigned long long flags, CUstream)— executes a single 3D copy operation described by aCUDA_MEMCPY3D_BATCH_OP, reusing theCUmemcpySrcAccessOrder/CUmemcpyFlagsmachinery already introduced forcuMemcpyBatchAsync(12.8) and consumed by#2636'sCopyOptions.
cuda.core has no API for 3D/strided copies with attributes today.
Underlying C API
cuMemcpy3DWithAttributesAsync (CUDA 13.2). Takes a single CUDA_MEMCPY3D_BATCH_OP:
src,dst:CUmemcpy3DOperand, each either a raw pointer operand (base pointer + row/depth pitch, matching a strided/3D buffer layout) or a CUDA array operand (CUarray/CUmipmappedArray+ subresource).extent:CUextent3D(width/height/depth), all three components must be nonzero.srcAccessOrder:CUmemcpySrcAccessOrder— the same three-way hint (STREAM/DURING_API_CALL/ANY) as the 1D and batched entry points.flags:CUmemcpyFlags(e.g.CU_MEMCPY_FLAG_PREFER_OVERLAP_WITH_COMPUTE), same as the other attribute-carrying copies.
Relation to #2636
Should reuse as much of #2636's plumbing as possible rather than duplicating it:
MemcpySrcAccessOrder/MemcpyOverlapMode(cuda.core._memory._copy_enums) — same enums, no new ones needed for those two fields.- The
DURING_API_CALLfallback hazard and its guard,_reject_unsupported_during_api_call— the pointer-operand fallback (if any) would have the same correctness hazard as the 1D/batched paths and needs the same treatment, not a silently-downgraded copy. - The
_with_attributes_available()-style CUDA 13.2 version gate, and the graph-capture /LEGACY_DEFAULT_STREAMrejection pattern established there.
Design sketch (draft — needs design-meeting review)
[!IMPORTANT]
Starting point only, not a settled design.
cuda.core already has StridedMemoryView with copy_to/copy_from, and separately texture.Array/texture.MipmappedArray for CUDA-array-backed storage. The pointer-operand case maps naturally onto a strided/3D buffer view; the array-operand case maps onto the existing texture array types. Open questions for the meeting:
- Does this become
StridedMemoryView.copy_to/copy_fromgaining anoptionskeyword (pointer operands), a new API for the array-operand case, or both under one entry point? - How do row/depth pitch and extent get derived: from
StridedMemoryView's shape/strides directly, or does the caller supply them explicitly? - Should mixed operand kinds be supported in one call (e.g. array → pointer), or should the initial version only cover the common pointer-to-pointer case?
- Version gating (CUDA 13.2+) and error behavior, matching the pattern established in #2636.
- Separate PR from any future
cuMemcpy3DBatchAsync(multiple 3D ops per call) work, if that ever gets tracked.
References
- Driver docs: https://docs.nvidia.com/cuda/cuda-driver-api/
- Parent issue: #2365
- Implements the 1D half: #2636
Guida per i contributori
Apri la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Direzione di ricerca
Inizia leggendo l’implementazione di #2636 e l’infrastruttura _memory.copy_enums esistente, quindi esamina StridedMemoryView e texture.Array/MipmappedArray. Risolvi prima dell’implementazione le questioni di progettazione relative ai tipi di operandi supportati, alla derivazione di pitch ed extent e al comportamento in base alla versione e in caso di errore. Il lavoro è completato quando è definita un’API per copie 3D con attributi, con il gating e i controlli di sicurezza corrispondenti di CUDA 13.2.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- python
- Ambito
- backend-api-design
- Tipo di issue
- Funzionalità
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Stato di attività
- Tranquilla
- Chiarezza
- Da chiarire
- Idoneità per principianti
- 25/100