NVIDIA / NVIDIA/cuda-python

cuda.core: support "3D copies with attributes" (follow-on to issue #2365)

Aperta
#2,660 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

cuda.core triage
Lingua principale
Cython
Stelle
3.4k
Fork
329
Merge medio
1g 21h
PR unite (30g)
113

Descrizione

Summary

#2365 tracked two CUDA 13.2 attribute-carrying copy entry points together. #2636 implemented the flat 1D case (cuMemcpyWithAttributesAsync, exposed via Buffer.copy_to/Buffer.copy_from's new options keyword). This issue splits off the remaining, unimplemented half:

  • cuMemcpy3DWithAttributesAsync(CUDA_MEMCPY3D_BATCH_OP* op, unsigned long long flags, CUstream) — executes a single 3D copy operation described by a CUDA_MEMCPY3D_BATCH_OP, reusing the CUmemcpySrcAccessOrder/CUmemcpyFlags machinery already introduced for cuMemcpyBatchAsync (12.8) and consumed by #2636's CopyOptions.

cuda.core has no API for 3D/strided copies with attributes today.

Underlying C API

cuMemcpy3DWithAttributesAsync (CUDA 13.2). Takes a single CUDA_MEMCPY3D_BATCH_OP:

  • src, dst: CUmemcpy3DOperand, each either a raw pointer operand (base pointer + row/depth pitch, matching a strided/3D buffer layout) or a CUDA array operand (CUarray/CUmipmappedArray + subresource).
  • extent: CUextent3D (width/height/depth), all three components must be nonzero.
  • srcAccessOrder: CUmemcpySrcAccessOrder — the same three-way hint (STREAM/DURING_API_CALL/ANY) as the 1D and batched entry points.
  • flags: CUmemcpyFlags (e.g. CU_MEMCPY_FLAG_PREFER_OVERLAP_WITH_COMPUTE), same as the other attribute-carrying copies.

Relation to #2636

Should reuse as much of #2636's plumbing as possible rather than duplicating it:

  • MemcpySrcAccessOrder / MemcpyOverlapMode (cuda.core._memory._copy_enums) — same enums, no new ones needed for those two fields.
  • The DURING_API_CALL fallback hazard and its guard, _reject_unsupported_during_api_call — the pointer-operand fallback (if any) would have the same correctness hazard as the 1D/batched paths and needs the same treatment, not a silently-downgraded copy.
  • The _with_attributes_available()-style CUDA 13.2 version gate, and the graph-capture / LEGACY_DEFAULT_STREAM rejection pattern established there.

Design sketch (draft — needs design-meeting review)

[!IMPORTANT]
Starting point only, not a settled design.

cuda.core already has StridedMemoryView with copy_to/copy_from, and separately texture.Array/texture.MipmappedArray for CUDA-array-backed storage. The pointer-operand case maps naturally onto a strided/3D buffer view; the array-operand case maps onto the existing texture array types. Open questions for the meeting:

  1. Does this become StridedMemoryView.copy_to/copy_from gaining an options keyword (pointer operands), a new API for the array-operand case, or both under one entry point?
  2. How do row/depth pitch and extent get derived: from StridedMemoryView's shape/strides directly, or does the caller supply them explicitly?
  3. Should mixed operand kinds be supported in one call (e.g. array → pointer), or should the initial version only cover the common pointer-to-pointer case?
  4. Version gating (CUDA 13.2+) and error behavior, matching the pattern established in #2636.
  5. Separate PR from any future cuMemcpy3DBatchAsync (multiple 3D ops per call) work, if that ever gets tracked.

References

Guida per i contributori

Apri la guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Direzione di ricerca

Inizia leggendo l’implementazione di #2636 e l’infrastruttura _memory.copy_enums esistente, quindi esamina StridedMemoryView e texture.Array/MipmappedArray. Risolvi prima dell’implementazione le questioni di progettazione relative ai tipi di operandi supportati, alla derivazione di pitch ed extent e al comportamento in base alla versione e in caso di errore. Il lavoro è completato quando è definita un’API per copie 3D con attributi, con il gating e i controlli di sicurezza corrispondenti di CUDA 13.2.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
python
Ambito
backend-api-design
Tipo di issue
Funzionalità
Difficoltà
5/5
Tempo stimato
Più di una settimana
Stato di attività
Tranquilla
Chiarezza
Da chiarire
Idoneità per principianti
25/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.