cuda.core: support "3D copies with attributes" (follow-on to issue #2365)
Nobody has claimed this yet.
- Dominant language
- Cython
- Stars
- 3.4k
- Forks
- 329
- Avg merge
- 1d 23h
- Merged PRs (30d)
- 116
Description
Summary
#2365 tracked two CUDA 13.2 attribute-carrying copy entry points together. #2636 implemented the flat 1D case (cuMemcpyWithAttributesAsync, exposed via Buffer.copy_to/Buffer.copy_from's new options keyword). This issue splits off the remaining, unimplemented half:
cuMemcpy3DWithAttributesAsync(CUDA_MEMCPY3D_BATCH_OP* op, unsigned long long flags, CUstream)— executes a single 3D copy operation described by aCUDA_MEMCPY3D_BATCH_OP, reusing theCUmemcpySrcAccessOrder/CUmemcpyFlagsmachinery already introduced forcuMemcpyBatchAsync(12.8) and consumed by#2636'sCopyOptions.
cuda.core has no API for 3D/strided copies with attributes today.
Underlying C API
cuMemcpy3DWithAttributesAsync (CUDA 13.2). Takes a single CUDA_MEMCPY3D_BATCH_OP:
src,dst:CUmemcpy3DOperand, each either a raw pointer operand (base pointer + row/depth pitch, matching a strided/3D buffer layout) or a CUDA array operand (CUarray/CUmipmappedArray+ subresource).extent:CUextent3D(width/height/depth), all three components must be nonzero.srcAccessOrder:CUmemcpySrcAccessOrder— the same three-way hint (STREAM/DURING_API_CALL/ANY) as the 1D and batched entry points.flags:CUmemcpyFlags(e.g.CU_MEMCPY_FLAG_PREFER_OVERLAP_WITH_COMPUTE), same as the other attribute-carrying copies.
Relation to #2636
Should reuse as much of #2636's plumbing as possible rather than duplicating it:
MemcpySrcAccessOrder/MemcpyOverlapMode(cuda.core._memory._copy_enums) — same enums, no new ones needed for those two fields.- The
DURING_API_CALLfallback hazard and its guard,_reject_unsupported_during_api_call— the pointer-operand fallback (if any) would have the same correctness hazard as the 1D/batched paths and needs the same treatment, not a silently-downgraded copy. - The
_with_attributes_available()-style CUDA 13.2 version gate, and the graph-capture /LEGACY_DEFAULT_STREAMrejection pattern established there.
Design sketch (draft — needs design-meeting review)
[!IMPORTANT]
Starting point only, not a settled design.
cuda.core already has StridedMemoryView with copy_to/copy_from, and separately texture.Array/texture.MipmappedArray for CUDA-array-backed storage. The pointer-operand case maps naturally onto a strided/3D buffer view; the array-operand case maps onto the existing texture array types. Open questions for the meeting:
- Does this become
StridedMemoryView.copy_to/copy_fromgaining anoptionskeyword (pointer operands), a new API for the array-operand case, or both under one entry point? - How do row/depth pitch and extent get derived: from
StridedMemoryView's shape/strides directly, or does the caller supply them explicitly? - Should mixed operand kinds be supported in one call (e.g. array → pointer), or should the initial version only cover the common pointer-to-pointer case?
- Version gating (CUDA 13.2+) and error behavior, matching the pattern established in #2636.
- Separate PR from any future
cuMemcpy3DBatchAsync(multiple 3D ops per call) work, if that ever gets tracked.
References
- Driver docs: https://docs.nvidia.com/cuda/cuda-driver-api/
- Parent issue: #2365
- Implements the 1D half: #2636
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reading the implementation from #2636 and the existing _memory._copy_enums plumbing, then review StridedMemoryView and texture.Array/MipmappedArray. Resolve the design questions for supported operand types, pitch and extent derivation, and version/error behavior before implementation. Done means a settled API for 3D attribute-carrying copies with matching CUDA 13.2 gating and safety checks.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- backend-api-design
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100