NVIDIA / NVIDIA/cuda-python

[FEA] cuda.core: memory-pool location routing and new pool attributes (CUDA 13.0/13.2)

Open
#2,362 0 comments 0 reactions 1 assignee View on GitHub

@Andy-Jost is already working on this.

Since Jul 23, 2026.

cuda.core feature triage
Dominant language
Cython
Stars
3.4k
Forks
329
Avg merge
1d 23h
Merged PRs (30d)
116

Description

Summary

CUDA 13.0 makes memory-pool routing location-keyed: pools are addressed by
(CUmemLocation, CUmemAllocationType) — covering device pools, host / host-NUMA pools, and
managed pools (CU_MEM_ALLOCATION_TYPE_MANAGED, 13.0) under one surface:

  • cuMemGetDefaultMemPool(CUmemoryPool* pool_out, CUmemLocation* location, CUmemAllocationType type)
  • cuMemGetMemPool(CUmemoryPool* pool, CUmemLocation* location, CUmemAllocationType type) (current pool)
  • cuMemSetMemPool(CUmemLocation* location, CUmemAllocationType type, CUmemoryPool pool)

CUDA 13.2 adds six CUmemPool_attribute members: CU_MEMPOOL_ATTR_ALLOCATION_TYPE,
EXPORT_HANDLE_TYPES, HW_DECOMPRESS_ENABLED, LOCATION_ID, LOCATION_TYPE, MAX_POOL_SIZE.

Current cuda.core state:

Relates to the memory-resource architecture work (#209, #528, #726) — none of which tracks
these APIs.

Underlying C APIs to cover

cuMemGetDefaultMemPool, cuMemGetMemPool, cuMemSetMemPool; CUmemPool_attribute members
ALLOCATION_TYPE, EXPORT_HANDLE_TYPES, HW_DECOMPRESS_ENABLED, LOCATION_ID,
LOCATION_TYPE, MAX_POOL_SIZE; CU_MEM_LOCATION_TYPE_{NONE,INVISIBLE} mapping decision.

Design sketch (draft — needs design-meeting review)

[!IMPORTANT]
Starting point only, not a settled design — review in the cuda.core design meeting.

  • Read-only properties on DeviceMemoryResource (and the host/managed resources) for the new
    pool attributes; max_pool_size may warrant a setter — TBD.
  • Device.default_memory_resource (and a host-side equivalent) backed by
    cuMemGetDefaultMemPool, returning the corresponding memory-resource wrapper.
  • cuMemSetMemPool is a process-global state mutation — either expose it as an explicit,
    loudly-documented function, or deliberately keep it unexposed initially.

Open questions for the meeting:

  1. Expose the global setter at all (footgun) vs. get-only in the first pass?
  2. How (location, allocation type) keying maps onto cuda.core's Device/Host location
    objects (incl. host-NUMA IDs).
  3. Does HW_DECOMPRESS_ENABLED belong here or with a future decompress feature?
  4. Public StrEnum mapping for NONE/INVISIBLE location types.

References

-- Leo's bot

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.