NVIDIA / NVIDIA/cuda-python

[BUG]: SMResource.split() dry_run=True fails with cuDevSmResourceSplit when smCount is non-zero

Offen
#2,765 0 Kommentare 0 Reaktionen 2 zugewiesene Personen Auf GitHub ansehen

@leofang arbeitet bereits daran.

Seit 08.9.2026.

bug
Vorherrschende Sprache
Cython
Sterne
3.4k
Forks
329
Ø Merge
1 T. 23 Std.
Gemergte PRs (30 T.)
116

Beschreibung

[BUG]: SMResource.split() dry_run=True fails with cuDevSmResourceSplit when smCount is non-zero

Type of Bug

Runtime Error

Component

cuda.core

Describe the bug

SMResource.split(..., dry_run=True) always raises an exception on devices using the general cuDevSmResourceSplit API path (CUDA 13.x, _can_use_structured_sm_split()=True), silently causing the greenContext sample to fail.

Root cause: _split_with_general_api passes result=NULL to cuDevSmResourceSplit while groupParams[i].smCount is non-zero. Per the CUDA driver API docs, result=NULL is only valid in discovery mode (smCount=0). Passing result=NULL with non-zero smCount causes the driver to return CUDA_ERROR_INVALID_RESOURCE_CONFIGURATION.

How to Reproduce

Found while running the new cuda_core example samples against the CI GPU pool in PR #2266 - Migrating cuda-python samples from cuda-samples.

CI failure on H100 NVL (132 SMs, sm_90, CUDA 13.3):

[Green Context Sample using CUDA Core API]
Device: NVIDIA H100 NVL
Compute Capability: sm_90
Total SMs:                 132
Min. SM partition size:    8
SM co-scheduled alignment: 8
Error: could not find an SM split that the driver accepts on this device (total SMs=132, min_partition_size=8).
       The driver enforces architecture-specific alignment rules beyond min_partition_size; try passing an explicit --split.

Internally, _driver_accepts_split calls sm.split(SMResourceOptions(count=(112, 16)), dry_run=True), which hits _split_with_general_apicuDevSmResourceSplit(result=NULL, smCount=[112,16]) → exception swallowed by except Exception: return False → all split candidates return False_find_working_split returns Nonesys.exit(1).

System information

  • GPU: NVIDIA H100 NVL, sm_90, 132 SMs
  • CUDA: 13.3.0 (local)
  • Driver: 596.36 (kernel-mode)
  • Python: 3.14t (free-threaded)

Beitragsleitfaden

Beitragsleitfaden öffnen

Erste Schritte

  1. Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
  3. Forke das Repository und arbeite in einem Branch.
  4. Öffne einen Pull Request, der die Issue-Nummer nennt.

Bewertung

Dieses Issue wurde noch nicht bewertet.

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.