[BUG]: SMResource.split() dry_run=True fails with cuDevSmResourceSplit when smCount is non-zero
@leofang がすでに取り組んでいます。
2026年9月8日 から。
評価
この issue はまだ評価されていません。
説明
[BUG]: SMResource.split() dry_run=True fails with cuDevSmResourceSplit when smCount is non-zero
Type of Bug
Runtime Error
Component
cuda.core
Describe the bug
SMResource.split(..., dry_run=True) always raises an exception on devices using the general cuDevSmResourceSplit API path (CUDA 13.x, _can_use_structured_sm_split()=True), silently causing the greenContext sample to fail.
Root cause: _split_with_general_api passes result=NULL to cuDevSmResourceSplit while groupParams[i].smCount is non-zero. Per the CUDA driver API docs, result=NULL is only valid in discovery mode (smCount=0). Passing result=NULL with non-zero smCount causes the driver to return CUDA_ERROR_INVALID_RESOURCE_CONFIGURATION.
How to Reproduce
Found while running the new cuda_core example samples against the CI GPU pool in PR #2266 - Migrating cuda-python samples from cuda-samples.
CI failure on H100 NVL (132 SMs, sm_90, CUDA 13.3):
[Green Context Sample using CUDA Core API]
Device: NVIDIA H100 NVL
Compute Capability: sm_90
Total SMs: 132
Min. SM partition size: 8
SM co-scheduled alignment: 8
Error: could not find an SM split that the driver accepts on this device (total SMs=132, min_partition_size=8).
The driver enforces architecture-specific alignment rules beyond min_partition_size; try passing an explicit --split.
Internally, _driver_accepts_split calls sm.split(SMResourceOptions(count=(112, 16)), dry_run=True), which hits _split_with_general_api → cuDevSmResourceSplit(result=NULL, smCount=[112,16]) → exception swallowed by except Exception: return False → all split candidates return False → _find_working_split returns None → sys.exit(1).
System information
- GPU: NVIDIA H100 NVL, sm_90, 132 SMs
- CUDA: 13.3.0 (local)
- Driver: 596.36 (kernel-mode)
- Python: 3.14t (free-threaded)
- 主要言語
- Cython
- スター
- 3.4k
- フォーク
- 329
- 平均マージ
- 1日 21時間
- マージ済み PR(30日)
- 113
コントリビューションガイド
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
NVIDIA/cuda-python のほかの issue
-
bug cuda.core
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
NVIDIA/cuda-python#2886 · コメント 1 件 ·
-
triage
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
NVIDIA/cuda-python#2717 ·
-
triage
難易度 1/5 1〜3時間 初心者へのやさしさ 90/100
NVIDIA/cuda-python#2712 ·
-
triage
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
NVIDIA/cuda-python#2646 · リアクション 1 件 ·
-
cuda.core triage
難易度 2/5 1〜3時間 初心者へのやさしさ 62/100
NVIDIA/cuda-python#2435 · コメント 1 件 ·