[PERF]: Epic for binding overhead improvements

オープン
#1,645 コメント 0 件 リアクション 0 件 担当者 1 名 GitHub で見る

@mdboom がすでに取り組んでいます。

2026年2月18日 から。

評価

この issue はまだ評価されていません。

説明

cuda.bindings P2 performance

This issue is tracking performance improvements and investigations to Python-to-C binding overhead, mostly driven by the benchmark of cuTensorMapEncodeTiled devised in #659. That is a useful benchmark because it is a function with an unusually high number of arguments (and therefore unusually high Python-to-C overhead).

Comparison to a more limited Cython binding

As an interesting experimental datapoint, a colleague provided a vibe-coded Cython binding for cuTensorMapEncodeTiled that runs about 4x faster than cuda-bindings official one. It is useful to see where some overheads may be reduced, but care should be taken looking at its raw performance: this wrapper accepts far fewer things as inputs than the CUDA bindings, and doesn't include developer niceties, like enums.

Merged or in-progress fixes

Timings below are per-iteration of the benchmark in #659. This includes /both/ binding overhead and some fixed amount of time in the actual CUDA call.

  • 4.80us Baseline time
  • 3.63us #1543
  • 2.73us #1545
  • 2.70us #1581
  • 2.59us #1616
  • 2.38us #1638
  • (no change on this benchmark) #1644

Under investigation

Issues in this category are theoretical findings to reduce the operations required for type conversion, but haven't necessarily yet been confirmed to have a measurable effect.

  • #1639
  • #1640
  • #1642

Deferred (effective, but high effort)

  • #1643

Rejected (ineffective)

  • #1605
  • #1649
  • #1637
主要言語
Cython
スター
3.4k
フォーク
329
平均マージ
1日 21時間
マージ済み PR(30日)
113

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

NVIDIA/cuda-python のほかの issue

NVIDIA/cuda-python の issue をすべて見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。