[PERF]: Epic for binding overhead improvements

未关闭
#1,645 0 条评论 0 个 reaction 已指派 1 人 在 GitHub 查看

@mdboom 已经在做这个了。

开始于 2026年2月18日。

评估

这个 Issue 还没有评估数据。

描述

cuda.bindings P2 performance

This issue is tracking performance improvements and investigations to Python-to-C binding overhead, mostly driven by the benchmark of cuTensorMapEncodeTiled devised in #659. That is a useful benchmark because it is a function with an unusually high number of arguments (and therefore unusually high Python-to-C overhead).

Comparison to a more limited Cython binding

As an interesting experimental datapoint, a colleague provided a vibe-coded Cython binding for cuTensorMapEncodeTiled that runs about 4x faster than cuda-bindings official one. It is useful to see where some overheads may be reduced, but care should be taken looking at its raw performance: this wrapper accepts far fewer things as inputs than the CUDA bindings, and doesn't include developer niceties, like enums.

Merged or in-progress fixes

Timings below are per-iteration of the benchmark in #659. This includes /both/ binding overhead and some fixed amount of time in the actual CUDA call.

  • 4.80us Baseline time
  • 3.63us #1543
  • 2.73us #1545
  • 2.70us #1581
  • 2.59us #1616
  • 2.38us #1638
  • (no change on this benchmark) #1644

Under investigation

Issues in this category are theoretical findings to reduce the operations required for type conversion, but haven't necessarily yet been confirmed to have a measurable effect.

  • #1639
  • #1640
  • #1642

Deferred (effective, but high effort)

  • #1643

Rejected (ineffective)

  • #1605
  • #1649
  • #1637
主要语言
Cython
星标
3.4k
派生
329
平均合并
1 天 21 小时
30 天内合并 PR
113

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

NVIDIA/cuda-python 的其他 Issue

查看 NVIDIA/cuda-python 的全部 Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。