[PERF]: Epic for binding overhead improvements
@mdboom y travaille déjà.
Depuis le 18/2/2026.
- Langage dominant
- Cython
- Étoiles
- 3.4k
- Forks
- 329
- Merge moyen
- 1 j 21 h
- PR mergées (30 j)
- 113
Description
This issue is tracking performance improvements and investigations to Python-to-C binding overhead, mostly driven by the benchmark of cuTensorMapEncodeTiled devised in #659. That is a useful benchmark because it is a function with an unusually high number of arguments (and therefore unusually high Python-to-C overhead).
Comparison to a more limited Cython binding
As an interesting experimental datapoint, a colleague provided a vibe-coded Cython binding for cuTensorMapEncodeTiled that runs about 4x faster than cuda-bindings official one. It is useful to see where some overheads may be reduced, but care should be taken looking at its raw performance: this wrapper accepts far fewer things as inputs than the CUDA bindings, and doesn't include developer niceties, like enums.
Merged or in-progress fixes
Timings below are per-iteration of the benchmark in #659. This includes /both/ binding overhead and some fixed amount of time in the actual CUDA call.
- 4.80us Baseline time
- 3.63us #1543
- 2.73us #1545
- 2.70us #1581
- 2.59us #1616
- 2.38us #1638
- (no change on this benchmark) #1644
Under investigation
Issues in this category are theoretical findings to reduce the operations required for type conversion, but haven't necessarily yet been confirmed to have a measurable effect.
- #1639
- #1640
- #1642
Deferred (effective, but high effort)
- #1643
Rejected (ineffective)
- #1605
- #1649
- #1637
Guide de contribution
Ouvrir le guide de contribution
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Évaluation
Cette issue n'a pas encore été évaluée.