NVIDIA / NVIDIA/cuda-python

[PERF]: Measure the performance potential of caching pre-converted arguments at the call site

Ouverte
#1,655 0 commentaires 1 réaction 1 personne assignée Voir sur GitHub

@mdboom y travaille déjà.

Depuis le 27/2/2026.

cuda.bindings experiment performance triage
Langage dominant
Cython
Étoiles
3.4k
Forks
329
Merge moyen
1 j 23 h
PR mergées (30 j)
116

Description

I first came across the idea presented here in this paper: https://drops.dagstuhl.de/storage/00lipics/lipics-vol313-ecoop2024/LIPIcs.ECOOP.2024.6/LIPIcs.ECOOP.2024.6.pdf

The idea is that function calls across a language boundary spend a lot of time converting data structures from one form to another. Particulary in dynamic languages, there are many different code paths to handle many supported data types, but in practice, the types are quite stable between calls even in as the values change. The general idea is that once a particular call gets "hot", you replace a general function that handles all types with one that is specialized for specific types.

If you suppose the following times:

  • T: The total time to convert any acceptable types in a dynamic and fully safe way
  • H: The time to hash the types of the arguments
  • F: The time to convert specific types, in a less safe way

In order for this to work, H + F should be significantly less than T, and the rate at which you see the same types at the same call site needs to be high enough that (T + H)x + (H + F)(1 - x) amortizes to less than T.

We can start by building a prototype to measure T, H, F, and x, and then decide whether it is worth the effort to proceed further.

The best solution is to add support for this "call site caching" to the CPython interpreter (though there are object lifetime issues that made that tricky). It is also possible to do this outside of the CPython interpreter by creating a call-site-to-cache mapping (at the expense of some performance).

Guide de contribution

Ouvrir le guide de contribution

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Évaluation

Cette issue n'a pas encore été évaluée.

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.