[PERF]: Measure the performance potential of caching pre-converted arguments at the call site

Offen
#1,655 0 Kommentare 1 Reaktion 1 zugewiesene Person Auf GitHub ansehen

@mdboom arbeitet bereits daran.

Seit 27.2.2026.

Bewertung

Dieses Issue wurde noch nicht bewertet.

Beschreibung

cuda.bindings experiment performance triage

I first came across the idea presented here in this paper: https://drops.dagstuhl.de/storage/00lipics/lipics-vol313-ecoop2024/LIPIcs.ECOOP.2024.6/LIPIcs.ECOOP.2024.6.pdf

The idea is that function calls across a language boundary spend a lot of time converting data structures from one form to another. Particulary in dynamic languages, there are many different code paths to handle many supported data types, but in practice, the types are quite stable between calls even in as the values change. The general idea is that once a particular call gets "hot", you replace a general function that handles all types with one that is specialized for specific types.

If you suppose the following times:

  • T: The total time to convert any acceptable types in a dynamic and fully safe way
  • H: The time to hash the types of the arguments
  • F: The time to convert specific types, in a less safe way

In order for this to work, H + F should be significantly less than T, and the rate at which you see the same types at the same call site needs to be high enough that (T + H)x + (H + F)(1 - x) amortizes to less than T.

We can start by building a prototype to measure T, H, F, and x, and then decide whether it is worth the effort to proceed further.

The best solution is to add support for this "call site caching" to the CPython interpreter (though there are object lifetime issues that made that tricky). It is also possible to do this outside of the CPython interpreter by creating a call-site-to-cache mapping (at the expense of some performance).

Vorherrschende Sprache
Cython
Sterne
3.4k
Forks
329
Ø Merge
1 T. 21 Std.
Gemergte PRs (30 T.)
113

Beitragsleitfaden

Beitragsleitfaden öffnen

Erste Schritte

  1. Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
  3. Forken Sie das Repository und arbeiten Sie in einem Branch.
  4. Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.

Mehr aus NVIDIA/cuda-python

Alle Issues in NVIDIA/cuda-python

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.