[PERF]: Measure the performance potential of caching pre-converted arguments at the call site
@mdboom がすでに取り組んでいます。
2026年2月27日 から。
評価
この issue はまだ評価されていません。
説明
I first came across the idea presented here in this paper: https://drops.dagstuhl.de/storage/00lipics/lipics-vol313-ecoop2024/LIPIcs.ECOOP.2024.6/LIPIcs.ECOOP.2024.6.pdf
The idea is that function calls across a language boundary spend a lot of time converting data structures from one form to another. Particulary in dynamic languages, there are many different code paths to handle many supported data types, but in practice, the types are quite stable between calls even in as the values change. The general idea is that once a particular call gets "hot", you replace a general function that handles all types with one that is specialized for specific types.
If you suppose the following times:
- T: The total time to convert any acceptable types in a dynamic and fully safe way
- H: The time to hash the types of the arguments
- F: The time to convert specific types, in a less safe way
In order for this to work, H + F should be significantly less than T, and the rate at which you see the same types at the same call site needs to be high enough that (T + H)x + (H + F)(1 - x) amortizes to less than T.
We can start by building a prototype to measure T, H, F, and x, and then decide whether it is worth the effort to proceed further.
The best solution is to add support for this "call site caching" to the CPython interpreter (though there are object lifetime issues that made that tricky). It is also possible to do this outside of the CPython interpreter by creating a call-site-to-cache mapping (at the expense of some performance).
- 主要言語
- Cython
- スター
- 3.4k
- フォーク
- 329
- 平均マージ
- 1日 21時間
- マージ済み PR(30日)
- 113
コントリビューションガイド
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
NVIDIA/cuda-python のほかの issue
-
bug cuda.core
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
NVIDIA/cuda-python#2886 · コメント 1 件 ·
-
triage
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
NVIDIA/cuda-python#2717 ·
-
triage
難易度 1/5 1〜3時間 初心者へのやさしさ 90/100
NVIDIA/cuda-python#2712 ·
-
triage
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
NVIDIA/cuda-python#2646 · リアクション 1 件 ·
-
cuda.core triage
難易度 2/5 1〜3時間 初心者へのやさしさ 62/100
NVIDIA/cuda-python#2435 · コメント 1 件 ·