python / python/cpython

Broader specialization in the Specializing Adaptive Interpreter for better JIT performance

未關閉
#143,732 13 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

還沒有人認領這個 Issue。

3.15 interpreter-core performance topic-JIT
主要語言
Python
星號
77.2k
分支
36k
PR 合併指標
PR 指標待擷取

描述

Until now, our choice of specialization in the SAI has been driven by performance of the interpreter alone https://github.com/python/cpython/blob/main/InternalDocs/interpreter.md#performance-analysis.

However, we now expect any further performance improvements to be provided by the JIT, not the interpreter.
This means that specializations other role, that of gathering type and branching information for the JIT, is at least as important as pure interpreter performance.

We should therefore seek to broaden specialization to gather more information, as long as it does not make interpreter performance worse, or at least no significantly so.

Using some old stats, by fraction of unspecialized bytecode executed, the top 10 were:
BINARY_OP 31.3%
FOR_ITER 19.4%
LOAD_ATTR 10.9%
STORE_SUBSCR 9.2%
BINARY_SLICE 7.3%
COMPARE_OP 7.0%
TO_BOOL 5.8%
CALL 2.5%
CONTAINS_OP 2.4%
SEND 1.7%

We should fully specialize most, if not all, of these.

In general, the above instructions have a matching __dunder__ method which determines the behavior of the operation. Recording the type of the operand(s) allows us to know what __dunder__ method is to be called.

We cannot specialize for all possible types, but we can ensure we have good inputs and type information for the JIT by adding the following two specializations for all families of instructions:

  • __dunder__ implemented in Python. Most of the above instructions have a matching __dunder__ method. These specializations should jump directly into the method. LOAD_ATTR_GETATTRIBUTE_OVERRIDDEN already does this for LOAD_ATTR. Other families should follow this template.
  • __dunder__ implemented in C. In practice, this is just the generic instruction with a bit more information recorded.

Three instructions need special casing:

  • BINARY_OP. Because the behavior depends on two types, we will need a table driven approach: https://github.com/python/cpython/issues/100239
  • BINARY_SLICE. This is supposed to avoid creating temporary slice objects for expressions like a[b:c] but has yet to be implemented properly. There is no corresponding __dunder__ method, so we would need to expose slicing methods to use.
  • SEND. There is no __send__ method. For iterators, __next__ is called if the value is None, otherwise .send() is called. Rather than try to replicate the specializations of FOR_ITER we should maybe look to combine SEND and FOR_ITER much like we did for CALL and CALL_METHOD
First step

Add two specializations for __dunder__ in Python and the fallback __dunder__ in C for:

  • FOR_ITER
  • LOAD_ATTR
  • STORE_SUBSCR
  • COMPARE_OP
  • TO_BOOL
  • CALL
  • CONTAINS_OP

For a total of 12 new instructions as LOAD_ATTR already has the specialization for the Python __getattribute__ and CALL already has the generic fallback.

Second step

Implement https://github.com/python/cpython/issues/100239

Third step

Handle BINARY_SLICE and SEND

Linked PRs
  • gh-148113
  • gh-148128
  • gh-148271
  • gh-148745
  • gh-148963
  • gh-156033

貢獻指南

開啟貢獻指南

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

研究方向

從 InternalDocs/interpreter.md 的效能分析部分開始,並查看 issue 中列出的相關 PR,以了解已在進行的工作。建議的第一步是為 FOR_ITER、LOAD_ATTR、STORE_SUBSCR、COMPARE_OP、TO_BOOL、CALL 和 CONTAINS_OP 新增由 Python 和 C 實作的 dunder 特化;最終完成還包括後續的 BINARY_OP、BINARY_SLICE 和 SEND 工作。

由索引模型根據 Issue 內容生成。

評估

技術堆疊
python
領域
compilers, performance
Issue 類型
功能
難度
5/5
預估耗時
一週以上
活躍度
停滯
描述清晰度
基本清楚
新手友好度
25/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。