Triple dispatching our way into a smaller interpreter
還沒有人認領這個 Issue。
- 主要語言
- Python
- 星號
- 77.2k
- 分支
- 36k
- PR 合併指標
- PR 指標待擷取
描述
Feature or enhancement
Proposal:
This is orthogonal to https://github.com/python/cpython/issues/148506. However, the problem remains the same. I'll repeat it again here: we have hit the limit of the computed goto/switch-case interpreter size such that we are seeing compiler bugs in all three: MSVC, Clang, and GCC. (I just fixed another compiler bug in Clang 22 due to interpreter size last week). The tail calling interpreter is the long-term solution. In the short term, we need another one. It is imperative that we reduce the size of the interpreter.
Instrumentation takes up 21 opcodes. This presents an opportunity for significant interpreter size reductions and recovery of opcodes. With this, we can remove all instrumented instructions. The INSTRUMENTED_X instructions become pseudo instructions with oparg > 255 . We store the instrumented pseudo-instruction in the instrumentation tools in get_tools_for_instruction instead of the bytecode, so that we can exceed the 255 limit.
To achieve instrumentation, we just swap out the interpreter dispatch table. The key observation is to repurpose the current TRACE_RECORD instruction to a VARIABLE_DISPATCH operation and change that single instruction to use call threading. This call threading instruction will support both JIT and instrumentation modes. We also move all instrumentation functions to small helper functions automatically using the cases generator. This is a small change, as we already refactored the interpreter to almost support call-threading due to the tail calling interpreter work, where each opcode is implemented as a function.
// For instrumentation
static void *instrumented_dispatch_table[256] = {
[FOR_ITER] = &&VARIABLE_DISPATCH,
}
static funcptr instrumented_targets_table[256] = {
[FOR_ITER] = &_CALL_FOR_ITER,
[INSTRUMENTED_FOR_ITER] = &_CALL_INSTRUMENTED_FOR_ITER,
}
// For JIT
static void *tracing_targets_table[256] = {
[FOR_ITER] = &&VARIABLE_DISPATCH,
}
inst(VARIABLE_DISPATCH, (--)) {
next_instr = this_instr;
if (dispatch_table_var == instrumented_dispatch_table) {
if (HAS_INSTRUMENTED_OPCODES[opcode]) {
opcode = get_instrumented_opcode(inst);
_CALL_ARGS = instrumented_targets_table[opcode](_CALL_ARGS);
DISPATCH();
}
else {
DISPATCH_NON_INSTRUMENTED();
}
}
else {
assert(dispatch_table_var == tracing_targets_table);
// Do what _TRACE_RECORD currently does in the JIT.
}
}
This will remove all instrumented opcodes from the main interpreter. Runtime instrumentation might be slower, but still faster than the old days of sys.settrace(). However instrumentation pauses will be significantly lower, as we won't have to scan bytecode to instrument, instead we just check the current tools and swap the dispatch table. Finally, the base interpreter should be faster too.
Has this already been discussed elsewhere?
No response given
Links to previous discussion of this feature:
No response
貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
研究方向
首先閱讀 TRACE_RECORD、VARIABLE_DISPATCH、get_tools_for_instruction 附近的直譯器分派邏輯,以及 proposal 中提到的 cases generator。追蹤 instrumentation 和 JIT 分派表目前如何運作。完成的標準是從主要直譯器中移除經過 instrumentation 的指令,同時保留 instrumentation 和 JIT 的行為,並減小直譯器的大小。
由索引模型根據 Issue 內容生成。
評估
- 技術堆疊
- c, python
- 領域
- compilers, performance
- Issue 類型
- 功能
- 難度
- 5/5
- 預估耗時
- 一週以上
- 活躍度
- 冷清
- 描述清晰度
- 基本清楚
- 新手友好度
- 35/100