python / python/cpython

Triple dispatching our way into a smaller interpreter

Đang mở
#148,543 1 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

interpreter-core type-feature
Ngôn ngữ chính
Python
Star
77.2k
Fork
35.9k
Chỉ số merge pull request
Chỉ số pull request đang chờ

Mô tả

Feature or enhancement

Proposal:

This is orthogonal to https://github.com/python/cpython/issues/148506. However, the problem remains the same. I'll repeat it again here: we have hit the limit of the computed goto/switch-case interpreter size such that we are seeing compiler bugs in all three: MSVC, Clang, and GCC. (I just fixed another compiler bug in Clang 22 due to interpreter size last week). The tail calling interpreter is the long-term solution. In the short term, we need another one. It is imperative that we reduce the size of the interpreter.

Instrumentation takes up 21 opcodes. This presents an opportunity for significant interpreter size reductions and recovery of opcodes. With this, we can remove all instrumented instructions. The INSTRUMENTED_X instructions become pseudo instructions with oparg > 255 . We store the instrumented pseudo-instruction in the instrumentation tools in get_tools_for_instruction instead of the bytecode, so that we can exceed the 255 limit.

To achieve instrumentation, we just swap out the interpreter dispatch table. The key observation is to repurpose the current TRACE_RECORD instruction to a VARIABLE_DISPATCH operation and change that single instruction to use call threading. This call threading instruction will support both JIT and instrumentation modes. We also move all instrumentation functions to small helper functions automatically using the cases generator. This is a small change, as we already refactored the interpreter to almost support call-threading due to the tail calling interpreter work, where each opcode is implemented as a function.

// For instrumentation
static void *instrumented_dispatch_table[256] = {
    [FOR_ITER] = &&VARIABLE_DISPATCH,
}

static funcptr instrumented_targets_table[256] = {
    [FOR_ITER] = &_CALL_FOR_ITER,
    [INSTRUMENTED_FOR_ITER] = &_CALL_INSTRUMENTED_FOR_ITER,
}

// For JIT
static void *tracing_targets_table[256] = {
    [FOR_ITER] = &&VARIABLE_DISPATCH,
}

inst(VARIABLE_DISPATCH, (--)) {
    next_instr = this_instr;
    if (dispatch_table_var == instrumented_dispatch_table) {
        if (HAS_INSTRUMENTED_OPCODES[opcode]) {
            opcode = get_instrumented_opcode(inst);
            _CALL_ARGS = instrumented_targets_table[opcode](_CALL_ARGS);
            DISPATCH();
        }
        else {
            DISPATCH_NON_INSTRUMENTED();
         }
    }
    else {
        assert(dispatch_table_var == tracing_targets_table);
        // Do what _TRACE_RECORD currently does in the JIT.
    }
}

This will remove all instrumented opcodes from the main interpreter. Runtime instrumentation might be slower, but still faster than the old days of sys.settrace(). However instrumentation pauses will be significantly lower, as we won't have to scan bytecode to instrument, instead we just check the current tools and swap the dispatch table. Finally, the base interpreter should be faster too.

Has this already been discussed elsewhere?

No response given

Links to previous discussion of this feature:

No response

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Hướng nghiên cứu

Bắt đầu bằng cách đọc logic dispatch của interpreter xung quanh TRACE_RECORD, VARIABLE_DISPATCH, get_tools_for_instruction và cases generator được đề cập trong proposal. Theo dõi cách instrumentation và các bảng dispatch của JIT hiện đang hoạt động. Công việc được xem là hoàn tất khi loại bỏ các instruction đã được instrument khỏi interpreter chính, đồng thời giữ nguyên hành vi của instrumentation và JIT và giảm kích thước của interpreter.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
c, python
Lĩnh vực
compilers, performance
Loại issue
Tính năng
Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức độ hoạt động
Ít trao đổi
Độ rõ ràng
Khá rõ ràng
Mức phù hợp với người mới
35/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.