Broader specialization in the Specializing Adaptive Interpreter for better JIT performance
Chưa có ai nhận issue này.
- Ngôn ngữ chính
- Python
- Star
- 77.2k
- Fork
- 35.9k
- Chỉ số merge pull request
- Chỉ số pull request đang chờ
Mô tả
Until now, our choice of specialization in the SAI has been driven by performance of the interpreter alone https://github.com/python/cpython/blob/main/InternalDocs/interpreter.md#performance-analysis.
However, we now expect any further performance improvements to be provided by the JIT, not the interpreter.
This means that specializations other role, that of gathering type and branching information for the JIT, is at least as important as pure interpreter performance.
We should therefore seek to broaden specialization to gather more information, as long as it does not make interpreter performance worse, or at least no significantly so.
Using some old stats, by fraction of unspecialized bytecode executed, the top 10 were:
BINARY_OP 31.3%
FOR_ITER 19.4%
LOAD_ATTR 10.9%
STORE_SUBSCR 9.2%
BINARY_SLICE 7.3%
COMPARE_OP 7.0%
TO_BOOL 5.8%
CALL 2.5%
CONTAINS_OP 2.4%
SEND 1.7%
We should fully specialize most, if not all, of these.
In general, the above instructions have a matching __dunder__ method which determines the behavior of the operation. Recording the type of the operand(s) allows us to know what __dunder__ method is to be called.
We cannot specialize for all possible types, but we can ensure we have good inputs and type information for the JIT by adding the following two specializations for all families of instructions:
__dunder__implemented in Python. Most of the above instructions have a matching__dunder__method. These specializations should jump directly into the method.LOAD_ATTR_GETATTRIBUTE_OVERRIDDENalready does this forLOAD_ATTR. Other families should follow this template.__dunder__implemented in C. In practice, this is just the generic instruction with a bit more information recorded.
Three instructions need special casing:
- BINARY_OP. Because the behavior depends on two types, we will need a table driven approach: https://github.com/python/cpython/issues/100239
- BINARY_SLICE. This is supposed to avoid creating temporary slice objects for expressions like
a[b:c]but has yet to be implemented properly. There is no corresponding__dunder__method, so we would need to expose slicing methods to use. - SEND. There is no
__send__method. For iterators,__next__is called if the value isNone, otherwise.send()is called. Rather than try to replicate the specializations ofFOR_ITERwe should maybe look to combineSENDandFOR_ITERmuch like we did forCALLandCALL_METHOD
First step
Add two specializations for __dunder__ in Python and the fallback __dunder__ in C for:
- FOR_ITER
- LOAD_ATTR
- STORE_SUBSCR
- COMPARE_OP
- TO_BOOL
- CALL
- CONTAINS_OP
For a total of 12 new instructions as LOAD_ATTR already has the specialization for the Python __getattribute__ and CALL already has the generic fallback.
Second step
Implement https://github.com/python/cpython/issues/100239
Third step
Handle BINARY_SLICE and SEND
Linked PRs
- gh-148113
- gh-148128
- gh-148271
- gh-148745
- gh-148963
- gh-156033
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Hướng nghiên cứu
Inizia dalla sezione sull’analisi delle prestazioni di InternalDocs/interpreter.md e rivedi le PR collegate elencate nell’issue per comprendere il lavoro già in corso. Il primo passo proposto è aggiungere specializzazioni dunder implementate in Python e C per FOR_ITER, LOAD_ATTR, STORE_SUBSCR, COMPARE_OP, TO_BOOL, CALL e CONTAINS_OP; il completamento include infine il lavoro successivo su BINARY_OP, BINARY_SLICE e SEND.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Đánh giá
- Công nghệ
- python
- Lĩnh vực
- compilers, performance
- Loại issue
- Tính năng
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức độ hoạt động
- Đình trệ
- Độ rõ ràng
- Khá rõ ràng
- Mức phù hợp với người mới
- 25/100