Improve macOS AArch64 JIT stencils by relaxing GOT loads
還沒有人認領這個 Issue。
- 主要語言
- Python
- 星號
- 77.2k
- 分支
- 36k
- PR 合併指標
- PR 指標待擷取
描述
Proposal:
The AArch64 JIT has support for optimising pairs of instructions that load an address through the global offset table (GOT).
On Linux, the JIT optimiser recognises assembly like:
adrp x9, :got:__PyRuntime
ldr x9, [x9, :got_lo12:__PyRuntime]
On macOS, Clang uses a different syntax recognised by the Mach-O linker:
adrp x9, __PyRuntime@GOTPAGE
ldr x9, [x9, __PyRuntime@GOTPAGEOFF]
We can add support for this using the existing patch_aarch64_33rx() function just by adjusting the regex which would enable the same optimisation we already use on Linux:
adrp x9, <__PyRuntime page>
add x9, x9, <page offset>
This is a natural followup to #157040. I have a small patch ready once jit development is unblocked.
Benchmark:
On my M5 MacBook Pro this shows a 1.16% improvement on the pyperformance suite though results are very noisy. Probably best to re-run with the official benchmarking runners.
Expand for full benchmarking results
Has this already been discussed elsewhere?
This is a minor feature, which does not need previous discussion elsewhere
Links to previous discussion of this feature:
#148598, #148501, #157040 /cc @diegorusso
貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
研究方向
從 Tools/jit/_optimizers.py 中的 patch_aarch64_33rx() 開始,將現有的 Linux GOT-load 模式與此處描述的 Clang macOS 語法進行比較。更新比對邏輯,使現有的 AArch64 最佳化能夠套用,然後執行 pyperformance 套件或官方 benchmarking runner 以檢查結果。
由索引模型根據 Issue 內容生成。
評估
- 技術堆疊
- macos, python
- 領域
- compilers, performance
- Issue 類型
- 功能
- 難度
- 3/5
- 預估耗時
- 1-2 天
- 活躍度
- 活躍
- 描述清晰度
- 描述清楚
- 新手友好度
- 72/100