python / python/cpython

`unpack_sequence` benchmark runs slower under JIT

オープン
#149,212 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

performance topic-JIT
主要言語
Python
スター
77.2k
フォーク
35.9k
PR マージ指標
PR 指標を取得中

説明

Bug report

Bug description:

The pyperformance unpack_sequence benchmark is the worst-performing JIT benchmark on https://www.doesjitgobrrr.com — geomean speedup 0.586× (JIT ~1.7× slower).

The benchmark is a single function with 400 inlined a,b,c,d,e,f,g,h,i,j = to_unpack statements inside a for loop. to_unpack = tuple(range(10)) is reused (refcount > 1, so _UNPACK_SEQUENCE_UNIQUE_TUPLE never fires).

Analysis of the slowdown

The JIT covers all 400 unpacks via ~22 sequential traces (~775 uops, ~19 unpacks each), linked tail-to-tail through _EXIT_TRACE → _START_EXECUTOR, totalling ~18 MB of JIT code. This is a consequence of trace length / fitness limits ([UOP_MAX_TRACE_LENGTH] https://github.com/python/cpython/blob/main/Include/internal/pycore_uop.h#L42), EXIT_QUALITY_* from gh-146073) and is arguably fine - 400 inline statements isn't realistic code. Inter-trace transition cost alone would not produce a 1.7× slowdown. Nevertheless, longer traces would help here.

Per unpack: ~33 uops vs tier 1's ~12 bytecodes — ~2.7× more dispatches. The trace recorder unconditionally emits a _CHECK_VALIDITY + _SET_IP pair before every source bytecode at Python/optimizer.c:902-906. Per unpack:

LOAD_FAST_BORROW t            (1)
UNPACK_SEQUENCE 10            (1)
STORE_FAST_STORE_FAST × 5     (5; pair-fused by the compiler)
                              ── 7 source bytecodes → 6 _CHECK_VALIDITY+_SET_IP pairs steady state

Tier 1 has no analog. Each _CHECK_VALIDITY issues a load+branch on current_executor->vm_data.valid; each _SET_IP issues a store to frame->instr_ptr. 12 such uops × 400 unpacks × 20000 iterations is the bulk of the regression.

A naive elimination pass cannot drop them because every gap contains at least one uop with HAS_ESCAPES_FLAG — typically _POP_TOP, conservatively flagged as escaping (its Py_DECREF could run __del__).

CPython versions tested on:

CPython main branch

Operating systems tested on:

Linux

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

調査の方向性

まず、pyperformance unpack_sequence の結果を再現し、Python/optimizer.c の 902-906 行付近にあるトレースレコーダーのロジックを読みます。Include/internal/pycore_uop.h の UOP_MAX_TRACE_LENGTH を確認し、レポートで説明されている生成されたトレースと妥当性チェックを調べます。この issue では具体的な修正方法も受け入れテストも定義されていないため、コードを変更する前に、意図された最適化とベンチマークベースの成功基準を確認してください。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
python
領域
compilers, performance
issue の種類
バグ
難易度
4/5
見積もり時間
3〜5日
活発さ
静か
明瞭さ
説明が足りない
初心者へのやさしさ
42/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。