python / python/cpython

Using the unit tests as the PGO task has problems

オープン
#130,701 コメント 0 件 リアクション 2 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

build performance type-feature
主要言語
Python
スター
77.2k
フォーク
35.9k
PR マージ指標
PR 指標を取得中

説明

Feature or enhancement

Proposal:

When Python is compiled with --enabled-optimizations, which turns on PGO (program guided optimizations), the build will run a subset of the unit tests as the "task" to generate profile information. Included in the profile is information like counts of the taken side of a CPU conditional branch instructions. To get the best optimization, your PGO task should match the branch taken behavior of your real workloads.

Using the unit tests has the advantage that we have good code coverage in terms of executing most branches and code paths. It also has the advantage of being available without any external dependencies. It has the disadvantage that the code executed during unit tests is likely quite atypical of what's executed during real applications. Running the ./python -X perf -m test --pgo under the "perf" tool, I see the following results:

Children Self Symbol
97.39% 21.16% _PyEval_EvalFrameDefault
34.37% 2.39% deduce_unreachable
24.82% 1.50% _PyGC_Collect
21.11% 1.32% gc_collect_region
20.95% 0.00% gc_collect
19.50% 0.00% py::gc_collect:/home/nas/src/cpython/Lib/test/support/init.py
14.98% 0.54% PyObject_Vectorcall
10.95% 1.48% dict_traverse
7.57% 0.00% _PyPegen_run_parser_from_string
7.46% 0.00% _PyPegen_run_parser
7.46% 0.00% _PyPegen_parse
7.33% 0.24% py::_make_iterencode.._iterencode_dict:/home/nas/src/cpython/Lib/json/encoder.py
7.18% 7.09% visit_reachable
6.43% 0.02% expression_rule
6.13% 0.00% PyRun_StringFlags
6.06% 1.90% _PyEval_Vector
5.99% 5.91% visit_decref
5.80% 0.00% builtin_eval
5.66% 0.04% disjunction_rule

This profile reveals a number of problems. First, a large fraction of time in spent in the cyclic GC. That's because the unit test framework calls test.support.gc_collect() before each test case. That function triggers three full GC collections. Other tests also call the GC explicitly. This is not behavior typical of a real program.

Also taking a lot of time are functions related to parsing and compiling Python code. Notice the builtin_eval() function, for example. I suspect that's mostly a result of using "doctest". Again, this would not be typical of real programs.

I think we should replace the PGO task with a program that more closely represents the behavior of real Python programs. There are at least two potential advantages: it could make the compiled Python binary faster for real programs, it could make our benchmark results less noisy since the compiler would be doing a better and most consistent job of generating optimal code.

Has this already been discussed elsewhere?

This is a minor feature, which does not need previous discussion elsewhere

Links to previous discussion of this feature:

No response

Linked PRs
  • gh-130702

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

調査の方向性

まず、現在の --enable-optimizations PGO タスクと ./python -X perf -m test --pgo コマンドを確認します。Lib/test/support/init.py の test.support.gc_collect と、Lib/json/encoder.py に関するプロファイリング結果も対象に含めます。代表的な置き換え用ワークロードを定義して検証し、そのプロファイルとベンチマークへの影響を既存のユニットテストタスクと比較できれば完了です。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
python
領域
build-system, performance
issue の種類
機能追加
難易度
5/5
見積もり時間
1週間以上
活発さ
停滞
明瞭さ
おおむね明確
初心者へのやさしさ
35/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。