python / python/cpython

Using the unit tests as the PGO task has problems

未關閉
#130,701 0 則留言 2 個 reaction 已指派 0 人 在 GitHub 檢視

還沒有人認領這個 Issue。

build performance type-feature
主要語言
Python
星號
77.2k
分支
35.9k
PR 合併指標
PR 指標待擷取

描述

Feature or enhancement

Proposal:

When Python is compiled with --enabled-optimizations, which turns on PGO (program guided optimizations), the build will run a subset of the unit tests as the "task" to generate profile information. Included in the profile is information like counts of the taken side of a CPU conditional branch instructions. To get the best optimization, your PGO task should match the branch taken behavior of your real workloads.

Using the unit tests has the advantage that we have good code coverage in terms of executing most branches and code paths. It also has the advantage of being available without any external dependencies. It has the disadvantage that the code executed during unit tests is likely quite atypical of what's executed during real applications. Running the ./python -X perf -m test --pgo under the "perf" tool, I see the following results:

Children Self Symbol
97.39% 21.16% _PyEval_EvalFrameDefault
34.37% 2.39% deduce_unreachable
24.82% 1.50% _PyGC_Collect
21.11% 1.32% gc_collect_region
20.95% 0.00% gc_collect
19.50% 0.00% py::gc_collect:/home/nas/src/cpython/Lib/test/support/init.py
14.98% 0.54% PyObject_Vectorcall
10.95% 1.48% dict_traverse
7.57% 0.00% _PyPegen_run_parser_from_string
7.46% 0.00% _PyPegen_run_parser
7.46% 0.00% _PyPegen_parse
7.33% 0.24% py::_make_iterencode.._iterencode_dict:/home/nas/src/cpython/Lib/json/encoder.py
7.18% 7.09% visit_reachable
6.43% 0.02% expression_rule
6.13% 0.00% PyRun_StringFlags
6.06% 1.90% _PyEval_Vector
5.99% 5.91% visit_decref
5.80% 0.00% builtin_eval
5.66% 0.04% disjunction_rule

This profile reveals a number of problems. First, a large fraction of time in spent in the cyclic GC. That's because the unit test framework calls test.support.gc_collect() before each test case. That function triggers three full GC collections. Other tests also call the GC explicitly. This is not behavior typical of a real program.

Also taking a lot of time are functions related to parsing and compiling Python code. Notice the builtin_eval() function, for example. I suspect that's mostly a result of using "doctest". Again, this would not be typical of real programs.

I think we should replace the PGO task with a program that more closely represents the behavior of real Python programs. There are at least two potential advantages: it could make the compiled Python binary faster for real programs, it could make our benchmark results less noisy since the compiler would be doing a better and most consistent job of generating optimal code.

Has this already been discussed elsewhere?

This is a minor feature, which does not need previous discussion elsewhere

Links to previous discussion of this feature:

No response

Linked PRs
  • gh-130702

貢獻指南

開啟貢獻指南

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

研究方向

先檢視目前的 --enable-optimizations PGO 工作和 ./python -X perf -m test --pgo 命令,包括 Lib/test/support/init.py 中的 test.support.gc_collect,以及涉及 Lib/json/encoder.py 的剖析觀察結果。完成的標準是定義並驗證一個具代表性的替代工作負載,然後將其剖析結果和基準測試影響與現有的單元測試工作進行比較。

由索引模型根據 Issue 內容生成。

評估

技術堆疊
python
領域
build-system, performance
Issue 類型
功能
難度
5/5
預估耗時
一週以上
活躍度
停滯
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。