Reduce some code sizes
まだ誰も着手していません。
- 主要言語
- Python
- スター
- 77.2k
- フォーク
- 35.9k
- PR マージ指標
- PR 指標を取得中
説明
In a release build without PGO or LTO, Python's .text segment is over 5 MB.
$ size ./python
text data bss dec hex filename
5726491 829744 468280 7024515 6b2f83 ./python
Much of this size is necessary, but some code can be reduced significantly with only small changes to code generation. Reducing the amount of code may improve the efficiency of CPU resources such as the instruction cache and branch prediction structures.
The following are the results of having Codex investigate the top 25 functions.
Largest functions in CPython
Top 25 functions in the rebuilt ./python on main, before merging the bytes and 1-byte Unicode decimal-output paths.
Sizes are function symbol sizes in bytes, obtained with readelf --symbols --wide ./python and sorted in descending order. They include inlined code but exclude separately stored data tables and out-of-line callees. The comments summarize the main reasons for the code size, based on the source and inline debug information.
| Rank | Function | Bytes | Why it is large |
|---|---|---|---|
| 1 | _PyEval_EvalFrameDefault |
61,487 | Main bytecode interpreter: many instructions and specialized variants, with inlined stack and reference-count operations. |
| 2 | add_ast_annotations |
49,036 | Generated code repeats type construction, dictionary insertion, cleanup, and error handling for every AST field. |
| 3 | obj2ast_stmt |
31,029 | Converts Python AST objects into internal statement nodes, with separate field validation and conversion for each statement type. |
| 4 | long_to_decimal_string_internal |
24,989 | Decimal digit output is duplicated for bytes and 1-, 2-, and 4-byte Unicode representations, with substantial SIMD expansion. |
| 5 | obj2ast_expr |
20,585 | The expression counterpart: many expression types, each with its own attribute checks and conversion logic. |
| 6 | type_ready |
20,027 | Many type-initialization helpers are inlined, especially the extensive slot-inheritance logic in inherit_slots. |
| 7 | sre_ucs2_match |
17,008 | Complete regex matching engine for 2-byte characters, including repetition, backtracking, and inlined character-set checks. |
| 8 | sre_ucs4_match |
16,341 | A separately compiled copy of the regex matching engine for 4-byte characters. |
| 9 | sre_ucs1_match |
16,288 | A separately compiled copy of the regex matching engine for 1-byte characters. |
| 10 | _PyConfig_Read |
15,804 | Command-line parsing, environment processing, encoding configuration, and other initialization helpers are extensively inlined. |
| 11 | ast2obj_stmt.part.0 |
15,590 | Builds Python AST objects from internal statements, with repeated attribute assignments and inlined list conversions. |
| 12 | _PyAST_Fini |
13,914 | Generated Py_CLEAR operations individually expand reference-count checks and cleanup for many AST state fields. |
| 13 | errno_exec |
13,753 | Registers many errno constants individually; the _add_errcode helper is inlined at repeated call sites. |
| 14 | _PyUnicode_ToNumeric |
13,106 | A generated switch maps a large, sparse set of Unicode code points to numeric values. |
| 15 | init_types |
12,346 | Generated initialization for all AST classes and operator singletons, plus inlined identifier initialization. |
| 16 | split |
12,249 | Inlines whitespace and separator splitting for ASCII and 1-, 2-, and 4-byte Unicode representations. |
| 17 | replace |
12,168 | Handles several replacement strategies and character-width conversions, with inlined search and counting routines. |
| 18 | r_object |
12,135 | Main marshal decoder: handles many object types, nested containers, code objects, and shared-reference reconstruction. |
| 19 | PyUnicode_Format |
11,388 | Inlines % format parsing, argument conversion, padding, width, and precision handling. |
| 20 | codegen_visit_expr_impl |
11,387 | Compiles many expression types, with inlined helpers for calls, method-call optimization, subscripts, and string construction. |
| 21 | _PyUnicode_InitGlobalObjects |
11,154 | Contains generated static-string interning calls and inlined initialization of the global table and interpreter intern dictionary. |
| 22 | simple_stmt_rule |
9,764 | Generated parser alternatives inline several statement rules, particularly imports, plus lookahead, backtracking, and syntax-error handling. |
| 23 | codegen_visit_stmt |
9,732 | Dispatches statement compilation and inlines helpers for imports, assignments, returns, and pattern matching. |
| 24 | compound_stmt_rule |
9,705 | Inlines substantial parsing logic for with, match, try, and for, including invalid-syntax rules. |
| 25 | _PyLexer_get_normal |
9,681 | Combines token recognition, indentation handling, and identifier validation with inlined character-reading and buffer-refill logic. |
Linked PRs
- gh-157393
- gh-157483
- gh-157531
コントリビューションガイド
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
調査の方向性
まず、記載されているリンク先のPR(gh-157393、gh-157483、gh-157531)と、再ビルドした ./python で報告されている上位関数を確認します。readelf --symbols --wide と size を使用して、どの関数がリリースビルドの大部分を占めているかを特定します。PGO や LTO を使わずに .text セグメントを削減し、既存の動作を維持できれば完了です。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- python
- 領域
- performance
- issue の種類
- リファクタリング
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 活発さ
- 停滞
- 明瞭さ
- 説明が足りない
- 初心者へのやさしさ
- 25/100