Reduce some code sizes
還沒有人認領這個 Issue。
- 主要語言
- Python
- 星號
- 77.2k
- 分支
- 35.9k
- PR 合併指標
- PR 指標待擷取
描述
In a release build without PGO or LTO, Python's .text segment is over 5 MB.
$ size ./python
text data bss dec hex filename
5726491 829744 468280 7024515 6b2f83 ./python
Much of this size is necessary, but some code can be reduced significantly with only small changes to code generation. Reducing the amount of code may improve the efficiency of CPU resources such as the instruction cache and branch prediction structures.
The following are the results of having Codex investigate the top 25 functions.
Largest functions in CPython
Top 25 functions in the rebuilt ./python on main, before merging the bytes and 1-byte Unicode decimal-output paths.
Sizes are function symbol sizes in bytes, obtained with readelf --symbols --wide ./python and sorted in descending order. They include inlined code but exclude separately stored data tables and out-of-line callees. The comments summarize the main reasons for the code size, based on the source and inline debug information.
| Rank | Function | Bytes | Why it is large |
|---|---|---|---|
| 1 | _PyEval_EvalFrameDefault |
61,487 | Main bytecode interpreter: many instructions and specialized variants, with inlined stack and reference-count operations. |
| 2 | add_ast_annotations |
49,036 | Generated code repeats type construction, dictionary insertion, cleanup, and error handling for every AST field. |
| 3 | obj2ast_stmt |
31,029 | Converts Python AST objects into internal statement nodes, with separate field validation and conversion for each statement type. |
| 4 | long_to_decimal_string_internal |
24,989 | Decimal digit output is duplicated for bytes and 1-, 2-, and 4-byte Unicode representations, with substantial SIMD expansion. |
| 5 | obj2ast_expr |
20,585 | The expression counterpart: many expression types, each with its own attribute checks and conversion logic. |
| 6 | type_ready |
20,027 | Many type-initialization helpers are inlined, especially the extensive slot-inheritance logic in inherit_slots. |
| 7 | sre_ucs2_match |
17,008 | Complete regex matching engine for 2-byte characters, including repetition, backtracking, and inlined character-set checks. |
| 8 | sre_ucs4_match |
16,341 | A separately compiled copy of the regex matching engine for 4-byte characters. |
| 9 | sre_ucs1_match |
16,288 | A separately compiled copy of the regex matching engine for 1-byte characters. |
| 10 | _PyConfig_Read |
15,804 | Command-line parsing, environment processing, encoding configuration, and other initialization helpers are extensively inlined. |
| 11 | ast2obj_stmt.part.0 |
15,590 | Builds Python AST objects from internal statements, with repeated attribute assignments and inlined list conversions. |
| 12 | _PyAST_Fini |
13,914 | Generated Py_CLEAR operations individually expand reference-count checks and cleanup for many AST state fields. |
| 13 | errno_exec |
13,753 | Registers many errno constants individually; the _add_errcode helper is inlined at repeated call sites. |
| 14 | _PyUnicode_ToNumeric |
13,106 | A generated switch maps a large, sparse set of Unicode code points to numeric values. |
| 15 | init_types |
12,346 | Generated initialization for all AST classes and operator singletons, plus inlined identifier initialization. |
| 16 | split |
12,249 | Inlines whitespace and separator splitting for ASCII and 1-, 2-, and 4-byte Unicode representations. |
| 17 | replace |
12,168 | Handles several replacement strategies and character-width conversions, with inlined search and counting routines. |
| 18 | r_object |
12,135 | Main marshal decoder: handles many object types, nested containers, code objects, and shared-reference reconstruction. |
| 19 | PyUnicode_Format |
11,388 | Inlines % format parsing, argument conversion, padding, width, and precision handling. |
| 20 | codegen_visit_expr_impl |
11,387 | Compiles many expression types, with inlined helpers for calls, method-call optimization, subscripts, and string construction. |
| 21 | _PyUnicode_InitGlobalObjects |
11,154 | Contains generated static-string interning calls and inlined initialization of the global table and interpreter intern dictionary. |
| 22 | simple_stmt_rule |
9,764 | Generated parser alternatives inline several statement rules, particularly imports, plus lookahead, backtracking, and syntax-error handling. |
| 23 | codegen_visit_stmt |
9,732 | Dispatches statement compilation and inlines helpers for imports, assignments, returns, and pattern matching. |
| 24 | compound_stmt_rule |
9,705 | Inlines substantial parsing logic for with, match, try, and for, including invalid-syntax rules. |
| 25 | _PyLexer_get_normal |
9,681 | Combines token recognition, indentation handling, and identifier validation with inlined character-reading and buffer-refill logic. |
Linked PRs
- gh-157393
- gh-157483
- gh-157531
貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
研究方向
首先檢視列出的相關 PR(gh-157393、gh-157483 和 gh-157531),以及從重新建置的 ./python 中回報的頂級函式。使用 readelf --symbols --wide 和 size 來確定哪些函式佔據 release build 的主要部分。在不使用 PGO 或 LTO 的情況下減少 .text 區段,同時保留現有行為,即視為完成。
由索引模型根據 Issue 內容生成。
評估
- 技術堆疊
- python
- 領域
- performance
- Issue 類型
- 重構
- 難度
- 5/5
- 預估耗時
- 一週以上
- 活躍度
- 停滯
- 描述清晰度
- 需要釐清
- 新手友好度
- 25/100