Reduce some code sizes
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 77.2k
- Forks
- 35.9k
- PR merge metrics
- PR metrics pending
Description
In a release build without PGO or LTO, Python's .text segment is over 5 MB.
$ size ./python
text data bss dec hex filename
5726491 829744 468280 7024515 6b2f83 ./python
Much of this size is necessary, but some code can be reduced significantly with only small changes to code generation. Reducing the amount of code may improve the efficiency of CPU resources such as the instruction cache and branch prediction structures.
The following are the results of having Codex investigate the top 25 functions.
Largest functions in CPython
Top 25 functions in the rebuilt ./python on main, before merging the bytes and 1-byte Unicode decimal-output paths.
Sizes are function symbol sizes in bytes, obtained with readelf --symbols --wide ./python and sorted in descending order. They include inlined code but exclude separately stored data tables and out-of-line callees. The comments summarize the main reasons for the code size, based on the source and inline debug information.
| Rank | Function | Bytes | Why it is large |
|---|---|---|---|
| 1 | _PyEval_EvalFrameDefault |
61,487 | Main bytecode interpreter: many instructions and specialized variants, with inlined stack and reference-count operations. |
| 2 | add_ast_annotations |
49,036 | Generated code repeats type construction, dictionary insertion, cleanup, and error handling for every AST field. |
| 3 | obj2ast_stmt |
31,029 | Converts Python AST objects into internal statement nodes, with separate field validation and conversion for each statement type. |
| 4 | long_to_decimal_string_internal |
24,989 | Decimal digit output is duplicated for bytes and 1-, 2-, and 4-byte Unicode representations, with substantial SIMD expansion. |
| 5 | obj2ast_expr |
20,585 | The expression counterpart: many expression types, each with its own attribute checks and conversion logic. |
| 6 | type_ready |
20,027 | Many type-initialization helpers are inlined, especially the extensive slot-inheritance logic in inherit_slots. |
| 7 | sre_ucs2_match |
17,008 | Complete regex matching engine for 2-byte characters, including repetition, backtracking, and inlined character-set checks. |
| 8 | sre_ucs4_match |
16,341 | A separately compiled copy of the regex matching engine for 4-byte characters. |
| 9 | sre_ucs1_match |
16,288 | A separately compiled copy of the regex matching engine for 1-byte characters. |
| 10 | _PyConfig_Read |
15,804 | Command-line parsing, environment processing, encoding configuration, and other initialization helpers are extensively inlined. |
| 11 | ast2obj_stmt.part.0 |
15,590 | Builds Python AST objects from internal statements, with repeated attribute assignments and inlined list conversions. |
| 12 | _PyAST_Fini |
13,914 | Generated Py_CLEAR operations individually expand reference-count checks and cleanup for many AST state fields. |
| 13 | errno_exec |
13,753 | Registers many errno constants individually; the _add_errcode helper is inlined at repeated call sites. |
| 14 | _PyUnicode_ToNumeric |
13,106 | A generated switch maps a large, sparse set of Unicode code points to numeric values. |
| 15 | init_types |
12,346 | Generated initialization for all AST classes and operator singletons, plus inlined identifier initialization. |
| 16 | split |
12,249 | Inlines whitespace and separator splitting for ASCII and 1-, 2-, and 4-byte Unicode representations. |
| 17 | replace |
12,168 | Handles several replacement strategies and character-width conversions, with inlined search and counting routines. |
| 18 | r_object |
12,135 | Main marshal decoder: handles many object types, nested containers, code objects, and shared-reference reconstruction. |
| 19 | PyUnicode_Format |
11,388 | Inlines % format parsing, argument conversion, padding, width, and precision handling. |
| 20 | codegen_visit_expr_impl |
11,387 | Compiles many expression types, with inlined helpers for calls, method-call optimization, subscripts, and string construction. |
| 21 | _PyUnicode_InitGlobalObjects |
11,154 | Contains generated static-string interning calls and inlined initialization of the global table and interpreter intern dictionary. |
| 22 | simple_stmt_rule |
9,764 | Generated parser alternatives inline several statement rules, particularly imports, plus lookahead, backtracking, and syntax-error handling. |
| 23 | codegen_visit_stmt |
9,732 | Dispatches statement compilation and inlines helpers for imports, assignments, returns, and pattern matching. |
| 24 | compound_stmt_rule |
9,705 | Inlines substantial parsing logic for with, match, try, and for, including invalid-syntax rules. |
| 25 | _PyLexer_get_normal |
9,681 | Combines token recognition, indentation handling, and identifier validation with inlined character-reading and buffer-refill logic. |
Linked PRs
- gh-157393
- gh-157483
- gh-157531
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reviewing the listed linked PRs (gh-157393, gh-157483, and gh-157531) and the reported top functions from the rebuilt ./python. Use readelf --symbols --wide and size to establish which functions dominate the release build. Done means reducing the .text segment without PGO or LTO while preserving the existing behavior.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- performance
- Issue type
- Refactor
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100