Reduce some code sizes
Dieses Issue hat noch niemand übernommen.
- Vorherrschende Sprache
- Python
- Sterne
- 77.2k
- Forks
- 35.9k
- PR-Merge-Kennzahlen
- PR-Kennzahlen ausstehend
Beschreibung
In a release build without PGO or LTO, Python's .text segment is over 5 MB.
$ size ./python
text data bss dec hex filename
5726491 829744 468280 7024515 6b2f83 ./python
Much of this size is necessary, but some code can be reduced significantly with only small changes to code generation. Reducing the amount of code may improve the efficiency of CPU resources such as the instruction cache and branch prediction structures.
The following are the results of having Codex investigate the top 25 functions.
Largest functions in CPython
Top 25 functions in the rebuilt ./python on main, before merging the bytes and 1-byte Unicode decimal-output paths.
Sizes are function symbol sizes in bytes, obtained with readelf --symbols --wide ./python and sorted in descending order. They include inlined code but exclude separately stored data tables and out-of-line callees. The comments summarize the main reasons for the code size, based on the source and inline debug information.
| Rank | Function | Bytes | Why it is large |
|---|---|---|---|
| 1 | _PyEval_EvalFrameDefault |
61,487 | Main bytecode interpreter: many instructions and specialized variants, with inlined stack and reference-count operations. |
| 2 | add_ast_annotations |
49,036 | Generated code repeats type construction, dictionary insertion, cleanup, and error handling for every AST field. |
| 3 | obj2ast_stmt |
31,029 | Converts Python AST objects into internal statement nodes, with separate field validation and conversion for each statement type. |
| 4 | long_to_decimal_string_internal |
24,989 | Decimal digit output is duplicated for bytes and 1-, 2-, and 4-byte Unicode representations, with substantial SIMD expansion. |
| 5 | obj2ast_expr |
20,585 | The expression counterpart: many expression types, each with its own attribute checks and conversion logic. |
| 6 | type_ready |
20,027 | Many type-initialization helpers are inlined, especially the extensive slot-inheritance logic in inherit_slots. |
| 7 | sre_ucs2_match |
17,008 | Complete regex matching engine for 2-byte characters, including repetition, backtracking, and inlined character-set checks. |
| 8 | sre_ucs4_match |
16,341 | A separately compiled copy of the regex matching engine for 4-byte characters. |
| 9 | sre_ucs1_match |
16,288 | A separately compiled copy of the regex matching engine for 1-byte characters. |
| 10 | _PyConfig_Read |
15,804 | Command-line parsing, environment processing, encoding configuration, and other initialization helpers are extensively inlined. |
| 11 | ast2obj_stmt.part.0 |
15,590 | Builds Python AST objects from internal statements, with repeated attribute assignments and inlined list conversions. |
| 12 | _PyAST_Fini |
13,914 | Generated Py_CLEAR operations individually expand reference-count checks and cleanup for many AST state fields. |
| 13 | errno_exec |
13,753 | Registers many errno constants individually; the _add_errcode helper is inlined at repeated call sites. |
| 14 | _PyUnicode_ToNumeric |
13,106 | A generated switch maps a large, sparse set of Unicode code points to numeric values. |
| 15 | init_types |
12,346 | Generated initialization for all AST classes and operator singletons, plus inlined identifier initialization. |
| 16 | split |
12,249 | Inlines whitespace and separator splitting for ASCII and 1-, 2-, and 4-byte Unicode representations. |
| 17 | replace |
12,168 | Handles several replacement strategies and character-width conversions, with inlined search and counting routines. |
| 18 | r_object |
12,135 | Main marshal decoder: handles many object types, nested containers, code objects, and shared-reference reconstruction. |
| 19 | PyUnicode_Format |
11,388 | Inlines % format parsing, argument conversion, padding, width, and precision handling. |
| 20 | codegen_visit_expr_impl |
11,387 | Compiles many expression types, with inlined helpers for calls, method-call optimization, subscripts, and string construction. |
| 21 | _PyUnicode_InitGlobalObjects |
11,154 | Contains generated static-string interning calls and inlined initialization of the global table and interpreter intern dictionary. |
| 22 | simple_stmt_rule |
9,764 | Generated parser alternatives inline several statement rules, particularly imports, plus lookahead, backtracking, and syntax-error handling. |
| 23 | codegen_visit_stmt |
9,732 | Dispatches statement compilation and inlines helpers for imports, assignments, returns, and pattern matching. |
| 24 | compound_stmt_rule |
9,705 | Inlines substantial parsing logic for with, match, try, and for, including invalid-syntax rules. |
| 25 | _PyLexer_get_normal |
9,681 | Combines token recognition, indentation handling, and identifier validation with inlined character-reading and buffer-refill logic. |
Linked PRs
- gh-157393
- gh-157483
- gh-157531
Beitragsleitfaden
Erste Schritte
- Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
- Forke das Repository und arbeite in einem Branch.
- Öffne einen Pull Request, der die Issue-Nummer nennt.
Rechercherichtung
Beginne mit der Durchsicht der aufgeführten verknüpften PRs (gh-157393, gh-157483 und gh-157531) sowie der gemeldeten Top-Funktionen aus dem neu erstellten ./python. Verwende readelf --symbols --wide und size, um festzustellen, welche Funktionen den Release-Build dominieren. Als abgeschlossen gilt die Aufgabe, wenn das .text-Segment ohne PGO oder LTO verkleinert wurde und dabei das bestehende Verhalten erhalten bleibt.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- python
- Bereich
- performance
- Issue-Typ
- Refactoring
- Schwierigkeit
- 5/5
- Geschätzter Aufwand
- Über eine Woche
- Aktivitätsstatus
- Veraltet
- Klarheit
- Muss geklärt werden
- Anfängerfreundlichkeit
- 25/100