python / python/cpython

Optimize int add/sub for wide exact ints

Aperta
#151,289 6 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

interpreter-core performance type-feature
Lingua principale
Python
Stelle
77.2k
Fork
35.9k
Metriche di merge delle PR
Metriche PR in attesa

Descrizione

CPython has a fast path for compact integers in binary add/sub, but wide exact ints still go through the generic long arithmetic path even when both operands fit in int64_t.

This issue proposes adding a separate fast path for exact PyLong operands that fit in signed 64-bit integers, while preserving the existing compact-int path.

Suggested implementation:

  • Keep the current compact-int specialization unchanged.
  • Add a separate wide-int path for exact ints that fit in int64_t.
  • Preserve current behavior for overflow, subclasses, and other non-exact-int cases.

Motivation:

  • Improve performance for wide integer add/sub without affecting the common compact-int hot path.
  • Avoid adding new opcodes in the compact-int path.
  • Fit within the current interpreter and specialization structure.

Benchmark evidence:

  • I prototyped this locally with a benchmark covering compact and wide add/sub cases.
  • Wide cases improved substantially, while compact cases remained effectively flat.
  • Representative interpreter-only results with JIT disabled:
    • add_wide: about 25% faster
    • sub_wide: about 35% faster
    • add_compact/sub_compact: effectively unchanged

Benchmark script used locally:

"""Microbenchmark compact vs wide int add/sub with pyperf.

Use this with PYTHON_JIT=0 and -S if you want a stable interpreter-only run:

    PYTHON_JIT=0 ./python.exe -S Tools/scripts/bench_wide_int_pyperf.py
"""

from __future__ import annotations

import pyperf


def bench_add_compact() -> int:
    a = 1
    b = 2
    return a + b


def bench_add_wide() -> int:
    a = 10_000_000_000
    b = 1
    return a + b


def bench_sub_compact() -> int:
    a = 1
    b = 2
    return a - b


def bench_sub_wide() -> int:
    a = 10_000_000_000
    b = 1
    return a - b


def main() -> None:
    runner = pyperf.Runner()
    runner.bench_func("add_compact", bench_add_compact)
    runner.bench_func("add_wide", bench_add_wide)
    runner.bench_func("sub_compact", bench_sub_compact)
    runner.bench_func("sub_wide", bench_sub_wide)


if __name__ == "__main__":
    main()
Linked PRs
  • gh-151290

Guida per i contributori

Apri la guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Direzione di ricerca

Inizia esaminando la PR collegata gh-151290, quindi analizza la struttura di specializzazione dell’interprete descritta nell’issue. Esegui Tools/scripts/bench_wide_int_pyperf.py con PYTHON_JIT=0 e -S per confrontare i casi add/sub compact e wide. Il lavoro è completato quando gli operandi wide compatibili con int64 usano un fast path, mentre il comportamento di compact, overflow, subclass e non-exact-int rimane invariato.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
python
Ambito
performance
Tipo di issue
Refactoring
Difficoltà
4/5
Tempo stimato
3-5 giorni
Stato di attività
Ferma
Chiarezza
Abbastanza chiara
Idoneità per principianti
25/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.