python / python/cpython

Performance regression for loops in 3.12 vs 3.11

Abierto
#123,540 5 comentarios 1 reacción 0 asignados Ver en GitHub

Nadie ha tomado este issue todavía.

3.12 interpreter-core performance
Lenguaje dominante
Python
Estrellas
77.2k
Forks
36k
Métricas de merge de PR
Métricas de PR pendientes

Descripción

Description

I believe I've found a performance regression for large loops in Python 3.12 vs. 3.11. This effect is more pronounced when the value being stored is non-constant – with the list comprehension changed to [0 for x in range(chunk_size)], the relative difference dropped to 1.03x (still in favor of 3.11). Additionally, if the number of rows being generated is small, e.g. 1000, the difference disappears, and both versions are dead even. The results shown below were with 5,000,000 rows.

Interestingly, the array.array() performance difference was massive, but only on MacOS. On Linux, it was approximately the same as with a list. I'm not sure if this is due to the Linux installations being compiled with optimizations, hardware differences, OS differences, etc.

Environment

  • macOS Sonoma 14.6.1 on Apple Air M1
  • Debian Bullseye 12 5.10.0-32-amd64 on Xeon E5-2650 v2
  • Python 3.11.9, 3.12.5
    • Installed via Homebrew on Mac
    • Built from source on Linux with --enable-optimizations --with-lto=full --with-pkg-config=yes

Results

Mac
List
❯ hyperfine -w 20 -r 100 "python3.11 test_loops.py --num-rows 5000000" "python3.12 test_loops.py --num-rows 5000000"
Benchmark 1: python3.11 test_loops.py --num-rows 5000000
  Time (mean ± σ):     136.3 ms ±   1.3 ms    [User: 109.2 ms, System: 25.8 ms]
  Range (min … max):   133.7 ms … 142.3 ms    100 runs

Benchmark 2: python3.12 test_loops.py --num-rows 5000000
  Time (mean ± σ):     144.3 ms ±   4.2 ms    [User: 119.1 ms, System: 23.7 ms]
  Range (min … max):   138.6 ms … 170.5 ms    100 runs

Summary
  python3.11 test_loops.py --num-rows 5000000 ran
    1.06 ± 0.03 times faster than python3.12 test_loops.py --num-rows 5000000
Array
❯ hyperfine -w 20 -r 100 "python3.11 test_loops.py --num-rows 5000000" "python3.12 test_loops.py --num-rows 5000000"
Benchmark 1: python3.11 test_loops.py --num-rows 5000000
  Time (mean ± σ):     176.5 ms ±   1.3 ms    [User: 169.8 ms, System: 5.4 ms]
  Range (min … max):   174.7 ms … 185.0 ms    100 runs

Benchmark 2: python3.12 test_loops.py --num-rows 5000000
  Time (mean ± σ):     277.4 ms ±   1.6 ms    [User: 270.7 ms, System: 5.4 ms]
  Range (min … max):   274.2 ms … 283.9 ms    100 runs

Summary
  python3.11 test_loops.py --num-rows 5000000 ran
    1.57 ± 0.01 times faster than python3.12 test_loops.py --num-rows 5000000
Linux
List
❯ hyperfine -w 20 -r 100 "python3.11 test_loops.py --num-rows 5000000" "python3.12 test_loops.py --num-rows 5000000"
Benchmark 1: python3.11 test_loops.py --num-rows 5000000
  Time (mean ± σ):     378.5 ms ±  22.1 ms    [User: 260.1 ms, System: 118.3 ms]
  Range (min … max):   356.1 ms … 484.0 ms    100 runs

Benchmark 2: python3.12 test_loops.py --num-rows 5000000
  Time (mean ± σ):     405.0 ms ±  27.1 ms    [User: 283.1 ms, System: 121.9 ms]
  Range (min … max):   379.2 ms … 527.1 ms    100 runs

Summary
  python3.11 test_loops.py --num-rows 5000000 ran
    1.07 ± 0.10 times faster than python3.12 test_loops.py --num-rows 5000000
Array
❯ hyperfine -w 20 -r 100 "python3.11 test_loops.py --num-rows 5000000" "python3.12 test_loops.py --num-rows 5000000"

Benchmark 1: python3.11 test_loops.py --num-rows 5000000
  Time (mean ± σ):     387.4 ms ±  26.1 ms    [User: 263.8 ms, System: 123.5 ms]
  Range (min … max):   356.0 ms … 478.5 ms    100 runs

Benchmark 2: python3.12 test_loops.py --num-rows 5000000
  Time (mean ± σ):     408.3 ms ±  27.9 ms    [User: 284.6 ms, System: 123.6 ms]
  Range (min … max):   371.0 ms … 540.5 ms    100 runs

Summary
  python3.11 test_loops.py --num-rows 5000000 ran
    1.05 ± 0.10 times faster than python3.12 test_loops.py --num-rows 5000000

Code

from array import array
import argparse
from typing import Iterable, List


def get_args() -> argparse.Namespace:
    parser = argparse.ArgumentParser()
    parser.add_argument("--chunk-size", type=int, default=250_000)
    parser.add_argument("--num-rows", type=int, default=1_000_000)

    return parser.parse_args()


def generate(num_rows: int, chunk_size: int) -> Iterable[List]:
    for i in range(0, num_rows, chunk_size):
        chunk_size = min(chunk_size, num_rows - i)

        yield generate_chunk(chunk_size)


def generate_chunk(chunk_size: int):
    return [x for x in range(chunk_size)]
    # alternate return type for testing
    # return array("I", (x for x in range(chunk_size)))


if __name__ == "__main__":
    args = get_args()

    # optionally remove the list creation and discard the results
    lst: List = []

    for chunks in generate(
        num_rows=args.num_rows,
        chunk_size=args.chunk_size,
    ):
        # optional if discarding results of generator
        # pass
        lst.append(chunks)

Linked Issues:

Guía de contribución

Abrir la guía de contribución

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Línea de trabajo

Comienza con el script de benchmark de Python incluido y reproduce los tiempos de Python 3.11 frente a 3.12 en los entornos de macOS y Linux indicados, incluidos los casos de listas y arrays. Lee la discusión enlazada de faster-cpython para conocer el contexto; se considera terminado cuando la regresión esté aislada y su causa o resolución estén documentadas.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
python
Área
performance
Tipo de issue
Error
Dificultad
4/5
Tiempo estimado
3-5 días
Estado de actividad
Estancado
Claridad
Bastante claro
Aptitud para principiantes
35/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.