Performance regression for loops in 3.12 vs 3.11
Dieses Issue hat noch niemand übernommen.
- Vorherrschende Sprache
- Python
- Sterne
- 77.2k
- Forks
- 36k
- PR-Merge-Kennzahlen
- PR-Kennzahlen ausstehend
Beschreibung
Description
I believe I've found a performance regression for large loops in Python 3.12 vs. 3.11. This effect is more pronounced when the value being stored is non-constant – with the list comprehension changed to [0 for x in range(chunk_size)], the relative difference dropped to 1.03x (still in favor of 3.11). Additionally, if the number of rows being generated is small, e.g. 1000, the difference disappears, and both versions are dead even. The results shown below were with 5,000,000 rows.
Interestingly, the array.array() performance difference was massive, but only on MacOS. On Linux, it was approximately the same as with a list. I'm not sure if this is due to the Linux installations being compiled with optimizations, hardware differences, OS differences, etc.
Environment
- macOS Sonoma 14.6.1 on Apple Air M1
- Debian Bullseye 12 5.10.0-32-amd64 on Xeon E5-2650 v2
- Python 3.11.9, 3.12.5
- Installed via Homebrew on Mac
- Built from source on Linux with
--enable-optimizations --with-lto=full --with-pkg-config=yes
Results
Mac
List
❯ hyperfine -w 20 -r 100 "python3.11 test_loops.py --num-rows 5000000" "python3.12 test_loops.py --num-rows 5000000"
Benchmark 1: python3.11 test_loops.py --num-rows 5000000
Time (mean ± σ): 136.3 ms ± 1.3 ms [User: 109.2 ms, System: 25.8 ms]
Range (min … max): 133.7 ms … 142.3 ms 100 runs
Benchmark 2: python3.12 test_loops.py --num-rows 5000000
Time (mean ± σ): 144.3 ms ± 4.2 ms [User: 119.1 ms, System: 23.7 ms]
Range (min … max): 138.6 ms … 170.5 ms 100 runs
Summary
python3.11 test_loops.py --num-rows 5000000 ran
1.06 ± 0.03 times faster than python3.12 test_loops.py --num-rows 5000000
Array
❯ hyperfine -w 20 -r 100 "python3.11 test_loops.py --num-rows 5000000" "python3.12 test_loops.py --num-rows 5000000"
Benchmark 1: python3.11 test_loops.py --num-rows 5000000
Time (mean ± σ): 176.5 ms ± 1.3 ms [User: 169.8 ms, System: 5.4 ms]
Range (min … max): 174.7 ms … 185.0 ms 100 runs
Benchmark 2: python3.12 test_loops.py --num-rows 5000000
Time (mean ± σ): 277.4 ms ± 1.6 ms [User: 270.7 ms, System: 5.4 ms]
Range (min … max): 274.2 ms … 283.9 ms 100 runs
Summary
python3.11 test_loops.py --num-rows 5000000 ran
1.57 ± 0.01 times faster than python3.12 test_loops.py --num-rows 5000000
Linux
List
❯ hyperfine -w 20 -r 100 "python3.11 test_loops.py --num-rows 5000000" "python3.12 test_loops.py --num-rows 5000000"
Benchmark 1: python3.11 test_loops.py --num-rows 5000000
Time (mean ± σ): 378.5 ms ± 22.1 ms [User: 260.1 ms, System: 118.3 ms]
Range (min … max): 356.1 ms … 484.0 ms 100 runs
Benchmark 2: python3.12 test_loops.py --num-rows 5000000
Time (mean ± σ): 405.0 ms ± 27.1 ms [User: 283.1 ms, System: 121.9 ms]
Range (min … max): 379.2 ms … 527.1 ms 100 runs
Summary
python3.11 test_loops.py --num-rows 5000000 ran
1.07 ± 0.10 times faster than python3.12 test_loops.py --num-rows 5000000
Array
❯ hyperfine -w 20 -r 100 "python3.11 test_loops.py --num-rows 5000000" "python3.12 test_loops.py --num-rows 5000000"
Benchmark 1: python3.11 test_loops.py --num-rows 5000000
Time (mean ± σ): 387.4 ms ± 26.1 ms [User: 263.8 ms, System: 123.5 ms]
Range (min … max): 356.0 ms … 478.5 ms 100 runs
Benchmark 2: python3.12 test_loops.py --num-rows 5000000
Time (mean ± σ): 408.3 ms ± 27.9 ms [User: 284.6 ms, System: 123.6 ms]
Range (min … max): 371.0 ms … 540.5 ms 100 runs
Summary
python3.11 test_loops.py --num-rows 5000000 ran
1.05 ± 0.10 times faster than python3.12 test_loops.py --num-rows 5000000
Code
from array import array
import argparse
from typing import Iterable, List
def get_args() -> argparse.Namespace:
parser = argparse.ArgumentParser()
parser.add_argument("--chunk-size", type=int, default=250_000)
parser.add_argument("--num-rows", type=int, default=1_000_000)
return parser.parse_args()
def generate(num_rows: int, chunk_size: int) -> Iterable[List]:
for i in range(0, num_rows, chunk_size):
chunk_size = min(chunk_size, num_rows - i)
yield generate_chunk(chunk_size)
def generate_chunk(chunk_size: int):
return [x for x in range(chunk_size)]
# alternate return type for testing
# return array("I", (x for x in range(chunk_size)))
if __name__ == "__main__":
args = get_args()
# optionally remove the list creation and discard the results
lst: List = []
for chunks in generate(
num_rows=args.num_rows,
chunk_size=args.chunk_size,
):
# optional if discarding results of generator
# pass
lst.append(chunks)
Linked Issues:
Beitragsleitfaden
Erste Schritte
- Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
- Forke das Repository und arbeite in einem Branch.
- Öffne einen Pull Request, der die Issue-Nummer nennt.
Rechercherichtung
Beginne mit dem eingebetteten Python-Benchmarkskript und reproduziere die Zeitmessungen von Python 3.11 gegenüber 3.12 in den angegebenen macOS- und Linux-Umgebungen, einschließlich der Fälle mit Listen und Arrays. Lies zur Einordnung die verknüpfte faster-cpython-Diskussion; abgeschlossen ist die Aufgabe, wenn die Regression isoliert und ihre Ursache oder Behebung dokumentiert ist.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- python
- Bereich
- performance
- Issue-Typ
- Bug
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Aktivitätsstatus
- Veraltet
- Klarheit
- Größtenteils klar
- Anfängerfreundlichkeit
- 35/100