python / python/cpython

Dramatic slowdown with random.random on free-threading build.

Open
#118,392 5 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

performance topic-free-threading
Dominant language
Python
Stars
77.2k
Forks
35.9k
PR merge metrics
PR metrics pending

Description

out

This slowdown was found with one of my favorite benchmarks, which is calculating the pi value with the Monte Carlo method.

import os
import random
import time
from threading import Thread

def monte_carlo_pi_part(n: int, idx: int, results: list[int]) -> None:
    count = 0
    for i in range(n):
        x = random.random()
        y = random.random()

        if x*x + y*y <= 1:
            count += 1
    results[idx] = count


n = 10000
threads = []
num_threads = 100
results = [0] * num_threads
a = time.time()
for i in range(num_threads):
    t = Thread(target=monte_carlo_pi_part, args=(n, i, results))
    t.start()
    threads.append(t)

while threads:
    t = threads.pop()
    t.join()

b = time.time()
print(sum(results) / (n * num_threads) * 4)
print(b-a)

Acquiring critical sections for random methods causes this slowdown.
Removing @critical_section from the method, which uses genrand_uint32 and then updating genrand_uint32 to use atomic operation makes the performance acceptable.

Build Elapsed PI
Default (with specialization) 0.16528010368347168 3.144508
Free-threading (with no specialization) 0.548654317855835 3.1421
Free-threading with my patch (with no specialization) 0.2606849670410156 3.141108
Linked PRs
  • gh-118393
  • gh-118396

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reading the random.random implementation and the genrand_uint32 path, including the @critical_section usage described in the issue. Reproduce the supplied multithreaded Monte Carlo benchmark, then review linked PRs 118393 and 118396; done means the free-threading slowdown is reduced without changing random-number correctness.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.