POSIX multiprocessing spawn performance becomes 10x slower from a certain pickle size
Personne n'a encore pris cette issue.
Évaluation
- Difficulté
- 4/5
- Temps estimé
- 3-5 jours
- Accessibilité débutants
- 38/100
- Type d'issue
- Bug
- Clarté
- Plutôt claire
- Activité
- À l'abandon
- Stack technique
- linux, python
- Domaine
- operating-systems, performance
Piste de recherche
Reproduisez la différence de temps avec mp_pipe_limits.py sur Python 3.9 ou 3.10, puis inspectez multiprocessing/popen_spawn_posix.py autour de la création du pipe et de reduction.dump. Comparez le comportement avec 65536 et 65537 octets et déterminez une correction qui évite la chute brutale des performances ; le travail est terminé lorsque la reproduction ne montre plus le ralentissement signalé sans dépendre d’une taille de pipe fixe dangereuse.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Description
Bug report
We are using multiprocessing with the spawn start method. On my 32-thread PC, starting all worker processes for my project used to take 2 seconds. At a certain point, it jumped straight to taking 20 seconds.
The slowdown appears as soon as more than 64 KB needs to be sent to a child process over the pipe.
Consider this minimal reproduction case:
#!/usr/bin/python3
import multiprocessing
import random
import sys
import time
class Container:
def __init__(self, size):
self.data = random.randbytes(size)
class ChildProcess(multiprocessing.Process):
def __init__(self, name: str, container):
super().__init__(name=name)
self.container = container
def run(self) -> None:
print("Running")
def run():
fixed_overhead_3_9 = 885
difference = int(sys.argv[1])
container = Container((64*1024) - fixed_overhead_3_9 + difference)
children = [ChildProcess(f"child-{i}", container) for i in range(0, 2)]
start_time = time.perf_counter()
for child in children:
child.start()
end_time = time.perf_counter()
print(f"Running took {int((end_time - start_time) * 1000)}ms")
for child in children:
child.join()
if __name__ == "__main__":
multiprocessing.set_start_method("spawn")
run()
I added some "instrumentation" in multiprocessing/popen_spawn_posix.py to print the buffer size:
try:
reduction.dump(prep_data, fp)
reduction.dump(process_obj, fp)
finally:
set_spawning_popen(None)
print(len(fp.getbuffer()))
parent_r = child_w = child_r = parent_w = None
Running the example results in:
$ python3.9 mp_pipe_limits.py 0
65536
65536
Running took 9ms
Running
Running
$ python3.9 mp_pipe_limits.py 1
65537
65537
Running
Running took 96ms
Running
Changing the pipe size with fcntl in multiprocessing/popen_spawn_posix.py restores performance:
parent_r = child_w = child_r = parent_w = None
try:
parent_r, child_w = os.pipe()
child_r, parent_w = os.pipe()
fcntl.fcntl(parent_w, 1031, 100000)
cmd = spawn.get_command_line(tracker_fd=tracker_fd,
pipe_handle=child_r)
Where 1031 is fcntl.F_SETPIPE_SZ, which is not in Python 3.9.
Rerunning the reproduction case after this change:
$ python3.9 mp_pipe_limits.py 1
65537
65537
Running took 9ms
Running
Running
Of course, changing the pipe size will only delay the onset of the problem. The real solution (if there is any) will probably be different. Blindly setting a pipe size might also not be safe as it depends on limits set in /proc.
The example above is a best case example, since it has very limited pickle overhead. We hit this limit without any data caches involved. It's just our Python objects that live after application initialization. They are slower to pickle. However, then things are still 10x slower, so not just a fixed 80ms as seen in the example.
We use spawn instead of fork on Linux to avoid troubles with objects that cannot be pickled on other OSs (Windows).
Your environment
- CPython versions tested on: 3.9 and 3.10
- Operating system and architecture: Arch Linux (kernel 5.19.9-arch1-1), x86_64
- Langage dominant
- Python
- Étoiles
- 77.2k
- Forks
- 36k
- Merge moyen
- 1 j 9 h
- PR mergées (30 j)
- 558
Guide de contribution
Ouvrir le guide de contribution
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Autres issues de python/cpython
-
docs pending
Difficulté 2/5 1-3 heures Accessibilité débutants 78/100
-
stdlib type-feature
Difficulté 2/5 1-3 heures Accessibilité débutants 78/100
-
stdlib type-feature
Difficulté 2/5 1-3 heures Accessibilité débutants 72/100
-
build type-bug
Difficulté 2/5 1-3 heures Accessibilité débutants 76/100
-
stdlib topic-email type-feature
Difficulté 2/5 1-3 heures Accessibilité débutants 70/100
Toutes les issues de python/cpython
Issues similaires
-
fix: inaccuracy ⚠️
Difficulté 2/5 1-3 heures Accessibilité débutants 72/100
uabrc/uabrc.github.io#1255 · 1 commentaire ·
-
Difficulté 2/5 1-3 heures Accessibilité débutants 84/100
ethereum-optimism/factory#64 ·
-
Difficulté 2/5 1-3 heures Accessibilité débutants 90/100
duckdb/duckdb-python#627 ·
-
Difficulté 2/5 1-3 heures Accessibilité débutants 68/100
-
Add link for tutorial Ouvertedocumentation
Difficulté 1/5 Moins d'une heure Accessibilité débutants 78/100
Qiskit/qiskit-addon-sqd#376 ·