POSIX multiprocessing spawn performance becomes 10x slower from a certain pickle size
まだ誰も着手していません。
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 38/100
- issue の種類
- バグ
- 明瞭さ
- おおむね明確
- 活発さ
- 停滞
- 技術スタック
- linux, python
調査の方向性
Python 3.9 または 3.10 で mp_pipe_limits.py を使って時間差を再現し、その後、パイプの作成と reduction.dump の周辺にある multiprocessing/popen_spawn_posix.py を調査します。65536 バイトと 65537 バイトでの動作を比較し、パフォーマンスの急激な低下を回避する修正を決定します。固定のパイプサイズを安全でない形で使用せずに、再現で報告された低速化が確認されなくなれば完了です。
索引モデルが issue の本文から書いたものです。
説明
Bug report
We are using multiprocessing with the spawn start method. On my 32-thread PC, starting all worker processes for my project used to take 2 seconds. At a certain point, it jumped straight to taking 20 seconds.
The slowdown appears as soon as more than 64 KB needs to be sent to a child process over the pipe.
Consider this minimal reproduction case:
#!/usr/bin/python3
import multiprocessing
import random
import sys
import time
class Container:
def __init__(self, size):
self.data = random.randbytes(size)
class ChildProcess(multiprocessing.Process):
def __init__(self, name: str, container):
super().__init__(name=name)
self.container = container
def run(self) -> None:
print("Running")
def run():
fixed_overhead_3_9 = 885
difference = int(sys.argv[1])
container = Container((64*1024) - fixed_overhead_3_9 + difference)
children = [ChildProcess(f"child-{i}", container) for i in range(0, 2)]
start_time = time.perf_counter()
for child in children:
child.start()
end_time = time.perf_counter()
print(f"Running took {int((end_time - start_time) * 1000)}ms")
for child in children:
child.join()
if __name__ == "__main__":
multiprocessing.set_start_method("spawn")
run()
I added some "instrumentation" in multiprocessing/popen_spawn_posix.py to print the buffer size:
try:
reduction.dump(prep_data, fp)
reduction.dump(process_obj, fp)
finally:
set_spawning_popen(None)
print(len(fp.getbuffer()))
parent_r = child_w = child_r = parent_w = None
Running the example results in:
$ python3.9 mp_pipe_limits.py 0
65536
65536
Running took 9ms
Running
Running
$ python3.9 mp_pipe_limits.py 1
65537
65537
Running
Running took 96ms
Running
Changing the pipe size with fcntl in multiprocessing/popen_spawn_posix.py restores performance:
parent_r = child_w = child_r = parent_w = None
try:
parent_r, child_w = os.pipe()
child_r, parent_w = os.pipe()
fcntl.fcntl(parent_w, 1031, 100000)
cmd = spawn.get_command_line(tracker_fd=tracker_fd,
pipe_handle=child_r)
Where 1031 is fcntl.F_SETPIPE_SZ, which is not in Python 3.9.
Rerunning the reproduction case after this change:
$ python3.9 mp_pipe_limits.py 1
65537
65537
Running took 9ms
Running
Running
Of course, changing the pipe size will only delay the onset of the problem. The real solution (if there is any) will probably be different. Blindly setting a pipe size might also not be safe as it depends on limits set in /proc.
The example above is a best case example, since it has very limited pickle overhead. We hit this limit without any data caches involved. It's just our Python objects that live after application initialization. They are slower to pickle. However, then things are still 10x slower, so not just a fixed 80ms as seen in the example.
We use spawn instead of fork on Linux to avoid troubles with objects that cannot be pickled on other OSs (Windows).
Your environment
- CPython versions tested on: 3.9 and 3.10
- Operating system and architecture: Arch Linux (kernel 5.19.9-arch1-1), x86_64
- 主要言語
- Python
- スター
- 77.2k
- フォーク
- 36k
- 平均マージ
- 1日 9時間
- マージ済み PR(30日)
- 558
コントリビューションガイド
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
python/cpython のほかの issue
-
docs pending
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
-
stdlib type-feature
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
-
stdlib type-feature
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
-
build type-bug
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
-
stdlib topic-email type-feature
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 74/100
-
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
PolicyEngine/policyengine-us#9559 ·
-
priority: p3
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
googleapis/librarian#7636 ·
-
from:qa priority:P2 reliability tech-debt
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
spec-kitty/spec-kitty#4874 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100