python / python/cpython

multiprocessing race condition on flushing stdout, deadlocks child on exit

Open
#91,776 12 comments 3 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

topic-multiprocessing type-bug
Dominant language
Python
Stars
77.2k
Forks
36k
PR merge metrics
PR metrics pending

Description

I experienced deadlocks when using the logging example for multiprocessing, where a QueueHandler is used.

I found that sometimes a multiprocessing process is not terminating, when it has put an element in a Queue,
even if the parent process runs a thread to empty the queue and successfully retrieved the item.

The worker looks like this:

def worker(q):
    q.put(current_process().name)
    return

and while this works:

workers = []
for i in range(15):
    p = Process(target=worker, args=(logq,), name=f"Worker {i+1}")
    workers.append(p)

for p in workers:
    p.start()
    time.sleep(.1)

removing the sleep leads to a very high propability of deadlocking when I then try to join the processes, e.g,:

for w in workers:
    print(f'trying to join on {w.name}, alive={w.is_alive()}, exitcode={w.exitcode}', w.name, w.is_alive(), w.exitcode)
    w.join()

The Process is still marked as alive, hitting Ctrl+C gives this:

trying to join on Worker 14, alive=True, exitcode=None Worker 14 True None

^CTraceback (most recent call last):
  File "/home/fls/pybug/deadlock.py", line 43, in <module>
    w.join()
  File "/usr/lib/python3.10/multiprocessing/process.py", line 149, in join
    res = self._popen.wait(timeout)
  File "/usr/lib/python3.10/multiprocessing/popen_fork.py", line 43, in wait
    return self.poll(os.WNOHANG if timeout == 0.0 else 0)
  File "/usr/lib/python3.10/multiprocessing/popen_fork.py", line 27, in poll
    pid, sts = os.waitpid(self.pid, flag)
KeyboardInterrupt

Can reproduce using Python 3.9.10 and 3.10.4 on Linux:

import time
import threading
from multiprocessing import Process, Queue, current_process

def logger_thread(q: Queue):
    while True:
        record = q.get()
        if record is None:
            break
        print("logrecord: ", record)

def worker(q):
    q.put(current_process().name)
    return

logq = Queue()

lp = threading.Thread(target=logger_thread, args=(logq,), daemon=True)
lp.start()

workers = []
for i in range(15):
    p = Process(target=worker, args=(logq,), name=f"Worker {i+1}")
    workers.append(p)

print("starting workers")

for p in workers:
    p.start()
    # no deadlock when added:
    # time.sleep(.1)

print("waiting a bit")
time.sleep(1)
print("trying to join workers")

for w in workers:
    print(f'trying to join on {w.name}, alive={w.is_alive()}, exitcode={w.exitcode}', w.name, w.is_alive(), w.exitcode)
    w.join()
    print(f'joined on {w.name}', w.name)

logq.put(None)
lp.join()

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by running the supplied multiprocessing reproduction on Linux with the reported Python versions, then inspect the multiprocessing Queue and child-exit path, including the multiprocessing/process.py and multiprocessing/popen_fork.py paths shown in the traceback. Done means the workers reliably join after queue items are retrieved, with regression coverage for the reproduced case.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
distributed-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.