Deadlock when shutting down ThreadPoolExecutor from inside OS Signal handler
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 77.2k
- Forks
- 35.9k
- PR merge metrics
- PR metrics pending
Description
Bug report
Bug description:
When running the code below and sending a SIGTERM to the process results in what looks like a deadlock on the ThreadPoolExecutor._shutdown_lock. I guess the issue is that the SIGTERM is being handled in main thread while in the submit function where the _shutdown_lock is already locked. When I try calling shutdown inside the signal handler it waits forever.
TBH I'm not sure if this is a bug or expected behavior but I wanted to file this anyways. Thanks!
from concurrent.futures import ThreadPoolExecutor
import signal
def main():
with ThreadPoolExecutor(max_workers=2) as executor:
def __exit(signal, frame):
executor.shutdown(wait=True, cancel_futures=True)
signal.signal(signal.SIGTERM, __exit)
while True:
executor.submit(lambda: 1 + 1)
if __name__ == "__main__":
main()
AFTER sending SIGTERM to the process, this is what pystack shows:
(v) rbost@fedora:~/code/CheckDuplicate$ pystack remote 3431749
Traceback for thread 3431751 (python) [] (most recent call last):
(Python) File "/usr/lib64/python3.12/threading.py", line 1030, in _bootstrap
self._bootstrap_inner()
(Python) File "/usr/lib64/python3.12/threading.py", line 1073, in _bootstrap_inner
self.run()
(Python) File "/usr/lib64/python3.12/threading.py", line 1010, in run
self._target(*self._args, **self._kwargs)
(Python) File "/usr/lib64/python3.12/concurrent/futures/thread.py", line 89, in _worker
work_item = work_queue.get(block=True)
Traceback for thread 3431750 (python) [] (most recent call last):
(Python) File "/usr/lib64/python3.12/threading.py", line 1030, in _bootstrap
self._bootstrap_inner()
(Python) File "/usr/lib64/python3.12/threading.py", line 1073, in _bootstrap_inner
self.run()
(Python) File "/usr/lib64/python3.12/threading.py", line 1010, in run
self._target(*self._args, **self._kwargs)
(Python) File "/usr/lib64/python3.12/concurrent/futures/thread.py", line 89, in _worker
work_item = work_queue.get(block=True)
Traceback for thread 3431749 (python) [] (most recent call last):
(Python) File "/home/rbost/code/CheckDuplicate/test.py", line 18, in <module>
main()
(Python) File "/home/rbost/code/CheckDuplicate/test.py", line 14, in main
executor.submit(lambda: 1 + 1)
(Python) File "/usr/lib64/python3.12/concurrent/futures/thread.py", line 175, in submit
f = _base.Future()
(Python) File "/usr/lib64/python3.12/concurrent/futures/_base.py", line 328, in __init__
def __init__(self):
(Python) File "/home/rbost/code/CheckDuplicate/test.py", line 9, in __exit
executor.shutdown(wait=True, cancel_futures=True)
(Python) File "/usr/lib64/python3.12/concurrent/futures/thread.py", line 220, in shutdown
with self._shutdown_lock:
CPython versions tested on:
3.12
Operating systems tested on:
Linux
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Reproduce the supplied SIGTERM example, then inspect concurrent/futures/thread.py around ThreadPoolExecutor.submit(), shutdown(), and _shutdown_lock. The issue does not define the expected signal-handler behavior or name a test; progress requires agreeing on that behavior and adding a regression test that demonstrates completion without the reported deadlock.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- backend, operating-systems
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100