python / python/cpython

Thread.start() can hang indefinitely if the new thread fails (MemoryError) during its initialization

未關閉
#140,746 10 則留言 4 個 reaction 已指派 0 人 在 GitHub 檢視

還沒有人認領這個 Issue。

stdlib type-bug
主要語言
Python
星號
77.2k
分支
35.9k
PR 合併指標
PR 指標待擷取

描述

Bug report

Bug description:

There is a race condition in the threading module where a parent thread calling Thread.start() can wait forever if the newly created thread crashes with a MemoryError during its internal bootstrap process.

Case Explanation

In case we have a "Serving Thread" that creates threads on demand (e.g., for each HTTP request):

  • When this Serving Thread calls Thread.start(), the OS-level thread (pthread_create()^1 - Linux) is successfully created. The parent thread then waits for the new thread to signal that it has started correctly by calling self._started.wait()^2.
  • The new thread starts, but before it can signal the parent thread (Serving Thread) that it is alive with self._started.set()^3 it encounters a MemoryError.
  • This MemoryError can occur at the C level during the PyObject_Call to the _bootstrap method or inside the _bootstrap_inner method before _started.set() is reached, often due to memory pressure from other threads (or a heap limit being reached).
  • This exception is caught by the C-level entry point thread_run(), which calls _PyErr_WriteUnraisableMsg^4 and prints "Exception ignored in thread started by: ..."

The new thread then exits without ever signaling the _started event, and the parent thread waits indefinitely on the _started.wait().
This also leaves the threading module in an inconsistent state, as the "zombie" thread object may not be correctly cleaned up from the _limbo dict.

How to Reproduce

This has been observed in high-concurrency server applications under heavy, sustained load, where heap memory can be rapidly consumed and exhausted by concurrent threads ^5.
This is a race condition that is difficult to reproduce reliably, as it requires triggering a MemoryError at a specific moment.

I found a (deterministic?) way to reproduce the issue by restricting the heap memory until it reaches a threshold where we can start a new thread, but this new thread won't get enough memory for its initialization.

On some machines (and depending of Python versions), it is sometimes necessary to tweak HARD_LIMIT_START / LIMIT_REDUCTION (I reproduced it on Ubuntu based machine with Python 3.11/3.12/3.13/3.14)

import resource
import threading
import gc

def handler():
    pass

def serving():
    # These should be tweak (depending of Python version + system)
    HARD_LIMIT_START = 30_000_000
    LIMIT_REDUCTION = 5_000

    for _ in range(500_000):
        gc.collect(2)  # Force getting back memory: seems to increase the determinism of the script

        # Limit the heap size available for this process
        resource.setrlimit(resource.RLIMIT_DATA, (HARD_LIMIT_START, HARD_LIMIT_START * 2))
        try:
            handler_thread = threading.Thread(target=handler)
            print(f'Start Thread: {handler_thread} - Heap size limit : {HARD_LIMIT_START}')
            handler_thread.start()
            handler_thread.join()
            HARD_LIMIT_START -= LIMIT_REDUCTION
        except RuntimeError as r:  # If Python refused to launch a new Thread
            print(f'RuntimeError: {r} - Cannot start the thread at all => error not detected.')
            return

serving_thread = threading.Thread(target=serving)
serving_thread.start()
serving_thread.join()
Expected Behavior

I am not sure if this is an "accepted" limitation of (CPython) Thread or not. IMO, Thread.start() shouldn't hang indefinitely if the low-level thread is dead.

I didn't take the time to try to fix it yet (if possible). I would prefer to get your opinions on this first.

CPython versions tested on:

3.12, 3.13, 3.14

Operating systems tested on:

Linux

Linked PRs
  • gh-140799
  • gh-144750
  • gh-153776

貢獻指南

開啟貢獻指南

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

研究方向

從 Lib/threading.py 中 Thread.start() 和 _bootstrap_inner 附近開始,接著檢查 Modules/_threadmodule.c 中參照的進入點,以及 Python/thread_pthread.h 中的 pthread 實作。在列出的 CPython 版本上執行隨附的 Linux 資源限制重現程式。當初始化失敗時 Thread.start() 不再無限期等待,且針對失敗路徑有回歸測試涵蓋時,即表示完成。

由索引模型根據 Issue 內容生成。

評估

技術堆疊
python
領域
operating-systems
Issue 類型
缺陷
難度
4/5
預估耗時
3-5 天
活躍度
停滯
描述清晰度
基本清楚
新手友好度
25/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。