Data race on the GC debug flag (gc.set_debug/get_debug) in free-threading builds
還沒有人認領這個 Issue。
- 主要語言
- Python
- 星號
- 77.2k
- 分支
- 36k
- PR 合併指標
- PR 指標待擷取
描述
Bug report
In a free-threading build (--disable-gil), the garbage collector debug flag
gcstate->debug is read and written without synchronisation:
- Write:
gc_set_debug_impl()(Modules/gcmodule.c) does a plain
gcstate->debug = flags. Unlikegc_set_threshold_impl(), which runs its
free-threading branch under_PyEval_StopTheWorld(),gc.set_debug()takes
no lock and does not stop the world. - Read:
gc_get_debug_impl()returnsgcstate->debugdirectly, and the
collector inPython/gc_free_threading.creadsgcstate->debug/
interp->gc.debugin several places while walking the graph.
So one thread calling gc.set_debug() concurrently with another calling
gc.get_debug() or triggering a collection is an unsynchronised read/write of
the same int. The flag is only an int, so the effect stays benign at the
Python level, but it is undefined behaviour under C11 and ThreadSanitizer
reports it as a data race.
This is the same class of issue already fixed for sys dlopenflags
(gh-151644) and gc.get_stats() (gh-151646). gc.enable() / gc.disable()
in the same file already access gcstate->enabled atomically; the debug flag
was missed.
ThreadSanitizer output
Built with ./configure --with-thread-sanitizer --disable-gil and stressed
with concurrent gc.set_debug() / gc.get_debug() plus a thread churning
cyclic garbage so the collector runs:
WARNING: ThreadSanitizer: data race
Write of size 4 at 0x...6c by thread T2:
#0 gc_set_debug gcmodule.c.h:186
Previous write of size 4 at 0x...6c by thread T1:
#0 gc_set_debug gcmodule.c.h:186
Location is global '_PyRuntime'
How to reproduce
./configure --with-thread-sanitizer --disable-gil && make- Run a script that starts a few threads calling
gc.set_debug(...)/
gc.get_debug()in a loop, plus a thread that builds reference cycles and
callsgc.collect(). - TSan reports the write/write race on
gcstate->debug.
Suggested fix
Access the flag with FT_ATOMIC_STORE_INT_RELAXED / FT_ATOMIC_LOAD_INT_RELAXED
in gc_set_debug_impl() / gc_get_debug_impl() and in the collector reads,
matching how gcstate->enabled and dlopenflags are already handled. Relaxed
ordering is correct for an independent int flag. These wrappers compile to a
plain load/store in the default (GIL) build, so there is no change there.
Linked PRs
- gh-153015
貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
研究方向
先檢查連結的 PR gh-153015,然後檢視 Modules/gcmodule.c 中的 gc_set_debug_impl() 和 gc_get_debug_impl(),以及 Python/gc_free_threading.c 中對 gcstate-debug 的讀取。使用 --with-thread-sanitizer --disable-gil 重新建置,並執行 gc.set_debug()、gc.get_debug() 和收集操作的並行重現程式;當回報的 race 不再出現時即表示完成。
由索引模型根據 Issue 內容生成。
評估
- 技術堆疊
- c, python
- 領域
- backend
- Issue 類型
- 缺陷
- 難度
- 3/5
- 預估耗時
- 1-2 天
- 活躍度
- 停滯
- 描述清晰度
- 描述清楚
- 新手友好度
- 25/100