Data race on the GC debug flag (gc.set_debug/get_debug) in free-threading builds
还没有人认领这个 Issue。
- 主要语言
- Python
- 星标
- 77.2k
- 派生
- 35.9k
- PR 合并指标
- PR 指标待抓取
描述
Bug report
In a free-threading build (--disable-gil), the garbage collector debug flag
gcstate->debug is read and written without synchronisation:
- Write:
gc_set_debug_impl()(Modules/gcmodule.c) does a plain
gcstate->debug = flags. Unlikegc_set_threshold_impl(), which runs its
free-threading branch under_PyEval_StopTheWorld(),gc.set_debug()takes
no lock and does not stop the world. - Read:
gc_get_debug_impl()returnsgcstate->debugdirectly, and the
collector inPython/gc_free_threading.creadsgcstate->debug/
interp->gc.debugin several places while walking the graph.
So one thread calling gc.set_debug() concurrently with another calling
gc.get_debug() or triggering a collection is an unsynchronised read/write of
the same int. The flag is only an int, so the effect stays benign at the
Python level, but it is undefined behaviour under C11 and ThreadSanitizer
reports it as a data race.
This is the same class of issue already fixed for sys dlopenflags
(gh-151644) and gc.get_stats() (gh-151646). gc.enable() / gc.disable()
in the same file already access gcstate->enabled atomically; the debug flag
was missed.
ThreadSanitizer output
Built with ./configure --with-thread-sanitizer --disable-gil and stressed
with concurrent gc.set_debug() / gc.get_debug() plus a thread churning
cyclic garbage so the collector runs:
WARNING: ThreadSanitizer: data race
Write of size 4 at 0x...6c by thread T2:
#0 gc_set_debug gcmodule.c.h:186
Previous write of size 4 at 0x...6c by thread T1:
#0 gc_set_debug gcmodule.c.h:186
Location is global '_PyRuntime'
How to reproduce
./configure --with-thread-sanitizer --disable-gil && make- Run a script that starts a few threads calling
gc.set_debug(...)/
gc.get_debug()in a loop, plus a thread that builds reference cycles and
callsgc.collect(). - TSan reports the write/write race on
gcstate->debug.
Suggested fix
Access the flag with FT_ATOMIC_STORE_INT_RELAXED / FT_ATOMIC_LOAD_INT_RELAXED
in gc_set_debug_impl() / gc_get_debug_impl() and in the collector reads,
matching how gcstate->enabled and dlopenflags are already handled. Relaxed
ordering is correct for an independent int flag. These wrappers compile to a
plain load/store in the default (GIL) build, so there is no change there.
Linked PRs
- gh-153015
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
调研方向
先检查链接的 PR gh-153015,然后检查 Modules/gcmodule.c 中的 gc_set_debug_impl() 和 gc_get_debug_impl(),以及 Python/gc_free_threading.c 中对 gcstate-debug 的读取。使用 --with-thread-sanitizer --disable-gil 重新构建,并运行 gc.set_debug()、gc.get_debug() 和收集操作的并发复现程序;当报告的 race 不再出现时即表示完成。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- c, python
- 领域
- backend
- Issue 类型
- 缺陷
- 难度
- 3/5
- 预计耗时
- 1-2 天
- 活跃度
- 停滞
- 描述清晰度
- 描述清楚
- 新手友好度
- 25/100