python / python/cpython

Data race on the GC debug flag (gc.set_debug/get_debug) in free-threading builds

Open
#153,014 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

extension-modules topic-free-threading type-bug
Dominant language
Python
Stars
77.2k
Forks
35.9k
PR merge metrics
PR metrics pending

Description

Bug report

In a free-threading build (--disable-gil), the garbage collector debug flag
gcstate->debug is read and written without synchronisation:

  • Write: gc_set_debug_impl() (Modules/gcmodule.c) does a plain
    gcstate->debug = flags. Unlike gc_set_threshold_impl(), which runs its
    free-threading branch under _PyEval_StopTheWorld(), gc.set_debug() takes
    no lock and does not stop the world.
  • Read: gc_get_debug_impl() returns gcstate->debug directly, and the
    collector in Python/gc_free_threading.c reads gcstate->debug /
    interp->gc.debug in several places while walking the graph.

So one thread calling gc.set_debug() concurrently with another calling
gc.get_debug() or triggering a collection is an unsynchronised read/write of
the same int. The flag is only an int, so the effect stays benign at the
Python level, but it is undefined behaviour under C11 and ThreadSanitizer
reports it as a data race.

This is the same class of issue already fixed for sys dlopenflags
(gh-151644) and gc.get_stats() (gh-151646). gc.enable() / gc.disable()
in the same file already access gcstate->enabled atomically; the debug flag
was missed.

ThreadSanitizer output

Built with ./configure --with-thread-sanitizer --disable-gil and stressed
with concurrent gc.set_debug() / gc.get_debug() plus a thread churning
cyclic garbage so the collector runs:

WARNING: ThreadSanitizer: data race
  Write of size 4 at 0x...6c by thread T2:
    #0 gc_set_debug gcmodule.c.h:186
  Previous write of size 4 at 0x...6c by thread T1:
    #0 gc_set_debug gcmodule.c.h:186
  Location is global '_PyRuntime'

How to reproduce

  1. ./configure --with-thread-sanitizer --disable-gil && make
  2. Run a script that starts a few threads calling gc.set_debug(...) /
    gc.get_debug() in a loop, plus a thread that builds reference cycles and
    calls gc.collect().
  3. TSan reports the write/write race on gcstate->debug.

Suggested fix

Access the flag with FT_ATOMIC_STORE_INT_RELAXED / FT_ATOMIC_LOAD_INT_RELAXED
in gc_set_debug_impl() / gc_get_debug_impl() and in the collector reads,
matching how gcstate->enabled and dlopenflags are already handled. Relaxed
ordering is correct for an independent int flag. These wrappers compile to a
plain load/store in the default (GIL) build, so there is no change there.

Linked PRs
  • gh-153015

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Review linked PR gh-153015 first, then inspect gc_set_debug_impl() and gc_get_debug_impl() in Modules/gcmodule.c and the gcstate-debug reads in Python/gc_free_threading.c. Rebuild with --with-thread-sanitizer --disable-gil and run the concurrent gc.set_debug(), gc.get_debug(), and collection reproducer; done means the reported race is absent.

Written by the indexing model from the issue text.

Assessment

Tech stack
c, python
Domain
backend
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Clearly specified
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.