python / python/cpython

Data race reading `_thread.RLock` recursion count in `repr()` under free-threading

オープン
#154,928 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

extension-modules type-bug
主要言語
Python
スター
77.2k
フォーク
35.9k
PR マージ指標
PR 指標を取得中

説明

Bug report

Bug description:

This is a follow-up to #153292 (data race in repr() of _thread.RLock), which was fixed by making rlock_repr read the lock's owner with an atomic load. The fix covered the owner (self->lock.thread) field, but rlock_repr still reads self->lock.level with a plain (non-atomic) load to compute the recursion count:

https://github.com/python/cpython/blob/22a6c51c94a4fde986b8964f1d36d5ec3ac20dcc/Modules/_threadmodule.c#L1287-L1305

self->lock.level is written by acquire / release / _acquire_restore, e.g.:

https://github.com/python/cpython/blob/22a6c51c94a4fde986b8964f1d36d5ec3ac20dcc/Modules/_threadmodule.c#L1198-L1210

So on a free-threaded build, repr(rlock) concurrent with acquire() or release() is still a data race, now on the level field rather than the owner field the earlier fix addressed.

Reproducer:

import _thread
from threading import Thread, Barrier

shared_rlock = _thread.RLock()
state = (5, _thread.get_ident())

def chain1_thread():
    for _ in range(20000):
        try:
            shared_rlock._acquire_restore(state)
        except Exception:
            pass

def chain2_thread():
    for _ in range(20000):
        try:
            repr(shared_rlock)
        except Exception:
            pass

N_C1 = 2
N_C2 = 4
barrier = Barrier(N_C1 + N_C2)

def _c1():
    barrier.wait()
    chain1_thread()

def _c2():
    barrier.wait()
    chain2_thread()

threads  = [Thread(target=_c1) for _ in range(N_C1)]
threads += [Thread(target=_c2) for _ in range(N_C2)]
for t in threads: t.start()
for t in threads: t.join()

TSAN Report :

==================
WARNING: ThreadSanitizer: data race (pid=655652)
  Read of size 8 at 0x7fffb6610540 by thread T6:
    #0 rlock_repr /cpython/./Modules/_threadmodule.c:1295:28 
    #1 PyObject_Repr /cpython/Objects/object.c:784:11 
    #2 builtin_repr /cpython/Python/bltinmodule.c:2677:12 
    #3 _PyEval_EvalFrameDefault /cpython/Python/generated_cases.c.h:2712:35  

  Previous write of size 8 at 0x7fffb6610540 by thread T1:
    #0 _thread_RLock__acquire_restore_impl /cpython/./Modules/_threadmodule.c:1210:22 
    #1 _thread_RLock__acquire_restore /cpython/./Modules/clinic/_threadmodule.c.h:537:20 
    #2 method_vectorcall_O /cpython/Objects/descrobject.c:476:24 
    #3 _PyObject_VectorcallTstate /cpython/./Include/internal/pycore_call.h:144:11 
    #4 PyObject_Vectorcall /cpython/Objects/call.c:327:12 
    #5 _Py_VectorCallInstrumentation_StackRefSteal /cpython/Python/ceval.c:768:11 
    #6 _PyEval_EvalFrameDefault /cpython/Python/generated_cases.c.h:1906:35

SUMMARY: ThreadSanitizer: data race /cpython/./Modules/_threadmodule.c:1295:28 in rlock_repr
==================
CPython versions tested on:

CPython main branch

Operating systems tested on:

Linux

Linked PRs
  • gh-155381

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

調査の方向性

Modules/_threadmodule.c の rlock_repr と、レポートで説明されている acquire、release、_acquire_restore の実装から始めます。ThreadSanitizer を使って free-threaded ビルドで提供された reproducer を実行し、その後関連するテストを調べ、repr() の同時実行と lock の更新で race が報告されなくなったことを確認します。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
c, python
領域
operating-systems
issue の種類
バグ
難易度
4/5
見積もり時間
3〜5日
活発さ
停滞
明瞭さ
明確に書かれている
初心者へのやさしさ
35/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。