Race condition in `itertools.islice` under free-threading
Dieses Issue hat noch niemand übernommen.
- Vorherrschende Sprache
- Python
- Sterne
- 77.2k
- Forks
- 35.9k
- PR-Merge-Kennzahlen
- PR-Kennzahlen ausstehend
Beschreibung
Bug report
Bug description:
islice_next reads and writes three fields -- lz->cnt, lz->next, and lz->it -- and do operations on them without a critical section.
Two threads calling next on the same islice object concurrently can race:
- Both read
lz->cnt = N < lz->next, both execute the skip-loop body for the same slot, advancing the underlying iterator twice for one step. In that case, the step arithmetic gets corrupted. - Both read
lz->cnt = N < stop, both pass the stop check, both calliternext(it)and return an item -- the total number of yielded items exceedsstop. - There's a race on
lz->cnt++andlz->next += step: both threads can read the same value, then both can increment locally, and both store the same result. After that, the counter is permanently incorrect (such that subsequent skip/stop decisions use a wrong baseline). - One of the threads reaches exhaustion (
goto empty) and callsPy_CLEAR(lz->it), which setslz->it = NULLand DECREF's the iterator, which frees it (since each thread has readyit = lz->it. and not INCREF'd it). Another thread can have readlz->itbefore the clear and be in the middle ofiternext(it)on the now-freed object in which case there is a use-after-free.
chain_next for example wraps its body in Py_BEGIN_CRITICAL_SECTION(op), which I believe is what islice should do as well.
Reproducer
import itertools
import threading
STOP = 100
NTHREADS = 8
data = iter(range(STOP + NTHREADS * 2))
sl = itertools.islice(data, STOP)
results: list[int] = []
lock = threading.Lock()
def consume() -> None:
while True:
v = next(sl, None)
if v is None:
break
with lock:
results.append(v)
threads = [threading.Thread(target=consume) for _ in range(NTHREADS)]
for t in threads: t.start()
for t in threads: t.join()
The reproducer, on a free-threaded build, gets reports like this one.
WARNING: ThreadSanitizer: data race (pid=49886)
Read of size 8 at 0x00030275f8b0 by thread T2:
#0 islice_next itertoolsmodule.c:1663 (python.exe:arm64+0x10042bce4)
#1 builtin_next bltinmodule.c:1770 (python.exe:arm64+0x10028a900)
#2 cfunction_vectorcall_FASTCALL methodobject.c:449 (python.exe:arm64+0x1001379dc)
#3 _PyObject_VectorcallTstate pycore_call.h:144 (python.exe:arm64+0x100090e80)
#4 PyObject_Vectorcall call.c:327 (python.exe:arm64+0x100090e80)
...
CPython versions tested on:
CPython main branch
Operating systems tested on:
macOS
Linked PRs
- gh-151410
Beitragsleitfaden
Erste Schritte
- Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
- Forke das Repository und arbeite in einem Branch.
- Öffne einen Pull Request, der die Issue-Nummer nennt.
Rechercherichtung
Beginne in Modules/itertoolsmodule.c bei islice_next und vergleiche die Behandlung des kritischen Abschnitts von chain_next. Führe den mitgelieferten Multithread-Reproducer auf einem freethreaded CPython-Build aus und untersuche die relevanten itertools-Tests. Die Aufgabe ist erledigt, wenn gleichzeitige Aufrufe von next() nicht mehr die gemeldete Race Condition auslösen oder die stop-Grenze von islice überschreiten.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- python
- Bereich
- backend
- Issue-Typ
- Bug
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Aktivitätsstatus
- Veraltet
- Klarheit
- Klar beschrieben
- Anfängerfreundlichkeit
- 35/100