pytest-dev / pytest-dev/execnet
safe_terminate() leaves the abandoned termfunc thread running and reports success anyway
Nessuno ha ancora preso questa issue.
- Lingua principale
- Python
- Stelle
- 102
- Fork
- 47
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Descrizione
safe_terminate() abandons the termfunc thread it spawned and then reports success while it is still running.
Versions: execnet 2.1.2, CPython 3.12, Linux. Reached via pytest-xdist 3.8.0 (NodeManager.teardown_nodes → Group.terminate(EXIT_TIMEOUT=10)), but the defect is in safe_terminate itself.
The code
# src/execnet/multi.py:331
def safe_terminate(execmodel, timeout, list_of_paired_functions) -> None:
workerpool = WorkerPool(execmodel)
def termkill(termfunc, killfunc) -> None:
termreply = workerpool.spawn(termfunc)
try:
termreply.get(timeout=timeout)
except OSError:
killfunc()
replylist = []
for termfunc, killfunc in list_of_paired_functions:
reply = workerpool.spawn(termkill, termfunc, killfunc)
replylist.append(reply)
for reply in replylist:
reply.get()
workerpool.waitall(timeout=timeout)
Two problems:
-
The abandoned
termfuncthread is never cancelled. Whentermreply.get(timeout=timeout)raisesOSError,termkillcallskillfunc()and returns — but the thread runningtermfuncis still blocked and keeps running. ForGroup.terminate()thattermfuncisjoin_wait, i.e.gw.join(); gw._io.wait(), so it sits insubprocess.Popen.wait().killfunc(gw._io.kill()) is issued, but nothing waits for it to take effect and nothing joins the thread. -
waitall's timeout result is discarded.WorkerPool.waitall()returnsbool(gateway_base.py:470,return my_waitall_event.wait(timeout=timeout)). Line 349 ignores it, sosafe_terminatereturns normally — and thereforeGroup.terminate()returns normally — even when the abandoned threads from (1) are demonstrably still running.
Relatedly, Group.terminate(timeout) has no overall deadline: reply.get() on line 348 has no timeout at all, so the call can block indefinitely if a killfunc does.
Reproducer
"""safe_terminate() reports success while the termfunc thread it spawned is still running."""
import sys, threading, time
import execnet
from execnet.gateway_base import get_execmodel
from execnet.multi import safe_terminate
print("execnet", execnet.__version__, "| python", sys.version.split()[0])
execmodel = get_execmodel("thread")
entered, release = threading.Event(), threading.Event()
finished = []
def termfunc(): # stands in for join_wait: gw.join(); gw._io.wait()
entered.set()
release.wait(60)
finished.append(True)
def killfunc(): # stands in for kill: gw._io.kill()
pass # a kill that does not immediately reap the child
t0 = time.time()
safe_terminate(execmodel, 1.0, [(termfunc, killfunc)])
elapsed = time.time() - t0
print(f"safe_terminate(timeout=1.0) returned after {elapsed:.1f}s")
print(f" termfunc entered : {entered.is_set()}")
print(f" termfunc finished: {bool(finished)} <-- still running, never cancelled")
release.set()
execnet 2.1.2 | python 3.12.13
safe_terminate(timeout=1.0) returned after 2.0s
termfunc entered : True
termfunc finished: False <-- still running, never cancelled
killfunc is a no-op here to stand in for a kill that does not reap the child promptly; with a real popen gateway the same state is reached whenever Popen.wait() has not returned by the time waitall's timeout expires.
Why it matters
These threads are started by ThreadExecModel.start via _thread.start_new_thread (gateway_base.py:152), so they are not threading.Threads and threading._shutdown() never joins them. Once safe_terminate has returned, the caller is free to finish and the interpreter to finalize while those threads are still executing Python — they are then stopped at an arbitrary GIL checkpoint.
We hit this under pytest-xdist: teardown routinely exceeds xdist's 10s EXIT_TIMEOUT, so Group.terminate() takes the kill path on essentially every run, and faulthandler dumps taken in that window consistently show the leaked threads:
Thread ...: File ".../execnet/multi.py", line 231 in join_wait # x4, in subprocess.wait()
Thread ...: File ".../execnet/multi.py", line 339 in termkill # x3, in Reply.waitfinish
Thread ...: File ".../execnet/multi.py", line 348 in safe_terminate # main, unbounded reply.get()
File ".../execnet/multi.py", line 237 in terminate
File ".../xdist/workermanage.py", line 117 in teardown_nodes
On one such run the process died with SIGSEGV in exactly this window, with the truncated dump ending inside Gateway.__repr__ (reached from kill() at multi.py:234) on a frame whose code object could no longer be resolved. I am not claiming execnet segfaults — a faulthandler.dump_traceback_later watchdog was walking those stacks concurrently and that race is a plausible cause on its own. But the leaked threads are what put live Python into that window in the first place.
Suggested directions
- After
killfunc(), wait ontermreplyagain with a short grace period so the abandoned thread is actually observed to finish. - Propagate the
waitalltimeout instead of discarding it, soGroup.terminate()can tell its caller that gateways did not come down. - Give
Group.terminate(timeout)a single overall deadline rather than the currenttimeoutat line 339 + unbounded wait at 348 +timeoutat 349.
Happy to send a PR if you have a preference on the shape.
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Direzione di ricerca
Leggi safe_terminate e Group.terminate in src/execnet/multi.py, quindi esamina WorkerPool.waitall in gateway_base.py e il percorso di avvio del thread in ThreadExecModel.start. Riproduci prima il percorso di timeout; done dovrebbe significare che il lavoro termfunc abbandonato viene contabilizzato, che il risultato di waitall non viene scartato silenziosamente e che la terminazione non attende indefinitamente.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- python
- Ambito
- distributed-systems
- Tipo di issue
- Bug
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Stato di attività
- Attiva
- Chiarezza
- Abbastanza chiara
- Idoneità per principianti
- 45/100