pytest-dev / pytest-dev/execnet
safe_terminate() leaves the abandoned termfunc thread running and reports success anyway
还没有人认领这个 Issue。
- 主要语言
- Python
- 星标
- 102
- 派生
- 47
- PR 合并指标
- 30 天内没有已合并 PR
描述
safe_terminate() abandons the termfunc thread it spawned and then reports success while it is still running.
Versions: execnet 2.1.2, CPython 3.12, Linux. Reached via pytest-xdist 3.8.0 (NodeManager.teardown_nodes → Group.terminate(EXIT_TIMEOUT=10)), but the defect is in safe_terminate itself.
The code
# src/execnet/multi.py:331
def safe_terminate(execmodel, timeout, list_of_paired_functions) -> None:
workerpool = WorkerPool(execmodel)
def termkill(termfunc, killfunc) -> None:
termreply = workerpool.spawn(termfunc)
try:
termreply.get(timeout=timeout)
except OSError:
killfunc()
replylist = []
for termfunc, killfunc in list_of_paired_functions:
reply = workerpool.spawn(termkill, termfunc, killfunc)
replylist.append(reply)
for reply in replylist:
reply.get()
workerpool.waitall(timeout=timeout)
Two problems:
-
The abandoned
termfuncthread is never cancelled. Whentermreply.get(timeout=timeout)raisesOSError,termkillcallskillfunc()and returns — but the thread runningtermfuncis still blocked and keeps running. ForGroup.terminate()thattermfuncisjoin_wait, i.e.gw.join(); gw._io.wait(), so it sits insubprocess.Popen.wait().killfunc(gw._io.kill()) is issued, but nothing waits for it to take effect and nothing joins the thread. -
waitall's timeout result is discarded.WorkerPool.waitall()returnsbool(gateway_base.py:470,return my_waitall_event.wait(timeout=timeout)). Line 349 ignores it, sosafe_terminatereturns normally — and thereforeGroup.terminate()returns normally — even when the abandoned threads from (1) are demonstrably still running.
Relatedly, Group.terminate(timeout) has no overall deadline: reply.get() on line 348 has no timeout at all, so the call can block indefinitely if a killfunc does.
Reproducer
"""safe_terminate() reports success while the termfunc thread it spawned is still running."""
import sys, threading, time
import execnet
from execnet.gateway_base import get_execmodel
from execnet.multi import safe_terminate
print("execnet", execnet.__version__, "| python", sys.version.split()[0])
execmodel = get_execmodel("thread")
entered, release = threading.Event(), threading.Event()
finished = []
def termfunc(): # stands in for join_wait: gw.join(); gw._io.wait()
entered.set()
release.wait(60)
finished.append(True)
def killfunc(): # stands in for kill: gw._io.kill()
pass # a kill that does not immediately reap the child
t0 = time.time()
safe_terminate(execmodel, 1.0, [(termfunc, killfunc)])
elapsed = time.time() - t0
print(f"safe_terminate(timeout=1.0) returned after {elapsed:.1f}s")
print(f" termfunc entered : {entered.is_set()}")
print(f" termfunc finished: {bool(finished)} <-- still running, never cancelled")
release.set()
execnet 2.1.2 | python 3.12.13
safe_terminate(timeout=1.0) returned after 2.0s
termfunc entered : True
termfunc finished: False <-- still running, never cancelled
killfunc is a no-op here to stand in for a kill that does not reap the child promptly; with a real popen gateway the same state is reached whenever Popen.wait() has not returned by the time waitall's timeout expires.
Why it matters
These threads are started by ThreadExecModel.start via _thread.start_new_thread (gateway_base.py:152), so they are not threading.Threads and threading._shutdown() never joins them. Once safe_terminate has returned, the caller is free to finish and the interpreter to finalize while those threads are still executing Python — they are then stopped at an arbitrary GIL checkpoint.
We hit this under pytest-xdist: teardown routinely exceeds xdist's 10s EXIT_TIMEOUT, so Group.terminate() takes the kill path on essentially every run, and faulthandler dumps taken in that window consistently show the leaked threads:
Thread ...: File ".../execnet/multi.py", line 231 in join_wait # x4, in subprocess.wait()
Thread ...: File ".../execnet/multi.py", line 339 in termkill # x3, in Reply.waitfinish
Thread ...: File ".../execnet/multi.py", line 348 in safe_terminate # main, unbounded reply.get()
File ".../execnet/multi.py", line 237 in terminate
File ".../xdist/workermanage.py", line 117 in teardown_nodes
On one such run the process died with SIGSEGV in exactly this window, with the truncated dump ending inside Gateway.__repr__ (reached from kill() at multi.py:234) on a frame whose code object could no longer be resolved. I am not claiming execnet segfaults — a faulthandler.dump_traceback_later watchdog was walking those stacks concurrently and that race is a plausible cause on its own. But the leaked threads are what put live Python into that window in the first place.
Suggested directions
- After
killfunc(), wait ontermreplyagain with a short grace period so the abandoned thread is actually observed to finish. - Propagate the
waitalltimeout instead of discarding it, soGroup.terminate()can tell its caller that gateways did not come down. - Give
Group.terminate(timeout)a single overall deadline rather than the currenttimeoutat line 339 + unbounded wait at 348 +timeoutat 349.
Happy to send a PR if you have a preference on the shape.
贡献指南
这个仓库没有索引到贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
调研方向
阅读 src/execnet/multi.py 中的 safe_terminate 和 Group.terminate,然后检查 gateway_base.py 中的 WorkerPool.waitall 以及 ThreadExecModel.start 中的线程启动路径。先复现 timeout 路径;done 应表示已对被放弃的 termfunc 工作进行处理,waitall 的结果不会被默默丢弃,并且终止不会无限期等待。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- python
- 领域
- distributed-systems
- Issue 类型
- 缺陷
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 活跃度
- 活跃
- 描述清晰度
- 基本清楚
- 新手友好度
- 45/100