agent-substrate / agent-substrate/substrate

Stalled SUSPENDING actors have no system-driven recovery

未關閉
#1,502 0 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
area/api area/api-machinery area/reliability kind/bug
主要語言
Go
星號
1.8k
分支
316
平均合併
2 天 43 分鐘
30 天內合併 PR
287

描述

Split out of #791 (the deferred "janitor / retry budget" item) so #791 can close.

### Problem:
If `SuspendActor` fails after the actor is marked `SUSPENDING`, the actor stays `SUSPENDING` until a client calls `SuspendActor` again. In that state it cannot be resumed.

A running-origin suspend at least has a graceful termination exit: if the worker pod goes away, `DeleteWorker` marks the actor `CRASHED`.

A paused-origin is pinned to a node name, and once the node is gone every retry fails with just internal errors.

Xref: #660 #791 #817 #798

貢獻指南

開啟貢獻指南

評估

這個 Issue 還沒有評估資料。

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。