agent-substrate / agent-substrate/substrate
Stalled SUSPENDING actors have no system-driven recovery
Ouverte
area/api
area/api-machinery
area/reliability
kind/bug
- Langage dominant
- Go
- Étoiles
- 1.8k
- Forks
- 316
- Merge moyen
- 2 j 43 min
- PR mergées (30 j)
- 287
Description
Split out of #791 (the deferred "janitor / retry budget" item) so #791 can close.
### Problem:
If `SuspendActor` fails after the actor is marked `SUSPENDING`, the actor stays `SUSPENDING` until a client calls `SuspendActor` again. In that state it cannot be resumed.
A running-origin suspend at least has a graceful termination exit: if the worker pod goes away, `DeleteWorker` marks the actor `CRASHED`.
A paused-origin is pinned to a node name, and once the node is gone every retry fails with just internal errors.
Xref: #660 #791 #817 #798
Guide de contribution
Ouvrir le guide de contribution
Évaluation
Cette issue n'a pas encore été évaluée.