agent-substrate / agent-substrate/substrate
atelet: a node that never hosted a golden can't serve a first restore
- 主要语言
- Go
- 星标
- 1.8k
- 派生
- 316
- 平均合并
- 2 天 43 分钟
- 30 天内合并 PR
- 287
描述
**Measured** on GKE 3-node, `1.35.6-gke.1710000`, gVisor, demos/counter. 3 workers on 3 nodes with one `ActorTemplate`, found that on a node that never ran a golden, `static-files` is **absent entirely**. As a result, resumes actor on workers of other 2 nodes cause livelock. Tested this hypothesis by moving all workers to the warm node gives HTTP 200 in ~1 s; then cordoning it so a new template's golden lands on a cold node flips that node to warm and it resumes instantly — same node, cache alone.
e2e misses it because every suite creates its own template, hence its own golden and a warm node, before resuming anything — and single-node kind shares one cache across all workers.
**The actor is also unreapable.** After minutes in `ACTOR_STATE_RESUMING`, `suspend` returns `MarkSuspending prerequisite not met (got: ACTOR_STATE_RESUMING, want ACTOR_STATE_RUNNING or ACTOR_STATE_PAUSED)` and `delete` returns `not in a deletable state`. The only move left is deleting the worker pod, which makes the actor `CRASHED`.
贡献指南
评估
这个 Issue 还没有评估数据。