agent-substrate / agent-substrate/substrate

atelet: a node that never hosted a golden can't serve a first restore

未關閉
#1,238 9 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
area/gvisor area/node kind/bug
主要語言
Go
星號
1.8k
分支
316
平均合併
2 天 43 分鐘
30 天內合併 PR
287

描述

**Measured** on GKE 3-node, `1.35.6-gke.1710000`, gVisor, demos/counter. 3 workers on 3 nodes with one `ActorTemplate`, found that on a node that never ran a golden, `static-files` is **absent entirely**. As a result, resumes actor on workers of other 2 nodes cause livelock. Tested this hypothesis by moving all workers to the warm node gives HTTP 200 in ~1 s; then cordoning it so a new template's golden lands on a cold node flips that node to warm and it resumes instantly — same node, cache alone.
e2e misses it because every suite creates its own template, hence its own golden and a warm node, before resuming anything — and single-node kind shares one cache across all workers.

**The actor is also unreapable.** After minutes in `ACTOR_STATE_RESUMING`, `suspend` returns `MarkSuspending prerequisite not met (got: ACTOR_STATE_RESUMING, want ACTOR_STATE_RUNNING or ACTOR_STATE_PAUSED)` and `delete` returns `not in a deletable state`. The only move left is deleting the worker pod, which makes the actor `CRASHED`.

貢獻指南

開啟貢獻指南

評估

這個 Issue 還沒有評估資料。

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。