agent-substrate / agent-substrate/substrate

[Bug]: Lifecycle finalizers can overwrite a concurrent worker-deletion crash

未關閉
#1,443 0 則留言 0 個 reaction 已指派 1 人 已被 @EItanya 認領 在 GitHub 檢視
area/api-machinery area/reliability kind/bug
主要語言
Go
星號
1.8k
分支
316
平均合併
2 天 43 分鐘
30 天內合併 PR
287

描述

### What happened?

When a worker disappears during ResumeActor or SuspendActor, DeleteWorker marks the actor
CRASHED and clears its worker assignment. The lifecycle workflow can subsequently re-read that
updated actor and unconditionally finalize it:

- Resume changes CRASHED to RUNNING, leaving a running actor without a worker.
- Suspend changes CRASHED to SUSPENDED, potentially without a valid snapshot.

Optimistic version checks do not prevent this because the finalizers read and update the new
CRASHED version.

### Expected Behavior

finalization should require the actor to remain RESUMING or SUSPENDING. If worker deletion has already crashed it, the workflow should preserve CRASHED and return FailedPrecondition.

### Steps to Reproduce

Regression tests can reproduce both cases by deleting the worker from the fake atelet
immediately after restore/checkpoint completes.

### Sandbox Runtime

Both / Runtime Agnostic

### Agent Substrate Version / Commit SHA

main

### Kubernetes Version & Environment

_No response_

### Host OS & Architecture

_No response_

### Relevant Logs and Diagnostic Output

```shell

```

### Additional Context

_No response_

### Confirmation

- [x] I have searched existing issues and verified that this is not a duplicate.
- [x] I have verified that this issue occurs on the latest commit on `main`.

貢獻指南

開啟貢獻指南

評估

這個 Issue 還沒有評估資料。

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。