agent-substrate / agent-substrate/substrate

[Bug]: Lifecycle finalizers can overwrite a concurrent worker-deletion crash

未关闭
#1,443 0 条评论 0 个 reaction 已指派 1 人 已被 @EItanya 认领 在 GitHub 查看
area/api-machinery area/reliability kind/bug
主要语言
Go
星标
1.8k
派生
316
平均合并
2 天 43 分钟
30 天内合并 PR
287

描述

### What happened?

When a worker disappears during ResumeActor or SuspendActor, DeleteWorker marks the actor
CRASHED and clears its worker assignment. The lifecycle workflow can subsequently re-read that
updated actor and unconditionally finalize it:

- Resume changes CRASHED to RUNNING, leaving a running actor without a worker.
- Suspend changes CRASHED to SUSPENDED, potentially without a valid snapshot.

Optimistic version checks do not prevent this because the finalizers read and update the new
CRASHED version.

### Expected Behavior

finalization should require the actor to remain RESUMING or SUSPENDING. If worker deletion has already crashed it, the workflow should preserve CRASHED and return FailedPrecondition.

### Steps to Reproduce

Regression tests can reproduce both cases by deleting the worker from the fake atelet
immediately after restore/checkpoint completes.

### Sandbox Runtime

Both / Runtime Agnostic

### Agent Substrate Version / Commit SHA

main

### Kubernetes Version & Environment

_No response_

### Host OS & Architecture

_No response_

### Relevant Logs and Diagnostic Output

```shell

```

### Additional Context

_No response_

### Confirmation

- [x] I have searched existing issues and verified that this is not a duplicate.
- [x] I have verified that this issue occurs on the latest commit on `main`.

贡献指南

打开贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。