agent-substrate / agent-substrate/substrate

Pause finalization can record an empty node name in NodeVmsWithLocalSnapshots, leaving the actor permanently unschedulable

Abierto
#397 0 comentarios 0 reacciones 0 asignados Ver en GitHub
area/api area/api-machinery kind/bug
Lenguaje dominante
Go
Estrellas
1.8k
Forks
316
Merge medio
2 d 43 min
PR fusionados (30 d)
287

Descripción

### Summary

When a `PauseActor` workflow finalizes after the assigned worker record has already been deleted (e.g. the worker pod was removed while the pause was in flight), `FinalizePausedStep` cannot determine the node that holds the local checkpoint. It logs a warning ("Node name not found during finalize pause") but still writes the unknown node name into the actor record:

```
LatestSnapshotInfo.Local.NodeVmsWithLocalSnapshots = [""]
```

### Impact

On the next `ResumeActor`, `AssignWorkerStep.findFreeWorker` treats a non-empty `NodeVmsWithLocalSnapshots` list as a hard placement restriction: a candidate worker must have `NodeName` equal to one of the listed entries. No worker ever has an empty node name, so **no worker can ever match** and every resume of the affected actor fails with:

```
FailedPrecondition: no free workers available
```

The actor is permanently unschedulable. The only escape is `ResumeActor(boot=true)`, which discards the actor's checkpointed state.

The generic error message also makes this very hard to diagnose — it looks like a capacity problem, not a poisoned locality record.

### Proposed fix

Two guards, both small:

1. **Write side** (`FinalizePausedStep`): only append the node name to `NodeVmsWithLocalSnapshots` when it is non-empty. If the node is unknown, record the local snapshot with an empty node list — the snapshot still exists on some node, and an empty list lets the scheduler place the actor anywhere (a shuffled retry can still land on the node that holds the checkpoint, since local checkpoints live on the node, not in the deleted worker pod).

2. **Read side** (`findFreeWorker`): ignore empty-string entries in the restriction list. This heals actor records already poisoned by (1) without requiring a data migration.

### Notes

This is a targeted fix. The underlying modeling issue — snapshot placement being a write-once field on the Actor with no feedback from the nodes that actually hold the snapshots — is a larger topic being handled separately.

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.