agent-substrate / agent-substrate/substrate

Pause finalization can record an empty node name in NodeVmsWithLocalSnapshots, leaving the actor permanently unschedulable

Aperta
#397 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
area/api area/api-machinery kind/bug
Lingua principale
Go
Stelle
1.8k
Fork
316
Merge medio
2g 43m
PR unite (30g)
287

Descrizione

### Summary

When a `PauseActor` workflow finalizes after the assigned worker record has already been deleted (e.g. the worker pod was removed while the pause was in flight), `FinalizePausedStep` cannot determine the node that holds the local checkpoint. It logs a warning ("Node name not found during finalize pause") but still writes the unknown node name into the actor record:

```
LatestSnapshotInfo.Local.NodeVmsWithLocalSnapshots = [""]
```

### Impact

On the next `ResumeActor`, `AssignWorkerStep.findFreeWorker` treats a non-empty `NodeVmsWithLocalSnapshots` list as a hard placement restriction: a candidate worker must have `NodeName` equal to one of the listed entries. No worker ever has an empty node name, so **no worker can ever match** and every resume of the affected actor fails with:

```
FailedPrecondition: no free workers available
```

The actor is permanently unschedulable. The only escape is `ResumeActor(boot=true)`, which discards the actor's checkpointed state.

The generic error message also makes this very hard to diagnose — it looks like a capacity problem, not a poisoned locality record.

### Proposed fix

Two guards, both small:

1. **Write side** (`FinalizePausedStep`): only append the node name to `NodeVmsWithLocalSnapshots` when it is non-empty. If the node is unknown, record the local snapshot with an empty node list — the snapshot still exists on some node, and an empty list lets the scheduler place the actor anywhere (a shuffled retry can still land on the node that holds the checkpoint, since local checkpoints live on the node, not in the deleted worker pod).

2. **Read side** (`findFreeWorker`): ignore empty-string entries in the restriction list. This heals actor records already poisoned by (1) without requiring a data migration.

### Notes

This is a targeted fix. The underlying modeling issue — snapshot placement being a write-once field on the Actor with no feedback from the nodes that actually hold the snapshots — is a larger topic being handled separately.

Guida per i contributori

Apri la guida per i contributori

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.