stacklok / stacklok/mecatl

Parked run is cancelled when its SSE client disconnects, so a pending approval never lands

Aperta
#653 0 commenti 0 reazioni 1 assegnatario Vedi su GitHub

@JAORMX ci sta già lavorando.

Dal 7/9/2026.

Lingua principale
Go
Stelle
152
Fork
16
Merge medio
14h 48m
PR unite (30g)
536

Descrizione

Symptom

A parked run — one waiting on a permission ask — is cancelled if its SSE client disconnects, even briefly. When the operator then approves, the approval doesn't land: the same tool call re-prompts on the next turn, repeatedly.

This isn't specific to any one client. It reproduces with any HTTP/SSE consumer of relayRunSSE whose connection idles out or drops while a run is parked awaiting a verdict — for example, a browser client with a 120s idle timeout on its SSE connection, but the underlying bug is client-agnostic.

Root cause

relayRunSSE in internal/adapter/server/http.go is the shared relay for both the prompt run-entry and the approve-rehydrate path. It installs a disconnect hook plus a deregister:

defer h.svc.deregister(id, run)
...
// If the client disconnects, cancel the run.
go func() {
	<-r.Context().Done()
	run.Cancel()
}()

A run parked on a permission ask emits no events while it waits on the operator, so any idle timeout on the client side is enough to drop the connection. That trips the following chain:

  1. request ctx is Done → run.Cancel() fires, and the defer deregister runs when the loop ends.
  2. Operator submits a verdict → POST /[REDACTED]/sessions/{id}/approve.
  3. ApproveRun's lock-free LookupRun fast path misses (the run was deregistered) → falls through to resumeFromAwaiting.
  4. resumeFromAwaiting rejects any non-StateAwaiting session with ErrNoActiveRun — and the session is now cancelled.
  5. The pending tool never runs, so the next turn re-issues the same tool call and re-asks.

This inverts the cloud-native Phase 2 intent, where awaiting is the only non-terminal state precisely so that a parked session survives a disconnected client.

Note that deleting the disconnect goroutine alone is not sufficient: fail() in the same loop also calls run.Cancel(), so the first write attempt after the client is gone would cancel the run anyway.

Proposed fix

Daemon-side (primary): when the client disconnects while the run is parked on a human verdict, switch to the existing drain-to-discard mode instead of cancelling. The run outlives its stream, stays registered, and the verdict resolves it in-process via the same-process LookupRun path. Effects land in the durable log, so any client can pick the result up on reattach/replay.

Because relayRunSSE is shared by both entry points specifically so their discipline "cannot drift," any change here must keep the prompt path's behaviour intact, and is guarded by internal/adapter/server/resume_awaiting_test.go and internal/app/approve_after_restart_test.go.

Client-side (complementary): a well-behaved client shouldn't let its own idle timeout abort a stream while its own approval panel/prompt is open — an open approval means the run is alive by definition.

Secondary re-ask vectors (independent, both real)

  • Learned rules are Exact: true + ScopeUser (LearnableRule in engine/governance/evaluator.go), so "Always allow" only ever covers the byte-identical canonical command.
  • learnablePattern deliberately refuses compound or substituted Bash (a && b, pipelines, $(...), backticks). Approving such a command legitimately learns nothing, and the UI gives no signal that it didn't.
  • The ADR-0062 waiver and the permstore are in-memory and session-keyed. A client that restarts mecated on every config write drops both. Phase 3b replayApprovals only re-derives from EvApproval{AllowAlways}, and only via loadAndReopen.

Verification status

The call chain above is read from source. What is not yet confirmed end-to-end is that an idle timeout is what ends the parked stream in practice, rather than the daemon ending it for another reason — a repro with daemon logs alongside a client-side network capture would settle that before the fix lands.

Guida per i contributori

Apri la guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.