Parked run is cancelled when its SSE client disconnects, so a pending approval never lands
@JAORMX arbeitet bereits daran.
Seit 07.9.2026.
- Vorherrschende Sprache
- Go
- Sterne
- 152
- Forks
- 16
- Ø Merge
- 14 Std. 48 Min.
- Gemergte PRs (30 T.)
- 536
Beschreibung
Symptom
A parked run — one waiting on a permission ask — is cancelled if its SSE client disconnects, even briefly. When the operator then approves, the approval doesn't land: the same tool call re-prompts on the next turn, repeatedly.
This isn't specific to any one client. It reproduces with any HTTP/SSE consumer of relayRunSSE whose connection idles out or drops while a run is parked awaiting a verdict — for example, a browser client with a 120s idle timeout on its SSE connection, but the underlying bug is client-agnostic.
Root cause
relayRunSSE in internal/adapter/server/http.go is the shared relay for both the prompt run-entry and the approve-rehydrate path. It installs a disconnect hook plus a deregister:
defer h.svc.deregister(id, run)
...
// If the client disconnects, cancel the run.
go func() {
<-r.Context().Done()
run.Cancel()
}()
A run parked on a permission ask emits no events while it waits on the operator, so any idle timeout on the client side is enough to drop the connection. That trips the following chain:
- request ctx is Done →
run.Cancel()fires, and thedefer deregisterruns when the loop ends. - Operator submits a verdict →
POST /[REDACTED]/sessions/{id}/approve. ApproveRun's lock-freeLookupRunfast path misses (the run was deregistered) → falls through toresumeFromAwaiting.resumeFromAwaitingrejects any non-StateAwaitingsession withErrNoActiveRun— and the session is now cancelled.- The pending tool never runs, so the next turn re-issues the same tool call and re-asks.
This inverts the cloud-native Phase 2 intent, where awaiting is the only non-terminal state precisely so that a parked session survives a disconnected client.
Note that deleting the disconnect goroutine alone is not sufficient: fail() in the same loop also calls run.Cancel(), so the first write attempt after the client is gone would cancel the run anyway.
Proposed fix
Daemon-side (primary): when the client disconnects while the run is parked on a human verdict, switch to the existing drain-to-discard mode instead of cancelling. The run outlives its stream, stays registered, and the verdict resolves it in-process via the same-process LookupRun path. Effects land in the durable log, so any client can pick the result up on reattach/replay.
Because relayRunSSE is shared by both entry points specifically so their discipline "cannot drift," any change here must keep the prompt path's behaviour intact, and is guarded by internal/adapter/server/resume_awaiting_test.go and internal/app/approve_after_restart_test.go.
Client-side (complementary): a well-behaved client shouldn't let its own idle timeout abort a stream while its own approval panel/prompt is open — an open approval means the run is alive by definition.
Secondary re-ask vectors (independent, both real)
- Learned rules are
Exact: true+ScopeUser(LearnableRuleinengine/governance/evaluator.go), so "Always allow" only ever covers the byte-identical canonical command. learnablePatterndeliberately refuses compound or substituted Bash (a && b, pipelines,$(...), backticks). Approving such a command legitimately learns nothing, and the UI gives no signal that it didn't.- The ADR-0062 waiver and the permstore are in-memory and session-keyed. A client that restarts
mecatedon every config write drops both. Phase 3breplayApprovalsonly re-derives fromEvApproval{AllowAlways}, and only vialoadAndReopen.
Verification status
The call chain above is read from source. What is not yet confirmed end-to-end is that an idle timeout is what ends the parked stream in practice, rather than the daemon ending it for another reason — a repro with daemon logs alongside a client-side network capture would settle that before the fix lands.
Beitragsleitfaden
Erste Schritte
- Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
- Forke das Repository und arbeite in einem Branch.
- Öffne einen Pull Request, der die Issue-Nummer nennt.
Bewertung
Dieses Issue wurde noch nicht bewertet.