block / block/buzz

[Bug] A provider-backed agent stays "online" with a live Shutdown button after it has shut down — remote liveness is derived from backend_agent_id, not presence (violates I3)

Open
#4,730 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Rust
Stars
32.7k
Forks
4.3k
Avg merge
1d 13h
Merged PRs (30d)
253

Description

## Summary

Pressing **Shutdown** on a provider-backed (remote) managed agent stops the agent, but the desktop keeps rendering it as **online / running** indefinitely, and the primary action button stays **"Shutdown"** forever — it never flips back to **"Deploy"**.

The agent is not at fault: it publishes both of the signals the spec asks for (kind:10100 roster entry with `"status":"offline"`, and kind:20001 presence `"offline"`), and the relay accepts and applies both. The desktop simply does not consult either one when it decides whether a *remote* agent is alive.

Root cause: for `backend != Local`, `build_managed_agent_summary()` derives the status **exclusively** from `record.backend_agent_id.is_some()`, and `backend_agent_id` is never cleared outside of agent deletion. Every liveness-shaped UI decision downstream is keyed off that one field.

This contradicts **invariant I3 ("Presence is the status")** in `docs/remote-agents.md:200`, which states that the deployment axis "is bookkeeping, not liveness", and it is not among the 8 entries in §Known Defects.

## Environment

- Buzz Desktop **0.5.4** (`desktop/src-tauri/tauri.conf.json`), macOS
- Agent deployed through a third-party backend provider implementing the documented provider protocol (`buzz-backend-*`, local-host substrate)
- All `file:line` references below verified against `feccf4eabc23fdba94ce3537a194357ed17b197c`

## Steps to reproduce

1. Create a managed agent with a **provider** backend and press Deploy. The agent comes up, appears online, and answers in a channel.
2. Open the agent's profile panel and press **Shutdown**. A toast appears: *"Shutdown command sent. Agent will stop shortly."*
3. The agent receives the owner `!shutdown`, drains, publishes its farewell (kind:10100 `status: "offline"`), publishes presence `offline` (kind:20001), and exits. It stops answering in the channel.
4. Wait — minutes, hours, a restart of the desktop.

### Expected

The agent's card / profile stops showing a live indicator, and the primary action becomes **Deploy** so the agent can be brought back.

### Actual

- The agent still renders with a green "online" dot.
- The Runtime tab still shows a green *Running* dot.
- The status badge still reads a green **deployed**.
- The primary action button still reads **Shutdown** and is still styled as live.
- Pressing **Shutdown** again *succeeds* (green toast, `!shutdown` re-sent into a channel of a process that no longer exists) — because the offline roster entry still carries `channel_ids`, so `resolveManagedAgentChannelId` still resolves.
- **There is no way to redeploy the agent from this surface.** `handleAgentPrimaryAction` (`desktop/src/features/profile/ui/useAgentLifecycleActions.ts:30`) branches on `isManagedAgentActive`, which is permanently true, so the `startManagedAgentWithRules` arm is unreachable.

## Evidence: the agent reported correctly, the desktop ignored it

Agent-side log at shutdown:

```
13:38:41.122Z info draining reason=owner_shutdown
13:38:43.518Z warn relay connection closed
```

Process gone. I then reconnected to the relay with the agent key (NIP-42 + NIP-OA auth tag) and queried its own events:

```
=== kind 10100 (roster) ===
created_at: 1785850721 → 2026-08-04T13:38:41Z (exactly the drain instant)
content: {"name":"…","agent_type":"agent","status":"offline","capabilities":[],
"channels":[],"channel_ids":["650ac09e-…","81f2eb5e-…"]}

=== kind 20001 (presence) ===
(none — the relay holds no presence for this pubkey)
```

Both axes are correct on the relay:

- the roster entry was replaced with `"status":"offline"`;
- presence was cleared, which is exactly what `crates/buzz-relay/src/handlers/event.rs:823-830` does on a kind:20001 whose content is `"offline"`.

So the desktop had everything it needed. It just never looks.

## Root cause

**1. Remote status is a function of one persisted field.**

`desktop/src-tauri/src/managed_agents/runtime.rs:148-169`:

```rust
let (status, pid, log_path) = if record.backend != BackendKind::Local {
// Two-axis status model for remote agents:
// …
// Live axis (relay presence, polled by frontend): online/away/offline.
// Shown as a PresenceDot next to the agent name. …
// After !shutdown the agent goes offline (presence) but stays "deployed"
// (infrastructure still exists). This is intentional …
let status = if record.backend_agent_id.is_some() {
"deployed".to_string()
} else {
"not_deployed".to_string()
};
(status, None, String::new())
} else { /* local: real pid probe */ };
```

The comment's two-axis model is a reasonable design. The bug is that **the frontend never implements the second axis** for the decisions that matter — it treats the control-plane axis as liveness.

**2. `backend_agent_id` is write-once.** Set at `desktop/src-tauri/src/commands/agents.rs:501` on a successful deploy; the only paths that clear it are record creation defaults and deletion. There is no `undeploy` (deferred to "v2" per the comment at `runtime.rs:163`), and `stop_managed_agent` explicitly refuses remote agents (`desktop/src-tauri/src/commands/agents.rs:1244-1248`). So `status` is pinned at `"deployed"` for the entire life of the record.

**3. Every liveness-shaped UI decision keys off it.**

| Surface | Location | Behaviour |
|---|---|---|
| Active predicate | `desktop/src/features/agents/lib/managedAgentControlActions.ts:34` | `status === "running" \|\| "deployed"` → permanently `true` |
| Primary action label | `.../managedAgentControlActions.ts:38-47` | provider + active → `"Shutdown"`, never `"Deploy"` |
| Button "live" styling | `desktop/src/features/profile/ui/UserProfilePanelSections.tsx:364-367` | `agentActionLive` = same predicate, inlined |
| **Green online dot** | `desktop/src/features/agents/ui/AgentRuntimeAvatarControl.tsx:130` | `` — **hardcoded literal**, gated only on `isActive`. Presence is never read here |
| Status badge | `desktop/src/features/agents/ui/AgentStatusBadge.tsx:27-31` | green `deployed`; the `"Starting…"` de-escalation branch tests `status === "running"`, so it can never fire for a remote agent |
| Runtime tab dot | `desktop/src/features/profile/ui/UserProfilePanelSections.tsx:129-148` | `deployed` → `"running"` → `bg-emerald-500` |

`ManagedAgentRow.tsx:230` and `ProfileHero` (`UserProfilePanelSections.tsx:485-497`) *do* render a real presence-driven `PresenceDot` — so the app disagrees with itself on the same screen once the two axes diverge.

## Contributing (secondary) issue: nothing is refetched after `!shutdown`

`handleStop` (`desktop/src/features/agents/ui/useManagedAgentActions.ts:240-267`) invalidates and refetches **nothing** for the provider branch, unlike `handleStart` (`:220-221`). The same is true of `useMembersSidebarActions.ts:173-188` and `useAgentLifecycleActions.ts:30-41`.

Meanwhile `useRelayAgentsQuery` polls kind:10100 on a **5-minute** interval (`desktop/src/features/agents/hooks.ts:337`), and `usePresenceQuery` backstop-polls at 60s (`desktop/src/features/presence/hooks.ts:94`).

So even the surfaces that *do* honour presence lag by up to a minute, and roster-derived status by up to five — right at the moment the user is watching for feedback on an action they just took. Fixing the root cause without also invalidating `["presence"]` / `relay-agents` / `managed-agents` after a successful `!shutdown` send would leave a visible multi-minute lie.

## Why this is a spec violation, not a design choice

`docs/remote-agents.md:200-206`, invariant **I3**:

> **(I3) Presence is the status.** `D` derives a remote agent's live state *exclusively* from relay presence events self-signed by the agent key: `online`/`away`/`offline` (kind:20001, ephemeral, WS-published). The deployment axis (`deployed`/`not_deployed`, from the stored `backend_agent_id`) **is bookkeeping, not liveness.**

The implementation inverts this for every decision listed in the table above. This is not in §Known Defects (`docs/remote-agents.md:1572-1681`, defects 1–8).

I3 also promises a *bounded* wrong dot (180s `PRESENCE_TTL_SECS`). Today the wrong dot is unbounded: it survives shutdown, app restart, and machine reboot, because it is not a presence dot at all.

## Suggested fix

Smallest correct change is to give remote agents the live axis the comment already describes, and let it drive the "is this thing alive" decisions:

1. **Introduce a live signal for remote agents.** Either
(a) have `build_managed_agent_summary()` fold relay presence into the returned status for `backend != Local`, or
(b) keep the Rust summary purely control-plane and add a frontend helper — e.g. `isManagedAgentLive(agent, presenceLookup)` — that returns `presence !== "offline" && presence !== undefined` for provider agents and falls back to `status` for local ones.
(b) is less invasive and keeps `deployed`/`not_deployed` honest as bookkeeping, per I3.
2. **Route the liveness-shaped call sites through it**: `getManagedAgentPrimaryActionLabel`, `agentActionLive`, `AgentRuntimeAvatarControl` (drop the hardcoded `status="online"` and pass the real presence through), `AgentStatusBadge`, `resolveRuntimeTabStatus`. Keep `isManagedAgentActive` for the genuinely control-plane questions (e.g. "should Delete warn about orphaning a remote deployment?" in `deleteManagedAgentWithRules`, which already reads `presenceLookup` and gets this right).
3. **Invalidate after a `!shutdown` send** — `["presence"]`, `relay-agents`, `managed-agents` — in `stopManagedAgentWithRules`' callers, so the transition is visible immediately rather than on the next 5-minute poll.
4. Once (1)–(3) land, an offline remote agent's button becomes **Deploy**, and `startManagedAgentWithRules` → `deploy` reaches the provider, which is already a converge-to-one-live-instance operation (§Deploy State Machine) and safely re-adopts or recreates. That closes the "cannot bring it back" half of this report without needing the v2 `undeploy`.

## Related

- #2798 — provider-backed config saves don't redeploy. Independently notes that v1 has no `shutdown` operation and that Shutdown is not a redeploy path; different defect (stale config, not stale status).
- #4605 — no management surface on the owner's other desktops; cites the same missing-`undeploy` deferral, but for orphaned deployments elsewhere.
- #3831, #2453 — same symptom class (UI says online, agent unusable) for *local* agents dying accidentally; different root cause.
- #4537 (PR) — presence reporting for relay-discovered agents; opposite direction (always offline), different files.
- #3034 — provider-backed agents display local config instead of the remote runtime; adjacent "remote agent card shows the wrong thing" family.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.