block / block/buzz

buzz-acp: agents accept messages but never dispatch to harness — dispatch loop appears to wedge minutes after startup (all harness types, v0.5.3)

Open
#4,408 4 comments 0 reactions 0 assignees View on GitHub
Dominant language
Rust
Stars
32.7k
Forks
4.3k
Avg merge
1d 13h
Merged PRs (30d)
253

Description

# Buzz bug report — OmniRoute Operator (custom ACP harness) stopped processing messages

**Buzz version:** v0.5.3 (Windows)
**Agent:** OmniRoute Operator, custom ACP harness (bridge validated independently — see "Harness is not the problem" below)
**Date symptoms began:** 2026-08-02 (worked correctly 2026-08-01)
**Scope update (2026-08-02 ~21:00):** the dispatch failure is NOT specific to the custom harness — see Bug C below; claude-agent-acp and codex-acp agents show identical pool-init-without-dispatch, and one agent gets stuck in setup-listener mode.

## Summary

The agent receives messages (relay subscription active, membership/channel events arrive, agent pool initializes in response to an incoming message) but never dispatches a prompt to the harness. No `agent_returned`, no error, no requeue — the harness's ACP stdio stays silent. The same harness processed owner DMs correctly the previous day.

## Timeline (all times UTC, from agent logs)

**2026-08-01 — WORKING**
- 12:07–13:22: Owner DMs processed end-to-end; two successful runs (~6 min each), replies posted in channel. Agent command at the time was incorrectly set to `claude-agent-acp.cmd` (unauthenticated) — expected "Authentication required" errors observed, then harness switched to custom bridge and everything worked.

**2026-08-01 13:39 — custom harness confirmed healthy**
- `agent_cmd=C:\WINDOWS\system32\wsl.exe bash -lc /home/webgive/stacks/crewai-orchestrator/bridge/acp_bridge.py`
- `model=crewai-fixed-lineup` picked up from `_meta.models` in initialize response.

**2026-08-02 10:43 — first silent failure**
- Owner (me) sent a DM. Log shows ZERO events at that timestamp — no channel event, no batch, nothing. Presence showed online.

**2026-08-02 14:27 — restart after config change**
- Restarted agent (stop/start). Startup banner still shows `respond_to=owner-only` despite "Who can send instructions" being set to "Selected people" and saved (see Bug B below).

**2026-08-02 15:02:39 — pool initializes but no dispatch**
- ~1 min after owner message, all 10 pool agents initialize (`agent initialized ... agent=0..9`, `agent_pool_ready agents=10`), then nothing. No session/prompt reaches the harness (no activity on its stdio), no `agent_returned`, no queue/requeue lines.

**2026-08-02 20:27 — @mention test, total silence**
- Owner sent `@OmniRoute Operator Create a file called world5.txt ...` in the original DM channel (the one that worked on 2026-08-01). Log shows no new lines at all in the following minutes.

## Bug A: messages accepted but never dispatched to harness

Symptoms as above. Notably the failure mode differs by send:
- Some sends produce zero log lines (20:27)
- One send produced pool init but no dispatch (15:02)
- All produce no user-visible response, no error surfaced anywhere in the UI

Hypotheses I cannot verify from outside: a wedged internal queue after 2026-08-01's auth-failure requeue loop (attempts 1–6 of 10 logged before agent stop), DM/mention routing change between 08-01 and 08-02, or stale channel state (the agent is subscribed to 4 channels: 2c7038a9, 84721bf3, b11e7898, d1821d0d — the original working DM is 84721bf3).

## Bug B: "Who can send instructions" does not propagate to runtime

- Set to "Selected people" (3 agents added via the people picker), clicked Save changes, restarted the agent 6+ times across the day.
- Every startup banner still shows `respond_to=owner-only`.
- Consequence: agent-to-agent delegation is silently dropped — another agent (Codex Builder) added this agent to a group chat and sent a brief at 14:42; only the `membership notification: subscribing to new channel` line appears, the message itself is never processed.

## Harness is not the problem

The custom ACP bridge is independently verified:
- 16/16 unit tests; live JSON-RPC transcript tests (initialize / session/new / session/prompt / session/cancel) all pass outside Buzz.
- Streams sessionUpdate chunks correctly; handled a full real run on 2026-08-01 with 2,422 streamed chunks.
- Responds to `session/set_config_option` (visible in logs) — so Buzz CAN reach it.
- The failure is upstream of the harness: no `session/prompt` ever arrives on its stdio during the silent windows.

## Environment

- agent_cmd: `C:\WINDOWS\system32\wsl.exe bash -lc /home/webgive/stacks/crewai-orchestrator/bridge/acp_bridge.py`
- model: `crewai-fixed-lineup` (advertised via `_meta.models`; also via `session/new` configOptions per ACP v1)
- Startup flags: `subscribe=Mentions dedup=Queue respond_to=owner-only permission_mode=bypassPermissions agents=10`
- Relay: `wss://cenenta.communities.buzz.xyz` (two WebSocket "Connection reset without closing handshake" events on 08-01, both auto-recovered)

## Reproduction

1. Create custom ACP harness agent, verify owner DM works (08-01 behavior).
2. Next day, send owner DM → silence.
3. Restart agent → send again → pool may initialize, still no dispatch.
4. Set "Who can send instructions" → "Selected people", save, restart → banner still `respond_to=owner-only`; agent messages still dropped.

## Bug C: identical dispatch failure on built-in ACP agents (claude-agent-acp, codex-acp)

The failure is not harness-specific — it is in buzz-acp's message→dispatch path. Observed same evening, all owner DMs, all `respond_to=owner-only` so owner messages should pass the filter:

- **Claude Dev** (`claude-agent-acp` v0.64.0, agents=10): owner DM at 20:46 → all 10 pool agents initialize (`agent_pool_ready agents=10`) → **no `session/prompt` dispatch, no reply**. Agent restarted at 20:51:10; a second owner message after restart → silence. NOTE: the 20:46 message also arrived during the previous run's `idle_timeout=900s` window (agent had been idle since 15:01, i.e. ~5h45m), so dispatch after long idle is also implicated.
- **Codex Builder** (`codex-acp` v1.1.7, agents=2): owner DM at 20:47 → pool re-inits (`agent_pool_ready agents=2`) → **no dispatch**. Earlier at 14:42 this agent DID process an owner message and replied — the only successful dispatch all day across all agents. (That success was ~3.5 min after a restart.)
- **Claude Dev setup-mode wedge (possibly separate):** across restarts at 09:42, 09:49, and 14:27 the agent entered `setup-listener mode ... requirements=1` and never exited it — subscribing to channels but processing nothing. Only the 15:01 restart (after re-authenticating Claude Code in the terminal) reached normal startup. If setup mode is supposed to be resolved by in-app auth UI, there was no user-visible prompt in any of those three windows.

Pattern: **the only successful dispatch of the day happened within ~4 minutes of an agent restart.** Every message sent to an agent that had been running for longer (minutes to hours) produced pool-init-only or total silence. This points at the event-processing/dispatch loop wedging shortly after startup, not at harnesses, auth, or channel subscriptions (which all initialize cleanly).

## Requested investigation

1. Why does an incoming owner message trigger pool init but no `session/prompt` dispatch (15:02, 20:46, 20:47 events)?
2. Why do some messages produce zero log lines at all (10:43, 20:27)?
3. Does the dispatch loop wedge after some time/idle period? (Only success today was ~4 min post-restart; failures at 14 min, ~1 min, ~5h45m post-start.)
4. Why doesn't the "Selected people" setting reach the runtime `respond_to` flag?
5. Why did Claude Dev sit in setup-listener mode across three restarts with no UI prompt to resolve it?
6. Is there a way to inspect/flush the agent's internal event queue (`dedup=Queue`)?

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.