block / block/buzz

`buzz-dev-mcp`: the 600s shell cap is silently clamped and unconfigurable — long agent jobs die mid-work with no signal

Open
#4,638 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Rust
Stars
32.7k
Forks
4.3k
Avg merge
1d 13h
Merged PRs (30d)
253

Description

## Summary

`crates/buzz-dev-mcp/src/shell.rs` caps every shell command at 600s and there is no way to raise it, no env var, and — the part that costs the most debugging time — **no report that the clamp happened**. A caller that asks for 3600s is silently given 600s, and on expiry the whole process group is killed.

## Where

```rust
// crates/buzz-dev-mcp/src/shell.rs
const DEFAULT_TIMEOUT_MS: u64 = 120_000; // :16
const MAX_TIMEOUT_MS: u64 = 600_000; // :17

let timeout_ms = p
.timeout_ms
.unwrap_or(DEFAULT_TIMEOUT_MS)
.min(MAX_TIMEOUT_MS); // :141-144

cmd.kill_on_drop(true); // :178
set_process_group(&mut cmd); // :179
```

`.min(MAX_TIMEOUT_MS)` is the whole story: `timeout_ms` is an accepted parameter of the tool (`:127 pub timeout_ms: Option`), so a caller reasonably believes it was honoured. It was not. Verified present at HEAD, not just on the v0.5.3 tag.

## Impact

Any agent-launched job longer than 10 minutes dies mid-work. Because `set_process_group` + `kill_on_drop` take out the entire group, the child's own output buffer goes with it, so what the agent sees is a truncated or empty log — not a timeout message.

Concretely, on a self-hosted relay driving an agent that fixes issues in a repo: four consecutive runs died at exactly 600s. Each had already done real work (written files, opened PRs, posted review comments) before the kill. The surviving artifact was a 15-byte log reading `Execution error`. The agent, given only that, concluded three separate times that its model backend was broken and proposed switching runners — a diagnosis that would have made things worse, since the same ceiling kills any runner identically. The ceiling is invisible from inside the sandbox; nothing in the failure points at it.

## Asks (either would be sufficient, in preference order)

1. **Make the cap configurable** — an env var (`BUZZ_DEV_MCP_MAX_TIMEOUT_MS`) or a config field. Operators running agents on their own hardware are already trusting them with a shell; a hard 10-minute ceiling is a policy decision that should belong to the deployment.
2. **At minimum, report the clamp.** If `p.timeout_ms > MAX_TIMEOUT_MS`, say so in the tool result, and on expiry return an explicit "killed at N s ceiling" rather than letting the caller infer it from a truncated stream. A silent clamp that a caller cannot detect is indistinguishable from a crash.

## Not covered by existing work

PR #935 ("feat(acp): agent timeout resilience") is adjacent and explicitly acknowledges the "600s max shell timeout" — it raises `DEFAULT_IDLE_TIMEOUT_SECS` 620 → 900 so the ACP idle timer stops firing during legitimate long tool calls. That fixes spurious *session* kills above the cap; it leaves the cap itself, and its silence, unchanged.

## Workaround for anyone hitting this

Launch through a `setsid` wrapper so the run lives in a session the group-kill cannot reach, and have the wrapper — not the agent — own the launch, since the detached run then outlives the MCP call that started it.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.