CommandCodeAI / CommandCodeAI/command-code

Background shell tasks return empty logs; provider errors (too_many_images, timeouts) and subagent failures stall long coding sessions

未關閉
#779 1 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

還沒有人認領這個 Issue。

主要語言
沒有語言資料
星號
4k
分支
350
PR 合併指標
30 天內沒有已合併 PR

描述

Long coding session crippled by tooling failures: empty background-task output, provider errors, lost subagent reports

Environment

  • Windows 11, Command Code CLI (npm, v1.39.x line at the time)
  • Large Android/Gradle project; long-running builds (1–5 min per gradle invocation)
  • One long single-session refactor task (~15 files edited + verification builds + CI screenshot step)

Summary

A large but well-scoped coding task stretched for hours almost entirely because of harness/infra failures, not task complexity. Four classes of failure, each reproducible within the session:

1. Background shell commands return empty output logs (most damaging)
  • shell_command with run_in_background=true running gradlew ... > file.log 2>&1 or piped through | powershell ... produced 0-byte output logs for 20–60+ minutes while gradle had actually failed in 40 seconds.
  • Exit status was never surfaced; the log stayed empty; the only way to learn the real failure was to manually read ~/.gradle/daemon/9.5.0/daemon-*.out.log (the daemon's own log contained e: file:///...kt:NN Unresolved reference... compile errors all along).
  • Piping task output through PowerShell (| powershell -Command "$input | Select-Object -Last N") also silently lost everything.
  • Cost: this failure mode alone consumed the majority of the session, with multiple redundant 5–10-minute polling loops staring at empty files.
  • Expected: background tasks should surface real exit codes/stderr (or at minimum, the streamed log should contain the process's output as it appears).
2. Repeated mid-session provider errors, each requiring a manual "continue"
  • Error: 200 Failed to process successful response
  • Error: 500 Cannot connect to API: Connect Timeout Error (172.65.90.20-23:443)
  • Error: 400 Invalid_request_error ... [too_many_images] GLM requests accept at most 8 inline PNG/JPEG/WEBP/GIF ... — a session that reviews screenshots (dev workflow!) becomes unsendable until the user manually compacts. The model can't fix this itself.
3. Subagent runs errored and lost their reports
  • One general subagent returned [sub-agent stopped early: the run errored] after 20 minutes, mid-task (had made real edits already).
  • A separate audit subagent died entirely with its report lost; had to be relaunched from scratch.
  • Expected: partial output preserved on subagent error, or automatic retry.
4. Tool-schema friction
  • search_tools repeatedly returned todo_write schema, but calling it kept failing/looping for several turns before it finally worked.

Trace IDs (from the error banners in one session)

  • 6b481a97ce2ad552cb4802dc72d0b0ef
  • 39a179d53c81ee36fea47ff7f116d066
  • 52c2bcc5f22d1a3815bb7b6a4ce81356
  • ca6c03357b892c7ca6e9042e6ecb6367
  • 2cc56255ce5d8e9aab41e69325ec4e81

Ask

  1. Surface real exit status + stderr of background and piped shell tasks; don't let empty logs masquerade as "still running."
  2. Handle the image budget proactively (auto-compact or drop stale images before the provider hard-fails the whole conversation).
  3. Preserve subagent partial reports when a run errors.

貢獻指南

這個儲存庫沒有索引到貢獻指南

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

研究方向

從 shell_command 的 CLI 入口開始,使用 run_in_background=true,重現 Gradle 和 PowerShell 的情況,並將擷取的輸出與 daemon 記錄和結束狀態進行比較。接著追蹤 provider 的 image-budget 錯誤和子代理錯誤處理,在可用時使用列出的 trace IDs。完成的標準是:背景失敗會暴露狀態和 stderr,provider 限制得到處理,發生錯誤的子代理會保留部分報告。

由索引模型根據 Issue 內容生成。

評估

技術堆疊
ai-infra-agents, cli
領域
ai, cli, tooling
Issue 類型
缺陷
難度
5/5
預估耗時
一週以上
活躍度
活躍
描述清晰度
需要釐清
新手友好度
25/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。