Autopilot task-completion enforcement can override explicit user instructions
まだ誰も着手していません。
- 主要言語
- Shell
- スター
- 11.2k
- フォーク
- 1.9k
- 平均マージ
- 14時間 16分
- マージ済み PR(30日)
- 6
説明
Describe the bug
In Copilot CLI autopilot mode, the task-completion enforcement behavior can cause the agent to continue taking action after the user has explicitly narrowed the task to research/explanation only.
In my case, the user explicitly instructed the agent to do nothing besides respond and asked for an explanation/research task. After the agent answered, a follow-up enforcement message appeared:
“You have not yet marked the task as complete using the task_complete tool. If you were planning, stop planning and start implementing.”
That message appears to push the agent toward implementation/action even when the active user request is intentionally non-implementation work.
Why this is a problem
This can lead to the agent making unintended workspace changes despite the user explicitly asking for diagnostic or explanatory help only. The issue is especially risky because the enforcement message is framed as higher-priority continuation guidance, so the agent may treat it as expanding the user’s requested scope.
Suggested improvement
Adjust the completion enforcement prompt/logic so it preserves the current user’s scope constraints. For example:
- If the user asks for research/explanation only, completing the answer should be considered sufficient.
- If the user explicitly says not to modify files or take action, the continuation message should not direct the agent to “start implementing.”
- The agent should be encouraged to call
task_completewith a concise summary when the requested non-code task is complete.
Notes
This may be related to autopilot/task-completion behavior introduced or changed recently. The problematic phrase above is included because it seems useful for locating the relevant enforcement path.
Affected version
GitHub Copilot CLI 1.0.77.
Steps to reproduce the behavior
- Start a Copilot CLI session with autopilot mode enabled.
- Ask for a research-only or explanation-only task, and explicitly prohibit action. For example:
Please help me understand this behavior. Do not edit files, run implementation steps, or make changes. Just explain what is happening.
- Let the agent provide a complete explanatory answer without calling
task_complete. - Observe that task-completion enforcement sends a follow-up message similar to:
“You have not yet marked the task as complete using the task_complete tool. If you were planning, stop planning and start implementing.”
- Observe that the agent may interpret this as a directive to take implementation/action-oriented steps, despite the user’s explicit “do not act” instruction.
Expected behavior
If the active user request is clearly research-only, explanation-only, or explicitly says not to modify anything, task-completion enforcement should allow the agent to finish by calling task_complete with a research summary, rather than nudging it toward implementation.
At minimum, the enforcement text should not say “start implementing” when the current task has no implementation component or when the user has explicitly prohibited changes.
Actual behavior
The agent was nudged to continue with implementation-oriented behavior after completing an explanatory/research response, despite the user saying to do nothing except respond.
Additional context
- Copilot CLI: 1.0.77
- Mode: autopilot enabled
- OS: Linux devcontainer
- Shell: bash
- Terminal emulator: Herdr-managed terminal session
コントリビューションガイド
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
調査の方向性
リポジトリを検索して、引用された「start implementing」メッセージと autopilot/task-completion の強制適用経路を見つけ、その後 issue 内の調査専用手順を使って挙動を再現する。完了とは、説明のみを求めるリクエストが、実装に向けたガイダンスを受け取ることも、ユーザーが明示した制約に反するアクションを引き起こすこともなく、task_complete の要約で終了できることを意味する。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- shell
- 領域
- ai, cli
- issue の種類
- バグ
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 活発さ
- 静か
- 明瞭さ
- おおむね明確
- 初心者へのやさしさ
- 45/100