Autopilot task-completion enforcement can override explicit user instructions
- 主要语言
- Shell
- 星标
- 11.2k
- 派生
- 1.9k
- 平均合并
- 14 小时 16 分钟
- 30 天内合并 PR
- 6
描述
### Describe the bug
In Copilot CLI autopilot mode, the task-completion enforcement behavior can cause the agent to continue taking action after the user has explicitly narrowed the task to research/explanation only.
In my case, the user explicitly instructed the agent to **do nothing besides respond** and asked for an explanation/research task. After the agent answered, a follow-up enforcement message appeared:
> “You have not yet marked the task as complete using the task_complete tool. If you were planning, stop planning and start implementing.”
That message appears to push the agent toward implementation/action even when the active user request is intentionally non-implementation work.
## Why this is a problem
This can lead to the agent making unintended workspace changes despite the user explicitly asking for diagnostic or explanatory help only. The issue is especially risky because the enforcement message is framed as higher-priority continuation guidance, so the agent may treat it as expanding the user’s requested scope.
## Suggested improvement
Adjust the completion enforcement prompt/logic so it preserves the current user’s scope constraints. For example:
- If the user asks for research/explanation only, completing the answer should be considered sufficient.
- If the user explicitly says not to modify files or take action, the continuation message should not direct the agent to “start implementing.”
- The agent should be encouraged to call `task_complete` with a concise summary when the requested non-code task is complete.
## Notes
This may be related to autopilot/task-completion behavior introduced or changed recently. The problematic phrase above is included because it seems useful for locating the relevant enforcement path.
### Affected version
GitHub Copilot CLI 1.0.77.
### Steps to reproduce the behavior
1. Start a Copilot CLI session with autopilot mode enabled.
2. Ask for a research-only or explanation-only task, and explicitly prohibit action. For example:
> Please help me understand this behavior. Do not edit files, run implementation steps, or make changes. Just explain what is happening.
3. Let the agent provide a complete explanatory answer without calling `task_complete`.
4. Observe that task-completion enforcement sends a follow-up message similar to:
> “You have not yet marked the task as complete using the task_complete tool. If you were planning, stop planning and start implementing.”
5. Observe that the agent may interpret this as a directive to take implementation/action-oriented steps, despite the user’s explicit “do not act” instruction.
### Expected behavior
If the active user request is clearly research-only, explanation-only, or explicitly says not to modify anything, task-completion enforcement should allow the agent to finish by calling `task_complete` with a research summary, rather than nudging it toward implementation.
At minimum, the enforcement text should not say “start implementing” when the current task has no implementation component or when the user has explicitly prohibited changes.
## Actual behavior
The agent was nudged to continue with implementation-oriented behavior after completing an explanatory/research response, despite the user saying to do nothing except respond.
### Additional context
- Copilot CLI: 1.0.77
- Mode: autopilot enabled
- OS: Linux devcontainer
- Shell: bash
- Terminal emulator: Herdr-managed terminal session
贡献指南
调研方向
在仓库中搜索引用的“start implementing”消息以及 autopilot/task-completion 强制执行路径,然后使用 issue 中仅限研究的步骤重现该行为。完成的标准是:仅请求解释的请求可以以 task_complete 摘要结束,且不会收到面向实现的指导,也不会触发违反用户明确约束的操作。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- shell
- 领域
- ai, cli
- Issue 类型
- 缺陷
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 活跃度
- 冷清
- 描述清晰度
- 基本清楚
- 新手友好度
- 45/100