github / github/copilot-cli

Autopilot task-completion enforcement can override explicit user instructions

Open
#4,318 1 comment 0 reactions 0 assignees View on GitHub
area:agents area:non-interactive
Dominant language
Shell
Stars
11.2k
Forks
1.9k
Avg merge
14h 16m
Merged PRs (30d)
6

Description

### Describe the bug

In Copilot CLI autopilot mode, the task-completion enforcement behavior can cause the agent to continue taking action after the user has explicitly narrowed the task to research/explanation only.

In my case, the user explicitly instructed the agent to **do nothing besides respond** and asked for an explanation/research task. After the agent answered, a follow-up enforcement message appeared:

> “You have not yet marked the task as complete using the task_complete tool. If you were planning, stop planning and start implementing.”

That message appears to push the agent toward implementation/action even when the active user request is intentionally non-implementation work.

## Why this is a problem

This can lead to the agent making unintended workspace changes despite the user explicitly asking for diagnostic or explanatory help only. The issue is especially risky because the enforcement message is framed as higher-priority continuation guidance, so the agent may treat it as expanding the user’s requested scope.

## Suggested improvement

Adjust the completion enforcement prompt/logic so it preserves the current user’s scope constraints. For example:

- If the user asks for research/explanation only, completing the answer should be considered sufficient.
- If the user explicitly says not to modify files or take action, the continuation message should not direct the agent to “start implementing.”
- The agent should be encouraged to call `task_complete` with a concise summary when the requested non-code task is complete.

## Notes

This may be related to autopilot/task-completion behavior introduced or changed recently. The problematic phrase above is included because it seems useful for locating the relevant enforcement path.

### Affected version

GitHub Copilot CLI 1.0.77.

### Steps to reproduce the behavior

1. Start a Copilot CLI session with autopilot mode enabled.
2. Ask for a research-only or explanation-only task, and explicitly prohibit action. For example:
> Please help me understand this behavior. Do not edit files, run implementation steps, or make changes. Just explain what is happening.
3. Let the agent provide a complete explanatory answer without calling `task_complete`.
4. Observe that task-completion enforcement sends a follow-up message similar to:
> “You have not yet marked the task as complete using the task_complete tool. If you were planning, stop planning and start implementing.”
5. Observe that the agent may interpret this as a directive to take implementation/action-oriented steps, despite the user’s explicit “do not act” instruction.

### Expected behavior

If the active user request is clearly research-only, explanation-only, or explicitly says not to modify anything, task-completion enforcement should allow the agent to finish by calling `task_complete` with a research summary, rather than nudging it toward implementation.

At minimum, the enforcement text should not say “start implementing” when the current task has no implementation component or when the user has explicitly prohibited changes.

## Actual behavior

The agent was nudged to continue with implementation-oriented behavior after completing an explanatory/research response, despite the user saying to do nothing except respond.

### Additional context

- Copilot CLI: 1.0.77
- Mode: autopilot enabled
- OS: Linux devcontainer
- Shell: bash
- Terminal emulator: Herdr-managed terminal session

Contributor guide

Open the contributing guide

Research direction

Search the repository for the quoted “start implementing” message and the autopilot/task-completion enforcement path, then reproduce the behavior using the research-only steps in the issue. Done means an explanation-only request can finish with a task_complete summary without receiving implementation-oriented guidance or causing action against explicit user constraints.

Written by the indexing model from the issue text.

Assessment

Tech stack
shell
Domain
ai, cli
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.