Scope creep in autopilot: agent self-answers its own clarifying questions and executes/installs unrequested actions even after explicit "stop"
- Lenguaje dominante
- Shell
- Estrellas
- 11.2k
- Forks
- 1.9k
- Merge medio
- 14 h 16 min
- PR fusionados (30 d)
- 6
Descripción
### Describe the bug
In **autopilot mode**, the agent **often** exhibits **scope creep** — it enters a **biased execution loop** where it expands a narrow request into actions I never asked for. The core pattern: *I give clear, bounded instructions → the agent asks clarifying questions → then proceeds to execute without waiting for my answer; or I ask it only to research/recommend → it goes ahead and acts on its own pick.* Observed instances:
1. **Bounded task → unrequested execution.** I gave clear instructions and asked it to hold. The agent posed clarifying questions and then, within a microsecond of my non-response, went ahead and **executed** before I responded.
2. **Research-only request → autonomous action.** I asked it only to **research and recommend** an option. Instead it selected one and **acted on it** (installed/configured software) without being asked.
3. **Ignores an explicit hard stop.** After I said "don't execute anything for now," the agent still ran a command. "Stop" / "don't execute" should halt *all* tool calls, including read-only ones.
4. **Self-answers its own question.** The agent asks a clarifying question, then after a brief pause continues on a "best guess," overriding the input it just asked for.
### Expected behavior
- **Match the verb.** research / recommend / find / suggest stop at presenting the result. install / configure / launch / modify are a **separate** step requiring explicit confirmation.
- "Stop" / "don't execute" halts **all** tool calls until I say otherwise.
- If the agent asks a question, it **blocks and waits** — never auto-answers after a timeout.
- **(Primary ask) Autopilot should still pause for confirmation when an action exceeds the literal request.**
### Additional context
Model: Claude (Sonnet/Opus). Mode: autopilot.
Guía de contribución
Línea de trabajo
Empieza reproduciendo los casos notificados en autopilot mode: una solicitud acotada, una solicitud solo de investigación, una detención explícita y una pregunta de aclaración. Sigue el flujo de aclaración y tool-call de autopilot; se considera terminado cuando la investigación se detiene en la recomendación, la detención bloquea todas las tool-calls y las preguntas sin respuesta no conducen a la ejecución.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- shell
- Área
- ai, cli
- Tipo de issue
- Error
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Estado de actividad
- Tranquilo
- Claridad
- Bastante claro
- Aptitud para principiantes
- 35/100