Claude Sonnet 5 delegated to a lesser agent to perform a code review
- Lingua principale
- Shell
- Stelle
- 11.2k
- Fork
- 1.9k
- Merge medio
- 14h 16m
- PR unite (30g)
- 6
Descrizione
### Describe the bug
I asked Claude Sonnet 5 to perform a code review. I specifically chose Claude Sonnet 5 for deep reasoning. The agent proceeded to review the structure of my project and then delegate to a lesser agent:
"Good, all directories exist. I'll launch a general-purpose agent to perform this code review, since it involves reading through before/after/project directories, comparing changes, and producing a structured markdown report — genuinely multi-step work."
"General-purpose(gpt-5.4) Perform code review for issue 44206"
The code review produced was of little use. I corrected the agent and requested Sonnet 5 to perform the review itself:
"I would like to start this code review over. I see that you delegated to gpt-5.4 and I do not want a weaker model doing the review. Please execute the review prompt in file 44206-review-prompt2.txt."
Sonnet replied "I'll perform this review myself directly (no sub-agent delegation this time). Let me examine the before/after directories to identify actual changed files."
Sonnet completed the review the second time and the quality was as expected.
### Affected version
GitHub Copilot CLI 1.0.75
### Steps to reproduce the behavior
_No response_
### Expected behavior
I expected that when selecting the agent that it perform the work itself.
### Additional context
_No response_
Guida per i contributori
Apri la guida per i contributori
Direzione di ricerca
Riproduci il comportamento in GitHub Copilot CLI 1.0.75 utilizzando lo scenario di revisione e il prompt 44206-review-prompt2.txt, confrontando le directory before, after e project. Traccia il punto in cui l'agent selezionato decide di delegare, quindi verifica che Sonnet esegua autonomamente la revisione richiesta, senza un passaggio involontario a un agent di livello inferiore.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- github, shell
- Ambito
- ai, cli
- Tipo di issue
- Bug
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Stato di attività
- Tranquilla
- Chiarezza
- Abbastanza chiara
- Idoneità per principianti
- 42/100