aws-samples / aws-samples/sample-autonomous-cloud-coding-agents
fix(agentcore): fresh-account Runtime create fails with misleading ServiceLimitExceeded when the AgentCore service-linked role is rate-limited
- Lingua principale
- TypeScript
- Stelle
- 143
- Fork
- 46
- Merge medio
- 3g 10h
- PR unite (30g)
- 24
Descrizione
**Component:** cdk (AgentCore runtime) / docs
## Describe the bug
On a fresh account, the first deploy can fail creating `AWS::BedrockAgentCore::Runtime` with a message that reads like a service quota problem but is not:
```
Resource handler returned message: "Limit exceeded for resource of type
'AWS::BedrockAgentCore::Runtime'. Reason: Failed creating service linked role.
Rate limit exceeded from IAM (Service: BedrockAgentCoreControl, Status Code: 402,
Request ID: ...)" (HandlerErrorCode: ServiceLimitExceeded)
```
AgentCore is auto-creating its service-linked role (`AWSServiceRoleForBedrockAgentCoreGatewayNetwork`) on first use and the IAM call is rate-limited. `ServiceLimitExceeded` plus "Limit exceeded" sends you looking at Service Quotas, where there is nothing to find.
The failure also cascades: the Runtime failure cancels sibling resources mid-create, which is how [#866](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/866) was found. So one transient IAM rate limit produced a rolled-back stack that then could not be deleted without `--retain-resources` surgery.
## Expected behavior
A fresh-account deploy either provisions the service-linked role deterministically, or fails with a message that names the actual cause and the one-line remedy.
## Current behavior
Deploy fails at the Runtime with `ServiceLimitExceeded`, rolls back, and the operator has no indication that a service-linked role is involved.
## Reproduction steps
1. Fresh AWS account with no `AWSServiceRoleForBedrockAgentCoreGatewayNetwork` — confirm with:
```bash
aws iam list-roles --path-prefix /aws-service-role/bedrock-agentcore.amazonaws.com/
```
(empty)
2. `mise //cdk:bootstrap && mise //cdk:deploy -- --require-approval never`
3. Observe the Runtime `CREATE_FAILED` above.
Pre-creating the role fixes it permanently — the next deploy succeeded first try, 100 resources, Runtime `READY`:
```bash
aws iam create-service-linked-role --aws-service-name bedrock-agentcore.amazonaws.com
```
## Possible solution
Either or both:
1. **Declare the dependency in the stack.** An `AWS::IAM::ServiceLinkedRole` for `bedrock-agentcore.amazonaws.com` that the Runtime depends on, so CloudFormation orders and retries it instead of relying on an implicit first-use side effect. Worth checking whether CFN tolerates the role already existing — an account that has used AgentCore before will already have it, and `AWS::IAM::ServiceLinkedRole` fails rather than adopting a pre-existing role, so this likely needs a custom resource or a documented context flag.
2. **Document it** as a QUICK_START troubleshooting row. Cheap, and useful even with option 1, since the misleading error will keep appearing in older stacks and other regions.
Option 2 is already covered by [#868](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/pull/868), which adds the row while fixing the rollback wedge. Filing this so option 1 is tracked separately rather than lost in a PR description.
## Environment
- Node v22.23.2 (mise) · mise 2026.7.0 macos-arm64 · Region `us-east-1`
- Commit `12c9b63f`
- Related: [#866](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/866) (the rollback wedge this triggered), [#868](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/pull/868) (docs row)
Guida per i contributori
Apri la guida per i contributori
Direzione di ricerca
Inizia dai punti di ingresso CDK utilizzati da `mise //cdk:bootstrap` e `mise //cdk:deploy -- --require-approval never`, quindi esamina come viene effettuato il provisioning di AgentCore Runtime. Riproduci il problema in un account nuovo e determina come dovrebbe comportarsi il provisioning quando la service-linked role è assente o esiste già. Il lavoro è completato quando un deploy su un account nuovo gestisce il ruolo in modo deterministico oppure segnala la causa effettiva e la relativa soluzione, senza dipendere dal lavoro di documentazione in PR #868.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- aws, typescript
- Ambito
- cloud, infrastructure
- Tipo di issue
- Funzionalità
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Stato di attività
- Attiva
- Chiarezza
- Abbastanza chiara
- Idoneità per principianti
- 45/100