aws-samples / aws-samples/sample-autonomous-cloud-coding-agents
agent: add deterministic secret and scope regex guards alongside Cedar
- Langage dominant
- TypeScript
- Étoiles
- 143
- Forks
- 46
- Merge moyen
- 3 j 10 h
- PR mergées (30 j)
- 24
Description
### Component
Agent (Python runtime)
### Describe the feature
Add a **fast, deterministic** PreToolUse guard layer in `agent/src/hooks.py` (or a dedicated `regex_guards.py`) that runs **before** Cedar evaluation for selected tool classes. Guards should be pure regex / JSON field checks with no LLM involvement.
Minimum guard set:
| Guard | Triggers on | Blocks |
|-------|-------------|--------|
| **Secret** | `Write`, `Edit`, bash heredocs, Gateway `push_files` payloads | AWS access keys (`AKIA…`), PEM blocks, `ghp_`/`gho_`/`github_pat_`, `sk-…`, Slack tokens |
| **Scope** | Gateway GitHub tools, git push MCP | `owner`/`repo`/`issue`/`branch` mismatch vs task context; fail-closed if task scope metadata missing |
| **Bash** | `Bash` tool | `rm -rf /`, `git push --force` to protected refs, `curl`/`wget` with `--data` exfil patterns, dumping `env`/`printenv` |
Emit `tool_decision` / progress events when a guard blocks (for observability parity with Cedar denies).
### Use case
Cedar policies express intent well but are not ideal for high-velocity pattern matching on raw file content or shell one-liners. A compromised or jailbroken model might still reach bash or write paths before policy catches edge cases.
Operators need defense-in-depth: Cedar for governance and HITL, regex guards for known credential and exfiltration shapes that must **never** reach the repo.
### Proposed solution
1. Implement guards as pure functions returning `ALLOW` / `DENY` + reason string.
2. Wire into existing `PreToolUse` hook path in `hooks.py` **before** `PolicyEngine.evaluate()`.
3. Load scope from `agent/src/config.py` task context (`repo`, `branch`, `issue_number`, etc.) — same fields hydration already sets.
4. Add `agent/tests/test_regex_guards.py` with table-driven cases (≥10 per guard): allowed benign cases + blocked malicious cases.
5. Document in `docs/design/SECURITY.md` under §Tool access control as "Layer 0: deterministic guards".
6. Optional: mirror critical patterns in `output_scanner.py` PostToolUse for defense in depth.
### Other information
- Complements existing Cedar hard/soft deny (`docs/design/CEDAR_HITL_GATES.md`) — does not replace it.
- Related: #225 (egress DLP on progress events / PR content), #390 (execution-layer hardening below tool policy). This issue is **PreToolUse input** blocking; #225 is egress; `output_scanner.py` today is PostToolUse only.
- Scope guard should align with per-task SessionRole tags (`repo`, `task_id`) semantics in `docs/design/SECURITY.md`.
- Consider `nosemgrep` only if semgrep flags intentional test fixtures.
### Acknowledgements
- [x] I may be able to implement this feature
- [ ] This might be a breaking change
Guide de contribution
Ouvrir le guide de contribution
Piste de recherche
Commencez par agent/src/hooks.py et le chemin PreToolUse existant, puis examinez agent/src/config.py pour l’hydratation de task-scope et l’appel d’évaluation de Cedar. Utilisez agent/tests/test_regex_guards.py pour les cas bénins et malveillants pilotés par table, et consultez docs/design/SECURITY.md ainsi que output_scanner.py pour les conventions existantes en matière de sécurité et d’observabilité. C’est terminé lorsque les guards s’exécutent avant Cedar, émettent des événements de blocage, disposent de la couverture demandée et sont documentées.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Évaluation
- Stack technique
- github, python
- Domaine
- security
- Type d'issue
- Fonctionnalité
- Difficulté
- 4/5
- Temps estimé
- 3-5 jours
- Activité
- Calme
- Clarté
- Plutôt claire
- Accessibilité débutants
- 48/100