aws-samples / aws-samples/sample-autonomous-cloud-coding-agents

agent: add deterministic secret and scope regex guards alongside Cedar

Ouverte
#549 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub
agent-runtime enhancement P1 security
Langage dominant
TypeScript
Étoiles
143
Forks
46
Merge moyen
3 j 10 h
PR mergées (30 j)
24

Description

### Component

Agent (Python runtime)

### Describe the feature

Add a **fast, deterministic** PreToolUse guard layer in `agent/src/hooks.py` (or a dedicated `regex_guards.py`) that runs **before** Cedar evaluation for selected tool classes. Guards should be pure regex / JSON field checks with no LLM involvement.

Minimum guard set:

| Guard | Triggers on | Blocks |
|-------|-------------|--------|
| **Secret** | `Write`, `Edit`, bash heredocs, Gateway `push_files` payloads | AWS access keys (`AKIA…`), PEM blocks, `ghp_`/`gho_`/`github_pat_`, `sk-…`, Slack tokens |
| **Scope** | Gateway GitHub tools, git push MCP | `owner`/`repo`/`issue`/`branch` mismatch vs task context; fail-closed if task scope metadata missing |
| **Bash** | `Bash` tool | `rm -rf /`, `git push --force` to protected refs, `curl`/`wget` with `--data` exfil patterns, dumping `env`/`printenv` |

Emit `tool_decision` / progress events when a guard blocks (for observability parity with Cedar denies).

### Use case

Cedar policies express intent well but are not ideal for high-velocity pattern matching on raw file content or shell one-liners. A compromised or jailbroken model might still reach bash or write paths before policy catches edge cases.

Operators need defense-in-depth: Cedar for governance and HITL, regex guards for known credential and exfiltration shapes that must **never** reach the repo.

### Proposed solution

1. Implement guards as pure functions returning `ALLOW` / `DENY` + reason string.
2. Wire into existing `PreToolUse` hook path in `hooks.py` **before** `PolicyEngine.evaluate()`.
3. Load scope from `agent/src/config.py` task context (`repo`, `branch`, `issue_number`, etc.) — same fields hydration already sets.
4. Add `agent/tests/test_regex_guards.py` with table-driven cases (≥10 per guard): allowed benign cases + blocked malicious cases.
5. Document in `docs/design/SECURITY.md` under §Tool access control as "Layer 0: deterministic guards".
6. Optional: mirror critical patterns in `output_scanner.py` PostToolUse for defense in depth.

### Other information

- Complements existing Cedar hard/soft deny (`docs/design/CEDAR_HITL_GATES.md`) — does not replace it.
- Related: #225 (egress DLP on progress events / PR content), #390 (execution-layer hardening below tool policy). This issue is **PreToolUse input** blocking; #225 is egress; `output_scanner.py` today is PostToolUse only.
- Scope guard should align with per-task SessionRole tags (`repo`, `task_id`) semantics in `docs/design/SECURITY.md`.
- Consider `nosemgrep` only if semgrep flags intentional test fixtures.

### Acknowledgements

- [x] I may be able to implement this feature
- [ ] This might be a breaking change

Guide de contribution

Ouvrir le guide de contribution

Piste de recherche

Commencez par agent/src/hooks.py et le chemin PreToolUse existant, puis examinez agent/src/config.py pour l’hydratation de task-scope et l’appel d’évaluation de Cedar. Utilisez agent/tests/test_regex_guards.py pour les cas bénins et malveillants pilotés par table, et consultez docs/design/SECURITY.md ainsi que output_scanner.py pour les conventions existantes en matière de sécurité et d’observabilité. C’est terminé lorsque les guards s’exécutent avant Cedar, émettent des événements de blocage, disposent de la couverture demandée et sont documentées.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
github, python
Domaine
security
Type d'issue
Fonctionnalité
Difficulté
4/5
Temps estimé
3-5 jours
Activité
Calme
Clarté
Plutôt claire
Accessibilité débutants
48/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.