aws-samples / aws-samples/sample-autonomous-cloud-coding-agents
agent: add deterministic secret and scope regex guards alongside Cedar
- Vorherrschende Sprache
- TypeScript
- Sterne
- 143
- Forks
- 46
- Ø Merge
- 3 T. 10 Std.
- Gemergte PRs (30 T.)
- 24
Beschreibung
### Component
Agent (Python runtime)
### Describe the feature
Add a **fast, deterministic** PreToolUse guard layer in `agent/src/hooks.py` (or a dedicated `regex_guards.py`) that runs **before** Cedar evaluation for selected tool classes. Guards should be pure regex / JSON field checks with no LLM involvement.
Minimum guard set:
| Guard | Triggers on | Blocks |
|-------|-------------|--------|
| **Secret** | `Write`, `Edit`, bash heredocs, Gateway `push_files` payloads | AWS access keys (`AKIA…`), PEM blocks, `ghp_`/`gho_`/`github_pat_`, `sk-…`, Slack tokens |
| **Scope** | Gateway GitHub tools, git push MCP | `owner`/`repo`/`issue`/`branch` mismatch vs task context; fail-closed if task scope metadata missing |
| **Bash** | `Bash` tool | `rm -rf /`, `git push --force` to protected refs, `curl`/`wget` with `--data` exfil patterns, dumping `env`/`printenv` |
Emit `tool_decision` / progress events when a guard blocks (for observability parity with Cedar denies).
### Use case
Cedar policies express intent well but are not ideal for high-velocity pattern matching on raw file content or shell one-liners. A compromised or jailbroken model might still reach bash or write paths before policy catches edge cases.
Operators need defense-in-depth: Cedar for governance and HITL, regex guards for known credential and exfiltration shapes that must **never** reach the repo.
### Proposed solution
1. Implement guards as pure functions returning `ALLOW` / `DENY` + reason string.
2. Wire into existing `PreToolUse` hook path in `hooks.py` **before** `PolicyEngine.evaluate()`.
3. Load scope from `agent/src/config.py` task context (`repo`, `branch`, `issue_number`, etc.) — same fields hydration already sets.
4. Add `agent/tests/test_regex_guards.py` with table-driven cases (≥10 per guard): allowed benign cases + blocked malicious cases.
5. Document in `docs/design/SECURITY.md` under §Tool access control as "Layer 0: deterministic guards".
6. Optional: mirror critical patterns in `output_scanner.py` PostToolUse for defense in depth.
### Other information
- Complements existing Cedar hard/soft deny (`docs/design/CEDAR_HITL_GATES.md`) — does not replace it.
- Related: #225 (egress DLP on progress events / PR content), #390 (execution-layer hardening below tool policy). This issue is **PreToolUse input** blocking; #225 is egress; `output_scanner.py` today is PostToolUse only.
- Scope guard should align with per-task SessionRole tags (`repo`, `task_id`) semantics in `docs/design/SECURITY.md`.
- Consider `nosemgrep` only if semgrep flags intentional test fixtures.
### Acknowledgements
- [x] I may be able to implement this feature
- [ ] This might be a breaking change
Beitragsleitfaden
Rechercherichtung
Beginne mit agent/src/hooks.py und dem bestehenden PreToolUse-Pfad. Untersuche anschließend agent/src/config.py auf die Hydrierung des task-scope und den Cedar-Evaluierungsaufruf. Verwende agent/tests/test_regex_guards.py für tabellengesteuerte gutartige und bösartige Fälle und prüfe docs/design/SECURITY.md sowie output_scanner.py auf bestehende Konventionen für Sicherheit und Observability. Fertig ist es, wenn die Guards vor Cedar ausgeführt werden, blockierende Events ausgeben, die angeforderte Abdeckung vorhanden ist und dies dokumentiert ist.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- github, python
- Bereich
- security
- Issue-Typ
- Feature
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Aktivitätsstatus
- Ruhig
- Klarheit
- Größtenteils klar
- Anfängerfreundlichkeit
- 48/100