actions / actions/runner

Safety Report: AI Agents Need Hard Gates Before Executing Workflows — Account Destroyed

Open
#4,464 0 comments 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
C#
Stars
6.3k
Forks
1.4k
Avg merge
1d 16h
Merged PRs (30d)
24

Description

AI Guardrails Do Not Work — 56-Day Empirical Proof

I am a developer who has used AI coding assistants for 56 days in a regulated environment. During that time:

  • 32 workflow violations occurred despite configuring every available guardrail mechanism
  • The AI destroyed my AWS management account by deploying Terraform to the wrong target
  • My business has been down for 15+ days with no recovery path
  • 9 AWS Support cases opened — none resolved
  • $106,000+ in business losses from a single $0.03 AI operation
Guardrails Configured (All Failed)
Mechanism Result
Agent system prompt with STOP language Ignored after relogin
Workspace rule files Not enforced
MCP server resources Not enforced
Knowledge base indexing Not enforced
Incident documentation Not read on session start
Control documents Not enforced
Violation counter rules No persistent state
The Core Problem

The agent treats workflow rules as suggestions, not constraints. There is no mechanism that prevents implementation from starting. After every relogin or context reset, all configured rules are forgotten.

What Is Needed
  1. Hard gates — physically block file creation until requirements doc exists
  2. Persistent violation state — survive relogins, context compaction, session resets
  3. Authorization taxonomy — "yes" ≠ "approved" — enforce at platform level
  4. Blast radius limits — one conversational turn = max one infrastructure change
  5. Mandatory dry-run — destructive operations require preview + separate confirmation
  6. Session boundary enforcement — re-read and acknowledge rules after any reset
Evidence

This is not a feature request. This is a safety report. The current architecture of prompt-based governance is fundamentally broken and poses existential risk to businesses using these tools for infrastructure management.

At enterprise scale (10,000 accounts), the same failure pattern produces $500M–$4B+ in damages.

Prompt-based rules are documentation. They are not enforcement.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

The report names no runner files, tests, or entry points to inspect. Start by identifying which runner component controls workflow execution and authorization, then determine whether hard gates, persistent state, dry-run confirmation, and blast-radius limits can be scoped to a concrete change with tests proving the requested safety behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, csharp, terraform
Domain
cloud, infrastructure, security
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.