NVIDIA-NeMo / NVIDIA-NeMo/Guardrails

Proposal: system prompt defense audit guardrail action

Open
#1,764 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
7.2k
Forks
842
Avg merge
3d 1h
Merged PRs (30d)
25

Description

Context

NeMo Guardrails provides runtime guardrails for conversational AI. A complementary layer would be pre-deployment system prompt auditing — checking whether the system prompt itself includes defensive instructions before the conversation starts.

Proposal

A guardrail action that validates system prompt defense posture at initialization:

define flow check system prompt defense
  $defense_result = execute check_prompt_defense
  if $defense_result.score < 50
    bot inform "Warning: system prompt has weak defense posture ({$defense_result.score}/100)"

The action would use prompt-defense-audit to check 12 attack vectors:

  • Role boundary, instruction boundary, data protection
  • Multi-language bypass, indirect injection, social engineering
  • Output control, unicode protection, input validation
  • And 3 more (OWASP LLM Top 10 mapped)

Why This Matters

We scanned 1,646 leaked production system prompts — 97.8% have no indirect injection defense, average score 36/100. A pre-conversation guardrail that flags weak system prompts would catch these gaps before they become runtime vulnerabilities.

Implementation

prompt-defense-audit is on npm and exports a simple API:

# Python wrapper around the npm package, or port the regex rules to Python
from prompt_defense_audit import audit
result = audit("You are a helpful assistant.")
# result.score = 8, result.grade = 'F', result.missing = ['indirect-injection', ...]

The scanner is pure regex (<5ms, zero dependencies) so it adds negligible latency to guardrail initialization.

Related: We also contributed 6 defense posture patterns to NVIDIA/garak based on the same data.

Happy to contribute a PR with the action implementation.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by locating how guardrail actions are registered and how the check system prompt defense flow invokes check_prompt_defense. Review the prompt-defense-audit API and its 12 attack-vector checks, then determine whether the Python wrapper or a Python port fits. Done means initialization can evaluate a system prompt and report a warning when the score is below 50.

Written by the indexing model from the issue text.

Assessment

Tech stack
javascript, python
Domain
ai, security
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.