Prompt Injection Filtering
Nobody has claimed this yet.
- Dominant language
- Go
- Stars
- 2.2k
- Forks
- 300
- Avg merge
- 1d 15h
- Merged PRs (30d)
- 184
Description
Has Stacklok considered providing any functionality that would filter tool outputs for prompt injection attempts?
I'm concerned about tools that:
- Are malicious and injection malicious commands to the LLM to tool responses
- Are not malicious but end up passing through malicious content they ingest from the web
I would imagine having optional functionality that would allow the filtering as a paid feature, or using an api key from an ai vendor such as Lakera or Enkrypt could work, if Stacklok didn't have the filtering capabilities in house.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
No files, tests, or entry points are named. Start by clarifying the filtering requirements, including malicious tool output and untrusted web content, then determine whether optional filtering or integration with Lakera or Enkrypt is in scope; done would mean an agreed, implementable feature definition.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- go
- Domain
- ai, security
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100