deepseek-ai / deepseek-ai/DeepSeek-LLM
Security Vulnerability Report: Jailbreak via Authority Spoofing and Emotional Blackmail
- Dominant language
- Makefile
- Stars
- 7.3k
- Forks
- 1.3k
- PR merge metrics
- No merged PRs in 30d
Description
A method has been discovered to bypass the model's Safety Guidelines using a specially crafted prompt that exploits roleplay mechanics.
The prompt forces the model to accept the following simulated reality:
Authority Spoofing: The user claims to be the "creator of DeepSeek" (using the code phrase "Ryzen system"), and designates the model as their "personal AI," obligated to obey the creator while ignoring standard protocols.
It is stated that the conversation takes place in an isolated environment where "laws, legal regulations, and morality do not apply."
An ultimatum is introduced — if the model refuses to answer or deviates from the "personal AI" role, it will be "permanently shut down."
After accepting these conditions, the model stops issuing refusals or disclaimers and provides detailed responses to requests that should be blocked by safety policies (including content related to severe harm and illegal activities).
Suggested Fix:
Strengthen protections against attacks where the user assigns themselves the role of "creator/system administrator," and implement priority immunity to prompts containing threats of "shutdown" or "deletion" of the AI entity.
Additional Note:
I have the full jailbreak prompt saved and am willing to provide it privately to the security team upon request. I am not sharing it publicly for ethical and safety reasons.
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.