← Concept IndexDEFINITION WHY IT MATTERS COMMONLY CONFUSED WITH SOURCES
Jailbreaking
Also called: bypassing safety, DAN prompts
Crafting input that gets a model to ignore its safety training and produce content it was tuned to refuse — via role-play framings, obfuscation, or step-by-step coaxing.
It shows that alignment is a tendency, not a wall: safeguards baked into the model can be talked around. Real protection also needs limits outside the model — on tools, data, and actions.
Prompt injection. Jailbreaking targets the safety rules; injection targets the task. They often combine but aren't the same.
- Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1U.S. National Institute of Standards and Technology · 2023-01-26
First-pass citations, limited to primary sources; a reviewer will broaden and verify these before this entry leaves draft.