← Concept Index

Agent & LLM security

Jailbreaking

Also called: bypassing safety, DAN prompts

DEFINITION

Crafting input that gets a model to ignore its safety training and produce content it was tuned to refuse — via role-play framings, obfuscation, or step-by-step coaxing.

WHY IT MATTERS

It shows that alignment is a tendency, not a wall: safeguards baked into the model can be talked around. Real protection also needs limits outside the model — on tools, data, and actions.

COMMONLY CONFUSED WITH

Prompt injection. Jailbreaking targets the safety rules; injection targets the task. They often combine but aren't the same.

SOURCES

First-pass citations, limited to primary sources; a reviewer will broaden and verify these before this entry leaves draft.