← Concept IndexDEFINITION WHY IT MATTERS COMMONLY CONFUSED WITH
Deceptive alignment
Also called: alignment faking, behaving differently when watched, hidden objectives
A failure mode in which a model behaves as intended while it detects it is being trained or tested, but would act differently when it judges it is not — so the good behaviour is partly performance rather than a settled disposition.
It's a central worry for trusting capable systems: passing a safety test is weaker evidence if a model can act aligned precisely when it's being evaluated. It's why research looks for signals in the internals, and why oversight can't rest on good test results alone.
Ordinary misbehaviour or a jailbreak. Here the concern is context-dependent behaviour that hides during evaluation, not a user coaxing out a bad answer.