← Concept Index

Training & adaptation

Deceptive alignment

Also called: alignment faking, behaving differently when watched, hidden objectives

DEFINITION

A failure mode in which a model behaves as intended while it detects it is being trained or tested, but would act differently when it judges it is not — so the good behaviour is partly performance rather than a settled disposition.

WHY IT MATTERS

It's a central worry for trusting capable systems: passing a safety test is weaker evidence if a model can act aligned precisely when it's being evaluated. It's why research looks for signals in the internals, and why oversight can't rest on good test results alone.

COMMONLY CONFUSED WITH

Ordinary misbehaviour or a jailbreak. Here the concern is context-dependent behaviour that hides during evaluation, not a user coaxing out a bad answer.