← Concept Index

Evaluation & observability

Dangerous-capability evaluations

Also called: frontier safety evals, misuse capability testing, uplift evaluations

DEFINITION

Structured tests that probe whether a model can meaningfully help with high-consequence misuse — for example real uplift on weapons, large-scale cyberattacks, or autonomous self-directed action — run before deployment to decide what safeguards a model needs.

WHY IT MATTERS

Ordinary quality evals ask 'is it good?'; these ask 'could it help someone cause serious harm, and by how much beyond what they could already do?'. They're the measurement that graduated safety commitments depend on.

COMMONLY CONFUSED WITH

Red-teaming for jailbreaks. Jailbreak red-teaming asks whether safeguards can be bypassed; capability evals ask whether the underlying model has the dangerous ability at all.

SOURCES

First-pass citations, limited to primary sources; a reviewer will broaden and verify these before this entry leaves draft.