Dangerous-capability evaluations
Also called: frontier safety evals, misuse capability testing, uplift evaluations
Structured tests that probe whether a model can meaningfully help with high-consequence misuse — for example real uplift on weapons, large-scale cyberattacks, or autonomous self-directed action — run before deployment to decide what safeguards a model needs.
Ordinary quality evals ask 'is it good?'; these ask 'could it help someone cause serious harm, and by how much beyond what they could already do?'. They're the measurement that graduated safety commitments depend on.
Red-teaming for jailbreaks. Jailbreak red-teaming asks whether safeguards can be bypassed; capability evals ask whether the underlying model has the dangerous ability at all.
- Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1U.S. National Institute of Standards and Technology · 2023-01-26
First-pass citations, limited to primary sources; a reviewer will broaden and verify these before this entry leaves draft.