Frontier risk and graduated safeguards
Also called: responsible scaling, capability thresholds, if-then safety commitments
An approach to governing the most capable ('frontier') models by tying safeguards to capability: as evaluations show a model approaching a defined threshold of dangerous capability, stronger deployment and security measures are triggered — an 'if it can do X, then we must do Y' commitment.
It's how several developers and emerging policies try to manage low-probability, high-consequence risks (such as uplift to weapons or large-scale cyberattacks) without waiting for harm to happen first. It makes capability evaluation, not just a release date, the thing that gates deployment.
The EU-style risk tiers for use cases. Those classify how a system is used; frontier safeguards classify what a model is capable of, and scale the developer's own controls accordingly.
- Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1U.S. National Institute of Standards and Technology · 2023-01-26
First-pass citations, limited to primary sources; a reviewer will broaden and verify these before this entry leaves draft.