AI Safety Institute Tutorial: Evals, Frontier Rules

TL;DRAI safety institutes and frontier rules aim to evaluate and govern the most capable models.

As models grew powerful, governments and labs stood up AI Safety Institutes (UK, US) and frontier-safety frameworks to test models for dangerous capabilities (cyber, bio, autonomy) before release, and to set thresholds for extra scrutiny. It's nascent and contested — how to test, who decides, voluntary vs mandatory — but reflects a shift: frontier AI is treated as something to evaluate and govern, not just ship.

Key points

Common mistakes

Try it: Name a dangerous-capability category these evaluations probe.

Example code

<!doctype html><html><head><meta charset="utf-8"></head>
<body style="background:#06040d;color:#e6e0ff;font-family:monospace;padding:20px"><pre>before release: test for cyber/bio/autonomy risks
thresholds → extra scrutiny
UK/US AI Safety Institutes</pre></body></html>
Open the interactive lesson →
Mechanistic Interpretability · Sae · Probing Agi · The Moving Target