freecoding.school100% FREE · NO SIGNUP
Tensor TownISSUE #114 of 120

ai safety institute · evals · frontier rules

NeuraVSThe Overfit Ogre
Neura saysAI safety institutes and frontier rules aim to evaluate and govern the most capable models.

As models grew powerful, governments and labs stood up AI Safety Institutes (UK, US) and frontier-safety frameworks to test models for dangerous capabilities (cyber, bio, autonomy) before release, and to set thresholds for extra scrutiny. It's nascent and contested — how to test, who decides, voluntary vs mandatory — but reflects a shift: frontier AI is treated as something to evaluate and govern, not just ship.

Power-ups you unlock

The Overfit Ogre attacks — common mistakes

Boss battleName a dangerous-capability category these evaluations probe.

Example code

<!doctype html><html><head><meta charset="utf-8"></head>
<body style="background:#06040d;color:#e6e0ff;font-family:monospace;padding:20px"><pre>before release: test for cyber/bio/autonomy risks
thresholds → extra scrutiny
UK/US AI Safety Institutes</pre></body></html>
▶ Open the interactive comic issue
‹ Mechanistic Interpretability · Sae · ProbingAgi · The Moving Target ›