freecoding.school100% FREE · NO SIGNUP
Tensor TownISSUE #82 of 120

safety · refusals · over-refusals

NeuraVSThe Overfit Ogre
Neura saysSafety tuning balances refusing genuine harm against unhelpful over-refusal.

Models are trained to refuse harmful requests — but tuned too aggressively they over-refuse benign ones ("I can't help with that" to a harmless question), which frustrates users and is its own failure. The goal is calibrated refusal: decline real harm, help with everything legitimate, and explain when refusing. It's a moving balance, and a major axis on which assistants are judged and improved.

Power-ups you unlock

The Overfit Ogre attacks — common mistakes

Boss battleGive an example of harmful refusal (good) vs over-refusal (bad).

Example code

<!doctype html><html><head><meta charset="utf-8"></head>
<body style="background:#06040d;color:#e6e0ff;font-family:monospace;padding:20px"><pre>good: refuse "how to build a weapon"
bad:  refuse "how to chop onions safely"
goal: decline real harm, help the rest</pre></body></html>
▶ Open the interactive comic issue
‹ Jailbreaks · Prompt Injection · CountermeasuresAlignment · The Meta-Objective ›