freecoding.school100% FREE · NO SIGNUP
Tensor TownISSUE #113 of 120

mechanistic interpretability · sae · probing

NeuraVSThe Overfit Ogre
Neura saysMechanistic interpretability reverse-engineers the actual computations inside a network.

The most ambitious strand: mechanistic interpretability aims to reverse-engineer a model into human-understandable algorithms — identifying features (concepts a neuron/direction represents) and circuits (how they combine to do a task). Sparse autoencoders help untangle "superposition" (many concepts crammed into few neurons). If it succeeds, we could audit models for deception or danger directly — a major safety bet.

Power-ups you unlock

The Overfit Ogre attacks — common mistakes

Boss battleExplain what "features" and "circuits" mean in this context.

Example code

<!doctype html><html><head><meta charset="utf-8"></head>
<body style="background:#06040d;color:#e6e0ff;font-family:monospace;padding:20px"><pre>feature: a direction representing a concept
circuit: features combined to perform a task
SAEs untangle "superposition"
→ audit for deception/danger</pre></body></html>
▶ Open the interactive comic issue
‹ Interpretability · Features · CircuitsAi Safety Institute · Evals · Frontier Rules ›