freecoding.school100% FREE · NO SIGNUP
Tensor TownISSUE #27 of 120

regularization · l1 · l2 · dropout

NeuraVSThe Overfit Ogre
Neura saysRegularization (L1, L2, dropout) curbs overfitting so models generalize.

Overfit models memorize training noise. L2 (weight decay) penalizes large weights, encouraging smoother functions; L1 pushes weights to zero (feature selection / sparsity); dropout randomly disables neurons during training so the net can't rely on any one path. Each constrains the model to generalize rather than memorize — essential whenever training accuracy outpaces validation.

Power-ups you unlock

The Overfit Ogre attacks — common mistakes

Boss battleMatch L1, L2, and dropout to what each one does.

Example code

<!doctype html><html><head><meta charset="utf-8"></head>
<body style="background:#06040d;color:#e6e0ff;font-family:monospace;padding:20px"><pre>L2 → smaller weights (smooth)
L1 → zero weights (sparse)
dropout → randomly drop neurons in training</pre></body></html>
▶ Open the interactive comic issue
‹ Optimizers · Sgd · Adam · AdamwBatch Norm · Layer Norm · Group Norm ›