freecoding.school100% FREE · NO SIGNUP
Tensor TownISSUE #53 of 120

scaling laws · params · data · compute

NeuraVSThe Overfit Ogre
Neura saysScaling laws: model quality improves predictably with more parameters, data, and compute.

Empirical scaling laws show loss falls as a smooth power law as you increase model size, dataset size, and compute together. This predictability — bigger + more data + more compute = better — drove the race to ever-larger models. The Chinchilla result refined it: for a compute budget, many models were under-trained on data; balance parameters and tokens. Scaling, not just clever ideas, drove recent leaps.

Power-ups you unlock

The Overfit Ogre attacks — common mistakes

Boss battleState the Chinchilla insight about balancing model size and data.

Example code

<!doctype html><html><head><meta charset="utf-8"></head>
<body style="background:#06040d;color:#e6e0ff;font-family:monospace;padding:20px"><pre>more params + data + compute → lower loss (power law)
Chinchilla: don’t over-size params, feed enough tokens</pre></body></html>
▶ Open the interactive comic issue
‹ Context Windows · How Much Fits At OnceEmergent Abilities · The Surprise Phase ›