Neura saysScaling laws: model quality improves predictably with more parameters, data, and compute.
Empirical scaling laws show loss falls as a smooth power law as you increase model size, dataset size, and compute together. This predictability — bigger + more data + more compute = better — drove the race to ever-larger models. The Chinchilla result refined it: for a compute budget, many models were under-trained on data; balance parameters and tokens. Scaling, not just clever ideas, drove recent leaps.
Power-ups you unlock
Loss falls as a power law with scale
Parameters + data + compute together
Predictability fueled the scale-up race
Chinchilla: balance params and tokens
The Overfit Ogre attacks — common mistakes
Adding parameters without enough data (Chinchilla lesson)
Assuming scaling improves everything equally
Expecting scaling laws to hold forever
Boss battleState the Chinchilla insight about balancing model size and data.