freecoding.school100% FREE · NO SIGNUP
Tensor TownISSUE #59 of 120

dpo · direct preference optimization

NeuraVSThe Overfit Ogre
Neura saysDPO aligns models directly from preferences — simpler than RLHF, no separate reward model.

Direct Preference Optimization skips the reward-model + RL loop. Given pairs of (preferred, rejected) responses, DPO directly adjusts the model to make preferred answers more likely and rejected ones less likely, via a clean loss function. It's simpler, more stable, and cheaper than RLHF while achieving comparable alignment — which is why it (and variants) became popular for open models.

Power-ups you unlock

The Overfit Ogre attacks — common mistakes

Boss battleContrast DPO and RLHF in pipeline complexity.

Example code

<!doctype html><html><head><meta charset="utf-8"></head>
<body style="background:#06040d;color:#e6e0ff;font-family:monospace;padding:20px"><pre>RLHF: rankings → reward model → RL (PPO)
DPO:  preference pairs → direct loss → done
simpler, stabler</pre></body></html>
▶ Open the interactive comic issue
‹ Rlhf · Reinforcement Learning From Human FeedbackConstitutional Ai · Self-Critique Training ›