Direct Preference Optimization skips the reward-model + RL loop. Given pairs of (preferred, rejected) responses, DPO directly adjusts the model to make preferred answers more likely and rejected ones less likely, via a clean loss function. It's simpler, more stable, and cheaper than RLHF while achieving comparable alignment — which is why it (and variants) became popular for open models.