CHAPTER 03 · Alignment, Turning a Text Predictor Into a Helpful Assistant · 3 / 5
Paper 2: DPO, the same goal with far less machinery
RLHF works, but it is complicated and fragile. You have to train and maintain a separate reward model, then run a delicate reinforcement learning loop that is hard to get right. The DPO paper asked: can we get the same result more simply?
The answer is yes, and the insight is captured by the paper's subtitle: your language model is secretly a reward model. The authors showed with math that you do not need a separate reward model and an RL loop at all. You can fold the whole thing into a single, direct training step on the human comparison data.
DPO works straight from preference pairs: a prompt, a preferred answer, and a rejected answer. It then trains the model with one simple objective: make the preferred answer more likely and the rejected answer less likely, while gently staying close to the original model (the same leash idea as before, baked right in).
flowchart LR
subgraph RLHF[RLHF: three moving parts]
a1[Reward model] --> a2[RL loop with PPO] --> a3[Aligned model]
end
subgraph DPO[DPO: one step]
b1[Preferred vs rejected pairs] --> b2[One direct training objective] --> b3[Aligned model]
end
A simple analogy
RLHF is like hiring a judge (the reward model), then having a coach (PPO) run an athlete through endless practice rounds in front of that judge. DPO removes the judge and the coach. It just shows the athlete pairs of "this was good, that was bad" and lets them learn directly. Fewer moving parts, less that can break.
Why DPO caught on
DPO is simpler to implement, cheaper to run, and more stable than full RLHF. For many teams it produces results just as good. It does not completely replace RLHF (the big labs still use reinforcement learning for the hardest cases, as you will see with DeepSeek-R1 in folder 02), but DPO made high-quality alignment accessible to far more people.