CHAPTER 03 · Alignment, Turning a Text Predictor Into a Helpful Assistant · 5 / 5
The one-sentence takeaway
A pre-trained model is just a powerful autocomplete, and alignment is the process of teaching it human preferences, either through the three-step RLHF pipeline (demonstrations, a reward model, then reinforcement learning) or through DPO, which reaches the same place with a single, simpler training step.
Next: Chapter 4, LoRA, where we learn to customize giant models without retraining all of them.