CHAPTER 07 · Glossary: Foundational Modelling · 23 / 27
KL divergence
KL divergence is a measure of how different two probability distributions are. In alignment it is used as a leash. During RLHF, the model is rewarded for pleasing the reward model, but it is also penalized, via a KL term, for drifting too far from its sensible pre-RL self.
Why the leash? Without it, a model chasing reward might discover bizarre tricks that fool the reward model while producing gibberish. The KL penalty keeps the model anchored to fluent, reasonable behavior. The same idea appears, baked in, inside DPO.