CHAPTER 07 · Glossary: Foundational Modelling · 21 / 27
Reward model
A reward model is a separate model trained to predict how much a human would like a given answer. It is built from preference data: humans rank several answers, and the reward model learns to reproduce those judgments, outputting a score for any answer it sees.
Its purpose is to bottle human judgment into automatic software, so that during reinforcement learning the main model can be graded on millions of its own attempts without needing a human in the loop each time.