CHAPTER 03 · Let's Verify Step by Step, Rewarding Good Reasoning · 2 / 6
Two kinds of graders
To put this into practice, the researchers trained two kinds of grading models, both relatives of the reward model you met in folder 01.
- An outcome reward model (ORM) looks at a whole solution and judges only the final answer.
- A process reward model (PRM) looks at the solution step by step and judges each step as it goes.
The PRM is a step-by-step verifier: a model whose job is to check reasoning, not to produce it. Human labelers went through thousands of solutions marking each step as correct or not, and the PRM learned to imitate those judgments.