Skip to slide
Chapter 6 · Glossary: Planning and Reasoning
51 / 53

CHAPTER 06 · Glossary: Planning and Reasoning · 10 / 12

Checking the work

Outcome supervision vs process supervision

This is the central distinction of Chapter 3. Outcome supervision judges a solution only by its final answer: right or wrong. Process supervision judges every individual step of the reasoning along the way.

Process supervision turns out to work better on hard problems, for two reasons. It gives more precise feedback (it pinpoints exactly which step failed, not just that the whole thing failed), and it rewards genuinely sound reasoning rather than lucky answers that were reached through flawed steps. In short, it cares about how you got there, not just where you ended up.

Verifier

A verifier is a model whose job is to check work rather than produce it. In Chapter 3, you generate many candidate solutions to a hard problem and use a verifier to score them and pick the best. A good verifier is hard to fool, which is what makes this approach powerful: even if most of a model's attempts are flawed, a strong verifier can reliably find the good ones.

Process reward model (PRM)

A process reward model is a verifier that grades a solution step by step, judging each reasoning step as correct or not (Chapter 3). Because it checks the whole path rather than only the final answer, it is much better at catching solutions that look right but contain hidden errors. The "Let's Verify Step by Step" paper showed that selecting solutions with a PRM beats selecting them with an answer-only grader.

Outcome reward model (ORM)

An outcome reward model is a verifier that grades a solution by its final answer alone (Chapter 3). It is simpler than a process reward model but weaker on hard problems, because it cannot tell a soundly reasoned correct answer from a lucky one, and it gives no information about where a wrong solution went astray.

Reward model

A reward model is a model trained to score how good an output is, standing in for a human judge so that scoring can be done automatically and at scale. You met it first in folder 01 as part of the alignment pipeline. In folder 02, the process and outcome reward models are specialized reward models aimed at grading reasoning. (More background in the folder 01 glossary.)


← → arrow keys work too