Skip to slide
Chapter 3 · Let's Verify Step by Step, Rewarding Good Reasoning
23 / 53

CHAPTER 03 · Let's Verify Step by Step, Rewarding Good Reasoning · 3 / 6

How a verifier makes a model smarter

Here is how this actually boosts performance. You let the main model generate many candidate solutions to a hard problem (say, dozens of them). Most will be wrong in various ways. Then the verifier scores them, and you pick the solution it rates most highly.

flowchart TD
    P[Hard problem] --> Gen[Model generates<br/>many candidate solutions]
    Gen --> V[Verifier scores each one]
    V --> Pick[Keep the best-scored solution]
    Pick --> Ans[Final answer]

The paper's key result: a process-based verifier picks the right solution far more reliably than an outcome-based one, especially on genuinely hard math. By checking the reasoning rather than just the answer, the PRM is much harder to fool with confident-sounding but flawed solutions.

← → arrow keys work too