CHAPTER 03 · Let's Verify Step by Step, Rewarding Good Reasoning · 3 / 6
How a verifier makes a model smarter
Here is how this actually boosts performance. You let the main model generate many candidate solutions to a hard problem (say, dozens of them). Most will be wrong in various ways. Then the verifier scores them, and you pick the solution it rates most highly.
flowchart TD
P[Hard problem] --> Gen[Model generates<br/>many candidate solutions]
Gen --> V[Verifier scores each one]
V --> Pick[Keep the best-scored solution]
Pick --> Ans[Final answer]
The paper's key result: a process-based verifier picks the right solution far more reliably than an outcome-based one, especially on genuinely hard math. By checking the reasoning rather than just the answer, the PRM is much harder to fool with confident-sounding but flawed solutions.