CHAPTER 03 · Let's Verify Step by Step, Rewarding Good Reasoning · 4 / 6
Why process supervision wins
Two reasons, both worth understanding:
First, precision of feedback. If a solution goes wrong at step 5 of 10, outcome supervision only knows "the whole thing is wrong." Process supervision pinpoints step 5. That precise signal is far more useful for both selecting good solutions and training better models, the way a teacher's targeted correction beats a bare "wrong" stamped on the page.
Second, safety and trust. Process supervision rewards reasoning that is actually sound, not answers that merely happen to be right. A model trained to reason correctly at every step is one you can trust more, because it is not being rewarded for lucky guesses or hidden leaps. This alignment of "good reasoning" with "reward" is part of why the idea matters beyond just raw scores.