CHAPTER 03 · Let's Verify Step by Step, Rewarding Good Reasoning · 5 / 6
Why this paper mattered
This work established a principle that shapes today's most advanced reasoning systems: the path to an answer is worth supervising, not just the destination. The idea of using a verifier to check reasoning and select the best of many attempts is now a standard tool. It connects directly to the next chapter, where DeepSeek-R1 uses reward signals to teach a model to reason, and to the recurring theme of this folder, spending extra effort, here in the form of generating and checking many solutions, to get harder problems right.