Skip to slide
Chapter 3 · Let's Verify Step by Step, Rewarding Good Reasoning
25 / 53

CHAPTER 03 · Let's Verify Step by Step, Rewarding Good Reasoning · 5 / 6

Why this paper mattered

This work established a principle that shapes today's most advanced reasoning systems: the path to an answer is worth supervising, not just the destination. The idea of using a verifier to check reasoning and select the best of many attempts is now a standard tool. It connects directly to the next chapter, where DeepSeek-R1 uses reward signals to teach a model to reason, and to the recurring theme of this folder, spending extra effort, here in the form of generating and checking many solutions, to get harder problems right.

← → arrow keys work too