Skip to slide
Chapter 3 · Let's Verify Step by Step, Rewarding Good Reasoning
20 / 53

CHAPTER 03 · Let's Verify Step by Step, Rewarding Good Reasoning

Let's Verify Step by Step, Rewarding Good Reasoning

Paper: Let's Verify Step by Step (2023)

Chapters 1 and 2 got the model to reason and act. But a hard question lurks underneath: when a model produces a long chain of reasoning, how do we judge it? The obvious answer is to check whether the final answer is right. This paper shows that the obvious answer is not the best one. It turns out that grading each step of the reasoning, rather than just the final result, makes models dramatically better at hard problems. This is a subtle idea with big consequences.

← → arrow keys work too