CHAPTER 04 · DeepSeek-R1, Learning to Reason Through Reinforcement Learning · 2 / 6
DeepSeek-R1-Zero: reasoning appears on its own
The first model, called R1-Zero, was trained with pure reinforcement learning, with no supervised fine-tuning warm-up at all, no human-written reasoning examples first. It was simply rewarded for correct answers, over and over.
What happened next is the heart of the paper. With nothing but this reward, the model taught itself to reason. It began, entirely on its own, to write longer chains of thought, to double-check its own work, to try an approach and then reconsider it. Nobody programmed these behaviors in. They emerged because they led to more correct answers, and correct answers were rewarded.
The researchers describe an "aha moment," where the model learned to pause and re-evaluate its approach mid-solution, spending more thinking time on harder problems. This is a striking demonstration that the strategy of careful reasoning can be discovered through reward alone, much as a game-playing AI discovers clever tactics by being rewarded only for winning.