CHAPTER 04 · DeepSeek-R1, Learning to Reason Through Reinforcement Learning · 3 / 6
The catch, and the fix: DeepSeek-R1
R1-Zero reasoned well but had rough edges. Because it was never trained on clean human examples, its output was often hard to read and sometimes mixed languages mid-answer. It was a brilliant thinker with messy handwriting.
To fix this, the team built the full DeepSeek-R1 with a multi-stage recipe. The key addition is a small amount of cold-start data: a curated set of clean, well-formatted reasoning examples used to gently warm up the model before the reinforcement learning begins. This gives the model good habits of presentation first, and then RL sharpens its reasoning.
flowchart LR
subgraph Zero[DeepSeek-R1-Zero]
z1[Base model] --> z2[Pure reinforcement learning] --> z3[Strong reasoning,<br/>messy output]
end
subgraph Full[DeepSeek-R1]
f1[Base model] --> f2[Cold-start clean examples] --> f3[Reinforcement learning] --> f4[More stages] --> f5[Strong reasoning,<br/>clean output]
end
The payoff: DeepSeek-R1 reached performance comparable to OpenAI's o1, one of the best reasoning models in the world at the time, but as an open model the whole community could study and use.