CHAPTER 02 · Scaling Laws and Chinchilla, How Big Should a Model Be? · 4 / 5
Putting the two papers together
These papers are not in conflict. The first one discovered that scaling works in a predictable way. The second one refined the recipe for how to scale, by balancing model size against data instead of just inflating the model.
| Question | Scaling Laws (2020) | Chinchilla (2022) |
|---|---|---|
| Does scaling help predictably? | Yes, it follows a smooth power law | Agrees |
| Where to spend extra compute? | Mostly on a bigger model | Split evenly between model size and data |
| Practical rule | Bigger is better | About 20 tokens of data per parameter |
| Lasting lesson | Forecast performance before training | Do not starve your model of data |