Skip to slide
Chapter 7 · Glossary: Foundational Modelling
66 / 74

CHAPTER 07 · Glossary: Foundational Modelling · 19 / 27

Reinforcement learning

Reinforcement learning, or RL, is a style of training where a model learns by trial and error guided by rewards, rather than by copying labeled examples. The model tries an action, receives a reward signal indicating how good the outcome was, and adjusts to earn more reward over time. It is how you train an agent to play a game: not by showing it the perfect moves, but by rewarding it when it wins.

In language models, RL is used to push a model toward responses that score well according to human preferences (see RLHF). It also powers the reasoning models in folder 02, such as DeepSeek-R1.

← → arrow keys work too