Skip to slide
Chapter 7 · Glossary: Foundational Modelling
63 / 74

CHAPTER 07 · Glossary: Foundational Modelling · 16 / 27

Training the model

Pre-training

Pre-training is the first and most expensive stage of building a language model. The model is shown enormous amounts of text from the internet and books and trained on one simple task: predict the next token. By doing this billions of times, it gradually absorbs grammar, facts, reasoning patterns, and styles.

A pre-trained model is knowledgeable but raw. It is a powerful autocomplete, not yet a helpful assistant. Turning it into an assistant is the job of fine-tuning and alignment (Chapter 3).

Fine-tuning

Fine-tuning means taking an already pre-trained model and training it a bit more on a narrower dataset to specialize it. Pre-training gives broad general ability; fine-tuning shapes that ability toward a specific goal, such as following instructions, writing in a company's voice, or answering medical questions.

Fine-tuning is far cheaper than pre-training because the model already knows language; you are only nudging it. Chapter 4 (LoRA) is about making fine-tuning cheaper still.

Loss function

The loss function is the number that measures how wrong the model is. During training, the model makes a prediction, the loss function compares it to the correct answer, and produces a score where lower means better. The entire goal of training is to make this number as small as possible.

For language models, the loss is essentially a measure of how surprised the model was by the true next word. Confident and correct gives low loss; confident and wrong gives high loss. The scaling-law curves in Chapter 2 are plots of this loss shrinking as models get bigger.

Gradient descent and backpropagation

These two together are how a model actually learns. The model makes a prediction and computes its loss. Backpropagation is the process of working backward through the network to figure out, for every parameter, which direction it should move to reduce the loss. Gradient descent is then the act of nudging each parameter a small step in that helpful direction.

Repeat this billions of times and the model gradually improves. A simple image: you are in fog on a hillside trying to reach the valley. Backpropagation tells you which way is downhill; gradient descent takes a small step that way. Do it over and over and you reach the bottom.

Overfitting

Overfitting is when a model memorizes its training data instead of learning general patterns. An overfit model performs great on examples it has seen but fails on new ones, like a student who memorized last year's exam answers but cannot solve a fresh problem.

The goal is always generalization: doing well on data the model has never seen. Much of the craft of training is about getting strong performance without overfitting.

Hyperparameter

A hyperparameter is a setting chosen by the people training the model, as opposed to a parameter, which the model learns on its own. Examples include how many layers to use, how big to make the model, and how large a step to take during gradient descent.

Hyperparameters are the dials the humans set before and during training. Choosing them well is part science, part experience. The scaling laws in Chapter 2 are partly about choosing two of the most important ones: model size and data size.

Perplexity

Perplexity is a common way to report how well a language model predicts text. It is closely tied to the loss function: low perplexity means the model is rarely surprised by the next word, which means it models the language well. You can read perplexity loosely as "on average, how many words is the model torn between when guessing the next one." Lower is better.


← → arrow keys work too