Skip to slide
Chapter 7 · Glossary: Foundational Modelling
64 / 74

CHAPTER 07 · Glossary: Foundational Modelling · 17 / 27

Scale and cost

Compute and FLOPs

Compute is the total amount of calculation used to train or run a model, and it translates directly into time and money. It is often measured in FLOPs, which stands for floating-point operations, basically a count of how many individual arithmetic steps were performed.

Training a frontier model can take an astronomical number of FLOPs, which is why it costs millions of dollars and requires thousands of specialized chips. Chapter 2 is all about spending a fixed compute budget wisely.

Scaling law

A scaling law is a predictable mathematical relationship showing how a model's performance improves as you increase its size, its training data, or its compute. The surprising discovery (Chapter 2) is that this improvement is smooth and forecastable across a huge range, often appearing as a straight line on the right kind of graph.

Scaling laws matter because they remove guesswork. A small experiment can predict how a much larger model will perform before you spend the money to build it.

Emergent ability

An emergent ability is a skill that a model does not have when small but suddenly displays once it crosses a certain size. Below the threshold the ability is essentially absent; above it, the ability appears. Examples historically included certain kinds of arithmetic and multi-step reasoning.

Emergence is part of what made scaling so exciting: simply making models bigger sometimes unlocked qualitatively new capabilities, not just smoother improvements.

In-context learning

In-context learning is a model's ability to learn a task from examples placed directly in the prompt, without any change to its parameters. You show it a few examples of what you want, and it picks up the pattern on the spot.

Two related terms: "zero-shot" means you give the model a task with no examples, just the instruction. "Few-shot" means you include a handful of examples first. The discovery that large models are strong few-shot learners (from the GPT-3 work) was a major moment, because it meant one general model could handle countless tasks just by being shown examples in the prompt.

Inference

Inference is the act of actually using a trained model to produce an output, as opposed to training it. Every time you send a prompt to a model and get a response, that is one inference.

Inference cost is hugely important in practice because it is paid every single time anyone uses the model, across millions of users. This is why a smaller or more efficient model (Chapters 2, 4, and 5) is so valuable: it lowers the cost of every interaction.


← → arrow keys work too