Skip to slide
Chapter 7 · Glossary: Foundational Modelling
73 / 74

CHAPTER 07 · Glossary: Foundational Modelling · 26 / 27

Efficiency tricks

Matrix rank (low-rank)

A matrix is just a grid of numbers, and the weights inside a neural network are stored in such grids. The rank of a matrix is a measure of how much genuinely independent information it contains. A "low-rank" matrix, despite possibly being large, can be reconstructed from a much smaller amount of information.

The practical payoff (Chapter 4): a big grid of numbers can sometimes be approximated by multiplying two much skinnier grids together, storing nearly the same information with a tiny fraction of the numbers. This is the mathematical foundation that makes LoRA possible.

LoRA

LoRA stands for Low-Rank Adaptation (Chapter 4). It is a cheap way to fine-tune a giant model. Instead of adjusting all of the model's parameters, LoRA freezes the original model and trains only a tiny pair of skinny low-rank matrices that capture the adjustment a new task needs.

This cuts memory and storage dramatically, lets one frozen base model carry many small swappable adapters for different tasks, and adds no slowdown when the model runs, because the adapter can be merged back in. LoRA is what made fine-tuning accessible to people without giant compute budgets.

Mixture of Experts (MoE)

A Mixture of Experts is a model design (Chapter 5) where each layer contains many sub-networks called experts, but only a few of them activate for any given token. A small router picks which experts handle each word.

The benefit is that the model can hold a very large amount of knowledge across all its experts, while only doing a small amount of work per word. You get the quality of a big model at the running cost of a much smaller one. The trade-off is higher memory, because all experts must be kept loaded even though only a few are used at a time. Mixtral (Chapter 5) is the well-known open example.

Router (gating network)

The router, sometimes called the gating network, is the small, fast component inside a Mixture of Experts model that decides which experts should handle each token. Continuing the hospital analogy from Chapter 5, the router is the receptionist who glances at a patient and sends them to the right specialists.

The router is learned during training, not hand-coded. The model figures out on its own how to divide work among the experts.

Sparse and dense models

A dense model uses all of its parameters for every token it processes. A sparse model, like a Mixture of Experts, uses only a fraction of its parameters for each token, activating just the relevant parts.

The distinction matters for cost: dense models pay their full size on every word, while sparse models can hold far more total parameters while only paying for a slice each time. "Sparse" essentially means "most of the model stays asleep for any given word."


← → arrow keys work too