Skip to slide
Chapter 4 · LoRA, Fine-Tuning Giant Models on a Budget
27 / 74

CHAPTER 04 · LoRA, Fine-Tuning Giant Models on a Budget · 2 / 6

The key insight: big changes can hide in small matrices

Here is the idea that makes LoRA work. When you fine-tune a model for a new task, the change you make to its weights turns out to be surprisingly simple. Even though the weights are a huge grid of numbers, the adjustment needed to specialize them can be captured by something much smaller. In technical terms, the update has a low rank.

Let us unpack rank with a picture. A model's weights live in big grids of numbers called matrices. A large matrix might be 1000 by 1000, which is one million numbers. But some large matrices can be reconstructed by multiplying two skinny ones together. A 1000 by 1000 matrix can be approximated by a 1000 by 8 matrix times an 8 by 1000 matrix. Count the numbers: that is 8000 plus 8000, which is 16000 numbers instead of one million. Almost the same information, a tiny fraction of the storage.

flowchart LR
    Big["One big update grid<br/>1000 x 1000 = 1,000,000 numbers"]
    Small["Two skinny grids<br/>1000 x 8 and 8 x 1000<br/>= 16,000 numbers"]
    Big -. can be approximated by .-> Small

LoRA bets that the adjustment a model needs in order to learn a new task fits comfortably into those two skinny grids.

← → arrow keys work too