Skip to slide
Chapter 4 · LoRA, Fine-Tuning Giant Models on a Budget
28 / 74

CHAPTER 04 · LoRA, Fine-Tuning Giant Models on a Budget · 3 / 6

How LoRA works

Instead of editing the original weights, LoRA does this:

  1. Freeze the original model completely. Not a single original parameter changes.
  2. Add a small pair of skinny matrices (the low-rank adapter) alongside the layers you want to adapt.
  3. Train only those small matrices on your new data. They learn the adjustment.
flowchart TD
    Frozen[Original giant model<br/>FROZEN, never changes] --> Combine
    Adapter[Tiny LoRA adapter<br/>the only thing that trains] --> Combine
    Combine[Add the adapter's adjustment<br/>on top of the frozen model] --> Output[Model specialized<br/>for your task]

Because you are training only the skinny matrices, the number of trainable parameters can drop by a factor of thousands. Memory needs plummet, training is fast, and the resulting adapter file is small, often only a few megabytes instead of many gigabytes.

← → arrow keys work too