Skip to slide
Chapter 1 · The Transformer, the Engine Inside Every Modern Model
08 / 74

CHAPTER 01 · The Transformer, the Engine Inside Every Modern Model · 2 / 6

The one big idea: attention

Here is the core move. Instead of reading word by word, the Transformer looks at all the words at the same time and lets each word decide which other words it should pay attention to.

A simple analogy. Imagine every word in the sentence is a person in a room, and they are all trying to understand their own role. The word "it" raises its hand and asks the room, "Which of you is relevant to me?" Every other word answers with a relevance score. "Trophy" might shout loudly, "suitcase" a bit softer, "the" almost silently. The word "it" then builds its understanding mostly from the words that answered loudest.

That asking-and-weighing process is self-attention. Doing it lets every word pull in context from every other word in a single step, no matter how far apart they are.

How attention actually works, gently

Every word is first turned into a list of numbers called an embedding. A list of numbers is called a vector, and you can think of it as the word's location on a giant map of meaning, where similar words sit close together.

For attention, each word creates three different versions of itself:

  • A Query: "Here is what I am looking for."
  • A Key: "Here is what I contain."
  • A Value: "Here is what I will hand over if you pick me."

To decide how much word A should attend to word B, the model compares A's Query with B's Key. A strong match means a high score. All the scores get squashed into percentages that add up to 100 using a function called softmax. Then each word builds its new, context-aware representation by mixing together the Values of the other words, weighted by those percentages.

flowchart LR
    W["Word: 'it'"] --> Q[Query: what am I looking for]
    Q --> M{Compare with the Key<br/>of every other word}
    M --> S[Scores become<br/>percentages via softmax]
    S --> Mix[Mix the Values,<br/>weighted by the scores]
    Mix --> R["New meaning of 'it'<br/>now carries 'trophy'"]

The beautiful part: the model is not told the rules of grammar. It learns what to attend to on its own, by practicing next-word prediction on huge amounts of text.

Many heads are better than one

A single attention pass captures one kind of relationship. But words relate in many ways at once: grammar, subject matter, tone, who-did-what-to-whom. So the Transformer runs several attention passes in parallel, each free to focus on a different pattern. These parallel passes are called attention heads. One head might track which noun a pronoun refers to, another might track verb tense, another might link adjectives to the things they describe. Their results are combined. The original paper used eight heads.

← → arrow keys work too