Skip to slide
Chapter 7 · Glossary: Foundational Modelling
61 / 74

CHAPTER 07 · Glossary: Foundational Modelling · 14 / 27

Transformer

The Transformer is the neural network architecture introduced in Attention Is All You Need (Chapter 1), built around self-attention. It processes all words in parallel and lets any word attend to any other, which made it fast to train and good at long-range understanding.

It is the foundation of essentially every modern large language model. When people say GPT, Claude, Llama, or Gemini, they are referring to very large Transformers trained on huge amounts of text.

← → arrow keys work too