CHAPTER 05 · Mixtral and Mixture of Experts, More Brain, Same Speed · 6 / 6
The one-sentence takeaway
A Mixture of Experts model stores the knowledge of a very large model but, for each word, a small router activates only a couple of expert sub-networks, so you get big-model quality at small-model running cost, paying for it with extra memory to keep all the experts on standby.
Next: Chapter 6, Judging models, where we ask how anyone can possibly measure whether one model is better than another.