CHAPTER 05 · Mixtral and Mixture of Experts, More Brain, Same Speed · 1 / 6
The tension we are trying to escape
Recall the trade-off from earlier chapters:
- A bigger model knows more and reasons better.
- But every time you use a normal model, all of its parameters do work, so a bigger model costs more for every single answer. This running cost is called inference.
A normal model is dense: the whole network fires for every word. MoE breaks this rule. It builds a model with a huge total number of parameters, but arranges things so that only a small slice of them activates for any given word. Lots of knowledge stored, little work done per word.