Skip to slide
Chapter 5 · Mixtral and Mixture of Experts, More Brain, Same Speed
35 / 74

CHAPTER 05 · Mixtral and Mixture of Experts, More Brain, Same Speed · 3 / 6

How Mixtral is built

Inside each layer of the Transformer, Mixtral replaces the single feed-forward network from Chapter 1 with eight of them, the experts. For every word, the router scores the experts and sends the word to just the top two. The other six do nothing for that word and cost nothing.

The numbers tell the story. Mixtral (often written 8x7B) holds about 47 billion parameters in total, so it has a large store of knowledge. But because only two of eight experts are active per word, only about 13 billion parameters actually do work for each word. So:

  • Capacity of a roughly 47 billion parameter model.
  • Running cost closer to a 13 billion parameter model.
flowchart LR
    Total["Total knowledge<br/>about 47B parameters"] --> Active["Active per word<br/>about 13B parameters"]
    Active --> Result["Big-model smarts,<br/>small-model running cost"]

The result reported in the paper: Mixtral matched or beat much larger dense models such as Llama 2 70B, and rivaled GPT-3.5, while being significantly cheaper and faster to run. That combination is why MoE has become a backbone of many frontier models.

← → arrow keys work too