Skip to slide
Chapter 5 · Mixtral and Mixture of Experts, More Brain, Same Speed
37 / 74

CHAPTER 05 · Mixtral and Mixture of Experts, More Brain, Same Speed · 5 / 6

What you give up

MoE is not free magic. The catch is memory. Even though only two experts work per word, all eight must be loaded and ready, because the router might call on any of them for the next word. So an MoE model needs enough memory to hold all its parameters, even if it only uses a fraction at a time. You are trading higher memory for lower compute per word. For services answering millions of requests, that trade is usually well worth it, because compute is the cost that scales with every user, while memory is paid once.

flowchart LR
    subgraph Dense[Dense model]
        D1[All parameters<br/>work every word] --> D2[High cost per word]
    end
    subgraph MoE[Mixture of Experts]
        M1[All parameters<br/>loaded in memory] --> M2[Only a slice<br/>works per word] --> M3[Low cost per word]
    end
← → arrow keys work too