Skip to slide
Chapter 0 · Start Here: Benchmarks, Explained Simply
03 / 39

CHAPTER 00 · Start Here: Benchmarks, Explained Simply · 2 / 5

The big picture

There is a clever way to organize almost every test we give a language model, and it comes from the Chatbot Arena paper itself. You can sort benchmarks along two questions. First, where do the questions come from: a fixed list written down in advance (static), or a fresh stream of new questions from real users (live)? Second, how is an answer scored: against a known correct answer (ground truth), or by asking which answer people like better (human preference)?

flowchart TD
    Q[How do we test a model?] --> S{Where do<br/>questions come from?}
    S -->|Fixed written list| Static[Static benchmark]
    S -->|Fresh from real users| Live[Live benchmark]
    Static --> M1{How is it scored?}
    Live --> M2{How is it scored?}
    M1 -->|Known correct answer| BB["BIG-Bench<br/>broad skills, 204 tasks"]
    M1 -->|Run the code, did tests pass| SWE["SWE-bench<br/>real GitHub bug fixes"]
    M2 -->|People vote on the better reply| CA["Chatbot Arena<br/>live human preference"]

Each of the three papers stakes out a different corner of this map, and each one fixes a weakness in the styles that came before it.

← → arrow keys work too