Skip to slide
Chapter 1 · BIG-Bench, Testing the Full Breadth of What a Model Can Do
12 / 39

CHAPTER 01 · BIG-Bench, Testing the Full Breadth of What a Model Can Do · 5 / 6

Why it mattered

BIG-Bench set the template for the modern broad evaluation. It showed that the right response to fast-improving models is not one clever test but a wide, collaborative, deliberately-hard suite, paired with a human baseline and defenses against contamination. Its findings about smooth versus breakthrough scaling shaped how the field thinks about emergent abilities, and its canary-string trick is now standard practice. Later benchmarks, including the two in this folder, inherited its core lesson: a benchmark is only as good as it is hard to fake and broad enough to be honest.

← → arrow keys work too