Skip to slide
Chapter 0 · Start Here: Benchmarks, Explained Simply
06 / 39

CHAPTER 00 · Start Here: Benchmarks, Explained Simply · 5 / 5

One idea to hold in your head

Every benchmark is a bet about what "good" means, and every benchmark can be gamed or can go stale. A fixed test can leak into training data so a model looks smart without being smart. A narrow test can reward a skill nobody actually needs. The single thread running through this folder is the constant push to make tests that are harder to fake and closer to reality: broader tasks (BIG-Bench), real work with real pass-or-fail checks (SWE-bench), and live human judgment that cannot be memorized in advance (Chatbot Arena). Once you see evaluation as an arms race between tests and the models trying to ace them, every paper here becomes a different move in that game.

Ready? Start with Chapter 1: BIG-Bench.

← → arrow keys work too