CHAPTER 04 · Glossary: Benchmarks
Glossary: Benchmarks
This is your shared reference for folder 04. Every term the chapters link to is explained here from scratch, in plain language, with enough depth to actually understand it rather than just recognize it. Read it straight through as a primer, or jump in whenever a chapter sends you here.
A few foundational terms live in the earlier folders and are linked across rather than repeated: what a benchmark is at heart, the Elo rating system, emergent ability, sparse and dense models, long context, and agent. This glossary focuses on the ideas specific to evaluation.
Terms are grouped by theme so related ideas sit together.
- Benchmark basics: Benchmark, Static vs live benchmark, Ground truth, Human preference, Task and task suite, Metric, Saturation, Contamination, Canary string
- Scoring and scaling behavior: Calibration, Aggregate score, Breakthrough behavior, Brittleness, BIG-Bench Lite
- Measuring real coding: Repository and codebase, Issue and pull request, Patch and diff, Unit test, Execution-based evaluation, HumanEval, SWE-Llama
- Ranking by human votes: Pairwise comparison, Crowdsourcing, Leaderboard, Elo rating, Bradley-Terry model, Confidence interval, Active sampling