CHAPTER 01 · BIG-Bench, Testing the Full Breadth of What a Model Can Do · 6 / 6
The one-sentence takeaway
BIG-Bench answered fast-improving models with a fast-broadening test, proving that the most honest way to measure a generalist is a huge, diverse, deliberately-hard suite of tasks measured against a human baseline.
Next: Chapter 2, SWE-bench, where the test stops using made-up tasks and starts using real bugs from real software.