Skip to slide
Chapter 4 · Glossary: Benchmarks
32 / 39

CHAPTER 04 · Glossary: Benchmarks · 5 / 12

Task and task suite

A task is one specific kind of problem you ask a model to solve, for example translating a sentence or answering a multiple-choice science question. A task suite is a bundle of many such tasks gathered into one benchmark. The reason large models are tested with suites rather than single tasks is that a generalist model can be strong in one area and surprisingly weak in another, and only a wide spread of tasks reveals the full shape of its abilities. BIG-Bench is a task suite of 204 tasks, deliberately spanning many different domains.

← → arrow keys work too