CHAPTER 06 · Judging Models, How Do We Measure Quality? · 2 / 7
Tool 1: MT-Bench, a hard conversation quiz
MT-Bench is a curated set of challenging, open-ended questions spanning writing, reasoning, math, coding, and more. Crucially, it is multi-turn: it asks a question, then a follow-up, to test whether the model can hold a coherent conversation rather than just answer one-off prompts. It is a fixed, repeatable test you can run any model against.