Skip to slide
Chapter 6 · Judging Models, How Do We Measure Quality?
45 / 74

CHAPTER 06 · Judging Models, How Do We Measure Quality? · 6 / 7

How the three pieces fit together

ToolWhat it isStrengthWeakness
MT-BenchFixed hard question setRepeatable, targetedLimited set of questions
Chatbot ArenaLive human voting with EloGold-standard human preferenceSlow, expensive, needs crowds
LLM-as-a-judgeA strong model grades answersFast, cheap, scalableCarries biases, needs care

Together they form a practical toolkit: use the fast LLM judge for everyday iteration, use MT-Bench for a consistent yardstick, and use the human Arena as the ultimate source of truth.

← → arrow keys work too