CHAPTER 06 · Judging Models, How Do We Measure Quality? · 6 / 7
How the three pieces fit together
| Tool | What it is | Strength | Weakness |
|---|---|---|---|
| MT-Bench | Fixed hard question set | Repeatable, targeted | Limited set of questions |
| Chatbot Arena | Live human voting with Elo | Gold-standard human preference | Slow, expensive, needs crowds |
| LLM-as-a-judge | A strong model grades answers | Fast, cheap, scalable | Carries biases, needs care |
Together they form a practical toolkit: use the fast LLM judge for everyday iteration, use MT-Bench for a consistent yardstick, and use the human Arena as the ultimate source of truth.