Skip to slide
Chapter 6 · Judging Models, How Do We Measure Quality?
43 / 74

CHAPTER 06 · Judging Models, How Do We Measure Quality? · 4 / 7

The big idea: let an LLM be the judge

Human voting is the gold standard, but it does not scale. You cannot summon thousands of humans every time you tweak a model. So the paper asks a bold question: can a strong model, like GPT-4, judge other models' answers in place of a human?

flowchart LR
    P[A question] --> M1[Model A's answer]
    P --> M2[Model B's answer]
    M1 --> J{Strong LLM judge<br/>reads both}
    M2 --> J
    J --> V[Picks the better one<br/>or scores each]

The appeal is obvious: an LLM judge is fast, cheap, and available around the clock. You can run thousands of automatic comparisons in minutes. But does it actually agree with human taste?

The key finding

Yes, to a striking degree. The paper found that a strong LLM judge agreed with human preferences about 80 percent of the time. That is roughly the same rate at which two humans agree with each other. In other words, the model judge is about as reliable as a typical human judge. This result is why "LLM-as-a-judge" is now a standard, everyday tool for evaluating AI systems quickly.

← → arrow keys work too