CHAPTER 06 · Judging Models, How Do We Measure Quality? · 5 / 7
Be careful: judges have biases
The paper is honest about the ways an LLM judge can be fooled, and knowing these is important if you ever rely on one:
- Position bias: the judge may favor whichever answer it sees first, regardless of quality. The fix is to swap the order and average.
- Verbosity bias: the judge tends to prefer longer answers, even when a shorter one is better. Length can masquerade as quality.
- Self-enhancement bias: a judge may favor answers written in its own style, subtly rating its own family of models more highly.
flowchart TD
J[LLM judge] --> B1[Position bias<br/>prefers the first answer]
J --> B2[Verbosity bias<br/>prefers longer answers]
J --> B3[Self-enhancement bias<br/>prefers its own style]
These biases do not make LLM judges useless. They make them tools to use carefully, with tricks like swapping answer order and watching out for padding.