CHAPTER 03 · Chatbot Arena, Letting Real People Pick the Winner · 4 / 6
A short worked example
Suppose a new model joins the Arena and, in its first matches, beats a model that everyone already agrees is excellent. Under an Elo-style or Bradley-Terry rating, that win counts for a lot, because defeating a strong opponent is strong evidence of strength, and the newcomer's rating jumps. If it had instead beaten a weak model, the rating would barely move, since that result was expected. Over thousands of such matches the ratings settle into an order that reflects real relative quality, with the confidence interval shrinking as more votes come in. This is why a single lucky win cannot crown a model: the system demands consistent wins against tough competition.