CHAPTER 06 · Judging Models, How Do We Measure Quality?
Judging Models, How Do We Measure Quality?
Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (2023)
We have built, scaled, aligned, and optimized our model. One question remains, and it is harder than it sounds: how do we know if it is any good? This final chapter of folder 01 is about measurement, which is the unglamorous but essential foundation of all progress. If you cannot measure quality, you cannot improve it.