
Hugging Face Blog
· 1 min read
Judge Arena: Benchmarking LLMs as Evaluators
LLM-as-a-Judge has emerged as a popular way to grade natural language outputs from LLM applications, but how do we know which models make the best judges?
We’re excited to launch Judge Arena - a platform that lets anyone easily compare models as judges side-by-side. Just run the judges on a test sample and vote which judge you agree with most. The results will be organized into a leaderboard that displays the best judges.
Judge Arena
Crowdsourced, randomized battles have proven effective at benchmarking LLMs. LMSys's Chatbot Arena has collected over 2M votes and is highly regarded as a field-test to identify the best language models. Since LLM evaluations aim to capture human preferences, direct human feedback is also key to determining which AI judges are most helpful.
How it works
- Choose your sample for evaluation:
- Let the system randomly generate a 👩 User Input / 🤖 AI Response pair
- OR input your own custom sample
- Two LLM judges will:
- Score the response
- Provide their reasoning for the score
Review both judges’ evaluations and vote for the one that best aligns with your judgment
(We recommend reviewing the scores first before comparing critiques)
After each vote, you can:
- Regenerate judges: Get new evaluations of the same sample
- Start a 🎲 New round: Randomly generate a new sample to be evaluated
- OR, input a new custom sample to be evaluated
To avoid bias and potential abuse, the model names are only revealed after a vote is submitted.
Selected Models
Judge Arena focuses on the LLM-as-a-Judge approach, and therefore only includes generative models (excluding classifier models that solely output a score). We formalize our selection criteria for AI judges as the following:
- The model should possess the ability to score AND critique other models' outputs effectively.
- The model should be prompt-able to evaluate in different scoring formats, for different criteria.
The Leaderboard
Early Insights
These are only very early results, but here’s what we’ve observed so far:
Original source
This story was published by Hugging Face Blog. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on huggingface.co


