Every time a new model ships, the announcement comes with the same magic phrase: "number 1 on the leaderboard". OpenAI says it, Google says it, Anthropic says it. And you, who just wanted to know which AI to use at work, are left with a fair question: number 1 according to whom?
Most of the time, the answer is a website where anyone can walk in, type a question and vote on which of two anonymous answers was better. Without knowing who wrote which. Millions of times a month.
That website is called Arena. And it became, almost by accident, the judge the entire industry agreed to trust.
A university project that ended up refereeing billions
Arena was born in 2023 as Chatbot Arena, a project by researchers at UC Berkeley. In 2024 it got its own domain as LMArena, in April 2025 it became an independent company and, in January 2026, it shortened its name to Arena.
The "university project" phase didn't last long. In May 2025 it raised US$100 million at a US$600 million valuation. In January 2026, another US$150 million, now valued at US$1.7 billion. Today it has more than 5 million monthly users across 150 countries and around 60 million conversations a month.
And the business model is curious: since September 2025, labs and companies pay to have their models evaluated by the Arena community. By December, that service was already running at a US$30 million annual pace. In other words: the referee also sells consulting to the teams it officiates.
The blind test no brand can buy
The mechanism is simple, and that's why it's powerful. You write a prompt, two anonymous models answer, you pick the better one and only then find out who was who. The votes feed a statistical system similar to chess Elo: beating a strong model is worth more than beating a weak one.
The difference from traditional benchmarks is significant:
| Traditional benchmark | Arena | |
|---|---|---|
| Who evaluates | A fixed answer key | Real people |
| Questions | Always the same | Endless and unpredictable |
| Risk of "memorizing the test" | High | Low |
| Brand bias | Not applicable | Removed by anonymity |
| What it measures | Technical accuracy | Human preference |
Fixed benchmarks age badly. The more famous they get, the more models end up trained (on purpose or not) to do well on them. On Arena, the test changes with every vote, because every person brings their own odd Tuesday afternoon question.
And anonymity solves the fan club problem. If you know you're judging ChatGPT against Claude, your brand loyalty votes with you.
When the scoreboard becomes the target
There's an old law in economics, Goodhart's Law: when a measure becomes a target, it stops being a good measure. Arena didn't escape it.
In April 2025, Meta launched Llama 4 and Maverick showed up at the top of the leaderboard. Except the version tested on Arena was different from the one released to the public, tuned to please voters. Arena changed its rules after the episode.
That same month, researchers from Cohere, Stanford, MIT, Princeton and other institutions published a 68-page study called The Leaderboard Illusion. The accusation: big labs privately tested dozens of variants before launch and published only the best one. In the most extreme case, 27 private Meta variants ahead of Llama 4.
A leaderboard can be honest and still be played best by whoever holds the most chips.
Arena replied that its private testing policy had been public since 2024 and open to any provider. Both sides have valid points. What became clear is that, with billions at stake, the scoreboard stopped being just science.
What the ranking really measures (and what it doesn't)
Here's the detail almost nobody reads: Arena measures which answer people prefer. Not which one is more correct.
Human preference is shaped by tone, formatting, answer length, confident writing. A polished wrong answer can beat a dry correct one, especially when the voter doesn't know the subject.
That doesn't make the ranking useless. It makes it specific. It answers the question "which AI do people enjoy talking to?" very well, and answers "which AI will get my company's tax calculation right?" poorly.
How to use it without falling into the trap
If you're choosing a model for your team, Arena is a great starting point and a terrible finish line. A few practical rules:
- Look at the categories, not just the overall top spot. Arena splits leaderboards by coding, vision, image, video and other tasks. The overall champion may not be the champion at what you need.
- Small gaps are ties. A few Elo points apart doesn't mean one model is better than the other in practice.
- Test with your own prompts. The site itself lets you pick models and compare them side by side. Use real questions from your day to day.
- Combine it with other sources. Technical benchmarks, cost per use, data privacy and integration with the tools you already use matter as much as the ranking.
The point
Arena solved a real problem: it took AI evaluation out of the exclusive hands of those who build AI and put it in the hands of millions of ordinary people. That's rare and valuable.
But no ranking decides for you. It tells you what the crowd prefers. You still need to know what your business needs.
If you like understanding technology before believing its marketing, subscribe to the macareno.net newsletter. Every week, a little less hype and a little more context.
Sources
- Arena (AI platform), Wikipedia
- LMArena is now Arena, Arena Blog
- LMArena lands $1.7B valuation four months after launching its product, TechCrunch
- The Leaderboard Illusion, arXiv
- LMArena Response to "The Leaderboard Illusion", LMArena
