The world's most trusted AI ranking is decided **by blind vote**
AIGenerative AIAnalysis

The world's most trusted AI ranking is decided by blind vote

Arena (formerly LMArena) ranks AI models using millions of anonymous human votes. How it works, why it became the industry's reference, and where the scoreboard misleads.

Macareno5 min read

Every time a new model ships, the announcement comes with the same magic phrase: "number 1 on the leaderboard". OpenAI says it, Google says it, Anthropic says it. And you, who just wanted to know which AI to use at work, are left with a fair question: number 1 according to whom?

Most of the time, the answer is a website where anyone can walk in, type a question and vote on which of two anonymous answers was better. Without knowing who wrote which. Millions of times a month.

That website is called Arena. And it became, almost by accident, the judge the entire industry agreed to trust.

A university project that ended up refereeing billions

Arena was born in 2023 as Chatbot Arena, a project by researchers at UC Berkeley. In 2024 it got its own domain as LMArena, in April 2025 it became an independent company and, in January 2026, it shortened its name to Arena.

The "university project" phase didn't last long. In May 2025 it raised US$100 million at a US$600 million valuation. In January 2026, another US$150 million, now valued at US$1.7 billion. Today it has more than 5 million monthly users across 150 countries and around 60 million conversations a month.

And the business model is curious: since September 2025, labs and companies pay to have their models evaluated by the Arena community. By December, that service was already running at a US$30 million annual pace. In other words: the referee also sells consulting to the teams it officiates.

The blind test no brand can buy

The mechanism is simple, and that's why it's powerful. You write a prompt, two anonymous models answer, you pick the better one and only then find out who was who. The votes feed a statistical system similar to chess Elo: beating a strong model is worth more than beating a weak one.

The difference from traditional benchmarks is significant:

Traditional benchmark Arena
Who evaluates A fixed answer key Real people
Questions Always the same Endless and unpredictable
Risk of "memorizing the test" High Low
Brand bias Not applicable Removed by anonymity
What it measures Technical accuracy Human preference

Fixed benchmarks age badly. The more famous they get, the more models end up trained (on purpose or not) to do well on them. On Arena, the test changes with every vote, because every person brings their own odd Tuesday afternoon question.

And anonymity solves the fan club problem. If you know you're judging ChatGPT against Claude, your brand loyalty votes with you.

When the scoreboard becomes the target

There's an old law in economics, Goodhart's Law: when a measure becomes a target, it stops being a good measure. Arena didn't escape it.

In April 2025, Meta launched Llama 4 and Maverick showed up at the top of the leaderboard. Except the version tested on Arena was different from the one released to the public, tuned to please voters. Arena changed its rules after the episode.

That same month, researchers from Cohere, Stanford, MIT, Princeton and other institutions published a 68-page study called The Leaderboard Illusion. The accusation: big labs privately tested dozens of variants before launch and published only the best one. In the most extreme case, 27 private Meta variants ahead of Llama 4.

A leaderboard can be honest and still be played best by whoever holds the most chips.

Arena replied that its private testing policy had been public since 2024 and open to any provider. Both sides have valid points. What became clear is that, with billions at stake, the scoreboard stopped being just science.

What the ranking really measures (and what it doesn't)

Here's the detail almost nobody reads: Arena measures which answer people prefer. Not which one is more correct.

Human preference is shaped by tone, formatting, answer length, confident writing. A polished wrong answer can beat a dry correct one, especially when the voter doesn't know the subject.

That doesn't make the ranking useless. It makes it specific. It answers the question "which AI do people enjoy talking to?" very well, and answers "which AI will get my company's tax calculation right?" poorly.

How to use it without falling into the trap

If you're choosing a model for your team, Arena is a great starting point and a terrible finish line. A few practical rules:

  • Look at the categories, not just the overall top spot. Arena splits leaderboards by coding, vision, image, video and other tasks. The overall champion may not be the champion at what you need.
  • Small gaps are ties. A few Elo points apart doesn't mean one model is better than the other in practice.
  • Test with your own prompts. The site itself lets you pick models and compare them side by side. Use real questions from your day to day.
  • Combine it with other sources. Technical benchmarks, cost per use, data privacy and integration with the tools you already use matter as much as the ranking.

The point

Arena solved a real problem: it took AI evaluation out of the exclusive hands of those who build AI and put it in the hands of millions of ordinary people. That's rare and valuable.

But no ranking decides for you. It tells you what the crowd prefers. You still need to know what your business needs.

If you like understanding technology before believing its marketing, subscribe to the macareno.net newsletter. Every week, a little less hype and a little more context.

Sources

Share article

Next business step

Connect this article with a relevant service and a real MacarenoNet case to move from insight to execution.

Recommended service

Digital Transformation Services

Consulting, implementation and continuous improvement on Microsoft stack.

View service

Recommended case

Project Dashboard

Real-world case with measurable execution and visibility gains.

View case