Arena, formerly LMArena and best known for the Chatbot Arena leaderboard, is a venture-backed company that runs a crowdsourced AI model ranking system and, since September 16, 2025, sells AI Evaluations to model labs and enterprises. The same organization says its public platform reached 3 million votes across 400-plus models by April 2025, then 5 million monthly users and 60 million monthly conversations by January 2026.
That combination is why Arena matters and why it draws scrutiny. Its public influence comes from a human-preference leaderboard that many labs cite as evidence of model quality, while its commercial business now sells evaluation services to the same industry it ranks; Arena does not publicly disclose pricing, customers, or revenue for AI Evaluations.
Arena started as a research-led public benchmark for head-to-head chatbot comparisons. On its methods page, the company says users compare two anonymous model outputs side by side, pick the better answer, and contribute data that feeds public rankings and feedback to developers. It now covers multiple modalities, with the company saying on its About page that it evaluates text, image, and coding systems.
The ranking system itself has changed over time. In a November 14, 2025, methods post, Arena said it had moved to Bradley-Terry ranking, a standard pairwise-comparison model that estimates how likely one model is to beat another from many head-to-head votes. The same post said Arena now publishes not just a raw rank but also rank spread and confidence intervals, a tacit acknowledgment that tiny position changes on a crowded leaderboard can look more meaningful than they are.
Arena’s public leaderboard scaled from 3M votes and 400+ models to 5M monthly users
Arena’s own numbers show a benchmark that grew from an academic curiosity into public infrastructure for model marketing. In its April 27, 2025, two-year update, the team said the platform had collected more than 3 million votes, hosted 400-plus models, and run 300-plus pre-release tests.
That post also said the system had generated more than 1.7 million unique battle pairings. In plain terms, Arena was no longer just showing a few headline models against each other; it had become a large matching engine for pairwise preference data, which is exactly the kind of data labs want when they are tuning new releases.
By January 28, 2026, after rebranding from LMArena to Arena, the company said the platform served 5 million monthly users in 150 countries and handled 60 million conversations per month. Those metrics were disclosed at different times, so they are not directly comparable one-to-one with earlier vote totals, but the direction is clear: the leaderboard became a very large consumer-facing evaluation surface.
Arena’s growth pitch to investors was even more explicit. In its January 6, 2026, Series A announcement, the company said it had raised \$150 million and described itself as building the “world’s most trusted AI evaluation platform,” adding that its community had grown 25x in the prior year.
The public leaderboard is also not Arena’s only ranking venue. The company’s broader evaluation logic now shows up across domains, including coding-model comparisons such as these Code Arena rankings, which reflect how quickly arena-style, pairwise human preference testing has spread as a product category.
AI Evaluations turned the leaderboard into a commercial service for labs and enterprises
Arena did not stumble into commercialization by accident; it said plainly that neutral evaluation was the business. In an April 17, 2025, post announcing that LMArena had started a company, the team wrote that AI companies were asking for third-party evaluations and that building sustainable infrastructure required a business behind the public platform.
“Many AI companies told us that they needed neutral third-party evaluations,” the team wrote in its April 2025 company announcement, adding that commercialization was how it planned to sustain the platform.
That matters because it pins down the timeline. LMArena said it had become a company by April 17, 2025, and Arena officially launched its paid AI Evaluations product on September 16, 2025. The later launch post said the service was for enterprises, model labs, and developers and positioned Arena as an external evaluator rather than just a public leaderboard host.
On the AI Evaluations launch page, Arena described products for model benchmarking, red-teaming, domain-specific assessment, and side-by-side human testing. That is a familiar AI-industry move: once a benchmark becomes influential enough, the benchmark operator can sell access to the machinery behind it. The underlying business logic is not very different from what you see in broader platform economics, including the pressure described in this look at the OpenAI vs. Anthropic business model: expensive model development creates demand for outside proof that a new release is actually better.
Arena’s About page now frames the company in exactly those terms, saying it provides evaluation infrastructure to enterprises and labs. What it does not disclose is basic commercial detail: the site does not list customers, pricing, or revenue.
Arena also says it shares feedback with model developers. On How Arena Works, it says user preferences and feedback can help improve models, and its April 2025 scale post highlighted 300-plus pre-release tests. That is not proof of misconduct; it is evidence that private testing by labs was already a substantial part of the system before the formal product launch.
Independent papers argue the ranking system can be gamed by private testing and clone submissions
The sharpest criticism is not that Arena’s math is unsound. It is that the incentives around the leaderboard can distort what the leaderboard measures.
A NeurIPS 2025 paper titled The Leaderboard Illusion analyzed historical Chatbot Arena data and argued that providers could benefit from private testing before public release, selective deployment, and treating the leaderboard as an optimization target rather than a neutral snapshot. The authors relied on public and scraped historical data, so their analysis may not capture Arena’s full internal controls or later policy changes.
The paper’s core point is simple: if a leaderboard strongly affects launch narratives, labs will adapt to the leaderboard. That can mean iterating privately until a model looks good under Arena-like comparisons, then releasing the strongest candidate while discarding weaker variants. A benchmark can still measure something real and still be gameable; those two facts are not mutually exclusive.
Arena’s methodology updates partly address presentation risk. In the ranking-method post, the company said it now reports confidence intervals and rank spread because close models may be statistically hard to separate. That is a useful improvement. It also undercuts the common habit of treating a one-place gap as a decisive win.
A later independent paper, Strategic Candidacy in Generative AI Arenas from 2026, made a different argument: model producers may gain an advantage by submitting multiple similar models or “clones” into arena-style systems. In that setup, a lab can increase the chance that one near-duplicate version lands higher, even if the underlying capability jump is modest. It is the leaderboard equivalent of buying more lottery tickets, except the tickets are model variants.
This is where the conflict-of-interest concern gets traction. Arena sells evaluations to the same class of customers whose models benefit from strong public rankings. The strongest version of that critique is still inferential rather than documentary; the available public sources do not show Arena manipulating rankings for paying clients. But the structure is hard to ignore: a company that monetizes evaluation has strong incentives to stay close to the labs seeking favorable evidence.
None of that means Arena’s leaderboard is useless. It means the leaderboard is best read as a large, valuable, but incentive-shaped measure of human preference under Arena’s own sampling and ranking rules. That is still informative. It is just not the same thing as an objective ground truth for overall model quality.
Arena has tried to make the system more legible. In its December 18, 2025, open-source ranking post, the company released the arena-rank package and said it was open-sourcing parts of the methodology along with a preference dataset. Greater transparency helps outsiders test the math. It does less to solve incentive problems created by who submits models, who tests privately, and who pays for services.
The short version is that Arena now plays two roles at once. It is a public referee for AI model popularity and a private vendor of evaluation services. Those roles are adjacent enough to be commercially sensible and close enough to raise credibility questions.
The next concrete milestone is whatever Arena discloses after its January 2026 Series A: additional methodology changes, more open data releases, or customer details for AI Evaluations.
Key Takeaways
- Arena, formerly LMArena, runs a crowdsourced AI model leaderboard and launched a paid AI Evaluations business on September 16, 2025.
- Arena said it had reached more than 3 million votes, 400-plus models, and 300-plus pre-release tests by April 2025, then 5 million monthly users and 60 million monthly conversations by January 2026.
- Arena said in April 2025 that labs wanted neutral third-party evaluations, then turned that need into a formal commercial product months later.
- Arena’s public rankings use Bradley-Terry pairwise modeling and now include confidence intervals and rank spread to show uncertainty.
- Independent papers in 2025 and 2026 argued that arena-style rankings can be strategically shaped by private testing and multiple similar submissions.
Further Reading
- About Arena | Crowdsourced AI Model Evaluation Platform, Arena’s overview of its mission, products, and claimed scale.
- New Product: AI Evaluations, Arena’s September 16, 2025, launch post for its commercial evaluations business.
- The Leaderboard Illusion, Independent analysis of incentives, private testing, and leaderboard dynamics in Chatbot Arena data.
- Strategic Candidacy in Generative AI Arenas, Independent paper on how multiple similar model submissions can affect arena rankings.
- Arena’s Ranking Method, Arena’s explanation of Bradley-Terry ranking, confidence intervals, and rank spread.
Frequently Asked Questions
What is Chatbot Arena?
Chatbot Arena is Arena’s public head-to-head voting system for AI models. Users see two anonymous model responses to the same prompt, pick the better one, and those pairwise outcomes feed the public leaderboard and model feedback pipeline described on How Arena Works.
When did LMArena become Arena?
LMArena said it had formed a company in an April 17, 2025, post. It later rebranded publicly as Arena on January 28, 2026.
How does Arena rank AI models?
Arena said in its November 2025 methods post that it uses a Bradley-Terry model, which estimates relative strength from many pairwise wins and losses. It also publishes confidence intervals, raw rank, and rank spread because models close together may not be meaningfully separable.
Why do people criticize the Arena leaderboard?
The main criticism is incentive-driven, not mathematical. Independent papers argued that if labs can privately test models or submit multiple near-duplicate variants, they can improve leaderboard outcomes without a proportionate real-world capability jump.
Does Arena disclose its evaluation customers or pricing?
No public page cited here lists named customers, pricing, or revenue for AI Evaluations. Arena describes the service and target users, but the business terms are not publicly disclosed on its site.
References
- Arena, 2025, LMArena is Growing to Support our Community Platform
- Arena, 2025, Celebrating Community Impact: 3M+ votes, 400+ models, and 300+ pre-release tests
- Arena, 2025, New Product: AI Evaluations
- Arena, 2025, Arena’s Ranking Method
- Arena, 2025, Arena-Rank: Open Sourcing the Leaderboard Methodology
- Arena, 2026, Fueling the World’s Most Trusted AI Evaluation Platform
- Arena, 2026, LMArena is now Arena
- Mok et al., 2025, The Leaderboard Illusion
- Zhou et al., 2026, Strategic Candidacy in Generative AI Arenas
Last reviewed: 2026-06
