In the frenzied race to build the most powerful artificial intelligence, one question has become increasingly urgent: who gets to decide which AI model is actually the best? As companies pour billions into developing large language models (LLMs), a quiet revolution in benchmarking has emerged from an unlikely source – a UC Berkeley PhD project that transformed into a $1.7 billion startup in just seven months. Arena, formerly LM Arena, has become the de facto public leaderboard for frontier AI models, influencing everything from funding decisions to product launches across the industry.
The Unlikely Arbiter
Founded by Anastasios Angelopoulos and Wei-Lin Chiang, Arena took a radically different approach to evaluating AI models. Instead of relying on static benchmarks that companies could optimize for, Arena uses crowd-sourced human comparisons – real people evaluating which AI responses are better in head-to-head matchups. This approach, they argue, creates a “leaderboard you can’t game” that more accurately reflects real-world performance. But here’s the twist: this supposedly neutral arbiter is funded by the very companies it ranks, including OpenAI, Google, and Anthropic.
The Neutrality Paradox
The question of structural neutrality becomes particularly relevant when examining the broader AI landscape. Consider the ongoing drama between Anthropic and the U.S. Department of Defense. After a $200 million contract collapsed because Anthropic refused to allow its AI systems to be used for mass surveillance or lethal targeting, the Pentagon designated the company as a “supply chain risk” and an “unacceptable risk to national security.” Meanwhile, OpenAI has been expanding its government footprint through deals with AWS to sell AI products to federal agencies.
This creates a fascinating tension: how can a benchmark platform maintain true neutrality when its funders are engaged in such high-stakes conflicts? Arena’s founders argue their funding structure – taking money from multiple competing companies – creates a balance that prevents any single player from exerting undue influence. But as AI becomes increasingly integrated into critical infrastructure and national security, the question of who controls the metrics of success takes on new urgency.
Beyond Chat: The Next Frontier
What makes Arena particularly interesting is its evolution beyond simple chat benchmarks. The platform is expanding to evaluate AI agents, coding capabilities, and real-world tasks through a new enterprise product. Currently, Anthropic’s Claude leads the expert leaderboard for legal and medical use cases – a significant achievement given the company’s recent government conflicts. This expansion reflects a broader industry shift: AI isn’t just about generating text anymore; it’s about performing complex tasks in professional environments.
The Gaming Problem Across Industries
The challenge of creating trustworthy benchmarks extends beyond AI models themselves. Consider the music streaming industry, where Deezer recently revealed that fraudsters are responsible for more than 80% of all streams of AI-generated music on its platform. Using tools like Suno and Udio that can create entire songs in seconds, these bad actors upload thousands of tracks and use bots to generate artificial plays, siphoning royalty payments away from legitimate artists. Deezer detected and tagged over 13 million AI-generated tracks in 2025 alone, with 60,000 new AI tracks being added daily.
This gaming of systems isn’t limited to entertainment. In recruitment, 89% of UK recruiters plan to use more AI in hiring this year, creating what some candidates describe as a “robotic” and “brutal” process. Law firm Mishcon de Reya turned to AI screening after receiving 5,000 applications for just 35 roles. As one early careers manager noted, “We’ve got more candidates using AI to write more applications… and it’s harder to tell the difference between those applications.”
The Compensation Question
As AI systems become more capable, the question of who gets paid for their training data has become increasingly contentious. Patreon CEO Jack Conte recently called AI companies’ fair use arguments “bogus,” noting that while they claim it’s fair to use creators’ work as training data, they simultaneously make multi-million dollar deals with rights holders like Disney and Warner Music. “If it’s legal to just use it, why pay?” Conte asked rhetorically. “Why pay them and not the millions of illustrators and musicians and writers whose work has been consumed by these models?”
The Human Factor
Despite the rush toward automation, human judgment remains crucial. In recruitment, Adecco CEO Denis Machuel acknowledges that while AI brings scale, it also creates frustration: “Before, you would reach out to 50 people, and out of that you will take one, so you will have 49 people frustrated. Now, if you reach out to 500 candidates, you create 499 people frustrated.” The solution, he suggests, is combining “the efficiency of AI with the judgement and human touch of people.”
This tension between automation and human oversight plays out across industries. In music streaming, Deezer focuses on AI detection to determine which tracks were made legitimately with AI tools. In government contracting, the Pentagon is developing its own LLMs after conflicts with commercial providers over ethical boundaries. And in benchmarking, Arena’s human evaluation system represents an attempt to inject human judgment into the process of evaluating machines.
The Bigger Picture
What emerges from these interconnected stories is a complex portrait of an industry in transition. The companies creating the most advanced AI systems are also funding the platforms that evaluate them. The same technology that can generate beautiful music is being used to defraud artists. The tools that promise to make hiring more efficient are creating frustrating experiences for job seekers. And the ethical boundaries that some companies insist on are causing conflicts with powerful institutions.
As Arena expands beyond chat to benchmark agents, coding, and real-world tasks, it’s not just evaluating AI models – it’s helping shape what we value in artificial intelligence. In a field where hype often outpaces reality, having trustworthy benchmarks matters. But as the industry matures, the question isn’t just which AI is best, but who gets to decide – and what values are embedded in those decisions.

