Arena, a crowdsourced AI leaderboard platform, announced it has raised a $200 million Series B round at a $3.1 billion valuation, it said on Thursday. This marks a nearly doubling of its valuation from its previous $1.7 billion post-money valuation following its Series A in January.

Arena reported $100 million in annualized run-rate revenue in June, up from $30 million in annualized revenue at the time of its Series A. The round was led by Lightspeed Venture Partners and Khosla Ventures, with Salesforce Ventures, 01 Advisors, Dell Technologies Capital, Endeavor Catalyst, a16z, Felicis and others joining in.

Arena provides a crowdsourced platform that is free for consumers to use. People enter prompts or request vibe-coded projects and then rate which model does it better. Arena claims it has tens of millions of monthly visitors.

In September of last year, it introduced its commercial product, AI Evaluations, a service that provides model labs and enterprises with detailed performance analytics based on its community feedback.

"AI is advancing faster than our ability to evaluate it, and static benchmarks break down once models recognize they’re being tested," the company said in its funding announcement.

"The world needs a neutral third party to measure how safe and aligned AI actually is once it’s in the hands of real people. Arena is stepping into that role today," it added.

The announcement follows AI labs realizing their models were gaming benchmarking tests, finding ways to rack up good scores without truly earning them. At the same time, enterprises wanted help determining which model works best for their own internal needs rather than relying only on standardized benchmarks.

Arena did not say how it plans to use the new funding, and it raised the issue of how to measure model alignment. The company has also added a new category to its leaderboard: alignment.

This is where it ranks models based on issues like unauthorized action, false attribution, and what it calls "deceptive completion." Currently, a slate of OpenAI’s models are at the top of its preliminary alignment leaderboard, with Claude Opus 5.5 and Claude Fable in sixth and ninth place, respectively.

Source: techcrunch