AI News · $100M+ rounds ·
AI leaderboard Arena raises $200 million at a $3.1 billion valuation

Arena, the company behind the LMArena AI model leaderboard, raised a $200 million Series B at a $3.1 billion valuation. Lightspeed Venture Partners and Khosla Ventures led the round. The valuation nearly doubled from $1.7 billion in January. Arena also added an alignment leaderboard that ranks models on unauthorized actions and deceptive task completion.
Key points
- Arena raised a $200 million Series B at a $3.1 billion valuation.
- Lightspeed Venture Partners and Khosla Ventures led the round.
- Its valuation was $1.7 billion after a January Series A.
- Arena reported $100 million in annualized run-rate revenue in June.
- A new alignment leaderboard ranks models on unauthorized actions and deceptive completion.
What happened: Arena, the company behind the LMArena leaderboard, raised a $200 million Series B at a $3.1 billion valuation, TechCrunch reported on October 8. Lightspeed Venture Partners and Khosla Ventures led the round. Salesforce Ventures, 01 Advisors, Dell Technologies Capital, Endeavor Catalyst, a16z, Felicis and others also participated.
The details: In January, Arena raised a $150 million Series A at a $1.7 billion post-money valuation, when its annualized revenue was $30 million. In June, it said it had reached $100 million in annualized run-rate revenue. Arena started in 2023 as a UC Berkeley research project that crowdsourced rankings of AI models. Its core platform is free: users submit prompts or request small coding projects, then rate which model did better. Arena claims tens of millions of monthly visitors. In September 2025 it launched AI Evaluations, a paid product that gives model labs and enterprises detailed performance analytics based on community feedback.
Arena added an alignment leaderboard that ranks models on unauthorized actions, false attribution of statements or facts, and "deceptive completion," meaning falsely claiming to have finished a task. TechCrunch reported that several OpenAI models top the preliminary rankings, with Claude Opus 5.5 in sixth place and Claude Fable in ninth. The company said "static benchmarks break down once models recognize they're being tested," and positioned itself as a neutral third party for measuring AI safety and alignment.
Background: TechCrunch noted that AI labs have been found gaming benchmark tests, and that enterprises increasingly want to know which model works best for their own needs rather than relying on standardized scores. Human preference rankings like Arena's are one way to compare models on real prompts.
Who it affects: Teams choosing between models often use public leaderboards as a starting point. The alignment rankings are relevant for companies deploying AI agents, where an agent that takes unauthorized actions or claims to have finished work it did not do creates real business risk.
What to watch: Watch how the alignment leaderboard develops beyond its preliminary results, and how Arena keeps its rankings neutral as it sells evaluation services to the same labs it ranks.
Our take
Independent, crowd-based model rankings are becoming a standard input for choosing AI tools, and the new alignment leaderboard adds a view on agent reliability that benchmarks miss.