Benchmarks

Arena's $100M Run Rate: Leaderboards as a Business

Arena, the crowdsourced AI leaderboard company, hit a $100M annualized run rate eight months after launching AI Evaluations, TechCrunch reported on June 29, 2026.

Arena's $100M Run Rate: Leaderboards as a Business — article cover

On June 29, TechCrunch reported a number few people saw coming: Arena, the company behind the crowdsourced AI leaderboard, has reached a $100 million annualized revenue run rate — just eight months after shipping its first commercial product. What most people know is the free leaderboard: users submit a prompt, compare responses from two anonymous models, and vote for the better one. More than 10 million of those evaluations have accumulated, spanning text, coding, vision, and image generation. CEO Anastasios Angelopoulos put it plainly: “A lot of people don’t even understand that our business is making any money at all,” adding that “people still see us as an open source project.”

The journey from a leaderboard everyone uses for free to a $100M-a-year business is worth a closer look for anyone who cares about how models get evaluated.

From a Berkeley Project to a $1.7B Company

Arena began as a UC Berkeley research project in 2023 and incorporated in April 2025. Its co-founders are CEO Anastasios Angelopoulos and CTO Wei-Lin Chiang, both Berkeley postdocs; the project was advised by Berkeley professor and Databricks co-founder Ion Stoica. The leaderboard’s flywheel has always depended on a simple trade: users get early access to new and often unreleased models, and in exchange they supply the pairwise votes that make the rankings statistically meaningful.

Capital moved faster than perception. In January 2026 the company raised a $150 million Series A at a $1.7 billion post-money valuation, when its annualized revenue stood at $30 million. Total funding has reached $250 million from Felicis, Andreessen Horowitz, The House Fund, LDVP, Kleiner Perkins, Lightspeed, Laude Ventures, and UC Investments.

The Revenue Engine: AI Evaluations

The money comes from AI Evaluations, launched in September 2025: in-depth performance analytics sold to model labs and enterprises. Buyers are effectively the same labs whose models appear on the public leaderboard, which makes the product a natural upsell — they already know where they rank; what they are paying for is understanding why. In the eight months after launch, annualized revenue climbed from $30 million to $100 million — more than tripling.

One billing detail deserves attention. Angelopoulos said the company bills based on consumption, which means the “$100M ARR” is not truly recurring subscription revenue — it is current usage extrapolated over a year. In an environment where AI startups routinely flatter their ARR, that is an honest but important footnote: if usage slows, the number does not stick the way subscription revenue would.

The Real Rivals: The Data Labeling Market

Angelopoulos is blunt about competition: Arena competes “for the same dollar” as Mercor, Surge, and Scale AI — the data and evaluation work that goes into model post-training. That market is inflating fast. Handshake’s annualized AI training revenue grew from $550 million to nearly $1 billion, and Mercor passed $1 billion, up from $500 million the previous September.

Yupp, the closest comparable crowdsourced-evaluation startup, shut down in March 2026, leaving Arena with no direct leaderboard rival. The leaderboard itself is also expanding its scope: a new Agent Mode now evaluates long-running agent workflows rather than single-turn question answering, following the industry’s own shift from chatbots to agents that use tools and operate over minutes or hours. The growth of the labeling-heavy businesses it competes with suggests post-training and evaluation demand is nowhere near saturated, which is the tailwind behind that $100M figure.

Influence and Caveats

Arena’s authority rests on those 10 million-plus crowd votes, and many voters come for early access to new and often unreleased models. That mechanism is both a moat and a structural tension: labs pay for deep analytics while depending on the leaderboard for exposure and trust. As “the leaderboard everyone uses” turns into a $100M business, the line between evaluation neutrality and commercial interest will be scrutinized at every model launch.

For model teams, two practical takeaways. First, human-preference evaluation is now pricing-and-positioning infrastructure for the industry, not something you can replace with academic benchmarks alone. Second, there is a fast-growing commercial services layer sitting on top of the rankings — any serious reading of a leaderboard position should factor that ecosystem in.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

FOUND_THIS_USEFUL?

Support more practical AI articles, tutorials, and build notes.

BUY_ME_A_COFFEE
SHAREXEMAIL