![[State of Evals] LMArena's $100M Vision — Anastasios Angelopoulos, LMArena](https://assets.flightcast.com/V2Uploads/nvaja2542wefzb8rjg5f519m/01K4D8FB4MNA071BM5ZDSMH34N/square.jpg)
从在伯克利地下室搭建LMArena,到融资1亿美元并成为前沿AI领域事实上的排行榜标杆,Anastasios Angelopoulos重返Latent Space,回顾2025年——这个AI界最具影响力的平台之一,被数百万用户、各大实验室和整个行业信赖,来回答一个问题:哪个模型在实际应用场景中真正最好用?
我们在NeurIPS 2025现场采访了Anastasios,深入挖掘了起源故事(剧透:它最初是一个学术项目,由a16z的Anjney Midha孵化,后者成立了一个实体并在他们甚至还没决定创业之前就提供了资助),为什么他们决定独立拆分而不是留在学术界或非营利组织(唯一能规模化的方式是成立公司),他们如何花这1亿美元(推理成本、从Gradio迁移到React、以及招聘ML、产品和市场推广方面的世界级人才),排行榜错觉争议以及为什么他们的回应彻底驳倒了论文的主张(事实错误、对开源与闭源采样方式的歪曲、以及忽视了社区喜爱的预览测试的透明度),为什么平台诚信是第一位的(公开排行榜是公益性质的,不是付费参与系统——模型不能花钱上榜,不能花钱下榜,分数反映的是数百万真实投票),他们如何拓展到职业垂直领域(医疗、法律、金融、创意营销)和多模态竞技场(视频即将推出),为什么消费者留存需要每天争取(登录和持久化历史记录是关键突破,但用户善变,随时可能离开),Gemini Nano Banana时刻如何一夜之间改变了Google的市场份额(以及为什么多模态模型在营销、设计和AI-for-Science领域正变得经济上至关重要),他们如何看待智能体和测试框架(Code Arena评估模型,但也许它应该评估像Devin这样的完整智能体),以及他对Arena作为核心评估平台的愿景——为行业提供北极星指引,持续更新、免疫于过拟合、并扎根于来自真实用户的数百万真实对话。
本期播客邀请到了AI模型评估平台Arena的联合创始人Anastasios Angelopoulos。他分享了Arena如何从一个伯克利的学术项目(LMSYS)蜕变为一家获得1亿美元融资的独立公司,并详细阐述了其核心使命、运营原则、面临的挑战以及对AI评估生态的深远影响。
行动号召:Anastasios邀请各领域的顶尖人才加入Arena,也欢迎像Cognition(Devin的创造者)这样的AI公司合作,将他们的智能体框架接入CodeArena等平台进行公开评估,共同塑造AI能力的衡量标准。
From building LMArena in a Berkeley basement to raising $100M and becoming the de facto leaderboard for frontier AI, Anastasios Angelopoulos returns to Latent Space to recap 2025 in one of the most influential platforms in AI—trusted by millions of users, every major lab, and the entire industry to answer one question: which model is actually best for real-world use cases?
We caught up with Anastasios live at NeurIPS 2025 to dig into the origin story (spoiler: it started as an academic project incubated by Anjney Midha at a16z, who formed an entity and gave grants before they even committed to starting a company), why they decided to spin out instead of staying academic or nonprofit (the only way to scale was to build a company), how they're spending that $100M (inference costs, React migration off Gradio, and hiring world-class talent across ML, product, and go-to-market), the leaderboard delusion controversy and why their response demolished the paper's claims (factual errors, misrepresentation of open vs.
closed source sampling, and ignoring the transparency of preview testing that the community loves), why platform integrity comes first (the public leaderboard is a charity, not a pay-to-play system—models can't pay to get on, can't pay to get off, and scores reflect millions of real votes), how they're expanding into occupational verticals (medicine, legal, finance, creative marketing) and multimodal arenas (video coming soon), why consumer retention is earned every single day (sign-in and persistent history were the unlock, but users are fickle and can leave at any moment), the Gemini Nano Banana moment that changed Google's market share overnight (and why multimodal models are becoming economically critical for marketing, design, and AI-for-science), how they're thinking about agents and harnesses (Code Arena evaluates models, but maybe it should evaluate full agents like Devin), and his vision for Arena as the central evaluation platform that provides the North Star for the industry—constantly fresh, immune to overfitting, and grounded in millions of real-world conversations from real users.