![[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor](https://assets.flightcast.com/V2Uploads/nvaja2542wefzb8rjg5f519m/01K4D8FB4MNA071BM5ZDSMH34N/square.jpg)
从伯克利机器人实验室和OpenAI 2017年Dota时代的实习,到在GPT-4o、o1和o3上实现强化学习突破,再到如今在Cursor领导模型开发,Ashvin Nair经历了这一切。
我们在NeurIPS 2025上采访了Ashvin,深入探讨了OpenAI推理团队的内幕故事(剧透:从十几个人扩展到300多人),为什么IOI金牌在2022年感觉触手可及,但当o1真正实现时却没有改变世界,强化学习为何无法泛化到训练分布之外(以及为什么这意味着你需要通过产品和模型的协同设计,将经济上有用的任务带入分布内),强化学习研究时代(2017-2022)的深层教训,以及为什么大部分研究没有成功——因为社区过度拟合了基准测试,Cursor如何独特地定位于大规模持续学习——每两小时更新一次策略,产品-模型协同设计让工程师保持在循环中,而不是切换到ADHD地狱般的上下文切换,以及他的赌注:下一个范式转变是带有无限记忆的持续学习——模型只经历一次某件事(一个bug、一个错误、一个用户模式)就永远不会忘记,将数百万部署token存储在权重中而不会过载容量。
好的,这是为您生成的播客内容摘要:
本次采访在NeurIPS 2024现场进行,嘉宾是前OpenAI O1/O3团队成员、现Cursor机器学习负责人Ashwin。他分享了自己从机器人学博士到投身大语言模型研究的独特路径,深入探讨了OpenAI内部如何孕育出O系列推理模型,并展望了AI智能体与机器人技术的未来。
行动号召:Ashwin代表Cursor发出邀请,欢迎对代码数据、奖励模型以及产品-模型共同设计感兴趣的人才加入。
From Berkeley robotics and OpenAI's 2017 Dota-era internship to shipping RL breakthroughs on GPT-4o, o1, and o3, and now leading model development at Cursor, Ashvin Nair has done it all.
We caught up with Ashvin at NeurIPS 2025 to dig into the inside story of OpenAI's reasoning team (spoiler: it went from a dozen people to 300+), why IOI Gold felt reachable in 2022 but somehow didn't change the world when o1 actually achieved it, how RL doesn't generalize beyond the training distribution (and why that means you need to bring economically useful tasks into distribution by co-designing products and models), the deeper lessons from the RL research era (2017–2022) and why most of it didn't pan out because the community overfitted to benchmarks, how Cursor is uniquely positioned to do continual learning at scale with policy updates every two hours and product-model co-design that keeps engineers in the loop instead of context-switching into ADHD hell, and his bet that the next paradigm shift is continual learning with infinite memory—where models experience something once (a bug, a mistake, a user pattern) and never forget it, storing millions of deployment tokens in weights without overloading capacity.