人工智能工程师世界博览会:技术洞察、模型评估与未来趋势 AI Engineer 2026-07-02

开场白

Announcer: 发射控制中心,我们准备就绪。收到。女士们先生们,欢迎来到人工智能工程师世界博览会 (AI Engineer World's Fair)。感谢您加入我们,共同度过激动人心的一周,这里有创新、技术见解以及塑造 AI 未来的对话。现在,请与我一起欢迎您的主持人,IBM 的开发者倡导者,Tjas Kuman。

Original English

Announcer: Launch control. We have a go. Roger. Ladies and gentlemen, welcome to the AI Engineer World's Fair. Thank you for joining us as we continue an exciting week of innovation, technical insights, and conversations shaping the future of AI. Now, please join me in welcoming your MC, developer advocate at IBM, Tjas Kuman.

欢迎来到第二天

Tjas Kuman: 早上好,AI 工程师们。我们来到了这里。我们成功了。我们在这里。这是第二天。非常荣幸和荣幸能在这里见到这么多人。这次大会打破了记录,对吧?去年的规模要小得多。今年,有 7000 人参加。太难以置信了。请给予热烈的掌声。就是这里。就是这里了。这就是发生一切的地方。听着,这里有各种公告。有大量收获。有涵盖 18 个赛道的内容。18 个赛道。有博览会环节。有分组讨论。有各种各样的事情,对吧?而且无可否认,我会这么说,这里有价值。是的。如果你觉得有价值,请在今天早上大声欢呼。绝对的。绝对的。

Original English

Tjas Kuman: Good morning, AI engineer. We are here. We made it. We are here. It is day two. It is such an honor and a privilege to see so many of you here today. This conference has broken records, right? Last year was way fewer. This year, 7,000 people. Incredible. Huge round of applause. This is it. This is it. This is where it happens. Listen, there's announcements. There's takeaways. There's content across 18 tracks. 18 tracks. There's expo sessions. There's breakouts. There's all kinds of things, right? And undeniably, I'll say this, there is value. Yes. If you've got value, make some noise this morning. Absolutely. Absolutely.

Tjas Kuman: 我从这里众多聪明的人身上学到了太多东西,毫无疑问你们也是。我们昨天举行了一场令人难以置信的主题演讲。昨天我们有太多主题演讲了,Swyx 在会议开始时谈论了循环 (loops)。主题是循环。这有什么好笑的?好吧,这不是个笑话,但我们随后有更多关于 AI 黄金时代的主题演讲,对吧?有一件事让我印象深刻,我相信对我们很多人来说也是,那就是在最上游将 Agent 接入意图 (intent) 之中,这真正能解锁更多的工作。当我们开始说明为什么事情很重要时,我们就能够解锁更多工作和高质量的工作。我们不仅是把任务交给它,而是告诉它:做这件事,这是原因,这是你验证的方式,这是你部署的方式。我们能完成更多的工作。

Original English

Tjas Kuman: I have learned so much from so many brilliant people here and I have no question that you have as well. We had an incredible keynote yesterday. We had so many keynotes yesterday where Swyx started the conference talking about loops. The theme was loops. Why is that funny? Okay, it wasn't a joke, but we had more keynotes after that about the golden age of AI, right? One thing that really stuck out to me and I'm sure many of us was wiring the agent into the intent upstream really unlocks more work. When we start to say why things are important, we're able to unlock more work and quality work. We don't just hand it the task, but we say do this and this is why and this is how you verify, this is how you deploy. We get so much more done.

Tjas Kuman: Teresa 谈到了可靠性,以及它是多么重要。她谈到了领导者和落后者之间有 30 倍的生产力差距,向我们展示了比起其他任何事情,这真正关乎的是可靠性。这次大会的一个巨大重点在于评估 (evals)。最后,我真的被 Daksh 昨天的演讲震撼了,他谈到了审查了 100 万个 AI 生成的 PR,并发现了一些令人难以置信的见解。如果你错过了那个,我强烈推荐去看视频和直播。太酷了。有一件引人注目的事情,Claude 生成代码里,大概是多了三倍的绕过漏洞的代码,很遗憾目前是这样,但从这之中得出的所有见解真是太酷了。

Original English

Tjas Kuman: Teresa talked about reliability, how important it is. She talked about the 30x productivity gap between leaders and laggers, showing us that really it's about reliability more than anything else. Huge focus about evals at this conference. And finally, I was really struck by Daksh yesterday who talked about reviewing 1 million AI generated PRs and found some incredible insights. If you didn't catch that, I highly recommend the videos, the live stream. So cool. One thing that stood out, Claude, code generates, what was it, three times more off bypass vulnerability code, unfortunately for now, but it's just so cool all the insights that come out of this.

Tjas Kuman: 今天我们有很多安排。这是行程满满的一天,我对此感到非常非常兴奋。如果你还没读新闻的话,这里有报纸。我们现在有一份模拟报纸,就是为了平衡一下你们脑海里的 AI。所以这里有一份供你们阅读的每日纸质报纸。还有直播观众。直播观众大家好。感谢你们的加入。有超过 100 个博览会合作伙伴。有人去过博览会了吗?这些博览会展位令人难以置信。我看到了很多很酷的东西。到处都有机器人。有太多的东西了。赞助商之一那里还有一个很酷的设备,我拿到了那个 B 设备,它是用于面对面会议的记录仪。总之,去看看博览会。它太棒了。我们有 3.5 天的博览会时间,还有四个博览会舞台。所以,敬请期待。

Original English

Tjas Kuman: Today we've got a lot of things. It is a jam-packed day and I'm very, very excited about it. There's the newspaper if you haven't yet read the news. We have a newspaper now analog just to balance you know the AI. So there's a daily print newspaper available for you. There's a live stream audience. Hello live stream. Thank you for joining. There is over 100 expo partners. Anyone been to the expo? These expo booths are incredible. I've seen so many cool things. There's robots lying around. So much stuff. There's also this cool device that I got the B one of the sponsors. It's a notetaker but for in-person meetings. Anyway, check out the expo. It's so incredible. We've got 3.5 days of expo and four stages as well, expo stages. So, look forward to that.

Tjas Kuman: 我们想向令人难以置信的赞助商表达巨大的感谢并致以热烈的掌声。老实说,如果没有赞助商的支持,这次大会是不可能举办的。所以,请大家一起为大会的赞助商鼓掌。我们有微软,冠名赞助商。继续鼓掌。我们有微软。我们有实验室和白金赞助商。你们得继续鼓掌。我们有金牌赞助商。我们有银牌和铜牌赞助商。我们有这么多赞助商。如果没有他们,这次大会真的不可能成行。所以,我们非常非常感激。

Original English

Tjas Kuman: We want to offer a huge thank you and a massive round of applause for the incredible sponsors. Honestly, this conference would not happen without the support of our sponsors. So, please everybody, your hands together for the sponsors of the conference. We've got Microsoft, the presenting sponsor. Keep it going. We got Microsoft. We've got the lab and platinum sponsors. You've got to keep it going. We've got the gold sponsors. We've got silver and bronze. We've got so many sponsors. And this conference genuinely would not be possible without. So, we're very very thankful.

介绍首位演讲嘉宾 Tariq Shihipar

Tjas Kuman: 现在我们要介绍……我们要拉开舞台的序幕。这太酷了。今天将是有着满满精彩议程的一天,我希望你们所有人都能参加想去的场次。我的意思是,有相当多的赛道,但不用担心。有直播,也有视频。我们要开始介绍我们的第一位演讲嘉宾。哦,我对这个很兴奋。谁看到了昨天关于 Fable 的公告?是的,让我们开始吧。这太令人兴奋了。所以,非常巧合的是,今天的第一场演讲被调换了。这个大会正在以 AI 的速度前进。这太酷了。我们的第一位演讲嘉宾,来自 Anthropic 的 Tariq。请为 Tariq 鼓掌。他来自 Anthropic。哦,我对这个很兴奋,我在后台和他聊天时说,“这将会是关于什么的内容?”这场演讲,如果我没弄错的话,我认为是它首次对外公布,将会教我们所有人如何使用新的 Mythos 模型类别,其中 Fable 即将推出。所以,请为 Tariq 献上你们最热烈的掌声。请欢迎 Anthropic 技术团队成员 Tariq Shihipar 上台。

Original English

Tjas Kuman: Now we get to introduce we get to open the stage. This is so cool. Today is going to be such an incredible jam-packed agenda and I hope all of you can make all that you want. I mean, there are quite a few tracks, but don't worry. There's a live stream, there's also videos. We're going to start introducing our first speaker. Oh, I'm excited about this one. Who saw the announcement about Fable yesterday? Yeah, let's go. This is so exciting. So, coincidentally, the first talk has changed today. This conference moves at the speed of AI. It's so cool. Our first speaker, Tariq comes to us from Anthropic. Give it up for Tariq. Comes to us from Anthropic. Oh, I'm excited about I was talking to him backstage and I said, "What's this going to be about?" This talk, I think the first time it's ever been given, if I'm not mistaken, is going to teach us all how to work with the new mythos class of models, of which Fable is going to be soonly available. So, your biggest round of applause for Tariq. Please welcome to the stage member of technical staff at Anthropic, Tariq Shihipar.

Tariq 上台分享

Tariq Shihipar: 大家好,我是 Tariq。我在 Anthropic 负责 Claude Code。在开始之前,我们在 Claude Code 有个传统,就是要在演讲前自拍。所以,如果你不介意的话,如果你能和我摆个姿势,我会在 AI 工程师大会上拍一张快速的自拍。好的。太棒了。好了,为了开启今天的环节,正如我们所说,Fable 回来了。我们将在今天晚些时候推出它。请继续关注确切的时间表。我和 Cat Woo 以及 Simon Wilson 将在 12:30 进行一场炉边谈话。那时我们可能会有一些更新消息给你们。

Original English

Tariq Shihipar: Hey everyone, I'm Tariq. I work at Anthropic on Claude Code. Before we get started, we have a tradition on Claude Code where we take a selfie before a talk. So, if you don't mind, if you strike a pose with me, I'll take a quick selfie at AI engineer. Okay. Incredible. Well, yeah, to kick things off, like we said, Fable is back. We're rolling it out later today. Keep stay tuned for exact timeline. Me and Cat Woo and Simon Wilson will be doing a fireside chat at 12:30. We might have some updates for you then.

Tariq Shihipar: 但 Fable 是一个我非常非常兴奋的模型。它是 Anthropic 的那些模型之一,你就是会……你就是会记住它。就像新的 Sonnet 3.5,Opus 4,Opus 4.5 一样。它是一个我怀有很多喜爱和期待的模型。对我来说,描述 Fable 最好的方式就像是……地图正在打开,你懂吗,就像你在玩一个角色扮演游戏,你一直待在新手村里,现在你到了……你知道的,开放世界开始的地方,对吧?你可以去做和探索很多事情。但同时它也有点令人畏惧和困惑,对吧?因为你能做的事情太多了。所以我想在这场演讲中做的是,给大家一份 Fable 的实践指南,对吧?你该如何使用这个新类别的模型?

Original English

Tariq Shihipar: But Fable is a model I'm just so so excited about. It's one of those Anthropic models where you just like you're just going to remember it. Like Sonnet 3.5 new, Opus 4, Opus 4.5. It's a model that I just have a lot of like affection and excitement for. And the best way to describe Fable to me is like the map is opening up, you know, like you are playing like an RPG and you've been on the tutorial and now you get to the point where the like, you know, the open world starts, right? And there's so much that you can do and explore. But there's also it's also a little bit intimidating and confusing, right? Because there's so much you can do. And so what I wanted to do in this talk is give you guys a field guide to Fable, right? How do you work with this new class of models?

Tariq Shihipar: 这部分我分为四个部分。我之前一直将其作为系列文章和博客文章来写。但是,当宣布 Fable 即将发布时,我想:好的,让我在演讲中一次性讲完所有这些——你知道的,来个速通。一共有四个部分:释放 (unhobbling) Claude,寻找你的未知,应对悲伤,以及变得不讲道理。

Original English

Tariq Shihipar: So I've got four parts to it. I've been working on this as a series of articles and blog post. But you know when we announced Fable was coming out I was like okay let me do all of this at once at the talk, you know, speedrun. So there are four parts unhobbling claude finding your unknowns dealing with the grief and being unreasonable.

Tariq Shihipar: 首先,释放 Claude。我认为我们经常说的一句话是,模型是培育出来的,而不是设计出来的,对吧?我们不会一觉醒来就说,我们需要在 su bench 上达到 99%,对吧?模型是……你知道,是我们精心培育的东西,我们给它数据、反馈和计算资源。但归根结底,它是……你知道这是有些有机的东西,我们通过使用模型,与它一起弄清楚并学习。因此,这也意味着束缚它们的是我们,对吧?我们给它们套上的安全带,以及我们向它们发出提示的方式,基本上就是我们对 Claude 理解的一个函数,对吧?我所说的释放它,指的是我们如何更好地理解 Claude 以释放它的潜力?而我们需要更多地了解 Fable。

Original English

Tariq Shihipar: So first unhobbling claude. I think something we say really often is that the models are grown not designed right we don't wake up and be like we need 99% on su bench right like the models are you know something we grow carefully we give it data and feedback and compute but ultimately it's you know something that we it's a little bit organic and we sort figure out and learn with the model as we use it. And so that what that also means is that what contains them is us, right? The harness we put them in and the way we prompt them is basically like a function of our understanding of Claude, right? And by unhobbling it, I mean how can we understand Claude better to unleash it? And we need to understand Fable more.

Tariq Shihipar: 所以我认为我的观点之一是,你知道的,我们仍然处于非常早期的阶段,我认为 Fable 还有很多需要理解的地方有待解锁。我想给大家举个简单的例子来说明模型是如何变得更聪明的,因为它有点违反直觉,对吧?像几周前我看到了这条病毒式传播的推文,大概是说:你知道为什么大语言模型说不出哪些宝可梦的名字是以 'aw' 结尾的吗?有一千多只宝可梦,对吧?结果发现有两只名字是以 'AW' 结尾的,crocodile (实际上应为 Krookodile) 和 dreadnot (实际上应为 Dreadnaw),对吧?而且事实证明,如果你问普通的聊天模型,它回答不上来,这有点让人困惑,因为像你……

Original English

Tariq Shihipar: So I think one of my points is that you know we're still so early and I think there's a lot more understanding in Fable to unlock and I think I'll give you a quick example about how models get smarter because it's a little bit unintuitive right like there I saw this viral tweet a couple weeks ago being like you know why can't LLM say which Pokemon end in aw there are a thousand Pokemon right and turns out there are two who whose names end in AW crocodile and dreadnot, right? And it turns out if you ask like a normal chat model, it can't answer it, which is kind of confusing because like you

Capability Overhang 与 Claude 的发展

Speaker: ……你知道,它肯定知道所有宝可梦的名字,对吧?但如果你问 Claude Code,它可以做到,对吧?因为它所做的是获取每一个宝可梦,并编写一个脚本来筛选出以 AW 结尾的,对吧?这就是我所说的“释放 Claude 的能力 (unhobbling Claude)”。我们把这称为能力悬垂 (capability overhang),对吧?Claude 以一种参差不齐 (spiky) 的方式变得更聪明。所以它不仅仅是记住每一个宝可梦并进行推理,而是如果你给它代码执行工具,它就能找到那两个以 AW 结尾的宝可梦,对吧?因此我认为,在使用 Fable 时,部分挑战在于弄清楚这种能力悬垂。现在什么是可能的?我觉得这就像是一次我很期待与大家一起进行的发现之旅。

Original English

Speaker: [..] know it definitely knows all the names of the Pokemon, right? But if you uh ask cloud code, it can, right? Because what it does is that it fetches every Pokemon and writes a script to filter for AW, right? And so this is what I mean by like unhobling claude. We call this capability overhang, right? Cloud gets smarter in spiky ways. So it doesn't just remember every Pokemon and reason through it, but if you give it the code execution tool, it can find the two Pokemons that end with AW, right? And so this is I think part of the challenge with Fable is figuring out this capability overhang. What is now possible? And I think this is like a discovery that I'm excited to go on with you.

Speaker: 为了让这一点更清楚,我将谈谈过去模型是如何发展的一些不同例子。其中一个很大的例子显然是类似聊天的功能。你知道,聊天模型过去必须被赋予上下文,对吧?比如,也许你把你的代码库粘贴进去,也许你会天真地以为,我们解决编程问题的方式就是上下文变得非常大,我可以直接粘贴我整个代码库。你知道,它将是一个拥有 1 亿上下文窗口的模型。但事实证明,相反,如果你给它提供“手臂”,比如给它 bash 工具以及与环境交互的方式,它就能构建和搜索自己的上下文,这差不多就是催生 Claude Code 的洞察,对吧?所以,这又是一个在“我们如何思考和使用模型”方面参差不齐的、类似新创新般的进步。最近我们推出了 Claude Tag,而解锁 Claude Tag 的是它能够主动工作和进行多人协作的能力。你知道,Claude Code 是你必须去提示它,它才会去工作的东西,对吧?而 Claude 能够自己醒来并开始工作的这种能力,是我们认为正在解锁新一波智能体的关键。

Original English

Speaker: Uh to make this a little bit clearer, I'm going to talk about a few different examples of how models have progressed in the past. Um one of the big examples obviously is like chat. You know the chat models were had to be given context, right? Like maybe you paste in your codebase and maybe naively you might have thought like you know the way we solve coding is by the context just gets really large and I can just paste in my entire codebase. You know it'll be a 100 million context window. But it turns out that instead if you give it arms like you give it the bash tool and ways to work with the environment it can build and search its own context and that's sort of like the insight that led to cloud code right and so again spiky like a new like innovation kind of right in how we think about and work with the model and then recently we've rolled out cloud tag uh and what's sort of unlocked cloud tag is its ability to work proactively and multiplayer uh cloud code, you know, is something that you have to prompt for it to do work, right? And uh this ability for cloud to wake itself up and do work is something that we think is unlocking the new wave of agents.

Speaker: 但这里还有更多内容。例如,我们最近删除了 Claude Code 中 80% 的系统提示词,对吧?这就是模型及其需求随时间变化的体现之一。所以最初,比如在 Sonnet 3.5 的时候,系统提示词的最佳实践是一个简短的系统提示词,加上少量工具和大量示例,对吧?然后随着模型变得更聪明,你可以给它们提供更多信息和更多指令,它们开始遵循这些指令,于是变成了一个包含大量示例和许多工具的更长的系统提示词,对吧?但最近我们发现,这类新模型想要更少的内容,想要更简短的系统提示词。示例往往会限制它,因为它实际上比我们给它的示例更有想象力。因此,我们试图给它提供上下文,而不仅仅是约束。我们真的尽量避免说“不要做这个”。这在以前的模型中是非常必要的。所以这反映了系统提示词正在改变的方式,并且可能还会继续改变。

Original English

Speaker: But there's there's more here. So for example, uh we recently removed 80% of the system prompt for cloud code, right? And this is one of the ways in which models, you know, and what they need uh changes over time. So originally like you know maybe back in Son of 3.5 new the best practices for a system prompt was a small system prompt few tools and lots of examples right and then as the models get smarter you can give them more information and more instructions and they start following them and so it's a larger system prompt with lots of examples and many tools right but most recently we found this new class of models want fewer want a smaller system prompt the examples tend to constrain it because it's actually more imaginative than the examples we give it. And so uh and we tried to give it con context and not just constraints. We really try and avoid being like do not do this. Um which is really necessary for the previous models. Um and so this is like a way that the system prompt is changing and and probably will continue to change.

Speaker: 另一个我非常喜欢的功能是“询问用户问题” (ask user question) 工具。这是我刚接手 Claude Code 时做的一个功能。当 Claude 在做计划或想问你一个问题时,它可以向你展示一个多选题对话框。对于 Opus 4,它几乎无法调用这个工具。我不得不去调整这个工具,以确保它能起作用,对吧?然后大概在 Opus 4.5 左右,我想,如果我让它就规范问我 40 个问题会怎样?它可以开始采访我,对吧?于是它提问的能力实现了飞跃,对吧?然后最近在使用 Opus 4.8 和 Fable 时,我现在可以构建一个完整的 HTML 报告,里面嵌入了这些问题。这就像是与 Claude 交互的一种全新方式,对吧?因此,Claude 从你这里获取信息的方式也在发生着这样的演变。

Original English

Speaker: Uh another feature I really like is the ask user question tool. This is something I worked on when I first got to cloud code and and it's uh when claude, you know, a is is planning or wants to ask you a question, it can show you a multiple choice dialogue. Uh for Opus 4, it could barely call it. I had to like really tweak the tool to make sure that it was uh that it would work, right? And then sometime opus 4.5, I was like, well, what if I asked it to like, you know, ask me 40 questions about the spec, it can start interviewing me, right? And so its ability to ask questions jumped, right? And then most recently with Opus 4.8 and Fable, I can now build a whole HTML report with the questions embedded inside of them. And uh it's just like a whole new way of interacting with uh with Claude, right? And and so this progression of like how Claude can get information from you is also changed.

Speaker: 说到这个,Markdown 和 HTML 也是我经常谈论的东西。起初,Markdown 对于模型来说是一个很好的输出格式。它可以展示一些丰富的资讯。然后有了计划模式 (plan mode),它开始变成让你去理解 Claude 准备要做什么。而现在,Claude 可以为你构建这些深度的 HTML 报告,对吧?所以这再次体现了模型以一种参差不齐的方式变得更聪明。我真的想强调,这更接近于生物学而不是物理学,对吧?它仍然非常依靠经验,非常有机的。我们并不知道所有的规则,但它背后有某种科学,对吧?就像构建它也需要一种直觉。所以我真的鼓励大家用这种方式来看待 Fable。我们在 Anthropic 写过的我最喜欢的一篇论文是关于大型语言模型的生物学的。我们所有的研究论文都旨在让具有不同技术专业程度的人阅读,但这是我最喜欢的一篇。所以,如果你想了解更多,建议你去看看。

Original English

Speaker: Um speaking of which, uh markdown and HTML is something I've also talked a lot about. um you know it turned initially markdown was a a good output for the model um you know it could show a little bit of rich information and then you know with plan mode it started to be for you like you could understand what cloud was about to do um and now you know claude can build you these in-depth HTML reports right and so again a way of this the models getting smarter in a spiky way I really like to emphasize that this is closer to a biology than a physics, right? It's still very empirical, very organic. Um, we don't know all the rules, but there is some sort of science behind it, right? Like there is an intuition to build as well. And so I really, you know, encourage you to treat Fable like that. Uh, one of my favorite papers uh that at Enthropic that we've written is on the biology of a large language model. Um, all of our research papers are meant to be read by, you know, people with various degrees of technical expertise, but this is one of my favorites. So, uh, if you're looking to learn a little bit more, suggest you check it out.

解开自身的束缚与发现未知

Speaker: 但是,我们谈论了解开 Claude 的束缚,结果发现在使用 Fable 时,你也需要解开自己的束缚,对吧?所以我经常思考的一点是:地图不等于疆域,对吧?当我在处理一个编程问题时,我脑海中的计划、提示词和规范就是地图,对吧?但疆域是实际的代码库、真实世界,以及 Claude 需要去应对的约束条件,对吧?每当 Claude 在疆域中遇到地图上没有的东西时,我称之为一个“未知 (unknown)”,对吧?Claude 必须弄清楚该怎么做。那是一个我没有说明的决策点。而 Fable 是最早让我感觉到必须去弄清楚我自己的未知事物的模型之一,因为如果不去弄清楚,它将会遍历如此大的一个区域,以至于它会遇到很多这样的情况。

Original English

Speaker: But, so, uh, yeah, we talked about unhobling Claude, but it turns out when youre working with Fable, you also need to unhobble yourself, right? And so, one of the things that I think a lot about is that the map is not the territory, right? When I'm working on a coding problem, the plan and prompt and spec that I have in my mind is the map, right? But the territory is the actual codebase, the real world, the constraints that Claude needs to navigate, right? And whenever Claude runs into something in the territory that's not in the map, I call that an unknown, right? Claude has to figure out what to do about it. It's a decision point that I haven't specified. And Fable is one of the first models where I felt that like I really have to figure out my unknowns because if not it's going to traverse such a large area that like it's going to run into a lot of them.

Speaker: 那么你如何弄清楚你的未知呢?在使用 Fable 时,我的瓶颈在于我将地图与疆域匹配以发现我的未知事物的能力。所以有几种思考这个问题的思路。我喜欢用一个矩阵来思考。对于任何一个问题,我有一堆“已知的已知 (known knowns)”。这通常是我写在提示词里的内容。我想要什么?对吧?然后我有“已知的未知 (known unknowns)”。比如我知道我还没真正了解、但我只是还没弄清楚的事情。然后我有“未知的已知 (unknown knowns)”。比如什么事情太明显了以至于我根本不会写下来,但当我看到它时我就知道,对吧?最后,是“未知的未知 (unknowns unknowns)”。我完全没有考虑过什么?我不知道什么?对吧?就像如果我知道了什么,就能改变我提示 Claude 的方式?幸运的是,你可以使用 Claude,你可以使用 Fable 来寻找你的未知事物。

Original English

Speaker: So how do you figure out your unknowns? Um it I fable bottleneck my abil by my ability to match the map and the territory to find my unknowns. So a few um few ways to think about this. I like to think of it in a a matrix. So like for any problem, I have a bunch of known knowns. This is usually like what I write in my prompt. What do I want? Right? Then I have known unknowns. Things that like I know I haven't don't really know yet, but I just haven't figured out yet. I can um uh yeah, then I've got unknown known. Like what's so obvious that I just wouldn't write it down, you know, but I I know it when I see it, right? And then finally, unknowns. Unknowns. What haven't I considered at all? What do I not know? right? Like what is something that if I knew could change how I prompt Claude? And and luckily you can use Claude, you can use Fable to find your unknowns.

Speaker: 接下来我将介绍几个关于我是如何用 Fable 做到这一点的例子。首先是我喜欢做所谓的盲点扫描 (blind spot pass)。我喜欢这样说:“嘿,我正在处理一个新的 OAuth 提供程序,我对它一无所知,比如在这个代码库里。你能做一次盲点扫描,帮我找出相关的未知的未知,并帮助我写出更好的提示词吗?”对吧?所以这可能会让 Claude 遍历 OAuth 模块,并发现,比如:“哦,你知道,这有点像一个经常出现的棘手的小死胡同。”它也许会搜索我的 git 差异或 Slack。我也许会告诉它哪里有上下文,对吧?这样我就能了解,你知道,所有的坑。你可以非常广泛地使用这个方法,对吧?你可以用它来教你新领域的知识。我最近在做视频剪辑的调色时就用过。因为我认为这非常强大,而 Fable 在这方面表现得令人难以置信。在很多方面,模型对几乎所有事情的了解都比我多,我只需要把它挖掘出来。

Original English

Speaker: So I'm going to go over a few examples of how I do that with Fable. Um the first is I like to do what I call a blind spot pass. So I like to say something like, "Hey, I'm working on a new O provider that I know nothing about uh like in this codebase. Can you do a blind spot pass to help me figure out my relevant unknown unknowns and help me prompt better? Right? And so this like might have Claude go through the the O module and figure out like, oh, you know, this is kind of like a hairy little uh dead end that comes up a lot. Maybe searches my git diff or slack. I might tell it where there's context, right? So that I can learn about, you know, all the gotchas. And and you can use this very broadly, right? You can use it to teach you about new fields. I I recently did this for color grading when doing video editing. Um because I think this is really powerful and and Fable is incredible at it. Um in many ways the model knows more about you know almost everything than I do. I just need to get it out of it.

Speaker: 然后我喜欢使用头脑风暴和原型。这能帮助我找出我的“未知的已知”,对吧?尤其是对我来说,设计就是那种“当你看到它时你就知道”的东西,对吧?所以,我也许会要求它创建一个仪表板,我会告诉它我没有视觉品味,给我做一个包含四种截然不同的设计决策的 HTML 页面,这样我就可以对它们做出反应,对吧?你知道,你可以随意调整这些,但是就像

Original English

Speaker: Um then I like to use brainstorms and prototypes. Uh this helps me figure out my unknown known right things like especially for design for me it's like know it when you see it, right? So, I might ask it to uh create a dashboard. Um, and I tell it I have no visual taste. Uh, make me an HTML page with four wildly different design decisions so I can react to them, right? And and you know, you tweak this as you want, but like the

利用 Fable 提升编程效率与重塑工作流

Speaker A: 核心想法是大致了解那些你无法用语言描述的东西,对吧?然后,和模型一起合作来帮助弄清楚那到底是什么。接下来是面试。一旦我大致知道了“这就是我想做的”,这里可能还是会有很多未知因素,对吧?我可能有些东西没有考虑到。我可能没有具体说明。所以我就会让 Clog(Claude)来面试我,对吧?我会在任何这些问题中给它提供多一点上下文。比如给它更多关于你、关于这项工作以及你目前所处阶段的上下文,比如,“嘿,是的,优先考虑那些会改变架构的问题”,这是非常有帮助的。

Original English

Speaker A: idea is to sort of get an idea of like what are the things that you uh you know, you can't describe in words, right? And uh like work with the model to help figure that out. Uh then f then interviews. So once I have an idea of like, you know, this is what I want to do. Uh there's probably still a lot of like uh unknowns here, right? Where I might not have considered something. I might not have specified it. And so I'll ask Clog to interview me, right? And I'll give it a little bit more context in any of these questions. Like giving it a little bit more context about you and the work and the stage you're at, like, hey, yeah, prioritize questions that would change the architecture is extremely helpful.

Speaker A: 接下来是参考资料。给 Claude 提供地图的最好方法之一就是给它另一张地图,对吧?所以,相比于我自己去写出规范说明,我可以直接说,“嘿,这是一些代表了我想完成的工作的代码,对吧?它可能是在不同的系统或语言里的。但是只要阅读这些代码,理解它,然后用它来开始你的工作就行了,对吧?”同样,这可以有很多种不同的方式。如果我在做一个 React 组件,我可能会有一个 HTML 模拟原型作为我的地图,对吧,我就把它作为参考传进去。我觉得这真的非常强大,而 Fable 在这方面简直不可思议。

Original English

Speaker A: Uh then references. One of the best ways to give Claude a map is to give it another map, right? So instead of me writing out the spec, uh I can just say, "Hey, here's some code that represents what I want to be done, right? It could be in a different uh system or language. Uh but just read this code, understand it, and then use that to start your work, right? And uh again, this can be in a lot of different ways. If I'm making a a React component, I might have an HTML mockup that is my map, right, that I pass in as a reference. I think this is really really powerful and Fable is really incredible at it.

Speaker A: 我还非常欣赏的另一点是实现笔记。如果在运行 Fable 时,它遇到了一个未知情况,让它记录下来,对吧?这样,你就可以看到偏差发生在哪里,然后你也可以找出原因。通常,我们会给你一些关于发生了什么的上下文。最后,我喜欢让 Fable 考考我发生的事情。只是为了确保我明白我在做什么,而且在我创建 PR 或者合并代码时,我能代表这份工作。这是一个确保你真正参与在 Fable 的循环之中的非常好的方法。我认为 Fable 最重要的部分之一就是保持在循环中,确保你得到了你想要的东西。以上就是我关于如何与 Fable 合作的一些提示。

Original English

Speaker A: Uh something else I've like really appreciated is implementation notes. So if uh while you're running Fable uh and it runs into an unknown, ask it to log it, right? So that um you uh you can see where the deviations happened and then you can sort of figure out why as well. you know, we'll usually give you some context about what happened. And then finally, I like to get a fable to quiz me about what happened. Uh, just to make sure I understand what I'm doing and I can represent this work, you know, when I'm creating a PR or merging it. Um, this is a really great way of like making sure that you're like really in the loop with Fable. And I think that's like one of the most important parts of Fable is like staying in the loop and making sure that you uh you get what you want. So um those are some of my tips for working with Fable.

从挣扎于失败到变得“不讲道理”

Speaker A: 我还想说,我第一次使用 mythosclass 模型……第一次使用 Fable 时,我同时感受到了一种巨大的获得感,但也有一种失落感,我想稍微谈谈这个,你知道,当我回想大语言模型出现之前的编程时,感觉就像是在一个异国他乡,你知道,就像我曾经运营过一家大概 30 人的 YC 初创公司,我们总是不断被迫做出权衡,因为写代码真的太难了,对吧?比如,我们可以让应用变得很快,或者我们可以尝试开发一个新功能的原型,而这可能需要一个月,或者那会需要两个月,所以我们不得不做出选择,那真的非常非常困难。几周前我又回到了那个代码库,我思考了一些我曾经想做的事情,那真的简单太多了。就好像那些过去需要花我几个星期的事情,我现在几个小时就能完成,你知道吗?到了某个时刻,这就像是,是啊,你怎么能不笑呢,但说实话,你怎么能不哭呢?这就像是那样的事情,我曾经非常非常喜欢编程和手写代码。我喜欢那种在脑海中看到代码库并旋转它的感觉,但我也记得那些熬夜试图调试、连续几周工作却不起作用的时光。我只记得自己在失败中挣扎。我只记得我参与过的大多数项目都失败了。大多数初创公司都破产了。你知道,我认为总体而言,编程和写代码是极其困难的,尽管我很享受那些高潮时刻,但我无法回去了,对吧?我在这里的反思是,唯一的出路就是穿过去,对吧?关于编程,依然还有很多东西要学。关于 Fable,也还有很多东西要学。但我认为,如果我们非常努力,如果我们保持在循环中,解除对它的束缚,我们就能到达那里,你知道,我们能从另一边出来,带着更多更多的东西。最后我想谈的一点就是这个“更多更多”的部分,对吧?我称之为变得“不讲道理”。

Original English

Speaker A: Uh I also want to say that the first time I used a mythosclass model uh used Fable I felt both a huge sense of like gain but also a sense of loss and I I wanted to talk a little bit about that you know um when I think about coding before LLMs it feels like a foreign country you know like I used to run a YC startup about 30 people and we were just constantly forced into trade-offs because of how hard code right? Like we could make the the app fast or we could try prototyping a new feature and and this might take a month or this would take two months and so we had to choose and it was just really really hard. Um and now I went back to that codebase a couple weeks ago and I thought about some of the things that I wanted to do and uh it was just way easier. It was like the things that would have taken me weeks I could do in hours, you know? And uh at some point it's like yeah like how can you not laugh but also how can you not cry honestly like it's like one of these things where um I really really loved programming and writing code by hand. I love the feeling of like seeing the codebase in my mind and like rotating it but I also remember just you know like staying up late nights trying to debug working on things for weeks without working right. I just remember swimming in failure. I just remember that like the most of the projects I've ever worked on have failed. Most startups go bankrupt. You know, I think just overall programming and coding is extremely hard and like as much as I enjoy those highs, I I can cannot go back, right? And uh the way my reflection here is like the only way out is through, right? There's still a lot to learn with the coding. There's a lot to learn with Fable. Uh but I think if we try really hard and if we like stay in the loop, we unhobble it, uh we can get there, you know, and we can come out on the other side, uh with just um so much more. And so the last bit I wanted to talk about is is the so much more part, right? I call this being unreasonable.

Speaker A: 我最喜欢 Anthropic 的一点是,我们相信权衡是不存在的。就像,我认为很多时候,比如在我之前的公司里,我非常习惯于变得“讲道理”。我会写下这个优先级列表,然后我会觉得,“好吧,我想我们可以优先考虑这个而不是那个,对吧?”而且,你知道,那听起来很合理。所以这将会是我们这个季度的优先级,但是如果你把所有的都做了呢?你知道,如果你迫使现实向你展示这种权衡呢,对吧?这是我在 Anthropic 的文化中非常看重的东西。我未来的反思是,我要变得更加“不讲道理”。我认为像 Claude 和 Fable 的数学模型真的改变了你思考权衡的方式。有很多权衡是你隐性地在脑海中做出的,对吧?比如好、快、便宜。现在变成了挑三个,对吧?我认为去做更有野心的工作,最好的方法就是去重构框架,并且让我们自己变得更有野心,因为我认为证明智能体(agents)有效的唯一方式,就是以前所未有的速度完成我们一生中最好的工作。

Original English

Speaker A: Um one of my favorite parts of anthropic is that we believe that trade-offs are not real. Um, like I think that very often I like in my previous company I was very used to being reasonable. So I'd like write down this list of priorities and I'd be like, "Well, I guess we can prioritize this against this, right?" Um, and uh, like, you know, that makes sense. So we'll we this will be our priority this quarter, but what if you uh just did all of it? You know, what if you forced reality to show you the trade-off, right? Um, this is something I've really valued at our culture and anthropic. And my reflection going forward is that I'm going to be a lot less reasonable. Um, I think one of this like the math of Claude and Fable really changes how you think about trade-offs. And there are so many trade-offs that you make implicitly in your head, right? Like good, fast, cheap. Now it's pick three, right? Um, I think that like the best way to like do more ambitious work is to uh like reframe and make big make ourselves more ambitious because I think the only way to prove that agents work is to do the best work of our lives faster than ever before.

Speaker A: 比如,我昨晚用 Fable 花了大概四个小时做了这个演示文稿。我觉得这是一个我非常喜欢的文稿,我非常享受这个过程,但同时我也做得非常快。我认为如果你在这里,你知道,作为一个 AI 工程师,世界正在看着你去证明 AI 是有效的,对吧?它不仅仅是一阵风潮,而是它能够让我们更有效率,同时也能节省我们的时间。这也是我今年的愿望,去工作,变得更高效,但要减少工作时间,花更多时间和我真正在乎的人在一起。我认为同样值得指出的是,构建变得更容易了,但创造价值依然很难。我认为这是我们作为 AI 工程师有时会遇到的问题,我们过多地思考构建的过程和我们的配置。但核心目的是创造价值,对吧?这就需要很多次的挥棒尝试。要找到有价值的东西需要很多次尝试。但这确实是我们的目标。这再次表明世界正在期待我们去证明 AI 能真正改变世界。所以在最后我只是想说,去探索,去实现它,并且,是的,变得更“不讲道理”一些。谢谢。

Original English

Speaker A: Um, you know, for example, I made this deck last night in about four hours with Fable. I feel like it's a it's a deck I really like and I I really enjoyed it, but I also um you know did it really fast. Uh and I think that if you're here, you know, at AI engineer, the world is kind of looking at you to prove that AI works, right? That it's not just like a fad or something, but that it can make us more productive and also save us time. And and that's my resolution for this year is to work, be more productive, but work less and spend more time with people I really care about. Uh, I think it's also worth calling out that building is easier, but generating value is still hard. And I think this is something that we run into, you know, as AI engineers sometimes where we think so much about the process of building and our our setups. Um, but the the point is to generate value, right? And uh there it takes a lot of swings. It takes a lot of tries to find the valuable stuff. Uh, but that really is the goal. And that's like you know again what the world is looking to us to prove that AI can really transform it. So to to end I just wanted to say like go explore make it real and uh yeah be less reasonable. Thank you.

AI 时代的代码验证与价值榨取

Announcer: 请和我一起欢迎 Sonar 的首席执行官 Tariq Shaukat。

Original English

Announcer: Please join me in welcoming the chief executive officer at Sonar, Tariq Shakat.

Tariq Shaukat: 大家早上好。你们喜欢刚才的演讲吗?太棒了。你们一定特别喜欢结尾那部分,关于“不讲道理”的部分。我觉得那很棒。我也想说,我正在试着计算作为早上头两个演讲环节,一个 Tar 接着另一个 Tar 上台的概率是多少。我觉得这个概率应该相当低,不过,今天非常激动能来到这里。正如刚才提到的,我就职于 Sonar。我们处于代码验证领域,今天我是来谈谈验证的。我想我们在座的各位,很大程度上是因为我们在某种程度上相信 AGI(通用人工智能)已经来了。它正在到来。我们刚刚听到的关于 Fable 的模型,现在世界上正在发生的事情真的不可思议。然而我们几乎是专门与世界各地的企业客户打交道。我们有更多对话是带着问号的版本:AGI 真的来了吗?

Original English

Tariq Shaukat: Morning everyone. Do you enjoy that last talk? That was amazing. Um, you particularly love the end, the being unreasonable part. I thought that was awesome. Um, I also want to just I'm trying to calculate the odds of Tar following Tar as the first two sessions in the morning. Uh, I think the odds are pretty low on this one, but uh, thrilled to be here today. Um, as as we just mentioned, I am with Sonar. We are in the code verification space and I'm here today to talk about verification. And I think we're all here uh in large part because we believe to some extent that AGI is here. It's coming. The models we just heard about Fable, it's really incredible what is going on in the in the world today. And yet we work almost exclusively with enterprises around the world. And the conversation that we have more is the question mark version. Is AGI here?

Tariq Shaukat: 为什么他们会问这些问题?那是因为你每天都能在新闻里读到。我不是想在这里点名批评,但如果你看看 KPMG,他们发布的报告因为幻觉而不得不撤回,安永(EY)也在做同样的事情,律师事务所因为编造的引用、编造的判例法之类的问题而惹上大麻烦。我认为我们真的可以开始质疑,我们如何从 AI 中获得价值?正如我们刚才所听到的,模型令人惊叹,但就像刚才那位演讲者所说,困难的部分是如何从中榨取价值。我们面临的困境是 AI 生成的劣质内容(slop)无处不在。我相信你们在内部一定都见过这种现象。

Original English

Tariq Shaukat: And why are they asking these questions? It's because you can read the news every day. And I'm not trying to name and shame here, but if you look at KPMG putting out reports that they have to uh uh retract because of hallucinations, uh EY doing the same thing, law firms getting into lots and lots of trouble because of madeup citations, madeup case law, things like this. I think we can really start to question how do we get value out of AI? The models are amazing as we just heard, but the hard part as the other target just said is getting value out of it. The struggle is that AI slop is everywhere. I'm sure you all see this inside

AI 在现实世界与软件开发中的挑战

Speaker: (这同样存在于)你们的组织之中。我确信你们在日常生活中已经看到了,AI 非常令人惊叹。这些模型在生成看似非常合理的输出方面表现得不可思议。它们在生成听起来正确的内容方面令人难以置信,但它们真的正确吗?你如何知道它们是否正确,这是一个大问题。正如我们所见,在专业服务领域这是一个大问题。在法律领域也是一个大问题。但说实话,我认为如果我们坦诚面对,在每个行业、每个领域,无论是营销还是金融,随便哪个领域,这都是一个大问题。你都有这样的疑问:你怎么真正知道它是否真实?你怎么知道它是好还是糟?而我们在编码领域,尤其是软件开发领域,面临的这个问题,也就是我们在与在座的许多人以及许多客户交谈时经常被问到的:软件开发难道不一样吗?我们可以看看关于这方面的数据,以及 Mythos 模型。这是来自 Meter 的数据。你可能已经看过这个 MER。编码智能体正以非常快的速度变得更好。它们以极快的速度获得了巨大的提升。你可以看到这里的进展,这条指数曲线。这张图表上显示的是模型完成人类所需耗时任务的能力。那么,它们能完成需要 1 小时、2 小时,或者任何时长任务的能力如何呢?根据大约一个月前在预览模式下进行的基准测试,最新的 Mythos 模型至少能达到 16 到 18 小时的水平。所以这些智能体实际上能够完成长时间运行的任务,这确实开始改变工作发生的方式。

Original English

Speaker: of your organizations. I'm sure you see this in your everyday life that AI is amazing. The models are incredible at generating very plausible output. They're incredible at generating things that sound correct, but are they correct? And how do you know that they're correct is a big problem. And it's a big problem in professional services as we saw. It's a big problem in legal. But really, I think if we're honest, it's it's a big problem in every sector, in every field, whether it's marketing or finance or you name it. You have this question of how do you actually know if it's true? How do you know if it's good or if it is slob? And the question that we we deal in the coding space in particular, we deal with software development. And the question we get as we talk to I'm sure many of the people here in the room and a lot of our customers is, isn't software development different? And we can look at the data on this and uh the mythos models. Um this is data from um meter. Uh you may have seen this MER um the coding agents are getting better uh very quickly. They're getting a lot better very quickly. And you can see uh the progression the exponential curve here. What this shows on this chart is how capable are the models at completing tasks that humans would take. So can they complete a task that takes 1 hour, 2 hours, whatever it is. is the latest Mythos model at least per the benchmarking which was done a month or so ago in the preview mode was you're getting to 16 to 18 hours. So they're actually able the agents are able to complete longunning tasks and it really is starting to transform how work is happening.

成功率、质量与基准测试

Speaker: 但是,当你阅读这些数据时,一个关键的注意事项是,这是在 50% 成功率的前提下。好。所以它再次表明能够完成任务,但它是否能够“正确地”完成任务,才是问题的关键。所以如果你开始观察,让我们把准确率提高,你把它提高到 80%。虽然仍有进展,但进展要慢得多。这就不是 18 小时了,而是在 3 个半小时左右,或者类似这样的水平。顺便说一句,这还是在 80% 准确率的情况下。当我在向我的一位大客户的首席技术官展示这一点时,他的回应是:“Betaric,如果有人在绩效评估中给我 80% 准确的信息,我可能还是会给他打个差评,对吧?这未必达到了企业级的要求。”问题在于模型本身,这里完全公开地说,我们还没有在 Fable 模型上进行过这种基准测试,显然是因为它们才刚刚发布。但是,当你观察这些模型时,你会发现它们变得越来越聪明,但它们仍然会产生大量有问题的代码。这是我们做的基准测试。我们给模型一系列超过 4000 个问题,然后基本上要求它生成对这些问题的响应,接着我们分析其功能正确性——这非常关键,而它们在功能正确性这个概念上都做得极其出色,对吧?嗯,然后我们会观察代码的复杂程度、代码的错误率、代码的安全性。而你看到的是,即使是最先进的模型,其复杂性仍然很高。正如你在这里看到的,它实际上是相当多变的。GPT55 在控制复杂性方面做得特别好。但它仍然会产生错误。它没有产生大量的错误,但仍然会产生错误,而且依然会引发安全问题。

Original English

Speaker: But the critical caveat when you read the data is this is at a 50% success rate. Okay. So it is again able to complete tasks but is it able to complete tasks correctly is the question. So if you start looking at let's dial up the accuracy rate you dial it up to 80%. And there's still progress but it is much slower progress. Instead of 18 hours you're at about 3 and a half hours or something along these lines. And by the way this is still at 80% accuracy. And as I was presenting this to the CTO of one of my uh large customers, his response was, "Betaric, I would still put someone who gave me 80% accurate information on a performance review probably, right? This isn't necessarily enterprisegrade. The problem is that the models themselves, and full disclosure, we have not yet uh done this benchmarking on the Fable models obviously because they are just being released. But as you look at the models, the models are getting smarter, but they still produce a lot of problem problematic code. This is benchmarking that we do. We give the models a series of over 4,000 problems and we basically ask it to generate the response to the problems and then we analyze both the functional correctness which is critical and they all do extremely well on this notion of functional correctness, right? Um, but then we look at how complex is the code, how buggy is the code, how secure is the code. And what you see with even the state-of-the-art models is that complexity is still high. It's actually quite variable as you can see here. Um, GPT55 is done particularly well on the complexity side of things. It still generates bugs. It doesn't generate massive amounts of bugs, but it still generates bugs and it still generates security issues.

技术债务与验证的必要性

Speaker: 所以这就是进入智能体工作流的模型输出。再强调一遍,我不是在(针对 AI,因为我)在 AI 工程师大会上。我不是说 AI 是假的或者是不正确的,而是试图解决这样一个问题:你如何真正地在生产环境中从 AI 中获取价值?卡内基梅隆大学进行了一项研究,它考察了你从使用 AI 编码智能体中看到的实际生产力效益究竟是多少。我认为你所看到的与我在市场上亲眼目睹的许多情况产生了强烈的共鸣,那就是你一开始会看到生产力的大幅提升,特别是速度方面。你看到的是生产力或速度有 3 到 5 倍的提升。但是这种提升在三个月后就会消散。在三个月结束时,它开始回落到你使用这些智能体之前的正常水平。如果你问为什么,那是因为这里标红的两个部分,你开始看到,速度增加了,但安全问题也增加了。可维护性问题增加了。可靠性问题增加了,而且复杂性也增加了。所以本质上,你积累技术债务的速度和你生成代码的速度一样快,甚至可能更快,这就产生了另一系列的工作。它造成了一个不同的瓶颈。所以对我们来说,现在这就是 AI 领域的关键问题,即在一个代码是可证明的世界里。有一些关于形式化方法和证明等内容的会议我其实非常期待去参加。代码是可证明的,但当你开始处理大型代码库时,软件却并非如此。它仍然非常复杂。它仍然非常混乱。其中有许多的依赖关系。在大多数代码库中已经存在了大量的技术债务。所以验证这个问题实际上是关键。我所要论证的是,你可以将验证作为一种事后的补充,或者你可以将验证融入到流程之中。如果你将其融入到生成代码、进行软件开发的流程中,与你将其视为一种事后的补充,或仅仅将其视为老式的代码审查相比,你实际上可以开始从编码智能体那里获得实质性更好的结果。

Original English

Speaker: So this is the output of the models that are going into the agentic workflows. And again, this is not, you know, I'm at the AI engineer conference. This is not me saying AI is fake or or um incorrect, but it is um trying to address this question of how do you really get value in a production setting out of AI? This is a study that was done in Carnegie Melon uh University and it looked at what is the actual productivity benefit that you see from the use of AI coding agents. And what you see I think really resonates with a lot of what I see firsthand in the market which is you have a initial just amazing boost of productivity of velocity in particular. what you see is a three to 5x boost in productivity or in in velocity. Um that dissipates in three months. At the end of three months, it starts to come back to the the normal before you were using the agents. And if you ask why, it is because of the two pieces in red here that you start to see there's an increase in velocity, but there's an increase in security issues. there's an increase in maintainability issues. There's an increase in reliability issues and there's an increase in complexity. So essentially you're building the technical debt as quickly as you are generating the code or maybe even more quickly and that creates a different set of work. it creates a different bottleneck. And so to us, this is now the critical question in AI, which is in a world in which code is provable. And there's sessions that um uh I'm actually very much looking forward to attending about formal methods and proofs and things like this. Code is provable, but when you start dealing with large code bases, software is not. It's still very complex. It is still very messy. there's lots of um dependencies. There's lots of uh technical debt already in most code bases. And so this question of verification is actually key. And what I'm going to be arguing is that you can treat verification as an afterthought or you can bake verification into the process. And if you bake it into the process of generating code, of doing software development, you can actually start to get materially better outcomes from the coding agents than if you view it as an afterthought, if you view it as just the old school code review.

以智能体为中心的开发周期 (AC/DC)

Speaker: 因此,当我们在思考这个问题时,我们基本上构建了一个框架,围绕这个问题有很多相互竞争的框架,但我只给大家讲讲我们的框架。我们称之为“以智能体为中心的开发周期”。简称我们有时称之为 AC/DC,这里的想法是,你如何将支持验证的智能体循环置于中心。目前有很多焦点放在代码生成部分,比如你如何真正让模型和智能体生成解决问题所需的代码。我们认为,你应该用正确的规范、正确的工具、正确的流程来包围它,以做好三件事:引导智能体,Tar 实际上就此的各个方面谈了很多。引导智能体,验证结果,然后解决问题。如果你想在 AI 世界中取得成功,你必须让这成为规范的一部分,流程的一部分,以及新的软件开发生命周期的一部分。

Original English

Speaker: So as we've been thinking through this, we basically have constructed a framework and there's lots of competing frameworks around this, but I'll just talk you through uh ours. We call it the agent centric development cycle. for shortand we call it AC/DC sometimes and the idea here is how do you get verification powered to Gentic loops at the center's a lot of focus on the code generation piece like how do you actually get the models and the agents to generate the code that you need to solve the problem and what we argue is that you should surround this with the right disciplines the right tools the right processes to do three things to guide the agents and tar is talking a lot about different aspects of this actually. Guide the agents, verify the outcomes and then solve the problems. And you have to make this part of the discipline, part of the process, part of the new software development life cycle if you want to be successful in the AI world.

引导:上下文与约束的作用

Speaker: 所以如果我深入探讨其中一些部分,我们所说的“引导”是什么意思?我们围绕引导做了很多实验。我想我们昨天刚发布了一款叫做 Sonar Vortex 的产品,开始涉足这个领域。我们发现至关重要的是要将“引导”视为“上下文”和“约束”,并且我们非常刻意地将上下文和约束分离开来,因为上下文指的是,你有你的代码库。我们如何让智能体更容易理解,让模型更容易理解你的代码库里有什么?如果你有一百万行代码,或者一亿行代码,甚至十亿行代码,如果智能体理解你的代码库,它们会工作得更好。那么,你如何赋予它架构意识?你如何提供语义导航地图,并帮助它们了解“领域”——借用 Tar 刚才说的词。我们发现提供约束也同样有价值,而且我认为这部分被谈论得不够。你有希望你的代码遵循的指导原则。你有你觉得可以使用的依赖项。你有你不允许存在的依赖项。你有编码标准。你有护栏。你有预期架构。我们花了很多时间讨论现有的架构。但你们想达到的目标架构呢?所以这种提供上下文和约束的想法,我们在测试中发现,这能极大提高智能体的有效性,并大幅减少 token 的消耗。在解决特定问题时,token 消耗降低了 30% 以上。如果你问为什么,那是因为你实际上让智能体的工作变得更轻松了。你在帮助……

Original English

Speaker: So if I double click on some of these pieces, what do we mean by guide? We've done a lot of experimenting around guide. We've just launched a product um yesterday I think called sonar vortex that starts to get into this area. What we find is critically important is to think about guide as context and constraints and we separate out context and constraints very deliberately because context is you have your code repositories. How do we make it easier for the agents to understand for the models to understand what is in your codebase? If you have a million lines of code, if you have a hundred million lines of code, you have a billion lines of code, the agents work better if they understand your codebase. So, how do you give it architectural awareness? How do you provide uh semantic navigation uh maps um and uh and help them understand the territory to borrow what Tar was just talking about and we find it equally valuable and I don't think this part is talked enough about to provide the constraints as well. You have guidelines that you want your code to follow. You have dependencies you are okay using. You have dependencies that you are not okay having. You have coding standards. You have guardrails. You have intended architecture. We spend a lot of time talking about existing architecture. But what about where you want to go? And so this idea of context and constraints uh we've found in our testing generates a massive improvement in agent effectiveness and a massive uh improvement in token consumption. O over 30% reduction in tokens being used to solve a given problem. And and if you ask why, it's because you're actually making the life of the agent easier. You're helping

Speaker A: “这样能更好地导航。然后我们进入问题的核心,我们真的把指导看作是先发制人的验证。你如何确保需要验证的内容更少,需要修复的内容更少,诸如此类的事情。接下来进入验证的核心。我们坚信且在实践中行之有效的,是‘零信任多层验证’的理念。零信任是因为每个模型都有偏见。每个模型生成的输出都有其特征和个性。因此,让我们确保使用不同的模型和不同的技术,来确保您的代码是安全的,确保它是可靠的,确保它是有保障的。而‘多层’则呼应了前面提到的观点,即软件是复杂的。软件非常杂乱,包含许多错综复杂的问题。因此我们相信(并且再次发现这非常有效),将数据流、控制流、已知模式、机密信息等算法验证的组合,与如今通过代理进行验证所能实现的意图、业务逻辑以及‘未知之未知’结合起来。其实,再次借用上一场演讲的话,这些事物的融合,即你所部署的精心设计的多层架构,能在生产环境中真正看到成果。因此,当我们观察那些使用多层验证方法的合作伙伴和客户时,他们报告称,由AI导致的生产环境中断故障比不使用该方法的人少了44%。所以你可以开始看到在可靠性、安全性和可维护性方面实质性的改善。”

Original English

Speaker A: it navigate better. So then we get into the heart of this and we really think of guide as preemptive verification. How do you make sure there's less to verify, less to fix, this sort of thing. Then you get to the heart of verification and what we believe quite strongly and what we've seen work in practice is this idea of zero trust multi-layered verification. Zero trust every model has biases. Every model produces has a character has a personality. So, let's make sure we use different models and different techniques to make sure your code is safe, to make sure it's reliable, to make sure it's secure. And multi-layered really speaks to the earlier point that software is complex. Software is very messy. Software has lots of intricacies involved with it. And so what we believe and again have found to be quite impactful here is that a combination of algorithmic verification looking at things like data flows, control flows, known patterns, secrets, these areas combined with what is now possible with agentic verification looking at intent, business logic, the unknown unknowns. Actually again to borrow from the last presentation the fusion of these things the deliberate multi-layered fabric that you put in place can actually you can see the results of this in production. So as we look at our partners and customers who use a multi-layered verification approach they are reporting AI derived production outages being 44% less frequent than the ones who do not. So you can start seeing a material improvement in reliability, in security and in maintainability.

Speaker A: “然后我提到的最后一点是,技术债务确实会爆炸式增长。对吧?当你生成代码时,技术债务也随之产生。我要重申,这并不是说要停止生成代码。而是要意识到这一点,并开始控制它。因此,我们发现非常有效的方法是,建立一个积极的流程,围绕代码维护建立一种积极的纪律,并思考如何进行‘经过验证的代码维护’。我不会在这里逐一讲解每个步骤,但代理——无论是修复代理,还是围绕验证的严格纪律——确实能保持你的代码库整洁。很多人问我:‘好吧,人类开发者在乎整洁的代码,但代理在乎吗?’我们再次发现,它们绝对在乎。因为如果代理要操作代码库,它们就必须理解它。所以这是一个一次性的视角。我们认为这是一种会产生复利效应的东西。如果你在典型的代码库上执行完全相同的代理任务,然后再在一个已经清理过的代码库上执行,你会发现,与典型代码库相比,那些更整洁的代码库所需的Token、推理和能源消耗等都有显著的减少。对吧?如果你让代理的工作变得更轻松,如果你维护好你的代码库,那么你实际上就会看到复利效应。”

Original English

Speaker A: And then the last point I mentioned is technical debt does explode. Right? As you generate code, technical debt is also generated. And again, this is not stop doing it. This is be aware and let's start controlling it. And so what we have seen be super effective is to have an active process to have an active discipline again around code maintenance and thinking about how you do verified code maintenance. Um I won't walk through every step of this but a the agents whether that is a set of remediation agents whether it's a strong discipline around verification does keep your codebase clean and a lot of people have asked me all right but do agents care about clean code human developers care about clean code do agents care about clean code and what we find again is they absolutely do because the agents have to understand the codebase if they're going to operate on it so this is is a oneshot view. Um we think this is something that compounds. But if you just do the exact same agentic tasks on a typical codebase and then one that has been cleaned, you see a material reduction in the amount of tokens, reasoning, energy, etc. needed for those cleaner uh code bases versus the typical code bases. Right? If you make the life of the of the agent easier, if you maintain your codebase, then you'll actually see compounding effects.

Speaker A: “现在,在我们的观念中,重要的事情是构建系统。这正是我一开始所说的:我相信大家都会进行代码审查,你们可能会使用静态分析工具,也可能会使用AI代码审查工具等等。而我们认为,你们必须将其置于一个系统中。重申一下,我们很乐意在楼下的展台详细探讨它的具体运作方式,但我们坚信,在AI时代的软件开发生命周期构建中,需要将‘指导、验证和解决’的概念嵌入其中。你需要通过三个循环来完成它,你也需要思考这三个循环。首先是代理循环(agentic loop),我认为这是这次大会的一个关键流行词。现在的问题是,如何在代理生成代码和执行工作的同时,为其提供上下文和约束,并通过内循环验证(inloop verification),让代理在工作时获得验证结果,以及如何解决问题——这就是这里的蓝色循环。我们所说的正是内部循环验证的部分。其次,是持续改进流程,即如何真正将算法的优势与代理结合,来生成你的拉取请求(pull request),审查代码,顺便说一句,这个过程的速度必须大幅提升。因此,要使用代理来审查代码,进行这种多层验证,然后你有你的评估(evals)。我想开场演讲者已经谈到过,‘eval’可能是本次大会的另一个流行词。你有你的评估和质量关卡,以检查你是否真正通过了标准。所以,你有了代码维护循环、代理循环和CI验证循环,将这些以验证为中心的循环进行精心设计,就构成了一个复利系统。”

Original English

Speaker A: Now the important thing in our mind is to construct the system. And this is how I started is saying, you know, I'm sure all of us do code reviews, you may use static analysis tools, you may use AI uh code review tools, a whole range of things. And we believe that you have to put this in a system. And again, uh, we're happy to in our booth downstairs talk through what this looks like, but we really believe that the construction of the software development life cycle in an AI world um, needs to embed this notion of guide, verify, and solve inside of it. And you need to do it in three loops. And you need to think about these three loops. There's the agentic loop, which I think is the key buzzword of the conference. Um now but how do you provide the agents as it's generating the code as it's doing the work with the context and constraints with the inloop verification so that the agent is getting verification as it's working and how do you fix problems that's that's the blue loop here what we what we talk about is the inner loop verification piece there's a second which is your continuous improvement process and how do you really combine the power of algorithmic and agentic to generate your your pull request, review the code and by the way the velocity of this has to go up massively. So to review the code using agents and to this multi-layered verification and then you have your evals and I think the opening speaker talked about how eval may be the buzzword of the um conference. You have your evals and you have your quality gates to check are you actually passing. So you have your your code maintenance loop, agentic loop, CI verification loop and deliberate design of these loops with verification at the center is a compounding system.

Speaker A: “这是一个自我强化的系统。它可以正向强化,也可以负向强化。我们已经看到一些客户在推出AI编程工具时,忽视了验证。他们忽视了代码质量和代码维护的理念,很快就陷入了恶性循环。卡内基梅隆大学的案例研究实际上表明,如果你不加控制,所有这些好处会开始消散,或者你会陷入这种负面的自我强化循环。我们在一家使用一些尖端(也是今天在场的许多人正在使用的)代理编码工具的大型银行中进行了一项测试。如果你在这些代理循环中实际采用这种‘指导-验证-解决’(guide verify solve)的方法,你可以减少92%的问题。我再次强调,这种效果是复利的。并不是说每个循环都改善了92%,而是在你花费几分钟甚至几小时解决问题的过程中,你实际上会看到复利的收益。这基本上就是我们看待其优势的角度,也是我们看待如何在企业环境中控制、有价值地使用AI的角度。我说的企业是指,那些已经拥有数百万行现有代码库的人。这里有代理循环,有CI验证循环,还有代码维护循环。我的营销团队要求我放上一个包含我们产品的演示版本。所以,这些就是我们的产品,大家之后可以来找我们。但最重要的一点是,我们的建议是采用这种‘以代理为中心的开发周期’(ACDC,agentcentric development cycle)。其核心是将精心设计的验证系统性地内置其中。如果您想了解更多信息,我们在楼下有一个红色的大展台,我们非常乐意与您进一步交流。我们还有一些深入探讨的会议即将举行,请务必参加,祝大家在大会上度过愉快的时光。谢谢大家!”

Original English

Speaker A: It's a system that reinforces itself and it reinforces itself in the positive and it reinforces itself in the negative. And we've seen customers who uh have kind of neglected as they've rolled out AI coding tools, they've neglected verification. they've neglected this idea of code quality, of code um maintenance, things like that, and you get into a downward spiral pretty quickly. This is what the Carnegie Melon uh case study uh or study actually shows is that you actually have all the benefits start to dissipate or you can get into this self-reinforcing loop. And one of the tests we did with one of the large banks who are using some of the cutting edge the folks who are all around here today um cutting edge agentic coding tools they can get a 92% reduction in issues if you actually take this guide verify solve approach inside of those agentic loops. If again this compounds it's not that each loop is 92% better. it's that as you go through solving the problem over minutes and hours that you actually see a compounding benefits. So that is uh essentially how we see the benefit here. The how we see the controlled um value creating use of AI in enterprise settings. And when I say enterprises, people with existing code bases, people with with you know millions of lines of code already. There's the agentic loop, there's a CI verification loop, there's the code maintenance loop. I am required by my marketing team to put up a version of this that has our products on here. So these are our products and you can come and see us later. But the most important thing is really to say our recommendation is this agent the ACDC agentcentric development cycle. The core part is deliberate verification built into the system. So if you'd like to learn more um we have a booth that's the big red booth downstairs. We'd love to talk more. We have some doubleclick sessions coming up. So please do uh join those and uh have a great conference. Thank you all.

AI Agent End-to-End Workflow Challenges

Speaker B: “现在有请亚马逊AGI实验室的技术人员Onjab上台。大家早上好!很高兴能回到AI工程师世界博览会(AI Engineer Worlds Fair)。就在一年前,难题还是如何让代理在屏幕上找到一个按钮并点击它,尤其是它以前从未见过的屏幕。现在,代理已经可以驱动浏览器了,并且也开始驱动桌面应用程序。但是我们发现,点击实际上是容易的部分。我们还没有解决的是‘实际的工作’。我这是什么意思呢?让我们举一个很简单的例子。周一有一位新团队成员入职,也许你的工作是为他们设置账户、将他们添加到Slack频道、预约同事间的介绍会议、订购笔记本电脑等。在公司里,并没有人真正端到端地拥有这个流程,它可能还涉及到五个不同的系统。现在,代理很可能可以执行这个工作流中的每一个独立步骤,但在端到端完成任务时依然困难重重,因为真正的工作存在于所有那些不同应用程序和各种不同步骤之间的接缝中。这也正是它们最容易崩溃的地方。代理可以使用你提供的每一个工具,但它仍然无法完成完整的工作。那么,我们为什么会看到这种差距呢?花一分钟想一想我们实际上构建了什么。我们教计算机如何使用计算机。我这是什么意思?我们是从构建基础操作开始的。我们教它们点击、滚动、打字、调用API、填写表单,我们让这些步骤变得非常可靠,并且您可以将它们串联成一个工作流。如今的代理在运行那些工作流方面已经相当出色了。”

Original English

Speaker B: Joining us on stage is a member of technical staff at Amazon AGI lab onjab. Good morning. It's so great to be back here at the AI Engineer Worlds Fair. Just a year ago, the hard problem was getting an agent to find a button and click it on a screen, especially screens it had never seen before. Now agents can drive browsers and they're starting to also drive desktop apps. But what we figured out click clicking was actually the easy part. What we didn't solve is the actual work. And what do I mean with this? Let's take a very simple example. A new team member starts on Monday. And maybe your job is to set up their accounts, add them to your Slack channel, book intros with colleagues, order the laptops, etc. And nobody really owns this end to end process in the company and it might be also touching five different systems. Now, agents can most likely perform each single individual individual step of this workflow, but agents still struggle to do this end to end because the real work lives within the seams of all of those different applications, of all of those different steps you have to take. And this is mostly where it all falls apart. The agent can use every single tool you give it, but it still can't do the full work. So why do we see this gap? Think about for a minute what we actually built. We taught computers to use computers. So what do I mean with this? We started building out the basics. We taught them clicking, scrolling, typing, calling an API, filling out a form, and we got those steps, these steps really reliable, and you can string them together in a workflow. And agents these days are fairly good at like operating those workflows.

Reliable Agents and Shared Context

Speaker: 那么,为什么你不能直接把更多工作交给它们,然后完全放手,相信它们能完成呢?我前面提到的所有事情,比如模型本身使用工具、工具调用、把多个智能体串联起来,这些都是能力,我们基本上已经弄清楚了如何为模型添加能力。现在下一个难点真正是可靠性,如果没有可靠性,我们无法真正建立对这些系统的信任。

Original English

Speaker: So, why can't you not just hand them more of your work and then literally just walk away and trust it to be completed? So all the things I talked about like using a tool models itself, tool use, stringing agents together, this is all capabilities and we mostly figured out how to add capabilities to models. Now the next hard part is really reliability and without reliability we cannot really build up trust in those systems.

Speaker: 这里做一个快速的直觉测试,也许大家都可以想象一下,一个智能体在端到端工作流中执行任务。你觉得现在它真正成功的频率有多高?大概 60% 或者 80% 吧。这听起来挺不错的,但如果你仔细想想,如果你的智能体每四次中就有一次删除了数据库,你就再也不会碰那个智能体了,对吧?所以,当你需要这种可靠性时,你需要它达到几个 9 的级别。你需要相信它确实能够成功完成工作。

Original English

Speaker: So here's a quick gut check and maybe all of you can just think about an agent doing work in an end to end workflow. How often do you think that actually succeeds these days? Maybe 60 maybe 80% of the time. And it sounds really fine, but if you look into this, if your agent one in four times deletes a database, you will never touch that agent again, right? So when you need this reliability, you really need to be it in the nines. You need to have the trust that it actually can do the work successfully.

Speaker: 不过,实际上在一个领域我们在可靠性和信任方面取得了巨大进展,那就是编程,对吧?想想编程演进得有多快。我还记得第一次它开始为你自动补全的时候,对吧?你只需点击自动补全。太棒了。不久之后,它开始编写函数。我们觉得这太神奇了。现在看看当今。编程智能体编写代码。它们自己发起 pull request。我们本周早些时候就看到了。代码不断地飞速生成。

Original English

Speaker: Now, there's actually one place where we made enormous progress on reliability and trust and this is coding, right? Think about how fast coding evolved. I still remember the first time when it started autocompleting for you, right? You just tapped autocomplete. Amazing. Then short time later, it started to write functions. And we thought that is amazing. And now look at these days. Coding agents write the code. They open up the pull requests themselves. And we had it earlier this week. Code keeps flying by.

Speaker: 曾经有一段时间,对于它生成的每一行代码,我们都觉得必须要仔细阅读,确保它正确。我想在座的大多数人仍然对此深有感触。而如今,我认为几乎没有人还在这么做了,我们甚至无法做到这一点,对吧?代码生成的节奏太快了。同时,编程实现了这种跨越,为什么呢?因为我们成功地让编程智能体从仅仅“有能力”,变成了真正“可靠”并值得“信赖”。

Original English

Speaker: So once in a time we were able to just every single line that it generated we felt like the urge we need to really read it and make sure it's correct right I think most in the audience here can still relate to that these days I think hardly anyone is still doing that like we cannot even do that right code is generated at such a pace at the same time coding made that jump so why is that because we were able to bring it from just being capable the coding agents to actually be reliable and then trusted.

Speaker: 那么这是为什么呢?为什么编程问题最先被解决?这是因为代码是可验证的。你可以运行它,你可以测试它,你可以检查它,并且可以确信它能运行。因此,可靠性首先出现在你可以真正验证答案的地方。

Original English

Speaker: So why is that? Why was coding first solved? It's because code is verifiable. You can run it, you can test it, you can check it and you can be for sure that it worked. So reliability showed up in the first place you can actually verify the answer.

Speaker: 但问题来了。如果你看看更广泛的知识工作领域,我们所做的大部分工作并非如此。知识工作是混乱的,而且说到底整个真实世界都是非常混乱的。我创建的报告达到效果了吗?这个设计符合品牌形象吗?它理解我的真实意图了吗?没有任何单元测试可以回答这些问题。因此,在大多数工作实际存在的地方,验证真的碰壁了。它存在于我们日常使用的所有这些应用程序的缝隙中。至今还没人真正攻克这一部分。当没有办法轻易验证答案时,你如何让一个智能体变得可靠?这仍然是一个完全开放的领域。

Original English

Speaker: But here's the catch. Most of the work we do if you look at the broader knowledge work areas is not like that. Knowledge work is messy and heck the whole real world is really messy. Did the report I created land? Is the design on brand? Did it get it what I actually meant? So there is no unit test that can answer those questions. So verification really hits the wall right where most of our work lives. It's living in the seams of all of those applications we're using on a day-by-day basis. And nobody really has cracked this part yet. How do you make an agent reliable when there's no way to verify the answer that easily? And that's a field that is still wide open.

Speaker: 那么,我们该如何解决这个问题呢?嗯,人类是如何处理混乱的工作的呢?我的意思是,我们在这方面很成功,对吧?我们每个人每天都在跨越不同的系统工作。我们设法弄清楚如何让新同事入职。我们是怎么做到的。嗯,我们是通过一起弄清楚事情来做到这一点的。你拉上一个同事,进入一个 Zoom 会议,你们讨论事情,审视要解决的问题,你们讨论,指着各个系统,可能两分钟后,你们就解决了。就搞定了。但这些工作实际上没有一个是能直接验证的。我们整天都在这样做。

Original English

Speaker: So, how can we solve this? Well, so how do humans handle messy work? I mean, we're successful at it, right? Each of us like every day we work across different systems. We manage out how to onboard a new colleague. We do this. Well, we're doing it by figuring things out together. You grab a colleague, you jump on a Zoom meeting, you're discussing things, you're looking at the problem to solve, you're discussing pointing at systems, and maybe two minutes later, you solved it. You're done. But none of this work is actually directly verifiable. And we do this all day.

Speaker: 所以,其中一件事是,我们主要看着同一个屏幕,对吧?如果你和同事开会,你们俩都看到同一个屏幕,你们实际上可以非常快地弄清楚需要做什么。这就是现在的智能体所缺乏的。你不一定需要一个更聪明的大脑。你需要的是这种共享的上下文。因为如果我和智能体看着同一个屏幕,我可能就不需要解释那么多了。

Original English

Speaker: So one of the things is we're looking mostly at the same screen, right? If you're jumping on a meeting with a colleague, you see the same screen, both of you, and you can actually like figure out really quickly what needs to be done. So this is what the agent these days is missing. You don't necessarily need a bigger brain. What you need is this shared context. Because if we're looking the agent and myself at the same screen, I probably have much less explanating to do.

Perception Agents

Speaker: 那么,为了实现这一点,我们真正需要构建什么样的智能体呢?如我所说,今天的智能体已经可以看到屏幕了,对吧?而且它们可以在屏幕上点击并采取操作。这部分是可行的。但如果它们触发了操作,它们通常做的是直接继续下一步。它们不会观察发生了什么,或者在某一步失败、事情出现偏差时进行恢复。我们需要一个实际上能像你一样工作、像人类一样工作的智能体。

Original English

Speaker: So what kind of agent do we really need to build to achieve this? And today's agent, as I said, they can already see a screen, right? and they can click and take actions in it. That part works. But if they fire off actions, what they usually do, they move on. They don't watch what happens or recover if one step didn't succeed or something goes sideways. And we need an agent that can actually work like you do, like humans work.

Speaker: 一个例子就是机器人技术。如果你花点时间看看机器人是如何做到的,机器人感知周围环境,计划做什么,然后采取行动。因此,从感知到计划再到行动的这个循环,实际上也是我们在屏幕上需要的。这里真正始于第一个词,即感知。智能体必须像你一样观察屏幕,不是抓取页面背后的代码,而是观察实际渲染出的内容、布局、状态、工作中刚刚发生了什么变化、我们正在做什么,然后再执行。

Original English

Speaker: And one example is robotics. If you just look for a moment as how robotics do it, a robot perceives what's around it and it plans what to do and then acts. So this loop here from perceiving to planning to acting, this is actually what we also would need on a screen. And it starts here really with the first word which is perceive. The agent has to take in the screen the way you do, not scrape the code behind the page, but what's actually rendered, the layout, the state, what just changed the work, what we're doing, and then do it.

Speaker: 而且它还必须保持实时跟进。想想我们人类是如何协同工作的。你加入进来,你在彼此发言的基础上做出反应。而今天的智能体仍然做不到这一点。我们现在的做法是发送一个提示,然后等待,它离开,过了一会儿智能体回来了,我们可能还需要来回沟通几个回合,对吧?因为智能体给出的结果并不完全是我们想要做的。所以我们又发送一个提示说:“嘿,回去,做这个,用不同的方式做这个。” 于是我们就有了这种我们在聊天机器人体验中已经非常习惯的长时间的来回沟通,以及这种轮流对话的节奏。

Original English

Speaker: And it would also have to keep up in real time. Think about how we as humans work together. You jump in, you react to build on top of what each other you say. And today agents can still don't do it. What we're doing is we're sending a prompt, we're waiting, it goes away and at one point the agent come back and we might have to take a couple of turns, right? Because what the agent come back with is not exactly what we might want to do. So we're sending another prompt say, "Hey, go back, do this, do this differently." And we have this long back and forth which we got so used to from our chatbot experience and from this rhythm taking turns.

Speaker: 但其实我们真正需要的,想一想,是一个可以在你仍在工作时做出反应的智能体。那不是很酷吗,对吧?就像你工作的同时,它也能提出建议,能帮助你,而且不需要等待。基本上就是一个能感知你所感知的内容并理解你意思的智能体。我们称它们为感知智能体。

Original English

Speaker: But what we actually would need, think about it, is an agent that can react while you're still working. Wouldn't that be really cool, right? Like at the same time you're working, it can also come up with suggestions, can help you, and there is no waiting time. So basically an agent that perceives what you perceive and understands what you mean. We call them perception agents.

Speaker: 那么为什么需要感知智能体呢?为什么它们很重要?首先,它们完成了计算机使用的闭环。再次强调,今天的智能体可以行动,它们可以点击,可以打字,可以滚动,但它们做不好的是查看结果,以及确认操作是否真的奏效了。感知智能体可以读取渲染出的屏幕,所以它可以确认自己的输出,而不是仅仅触发那些操作然后去祈祷。

Original English

Speaker: So why perception agents? Why do they matter? So first they complete the loop on computer use. Today's agents again they can act, they can click, they can type, they can scroll, but what they can't do well is looking at the results and whether it actually worked out. A perception agent can read the rendered screen so it can confirm its own output instead of just firing off those actions and then hoping.

Speaker: 其次,它不需要 API 或后端进程。这很重要,因为它是基于渲染出的界面工作的。它看到的是和你看到的相同的像素和结构。而人们每天使用的大多数当今软件根本不暴露 API。

Original English

Speaker: Second, it doesn't need an API or backend process. And that's important because it works off the rendered interface. It sees the same pixels and the structure you see. And most of today's software people use every day don't expose APIs at all.

Speaker: 第三,信息的输入也可以是反向的。与其写一大段文字来描述你想要修改的内容,比如假设你正在开发一个网站,你想描述所有需要应用的修改。与其写这段很长的描述,如果你可以直接指着它说:“嘿,这里这个标题需要改一下。嘿,你能更新这个部分吗?”那该有多好。这是一个比文本更精确、损耗更少的信号,智能体可以完全根据你标记的内容采取行动。

Original English

Speaker: And then third, the input also goes the other way here. Instead of writing a long paragraph to describe what you want to change, let's say you're working on a website and you want to describe all the changes you want to apply. Instead of writing this really long description, wouldn't it be great if you can just point to it and say, "Hey, here this heading needs to change. Hey, can you update this section?" This is a much more precise signal and less lossy than text. and the agent can act exactly on what you marked.

Speaker: 这就是我们的起点,我很高兴地分享,我们最近刚刚开源了我们的感知智能体测试工具的前两部分。这里包含两个部分。一个是标注,你可以用它来告诉智能体你想要什么。然后第二部分,即验证部分,赋予了智能体检查自己工作的能力。

Original English

Speaker: So this is where we started and I'm happy to share that we just recently launched the first two pieces of our perception agent harness open source. There's two pieces. There is annotation which you can use to tell it what you want. And then the second piece, the verification part gives the agent the capability to check its own work.

Annotation Tool Demo

Speaker: 让我给你们展示第一个部分。这是一个关于我们标注工具的非常简短的演示。这是一个 Chrome 扩展程序,所以使用起来非常简单。我将在这里播放这个简短的视频演示。安装扩展程序后,你就可以直接在屏幕上选择不同的元素。在这个例子中,我们就在标题周围画个圈,标记出这个区域。也许你想改变它。为什么不呢?我们把它变成红色吧。

Original English

Speaker: So let me show you the first one. So here's a very quick demo on our annotation tool. This one is a Chrome extension, so it's super easy to use. And I'm going to play here this quick video demo. So you have the extension installed and then you can just select different elements on a screen. So this example, we're just drawing around the heading there, marking the section. And maybe you want to change it. Why not? Let's change it to red.

Speaker: 你也可以选择这个页面上的元素。你看,如果我把鼠标悬停在它上面,它就能找到正确的元素。你点击它,选中它,然后说出比如“将字体大小加倍”之类的话。你还可以看到这里的智能体是如何在屏幕上准确捕捉反馈、位置、样式元素的,然后它生成了这段

Original English

Speaker: You could also select the elements on this page. You see how if I hover over it, finds the right element. You click it, you select it, and say something maybe double the font size. And you see also how the agent here captures on the screen exactly the feedback, the location, the style elements and it creates this

Automated Verification and Testing

Speaker A: (……一份)完整的总结,然后你就可以使用它,并交给你的智能体去实现。因此,这里不再需要反复沟通,因为你精确捕获了在屏幕上看到的内容,而智能体也能看到同样的东西。现在让我们非常简短地看一下第二部分:验证(verification)。验证的理念是,你可以描述——让我们继续以 Web 开发为例。你可以在一个 design.md 文件中描述你的设计规则是什么。如果我播放这里的这段视频,接下来会发生什么呢?智能体实际上可以根据这些设计规范来检查它自己的工作。

Original English

Speaker A: complete summary which you can then use and then give your agent to implement. So there is no back and forth anymore because you captured exactly what you saw on screen and the agent can see the same thing. Now let's have a very brief look at the second one at verification. So the idea of verification is that you can describe let's stay in this case of the web development. You can describe in a design MD file what your design rules are for this. And then what happens if I play this video here, the act the agent can actually check its own work against those design specs.

Speaker A: 于是它会获取你定义的内容,包括颜色、组件、你的布局,如果你还没有事先写下来,它会将其转化为那些规则。它会进行两种类型的检查。首先,它会进行视觉检查,这非常酷。例如,检查所有元素是否符合品牌要求,布局是否正确。另一部分则是检查用户流程。它在那里所做的是,它实际上会像真实用户一样在应用中走完整个体验流程,比如根据可用的任务。它可能会添加一个任务,也可能会删除一个任务。所以,它还可以帮助你以自动化的方式走完那些用户流程。

Original English

Speaker A: So it will take what you defined, the colors, the components, your layout, and it turns it into those rules if you don't have it written before yet. And it does two kinds of checks. Then it does a visual check, which is really cool. So everything is on brand, for example. it's the right layout. The other part is also checking user flows. So what it does there, it actually walks through this experience through the app for example depending on the tasks available. It might add a task, it might delete a task like a real user would. So it helps you walk through those user flows as well in an automated fashion.

Speaker A: 当它完成后,它会编写一份你可以审查的报告,并且会指出哪些测试通过了,还会告诉你任何未通过的地方。所以最终,你不必在一天结束时的半夜还在手动点击测试这一切,因为这很棒。智能体已经为你完成了这项工作。但是,可能并不总是有一块屏幕,对吧?所以我刚才讲了很多,我称之为感知(perception)。我谈到了智能体会看到你在屏幕上看到的东西。但是,在你的一天中有些时候你是没有屏幕的。也许你在办公室里。你正走进一个与同事的会议。

Original English

Speaker A: And then once it's done, it's writing a report which you can review and it's going to call out which tests passed and it's going to tell you anything that didn't. So ultimately, you're the one that doesn't have to click through this at midnight at the end of the day because great work. The agent already did this job for you. Now there might not always be a screen, right? So I talked a lot right now. I called it perception. I talked about the agent sees what you see on a screen. But there are times in your day where you don't have a screen. Maybe you're in the office. You're walking into a meeting with a colleague.

Multi-modal Perception Beyond the Screen

Speaker A: 昨天我在这里的会议上做了一个有趣的实验。我拉上了我的同事 Giovanni,他也在现场,实际上在二楼有一个很棒的小会议隔间。我们也是碰巧找到它的。于是我们走进去,并在那里开了我们的设计会议。这里的目标其实就是想向你们展示,感知远不仅限于视觉部分。所以在这个例子中,我们想向你们展示的是,感知还可以是倾听会议室里你们在讨论什么。你可以从照片上看到,我们俩都戴着我们的 B 品牌设备。非常感谢 B 赞助这些设备。

Original English

Speaker A: So I did a fun experiment yesterday at the conference here. So I grabbed my colleague Giovanni who is also here and actually on the second floor there's a great like little meeting booth. We found that by coincidence. So we went in there and we had our design meeting. And the goal here is really kind of show you how perception is so much more than just the visual part. So in this example, what we want to show you is perception can also be listening in the room to what you're discussing. And you can see here on the picture, both of us are wearing our B devices. Big shout out to B for sponsoring these.

Speaker A: 嗯,所以我们坐在那里。我们佩戴的 B 品牌设备可以生成文字记录(transcript)。它们在倾听我们所说的话。然后我们进行了这场设计会议。关于如何修改这个网站,我有几个很棒的想法。嗯,你马上就会在这里看到它们。让我们快速看一下,如何使用这个设备在这个网站上改变同样的工作流程。我们进行了讨论,B 设备生成了文字记录,你可以看到在这里的右侧,我们直接拉取了这份会议记录,里面有整个会议的详细摘要。里面有我们讨论的内容,并且它基本上捕获了那些洞察。我们直接把它们放在这里,然后我们可以点击“应用”(apply)。

Original English

Speaker A: Um, so we're sitting there. We have our B devices that can do a transcript. They're listening to what we're saying. And then we have this design meeting. And I had a couple of great ideas how to change this website. Um, you will see them in a in a second here. So let's have a quick look how this changed the same workflow on this website using this device. So we had the discussion the be did the transcript and you can see here on the right we're pulling this meeting transcript right in there is a whole detailed summary of the meeting. There is what we discussed and then it basically captures those insights. We have them right here and we can click apply.

Speaker A: 这个“应用”按钮所做的就是直接把它发送给智能体。你可以在这里看到我那些疯狂的想法——把背景变成黄色,把标题变成红色,并且还更改了一个表情符号——这些都被直接应用了。并且它还立即直接触发了验证过程。于是它生成了这份报告,幸运的是,这个配色方案显然是在已批准的规则之内的,否则这看起来就像是你在这里做了一些奇怪的事情。但是同样地,如果你不想要黄色背景,你可以更改这些规则,它会确保我们仍然遵守这些指南。它会标记出任何偏离规范的东西。因此,如果你想更新设计规范(因为你其实喜欢黄色),或者你采取行动说“不,修复这个违规行为”,这个决定权在你手里。

Original English

Speaker A: So what this apply button does is it sends it straight to the agent. And you can see here my crazy ideas to turn the background to yellow, turn the heading to red, and also change an emoji directly applied. And it also straight kicks off the verification right away. So it creates this report and and luckily this color scheme was apparently into in the approved rules otherwise this would have looked like you did some weird things here. But again you could change those rules if you don't want to have yellow backgrounds and it will make sure um that we still adhere to those guidelines. It would flag anything that's off. So you have the judgment call if you want to either update the design specs because you actually like yellow or you take an action and say no um fix this violation.

Open Sourcing AI Agent Patterns

Speaker A: 但这真的只是非常初步的第一步。这两个部分仅仅是最早期的开端,我们将会在开源环境中构建剩下的部分,因为只有当更多的人使用这些模式、在它们之上进行构建、甚至打破事物时,这些模式才能变得更好。所以我在这里对你们的请求是:去尝试一下它们吧。它们都在我们的 GitHub 仓库里。告诉我们我们遗漏了什么。给我们反馈,告诉我们你们希望看到什么,以及这接下来应该走向何方。因为最终,我们当中没有任何一个人可以单独变得聪明,而这正是整个事情的意义所在。我们希望构建能够让我们所有人一起变得更聪明的 AI。

Original English

Speaker A: But this is really the very first step. These two pieces are the very first beginning and we're building out the rest in the open because these patterns can only get better if more people are using them, building on top of them, breaking things. So my ask here to you is go and try them out. They're on our GitHub repos. Tell us what we're missing. Give us the feedback what you would like to see where this should go next. because ultimately none of us get smart alone and that's the whole point. We want to build AI that makes all of us smarter together.

Speaker A: 那么,如果你对人机交互(human-agent interactions)以及我们如何看待这些模式的变化更感兴趣的话,我强烈推荐我同事 Danielle Persik 制作的这个播客。她是一位认知科学家,也是我们在实验室的 AGI ACI 团队的负责人,并与行业专家探讨了很多关于人机交互模式的话题。你可以在主流的播客平台上找到这个播客。本周我们还有更多的会议环节。嗯,所以去看看吧。我们在下面有一个展位。我们有专家演讲。我的同事 Gav Mishra 也将在 1:30 于“计算机使用(computer use)”轨道带来另一场讲座。强烈推荐去听听他的演讲《从 RL 到 IRL(从强化学习到现实生活)》。

Original English

Speaker A: Now, if you're interested in a little bit more on human agent interactions and how we see those patterns changing, I would highly recommend this podcast by my colleague Danielle Persik. She is a cognitive scientist and runs our AGI ACI team at the lab and discusses a lot about human computer interaction patterns with experts in the industry. You can find the podcast on on a popular podcast platform. We also have more sessions this week. Um so check them out. We have a booth down there. We have expert talks. We also have another computer use track talk coming up with my colleague Gav Mishra at 1:30 in the computer use track. Highly recommend checking out his talk from RL to IRL.

Speaker A: 最后,欢迎来找我们。我们在展厅有很大的展示区。我们很乐意继续与大家交流。如果你不能亲自来到现场,你也可以在我们的 GitHub 仓库查看我们的代码,并访问我们的网站。那么,非常感谢大家。请欢迎 Google DeepMind 研究副总裁,Benois Schillings 上台。

Original English

Speaker A: And then ultimately come find us. We have a huge presence down at the expo hall. We would love to continue the conversation with you all. If you're not here in person, you can also check out our code on our GitHub repo and check out our website. And with that, thank you very much. Please welcome to the stage the vice president of research at Google DeepMind, Benois Schillings.

Scaling Up Models and Reasoning

Benois Schillings: 好的,早上好。呃,能来到这里并有机会与大家交谈,真的非常激动。我的名字是 Benois Schillings。其实在机器学习方面,我算是个新手。直到一年半前,我还在 Google X 工作,你们中的一些人可能知道它。我们做过像 Waymo 这样的项目,它现在似乎出现在每个街角。呃,我们还做过像 Glass(谷歌眼镜)这样的产品。所以你知道,我们经历了成功和失败的交织,但在许多方面,这段经历对我来说是一次有趣的成长,让我知道如何在一个像 DeepMind 这样的地方带领一支研究团队。

Original English

Benois Schillings: All right, good morning. Uh this is really quite exciting to be here and have a chance to to speak with all of you. Uh my name is Benois Shellings. I'm actually a bit of a noob when it comes to to machine learning. Till a year and a half ago, I was working for Google X which some of you may know. We've done things like Whimo which seems to be at every street corner now. Uh we also do things like Glass. So you know we we had a mix of hit and success but in many ways this was for me an interesting formative experience on how to run a research team in a place like deep mind.

Benois Schillings: 我确实拥有一支令人难以置信的团队。在 DeepMind,我团队的目标基本上是开发任何能够让 Gemini 在未来一个月到一年内变得不可思议的技术。之所以是一个月,是因为如果你开始研究一周后需要什么,那就是一种完全不同类型的工作;而之所以是一年,是因为我不认为有人能真正预测那么长远的事情。所以在我的观点中,去思考未来一年会发生的事情本身就已经相当有野心了。

Original English

Benois Schillings: I do have an incredible team. Uh my team goal in deep mind is basically to develop whatever technology will be needed to make Gemini incredible between one month and one year from now. So one month because if you start to work on what is needed in one week, that's a very different type of job and one year because I don't think anybody can really predict anything that far. So that's already pretty ambitious in my opinion to think about things that would happen one year in the future.

Benois Schillings: 在这个角色下我们做很多事情。很大一部分与代码有关,这也是我今天演讲的主要主题。但我们也做了大量关于模型推理演进的研究,例如,我们还做了拓扑学研究,探索哪些新型网络结构可能会带来更好的性能。我们在强化学习(reinforcement learning)科学方面做了基础性的工作,这对于我们今天使用机器学习所做的事情来说是非常根本的。让我们来讲一点起源故事。

Original English

Benois Schillings: We do many things under that role. Uh a lot of it is related to code which will be the main subject of my talk today. uh but we also do a lot of research on what is the evolution of reasoning for models for instance or we do topology research what are new type of network that might bring better performance uh we do fundamental work in the science of reinforcement learning which is so fundamental to what we're doing today with ML let's do a bit of an origin story

AI for Code: The Origin of Pitchfork

Benois Schillings: 嗯,我们在 2018 年在 X 实验室启动了一个名为 Pitchfork 的项目,旨在探索机器学习如何真正改善代码编写的方式。这非常有趣,因为在 2018 年我们向 Google 展示这个项目时,老实说,根本没有人理睬我们。当时有一种观点认为:你为什么会需要机器学习来写代码呢?在同时,我认为我们完全低估了这种发展的速度。

Original English

Benois Schillings: Um, we started the project at X named Pitchfork in 2018 which was aimed at looking at how ML could really improve the way code is being written. And this was very interesting because in 2018 when we presented that at Google honestly nobody would give us the time of day. uh there was that point like why would you ever need ML to to write code? Um at the same time I think that we totally underestimated how fast this could go.

Benois Schillings: 当我们最初做那个项目时,想法是看看我们如何能够加速一段代码的演进。我们如何才能执行那些会拖慢代码开发速度的大量微小修改?你知道,就是那种仅仅需要三天审查周期的微小编辑,以及我们如何压缩这个周期。有些人谈论起了用自然语言(英语)编写代码的“感觉编程(vibe coding)”,当时老实说我完全不以为然,我的态度是:这就是我们为什么需要编程语言,英语不是一种编程语言。好吧,我想在这个问题上我错得离谱,但当时我们感受到的阻力让我想起了我自己的职业生涯对变化的抗拒。我已经写了 45 年代码了。我最早开始编写……

Original English

Benois Schillings: When we did that project originally the idea was to look at how we could speed up the evolution of a piece of code. How could we make many of those small changes which slows down code speed development? you know the small edit which requires a review that takes three days and how we could compress that cycle. Some people were talking about vibe coding writing code in English and at the time honestly I totally dismissed that I was that's why we have programming language English is not a programming language. Well, I I I guess I was pretty wrong on that front, but the resistance we felt at the time reminded me of how my own career was pretty resistive to to change. Um, I've been writing code for 45 years. Uh, I started by writing

早期编程时代与汇编语言

Speaker:为 Apple 2 和 Commodore 64 编写电子游戏。所以,我的起点是写汇编语言。当你在写汇编语言上花了很长时间后,你会对编译器产生很大的怀疑,对吧?那些东西真的在正确运作吗?然后,当你转向 C++ 并使用编译器时,你会把垃圾回收语言看作“这不是真正的编程。你需要自己管理内存”。好吧,如今我也使用 Python 进行编程。所以即便是老狗也能学到新把戏。因此,我确实明白这里面发生了什么。

Original English

Speaker: video game for Apple 2 and Commodore 64. So, my formation was to write assembly language. And when you spend a long time writing assembly language, you look at compilers with a lot of suspicion, right? Are those things really working correctly? And then when you switch to C++ and use compiler, you look at garbage collected languages as this that's not real programming. You need to manage your memory. Well, today I use Python and VIP coding. So even old dogs can learn new tricks. So I do understand what happened there.

软件时代的变迁

Speaker:我认为,软件的发展经历了几个时代。第一个时代,也就是我开始写代码的那个时代,其根本限制其实在于机器本身。我们需要做大量工作去榨取那些机器的最后一丝性能,那就是汇编语言的时代,在那时你写代码的方式必须要极其精确。后来,计算成本变得便宜得多,我们就转向了现代的云时代,在这个时代,获得极致性能不再是最关键的方面。你实际上可以通过暴力破解来解决许多问题,但真正的限制因素变成了我们进行模块化设计的能力。你知道,在那个时代,软件是只写一次的,所有的核心理念都在于:你要如何构建代码库?你要如何编写函数?你要如何将那个问题拆解为长期可管理的东西?

Original English

Speaker: I think that we have a number of eras in what happened with software and the first one was you know the one where I started writing code where the fundamental limit was really the machine and there was a lot of work to go and extract the last ounce of power out of those machine and that was the days of assembly language where you really needed to be incredibly accurate in the way you were writing code. Computing became much cheaper and we switched to the modern cloud era where getting the best performance is not the most critical aspect. You can actually brute force many problems but really what became the limiting factor was the ability for us to design in a modular way. You know this was the era where software was write it only once and this was this whole idea of how are you going to build libraries? How are you going to write functions? How are you going to break down that problem into something that is long-term manageable?

人类认知的局限与人工智能前沿

Speaker:那里的局限性,并且在很大程度上决定了我们的软件流程是如何运作的,其实在于人类的大脑。一个典型的人类能够获取的上下文在 7 到 9 个词元(tokens)之间。我的意思是,我们拥有非常丰富的词元,但你将其与现代机器学习相比,后者很快就会具备几乎无限的上下文。人类的那种根本局限性,在很大程度上决定了软件是如何被编写出来的。但这已经结束了,我们现在正在转向那个 AI 前沿时代,在这个时代,真正编写代码已经不再是挑战了。

Original English

Speaker: The limitation there and that determine a lot of how our software process are working where actually the human brain. A typical human is able to get the context between seven and nine tokens. I mean we have very rich tokens but you compare that to modern ML where the context is basically going to be infinite pretty soon. That fundamental limitation of human determined a lot of how software was being written. This is over and we're switching now to that AI frontier where really writing the code is not the challenge anymore.

架构与设计能力的价值凸显

Speaker:我稍后会对此多谈一点,但现在的瓶颈其实在于,你要如何确保那段代码正是你真正想要的?因为写代码很容易,但要为特定问题准确界定出到底需要什么,可能会困难得多。所以在至少可预见的未来,人类将扮演架构的角色,或者去思考:我让机器学习模型去设计的那段代码,其真正的潜在影响到底是什么?归纳性思维是我认为人类仍然拥有非常明显优势的另一个领域,那就是在一个广阔得多的上下文中审视一个系统,并能够发现各种模式,然后根据这些模式做出决定。

Original English

Speaker: I'll speak some more about it but the bottlenecks are really how do you ensure that that code is what you really wanted because writing the code is easy but getting what is needed for a specific problem can be much harder to specify so humans at least in the near future will be that role of architecture or thinking of what are really the implication of that piece of code I'm getting the ML to design. Inductive thinking is another category where I think humans still have a very clear edge which is to look at a system in a much wider context and to be able to detect patterns and from those pattern take some decision.

超越人类的语法生成

Speaker:那么我们今天处于什么阶段呢?超越人类的语法生成。我上一次让 Gemini 帮我写个函数,然后我看着那个函数心想“我能写得更好”,那是多久以前的事了?那样的时代已经结束了。我认为在编写代码的繁枝末节上,我的意思是,你可以反驳,可以争论,也可以找到反例,但那个时代确实已经过去了。我们目前仍有许多工作要做的领域,是多步骤的代码库。软件工程不在于写代码,软件工程是指:当你第一次加入一家公司,你发现代码库里有 3500 万行 PHP 代码,而你需要做出一些修改——那一天你才会明白什么是软件工程。而那也正是我们今天的模型或者说前沿模型正在取得进展的领域,但是,管理那种极端的复杂性并将其拆解为可管理模块的能力,仍是前沿领域正在不断探索的方向。

Original English

Speaker: So where are we today? Superhuman syntax generation. When is the last time I built Gemini to write a function for me and I looked at the function and I was like I can do that better. It's over. I think that the minutia of code writing I mean you can fight you can argue you can find counter example but that time is gone where we still have a lot of work to do is multi-step code base. Software engineering is not about writing code software engineering is the first time you join a company and you realize that there are 35 million lines of PHP in the codebase and that you need to make some changes that's the day you understand what software engineering is and that's a place where our models today or frontier models are progressing but this ability to manage that extreme complexity and break it down into manageable pieces is a place where the frontier is still moving.

架构层面的深远影响

Speaker:这甚至一路延伸到架构层面。你看看,比如说谷歌的架构,感谢上帝我们有杰夫·迪恩(Jeff Dean),你知道他是那里的首席架构师,但那就是那种级别的思考,它具有诸多深远影响,其范围可以从“你如何进行硬件优化”、“你如何管理安全”,一直到“你如何构建一个系统,以至于十年后你不会感到后悔连连”。我认为,这确实是我们今天正在努力推进的进展范围。所以代码编写的时代结束了,但还有大量工作要做,还有很多进展需要去实现。

Original English

Speaker: It goes all the way to architecture you look at I don't know the Google architecture thanks god we have Jeff Dean which was you know the key architect there but that's the level of thinking which has many implication which can go from how do you do hardware optimization how do you manage security how do you build a system so that 10 years later you're not full of regrets and I think this is really the range of progress we are working on today so code is over but there's plenty to do there's plenty of progress to be made.

为什么代码对于 AI 训练如此独特

Speaker:如今,代码是一个非常独特的问题,从某种程度上说,这也是我们在这方面进行深度推进(pitchfork)的原因。首先,代码就是海量的数据。在其他一些领域,你也能找到大量数据来训练你的模型,但代码领域简直不可思议。你可以直接去 GitHub 上开始抓取数据。因此,这就是那种训练数据量处于非常独特状况的问题之一。这也是一个进行验证十分合理的领域。你可以运行一段代码,可以编译它,也可以进行单元测试。所以,要弄清楚模型生成的内容是否正确,是一件非常合情合理、完全可行的事情。这造就了我们今天的成就。

Original English

Speaker: Now code is a very unique problem and in some way that's the reason we did pitchfork on this. First of all, code is a lot of data. There are other domains where you can find a lot of data to train your model, but code was so incredible. You could go and go on GitHub and start to scrape GitHub. So this was one of those problem where the amount of training data was a very unique situation. It is also a domain where doing verification is reasonable. You can run a piece of code, you can compile it, you can have unit test. So the ability to figure out is the model generating something correct was something that was pretty reasonable to do. That brought us where we are today.

训练数据的枯竭与自我对弈的崛起

Speaker:但今天发生的情况是,我们的训练数据已经耗尽了。我认为今天添加到 GitHub 上的新代码中,有 80% 是机器生成的。因此,由人类带来可用于挖掘和训练模型的知识的这种概念,正在走向终结。但好消息是我们可以进行自我对弈(selfplay),而自我对弈是我们 DeepMind 一直非常喜欢的方法。我想大家都知道 AlphaZero。AlphaZero 在没有任何人类知识的情况下,仅仅通过与自己对弈,就成为了一名超越人类的围棋和国际象棋棋手。我们现在就处于这样一个阶段:用于编写代码的前沿模型也能做同样的事情,它们可以创建自己的挑战。它们可以判断答案的有效性。它们甚至可以在某种程度上评判架构。因此,进行这种数亿小时编写代码自我对弈的能力,正是将我们带入下一个层级的东西。

Original English

Speaker: But today what happened is that we ran out of training data. I think that 80% of the new code added to GitHub today is machine generated. So the notion of human bringing some knowledge that can be used for mining and to train model is reaching an end. But the good news is that we can do selfplay and selfplay is something we always liked a lot at deep mind. I suppose all know alpha zero. Alpha zero became a superhuman go and chess player without any human knowledge just by playing against itself. We are now at that stage where frontier model for code are able to do the same where they can create their own challenge. They can judge the validity of the answer. They can even to some extent judge the architecture. So that ability to do those hundreds of millions of hours of selfplay writing code is the thing that will bring us to the next layer.

培育超级程序员的实验

Speaker:你知道这很有趣。我们来做个实验。找一个才华横溢的软件工程师,把他或她关在一个房间里两年,只提供披萨,并给他们一个任务:你需要成为一名更好的软件工程师。作为一个人类,你会怎么做?你会给自己设定一些挑战。那些你可以验证的挑战,然后你不断地在这些挑战上工作和编写代码。我们在这里也能做同样的事情。所以这变成了一个我们能拥有多少算力、多少自我对弈时间的问题,但这将拓展我们在超人类编程领域能走多远的视界。

Original English

Speaker: You know it's interesting. Do the experiment. Take a brilliant software engineer, lock him or her in a room for two years and feed pizza and give the mission you need to become a better software engineer. What do you do as a person? You give yourself some challenges. Challenges that you can verify and you keep working and coding on those challenges. We can do the same here. So this is an issue of how much compute, how much selfplay time we can have, but that will bring the horizon of how far we go in superhuman coding.

软件经济学的巨变与安全挑战

Speaker:因此,代码的经济学正在发生戏剧性的变化。你知道,正如我所说,我们发展出了一整套软件工程的文化、基础设施以及一系列公司,全都是基于这样一个假设:写代码是困难的部分,是昂贵的部分。而我们现在正处于一个编写代码是免费或几乎免费的世界。这就是为什么我在这里画了个波浪号(~)。这意味着我们将看到产生的代码数量会呈现爆炸式增长。这其中伴随着一些艰难的影响。首先是设计与充分性的问题。面对那堆将被写出或动态生成的代码山,我们如何保持系统在微观层面上正常运作且可靠运行?这对人类来说是个大显身手的角色。此外还有一个问题,那就是你知道的,我们现在写很多代码,但我们不再过多地去阅读它了。我的意思是,我知道我们仍有代码审查,但我预测在一年内,我们会让 Gemini 或其他模型生成代码,而且实际上没有人会去查看它。你知道这就像编译器一样——现在谁还会去检查他们的编译器输出的汇编代码?也许还有人会看,但那可能已经是强弩之末了。所以同样的事情也将发生在代码上,这就带来了一些问题:我们需要实施哪些新的流程才能使其保持在可控范围内?这就是为什么我列出了一小份主动护栏(active guard rails)的清单。我的意思是,你们都看过 Mythos 查看一段代码并发现其中存在数量多得不合理的漏洞的新闻。人们会争先恐后地去修补那些漏洞,但我认为这将是一个永无止境的过程:我们将会有一层由模型发现的漏洞,我们会修复它们,然后模型会变得更聪明,它们会更加深入地挖掘,发现更加难以察觉的漏洞。所以我认为,第一方面的要求是:相较于编写代码本身,我们需要对代码安全以及一段代码的深远影响投入至少同样多的思考,而终极目标(grail)以及,你知道,我的……

Original English

Speaker: So the economics of code are changing dramatically. You know, as I say, we developed a whole software engineering culture and infrastructure and set of companies based on the assumption that writing code was the hard part, that this was the expensive part. We're now in a world where writing code is free or nearly free. That's why I've got the tilda there. That means that the amount of code that we're going to see produced is going to explode. And there are some hard implications to that. First is the question of design and adequacy. How in front of that mountain of code which will be written or written dynamically, how do we keep systems which works and are reliable at the microscopic level? Great role for human. It is also the issue that you know we're writing code and we're not reading it very much anymore. I mean I know we still have code review but I would predict that in one year we'll let Gemini or other model generate the code and nobody will actually look at it. You know it's similar to compilers who still check the assembly output of their compiler and maybe someone there but that's probably the end of it. So the same thing is going to happen to code and that brings some question of what are the new process that we need to put in place to keep that manageable and that's where I've got a bit of a list active guard rails. I mean you've all seen the news of mythos looking at a piece of code and detecting a unreasonable number of vulnerability in that code. There is a rush to go and patch those vulnerability but I think that's going to be a never ending process you know we're going to get a certain layer of vulnerability discovered by models we're going to fix those models will get smarter they will go a bit deeper and find even more subtle vulnerability so I think that the first aspect is that we need to think at least as much about code security and the implication of a piece of code than on the code writing itself and the grail and you know something my

机器学习与代码生成的新前沿

Speaker A: 团队正在积极研究的是,与其检测漏洞然后再建议一些修复,不如从一开始就教模型编写正确的东西。这非常非常难做到,因为它非常依赖于上下文。另一方面是,这就是我所说的归纳式架构。呃,我认为今天的模型在转移知识方面仍然不是很好,也就是把知识从一个领域应用到另一个领域,或者把两个概念结合起来找到那些上下文的交集,从而能够进行演绎思考。如果我们真的想使用机器学习来编写那些非常复杂的软件系统,这是一项我们需要教给模型的技能,而且你知道其中一个方面就是真正教模型如何在问题面前进行正确的规划。你如何看待一个非常复杂的问题,并决定该问题的正确分解方式,从而为问题带来最佳的清晰度或正确性。我们还需要改变我们进行评估(evaluation)的方式。我的意思是,在我的看法中,threebench 是声名狼藉的,因为 threebench 只验证一段代码是否能运行并产生正确的输出。正如我之前提到的,那只是代码工程的一小部分。因此,比如说,我认为我们需要在我们使用的那些基准测试中加入更多开放式问题。我,我来举个例子。呃,我很喜欢文本压缩这个问题。每个字符你需要多少比特,你能做到多极致?所以那是一个写起来非常简单的评估。你只需拿一段 10 兆字节的代码,然后告诉模型:写出你能写出的最好的无损压缩器,而在那种情况下的损失函数将会是,你知道的,压缩文件的大小加上源代码的大小,那是永无止境的。我的意思是,我认为正是这些问题将迫使这些模型去做一些新颖的事情,比如创造出完全新的算法。而且我,我认为我们现在正在进入那个阶段。

Original English

Speaker A: team is working actively on is instead of detecting the vulnerability and then suggesting some fix how about teaching model to write correct things from the start and that is very very hard to do because it is very context dependent the other aspect is that you know that's what I call inductive architecture uh I think that models today are still not very good at transferring knowledge of taking knowledge from one domain and applying it to another one or taking two concepts and finding the intersection of those context to be those context to be able to do deductive thinking. If we really want to write those very complex software system using ML that is a skill that we need to teach and you know one aspect of that is to really teach models how to do correct planning in front of a problem. How do you look at a very complex problem and decide what is the right decomposition of that problem that will bring the best clarity or correctness to the to the problem. We also need to change the way we do evaluation. I mean u threebench is infamous in in my book because threebench verifies if a piece of code runs and produce the right output. That's only a small part of as I mentioned earlier of code engineering. So for instance, I think that we need some problems much more in those benchmarks that we use which are open-ended problem. I I'll give an example. Uh I love the question of text compression. How many bits per character do you need and how far can you go? So that's a very simple eval to to write. You just take a piece of 10 megabyte of code and you tell the model write the best compressor you can that is lossless and the loss function in that case will be you know the size of the compressed file plus the size of the source code that's never ending I mean those problems are I think what's going to force those model to do novel things like creating totally new algorithmic for instance and I I think we're now getting to that stage

Speaker A: 编写代码或者做软件工程不仅仅是像一串 token 那样去思考。如今的思考和推理是思维链(chain of thought),你知道这非常成功并且极大地改进了模型。但是,人类在思考问题的方式上当然要复杂得多。我一直认为编写代码是一项非常视觉化的活动,那可能是,我不知道,可能是你正在做的事情的框图,或者是数据流经代码的走向。呃,但如果说代码仅仅是你输出的一组 token 然后就成了代码,我认为这只能适用到一定程度。对于我们在 Google Gemini 所做的工作来说,这是一个非常有趣的方面,我们从一开始就做出了选择,这将是一个多模态模型,你知道,文本只是 Gemini 能够应用的一种模态。我们开始看到,你知道,模型如何能够开始以空间或动态表示的形式来思考以解决问题,我认为这将成为必备能力。另一个有趣的问题是,现在是否是时候为模型创造一种新语言了?Python 或你叫得出来的名字都是为人类发明的,而这些语言在编写安全或可靠的代码方面并不是很好。我的意思是,它们用来写代码很棒,但它们绝对不是,不是最好的东西。我认为,我们正在达到这样一个阶段:既然编写代码的痛苦已经不复存在了。那么我们为何不让编写代码变得更困难一些呢?比如通过引入,你知道的,非常强类型的语言,或者,你知道的,从 Lean 语言中汲取一些灵感,看看如何编写在设计上就注定不可能是完美的代码。我的意思是,程序是有局限性的,但至少要把正确性的重担放在模型上。所以我不知道这里是否有语言设计师,但我,我,我认为在那个领域确实有一些事情可以做,而且它不需要是人类可读的。我,我不认为那还有什么要紧的。

Original English

Speaker A: Writing code or doing software engineering is not thinking as a chain of tokens. Thinking and reasoning today is chain of toad which has been you know very successful and improve models a lot. But humans of course are much more complex in the way they think about problems. I always think that code writing is a very visual activity and that can be I don't know the block diagram of what you're doing or the flow of data through your code. uh but saying that code will be just a set of token that you emit that are going to be the code I think goes only up to a certain point that's a very interesting aspect to what we do at Google Gemini we made that choice from the onset that this would be a multimodal model that you know text was only one of the modality that Gemini would be able to apply and we're starting to see you know how can a model start to think in term of spatial or dynamic representation to to solve problem and I think that's going to become a must have another interesting question is is this time to create a new language for models Python you name it have been invented for humans and those language are not very good to write safe or reliable code I mean they're great to write code but they're certainly not the the best thing I think We're getting to the point where since the pain of writing the code does not exist anymore. How about we make writing the code much harder by having you know very strongly typed languages or you know some inspiration from lean on how to write code that by design it's not going to be perfect. I mean program is something which has some limits but at least putting the burden of correctness on the model. So I don't know if we have some language designers here but I I I think there's something really to be done there and it doesn't need to be human readable. I I don't think that that will matter anymore.

Speaker A: 所以,除了代码之外……代码是解决问题的通用语言。我认为我们开始看到,这种在代码中极快地进行实验的能力,正以极快的速度影响着其他领域,因为进行实验基本上变得没有成本了。因此,我认为,着眼于代码编写与原子或科学的交汇点是我们正在开辟的另一个大前沿,那将是真正新颖事物出现的地方。对我来说尤其令人兴奋的两个领域之一是化学。嗯,你知道,作为人类,我们并不懂化学,或者说我们只懂化学中非常、非常小的一部分。一旦你的分子中超过了 20 个原子,情况就像是“哇,我们不知道那个东西会产生什么反应。”我认为我们将看到从那里涌现出不可思议的事物。我的意思是,一旦你能够把一万个原子放在一起,那开始看起来就像是生命了。所以,用一万个原子你还能做所有其他什么事?生物学。你们可能听过很多关于它的事,但你知道,生物学是自然界做了一项令人难以置信的工程工作,但在文档记录上做得糟糕透顶。嗯,但我们现在可以攻克它了。模型能够找到那些可能让我们觉得难以捉摸的关系。所以我认为这将打开一扇不可思议的大门。然后就是我称之为“我们看不见的黄金”的东西。人类在感觉什么是正确解决方案时有着令人难以置信的偏见。我的意思是,我们是那种为了帮助我们在丛林中生存下来的进化训练的结果,对吧?而不是为了做量子计算。因此,我认为,尽管我们可以是聪慧且有创新精神的,但有一大堆的进步和突破是可以实现的,只是我们看不见或者感知不到罢了。如果我有更多时间,我会举一些例子。我认为在那些方面中,机器学习对许多这类问题提供了一个如此不同的视角,以至于我们将会产生“天哪,这东西一直就在我们眼前,而我们居然没看出来”的感叹。所以,令人兴奋的时代就在前方。非常感谢。

Original English

Speaker A: So beyond code um code is a universal language to solve problems. I think that what we're starting to see is this ability to experiment very quickly in code is impacting other domain very quickly because doing experiment becomes basically free. So I think that looking at that intersection of code writing and atoms or science is another big front that we are opening that is the place where true novelty is going to appear. two which are especially exciting for me is chemistry. Um you know as humans we do not understand chemistry or we understand a very very small sliver of chemistry. Once you have more than 20 atoms in your molecule it's like wow we don't know what that thing is going to do. I think we're going to see incredible things emerging out of that. I mean once you are able to put 10,000 atom together that starts to look like life. So what are all the other things you can do with 10,000 atoms? Biology. You probably heard plenty about it, but you know, biology is the case of nature did an incredible engineering job and terrible job at documentation. Um, but we can crack through that now. Models are able to find those relationship that might be elusive for us. So I think that that is something that will open incredible door. And then there is what I call the gold we cannot see. Humans are incredibly biased in what we feel is the correct solution. I mean, we're the result of an evolutionary training that help us survive in the jungle, right? Not doing quantum computing. So, I think that even though we can be brilliant and innovative, there are a whole bunch of progress and breakthrough that can be done which we just cannot see or perceive. If I had more time, I would give some examples. I think that's one of the thing where ML is such a different viewpoint on many of those problems that we're going to get the oh my god this was in front of us the whole time and we could not see it. So exciting times ahead. Thank you very much.

过渡环节与模型评估(Evals)的未来

Announcer: 女士们,先生们,随着我们今天日程的继续进行,请欢迎你们的司仪,IBM 的开发者倡导者,Tea Scamar 再次上场。

Original English

Announcer: Ladies and gentlemen, as we continue today's program, please welcome back your MC developer advocate at IBM, Tea Scamar.

Tea Scamar: 今天真是个不可思议的开局。哇呼!大家正在离场。从这里看真是太棒了。嗯,在我们休息或者在那之后,嗯,让我们花点时间感谢一下赞助商。说实话,没有他们这不可能实现。我们要把幻灯片放上来。听着,你们需要给他们最热烈的掌声。我的意思是,这真的太酷了。谢谢你们。谢谢你们。谢谢。感谢微软。感谢这里所有的其他赞助商。没有他们,这场活动就不可能举办。还有很多其他的事情正在其他舞台上进行,但毫无疑问,嗯,评估(evals)在 AI 领域是一件大事。事实上,它们就是质量的把关门槛,对吧?嗯,我们可以发布很多东西,但如果它们没有经过评估,我们就会发布出一堆垃圾。因此,呃,我们的下一场讨论,我们的下一场会议,将由来自 Arise 的 Aparna Dinakaran 带来,她会跟我们多聊聊有关 EVAS 的事。请将你们最热烈的掌声送给 Aparna。

Original English

Tea Scamar: What an incredible start to the day. Woo! Everybody's leaving. This looks amazing from here. Um before we break off uh or after um let's take a moment and acknowledge the sponsors. Honestly, this would not be possible without them. We're going to get the slides up. Listen, you need to give them your biggest round of applause. I mean, it is so cool. Thank you. Thank you. Thank you. Thank you, Microsoft. Thank you to all the other sponsors here. This event would not be possible without them. There's plenty of other things happening um in the other stages, but there's no doubt that um evals are a huge deal in AI. In fact, they're the gate of quality, right? Um we can ship a lot of things, but if they're not eval, we ship a lot of slop. And so, uh our next discussion, our next session is going to be from Aparna Dinakan from Arise, who's going to talk to us a little bit about EVAS. Please, your biggest round of applause for Aparna.

Announcer: 请欢迎 Arise 联合创始人兼首席产品官 Aparna Dinakaran 登台。

Original English

Announcer: Please welcome to the stage co-founder and chief product officer at Arise Aparna Dinakaran.

Aparna Dinakaran: 大家好,你们都能听到我说话吗?好的,我们开始吧。哦,我往回翻一页。太棒了。那么,大家好。我叫 Aparna,是 Arise 的创始人之一。我们与一些出色的团队合作,帮助他们构建评估(evals)。嗯,今天我们在 Evals 分会场为大家准备了令人难以置信的演讲阵容。嗯,它就在 2005 房间举行,在这之后将会有来自 Turnbench、Uber 和 Snorkel 等公司的杰出讲者发表演讲。嗯,但今天我站在这里是为了和大家探讨评估的未来。评估已经从每个产品经理和每个 AI 工程师必须学习的新技能,变成了每一个严肃认真的 AI 团队所押注的东西。我们非常幸运能够与世界上最顶尖的一些 AI 团队合作。所以,我们在前排的位置不仅能看到他们在构建实际智能体时以及在实际发布之前发生了什么,还能真正看到评估团队是如何通过追踪记录在他们实时生产环境中的智能体上运行的。给大家分享一些统计数据。我们每个月运行超过一亿次评估。平均每个团队运行大约 12 个不同的评估任务,而最顶尖的团队运行的评估器超过 3800 个。离线评估、在线评估,它们各自都有用武之地。但今天,我要和你们谈论的,其实是那些正在——

Original English

Aparna Dinakaran: Hey everyone, can you all hear me? All right, let's go. Oh, let me go one back here. Awesome. Well, hey everyone. My name is Aperna, one of the founders of Arise. We work with some amazing teams to help them build evals. Um, and we have an incredible lineup of talks for you all today at the Evals track. Um, it's happening in room 2005 and there's going to be amazing speakers from Turnbench and Uber and Snorkel kind of all happening after this. Um, but today I'm here to talk to you about the future of evals. Evals have gone from the new skill that every PM and every AI engineer has to learn to the thing that every serious AI team is betting on. We've been really fortunate to get to work with some of the best AI teams in the world. So we get a front row seat into not just what's happening when they're building their actual agents and before they actually ship, but actually the eval teams are running on their live production agent via their traces. Little bit of some stats for you guys. We run over a 100 million evals every month. The average team runs about 12 different eval jobs with the top teams running over 3,800 different evaluators. And offline evals, online eval, they each have their own place. But today, what I'm actually going to talk to you about is the teams that are

Agent as a Judge 评估新范式

Speaker A: 在 traces(追踪记录)上运行 evals(评估)。这实际上能帮助团队弄清楚哪些方法有效、捕捉他们的故障,而这正是推动持续学习循环所需的数据类型。

Original English

Speaker A: running evals on their traces. This is actually what's helping teams figure out what's working, catch their failures, and that's the type of data you need to fuel your continual learning loops.

Speaker A: 整个行业基本对此达成了共识。我的意思是,Anthropic、OpenAI 的所有 CPO,还有你熟悉的 GDB,以及 Gary Tan 都在说 eval 就是你需要的一切。整个行业基本对此达成了共识。所以我们加入了 evals。它们能捕捉所有的故障。对吧?

Original English

Speaker A: And the industry kind of agrees. I mean all the CPOS of Anthropic, OpenAI, all you know GDB, you have Gary Tan saying eval are everything you need. And the whole industry kind of agrees. So we added evals. They catch all the failures. Right?

Speaker A: 问题在这里。当我们在构建所有这些第一代 evals 时,我们实际评估的对象在不知不觉中发生了变化。在 2023 年,一切还只是关于回答一个 prompt(提示词)。到了 2024 年,我们开始看到各种不同的前沿模型。它们加入了 tool calls(工具调用)。它们加入了推理。它们加入了深度研究。现在我们面对的是,团队在真实世界的数据上运行循环,并启动 sub agents(子智能体)来处理长周期的任务。这些任务中的每一项,实际上都是复杂性上的巨大跨越。我们不仅把问题变得更难了,而且实际上遇到了一个根本不同类型的问题。

Original English

Speaker A: Here's the problem. While we were building all of these firstgen evals, the thing that we were actually evaluating has changed underneath us. In 2023, it was about just answering a prompt. In 2024, we started to see all different frontier models. They've added tool calls. They've added reasoning. They've added deep research. Now, what we have is teams running loops on real world data with sub agents kicked off on um long horizon tasks. Every one of these was actually a massive jump in complexity. And we didn't just make the problem harder. we actually got a fundamentally different type of problem.

Speaker A: 这意味着随着这些系统变得越来越复杂,它们发生故障的方式也随之变得更复杂。我们非常幸运,因为我们在 UI 中构建了自己的智能体 Alex,而且我们能亲身体会到这种痛苦。每次前沿实验室(Frontier Labs)添加新功能,我们都会把它添加到我们的智能体中。现在 Alex 拥有了长得多的记忆。它具备创建动态 UI 的能力。它可以在海量的 traces 中进行搜索。但我们也意识到它会遗忘上下文。它不知道什么时候事情已经完成。有时候它就会卡在这些循环里。这里的关键是,在座各位可能写过的那些传统的“LLM 作为评判者”(LLM as a judge)的 evals,已经不足以让我们捕捉到我们所经历的所有类型的故障。我的意思是,这根本就不一样,对吧?你们以前拥有的是确定性的流程,而现在的情况是,几乎每次用户与 Alex 互动,它都会创建一个新的 UI,这是一个根本上完全不同的运行轨迹。所以这引发了我们一个非常重大的启示:如果评估一个智能体的最佳方式,实际上是用另一个智能体呢?这并不意味着我们以前使用的所有确定性 evals 和“LLM 作为评判者”的传统 eval 不再重要了,这只是意味着我们需要一种不同类型的工具来解决不同类型的问题。“智能体作为评判者”(Agent as a judge)是关于自适应的动态分析。而“LLM 作为评判者”只是给你一个带有这些固定分数的固定评分标准。这是大家都在做的事情。但是,当你的智能体在每次用户输入数据时都进行完全不同的运行轨迹时,这就意味着你需要一种根本上不同的 eval。

Original English

Speaker A: What that meant is that as these systems got more complex, so did the way that they actually fail. We're really lucky because we have our own agent that we've built, Alex, that lives in our UI and we get a kind of get to feel this pain ourselves. Every time the Frontier Labs added new functionality, we added it to our agent. And now Alex can has much longer memory. It has the ability to create dynamic UIs. it can go search across an enormous volume of traces. But we also realized that it would forget context. It wouldn't know when something was done. Um, sometimes it would just get stuck in these loops. And the key thing here is that the classical LLM as a judge evals that probably many of you have written in this room just weren't enough for us to be able to catch all the types of failures that we were experiencing. I mean, it's just fundamentally different, right? you have a deterministic flow and now what we have is literally every time a user interacted with Alex it would create a new UI that's a fundamentally different trajectory so this led to our really big revelation what if the best way to evaluate an agent was actually with an agent doesn't mean that all of the ways that we did eval with deterministic evals with LLM as a judge classic eval doesn't matter anymore but it just means that we have a different type of tool to solve a different type of problem. Agent as a judge is about adaptive dynamic analysis. LM as a judge just gives you a fixed rubric with these fixed scores. It's what everyone's doing. But when your agent's doing completely different trajectories every time a user puts in data, it just means that you need a fundamentally different type of eval.

Speaker A: 我的看法是,如今大多数团队都在做前两种,但 eval 的未来实际上是兼顾这三种。今天我非常激动地分享,我们发布了“智能体作为评判者”来帮助我们的团队进行他们的 eval 旅程。我们发布了 Signal。Signal 实际上是一个长期运行的智能体,它能够读取发送进来的 traces 并发现问题的模式。它能找出那些传统的“LLM 作为评判者” eval 通过那些确定性的评分标准永远无法发现的问题类型。它帮助我们找出了那些你甚至想不到的非常细微的故障,比如在循环中反复发生多次的问题。它在很长一段时间内重复调用同一个工具。运行轨迹非常低效。而它实际上能做的是,因为它掌握了所有的分析结果,它能够提交一个 PR(Pull Request)并提交修复代码。所以,如果你们想了解更多,请来我们的展位。我们就在 OpenAI 的展位旁边。我们会给你们做一个演示。我们会向你们展示更多关于它的内容。就像我说的,我们还将接管 eval 专场。所以,请来 205 房间。我们将探讨很多关于 evals 的未来以及它们的样貌。如果你们只是想和我们的团队聚一聚,我们今晚会举办一场美国队世界杯比赛的观赛派对。所以,去看看 Luma 并注册来加入我们。太棒了。非常感谢大家。

Original English

Speaker A: My take is that most teams today are doing the first two, but the future of eval is actually having all three. And today I'm actually excited to share we've released agent as a judge um to help our teams on their eval journey. We've released signal. Signal is actually a longunning agent that can read traces sent in discover patterns of issues. Um, it can figure out types of problems that a classical LLM as a judge eval just would never be able to do with these deterministic rubrics. It's helped us figure out um very subtle failures that you wouldn't even think of doing such as something going on in a loop for multiple times. It was calling the same tool uh for repeatedly long time. The trajectory was inefficient. And actually what this does is because it has all that analysis, it can go put up a PR and put up a fix. So, if you want to learn more, come to our come to our booth. We're right by the OpenAI booth. We'll give you a demo. We'll show you a bit more about it. Um, we're also, like I said, taking over the eval track. So, come to room 205. We're going to be talking a lot about the future of evals and what they look like. And if you just want to hang out with our team, we're throwing a viewing party for the USA World Cup uh game tonight. So, uh check out the Luma and register to come join us. Awesome. Thank you all so much.

OpenGV 智能体生产实践介绍

Gabe Dees Mesa: 讲述这一切是如何发生的。我们将探讨 OGs 在 Effect 上的重大押注,以及我们核心智能体循环的一点内容。我们将探讨 A2A 协议、eval。我们将探讨如何管理长上下文。

Original English

Gabe Dees Mesa: story of how this all kind of came to be. Uh we're going to talk about OGs uh big bet on effect uh a little bit into our core agent loop. Uh we're going to talk about the A2A protocol, eval. We're going to talk about how we manage long context.

Gabe Dees Mesa: 大家好,我叫 Gabe Dees Mesa。我是 OpenGV 的一名工程师,今天我们要讨论的是生产环境中的智能体,特别是 OpenGV 是如何构建和扩展 OG Assist 的。这次演示将会充满非常多干货。我们将探讨 AI 智能体。我们将探讨我们的 harness(运行工具)。我们将探讨 eval、observability(可观测性)和 traces(追踪)。我们将探讨 tools(工具)和 skills(技能)。这里会有非常多精彩内容。我们将向大家介绍我们在 OpenGV 做了什么,以及我们在生产环境中如何以现有的规模进行运营。所以你们将能看到一个带有 AI 智能体的真实用例和工作负载。话不多说,我们开始吧。

Original English

Gabe Dees Mesa: Hi everyone, my name is Gabe Dees Mesa. I'm an engineer here at OpenGV and today we're going to be talking about agents in production, specifically how OpenGV built and scaled OG Assist. Uh so um this presentation is going to be jam-packed with just so much good stuff. Uh we're going to talk about uh AI agents. We're going to talk about our harness. We're going to talk about um eval observability traces. We're going to talk about um tools and skills. Um it's there's going to be a lot of good stuff in here. We're going to talk to you guys about uh what we do at OpenGV and how we operate at the scale that uh we operate at um in production. So you'll be able to see a real use case and workload uh with AI agents. Um so without further ado, let's get started.

Gabe Dees Mesa: 好的,议程。我们将快速过一遍今天的高层面话题。我将向大家简单介绍一下 OG Assist 以及 Open Gov 是做什么的。我将告诉大家这一切是如何诞生的起源故事。我们将探讨 OG Assist 在 Effect 上的重大押注,并深入了解我们的核心智能体循环。我们将探讨 A2A 协议、eval。我们将探讨如何管理长上下文。我们将探讨监控和可观测性,如何收集反馈,以及如何对这些反馈进行迭代。最后,我们还将探讨工具和技能,以及在 Open Gov,我们如何不仅将 AI 作为向客户提供的外部服务,还在内部使用它来改善我们的开发工作流。

Original English

Gabe Dees Mesa: Okay, agenda. So just really quickly going to go through uh high level what we're going to talk about today. Uh I'm going to tell you guys a little bit about OG Assist and what uh Open Gov is. I'm going to tell you guys the origin story of how this all kind of came to be. Uh we're going to talk about OG Assist's uh big bet on effect uh a little bit into our core agent loop. Uh we're going to talk about the A2A protocol, eval. We're going to talk about how we manage long context. We're going to talk about um monitoring observability, how we collect feedback uh and how we iterate on that feedback. We're gonna lastly uh also talk about tools and skills and how at open gov uh we use um AI not only externally uh that we uh serve to customers but also internally to improve our development workflows.

Gabe Dees Mesa: 在继续之前,先简单介绍一下我自己。我叫 Gabe。我是 OpenGV 的一名软件工程师。我在 AI 智能体团队工作,而且我是参与构建 OG Assist 以及你们今天将看到的一些系统的研发人员之一。

Original English

Gabe Dees Mesa: Just a little bit about me before we go any further. My name is Gabe. I'm a software engineer here at OpenGV. I work on the AI agents team and uh I'm one of the folks that helped build uh OG Assist and some of the systems that you guys will be seeing today.

Gabe Dees Mesa: 简单介绍一下 OpenGV。OpenGV 是一家软件公司,其使命是赋能更高效、更负责任的政府。OpenGV 销售 ERP(企业资源规划)软件。包括预算编制、采购、资产管理和许可证办理等业务。我们大约是 14 年前成立的。很酷的一点是,我们有一个叫 OG Assist 的东西。OG Assist 是我们在所有产品导航栏顶部的一个小按钮。很酷的是,我们所有的产品套件和产品团队都构建了工具和技能来驱动这个按钮。比如,如果我点击这个按钮并打开 OG Assist,它会说:“嘿,我要问关于 rate codes(费率代码)的问题,这是非常针对 utility billing(公用事业计费)的”,这正是我当前所在的产品。你可以看到,在这个聊天界面里,我能和一个智能体对话,而智能体能够进行工具调用,以便在该套件内查询数据信息。所以,能够通过我们构建的这个叫做 OG Assist 的功能,第一方地创造出这样的体验,真的非常酷。

Original English

Gabe Dees Mesa: So, a little bit about OpenGV. OpenGV is a software company uh on a mission to power more effective and accountable government. Um so, OpenGV sells ERP software. That's things like budgeting, procurement, asset management, and permitting. And um we were founded about 14 years ago. And what's cool is um we have this thing called OG Assist. And OG Assist is this little button on the top of all of our products in the in the navigation bar. And what's cool is um all of our product suites and product teams um have built tools and skills in order to power this button. So, for example, if I open up uh this this um if I click this button and I open up OG Assist, it says, "Hey, um I'm going to ask about rate codes, which is very specific to utility billing, the current product that I'm in." And you can see that inside of this kind of chat interface, I'm able to speak to an agent, and the agent is able to make tool calls in order to um look up information against data inside of that suite. So, it's really cool um to be able to kind of first party create these experiences uh through the capability that we've built called OG Assist.

Gabe Dees Mesa: 好的,简单说一下这一切是如何诞生的故事。不久前,我们看到 AI 开始真正腾飞,于是我们的一位首席工程师(principal)组建了这个名为 AI 智能体的新团队,并邀请我加入。我立刻答应了,于是 OG Assist 开始成长。我们开始将 OG Assist 整合到我们所有的产品中,不仅整合到后端功能中,还包括前端功能。因此,你会看到我们赋予智能体的能力之一,就是它能够看到屏幕上的内容,并对其采取行动。所以,你可以看到我在这里问智能体:“嘿,屏幕上都有什么?你能帮我高亮显示一下我接下来可以采取的一些步骤吗?”你可以看到智能体在这里正在思考。它在想:“好吧,我有哪些可用的工具?嘿,让我去……”

Original English

Gabe Dees Mesa: Okay, so just a quick story about how this all came to be. So, um, a little while back, we we we saw that AI was really starting to take off and a principal, uh, spun up this new team called the AI agents team and asked me to join and, um, instantly I said yes and OG Assist started to to grow and we started to integrate, uh, OG Assist into all our products and, uh, not only our back-end capabilities, but also our front-end capabilities as well. So, you'll see that one of the capabilities that we give the agent is it's able to um see what's on the screen and and see and and and take action on what's on the page. So, you could see that um I'm asking the agent here, hey, hey, what's on the screen? Can you maybe highlight uh some of the next steps that I could take? So, you can see that the agent here is thinking. It's saying, okay, what tools do I have available to use? And hey, let me go and

Effect 代理循环与基础设施架构

Speaker A: 高亮显示一些你可以实际点击的元素,它随后会向你展示更多关于该内容的详细信息。所以,这仅仅是 OG Assist 所具备的另一项功能特性。在这里,我想分享一个简短的小故事,讲讲这一切功能和特性究竟是如何应运而生、最终形成的。这就要说到我们对 Effect 框架的巨大押注了。我之所以非常想把这一页幻灯片加进演讲中,是因为在我们代理(Agents)团队内部,我们下定决心做出了一个巨大的赌注——那就是把注押在 Effect 框架上。现在可以毫不夸张地说,这个决定已经为我们带来了极其丰厚的回报。在日常开发中,我们广泛使用 Effect 来编写代码。Effect 究竟是什么呢?它是一个专为 TypeScript 语言量身打造的高级开发库。它是完全开源的,其核心目的就是为了帮助开发者编写出质量更高、更健壮的 TypeScript 代码。要知道,这个库的内部已经内置了大量非常实用且强大的功能,例如,它拥有类似于 Zod 的 schema(模式定义与验证)功能——如果你曾经在项目中使用过 Zod,你就会很熟悉这种体验。除此之外,它还内置了针对错误处理(error handling)的完备机制,提供了用于系统日志记录(logging)以及执行追踪(traces)的工具。它的功能实在是太丰富了,包罗万象。它确确实实能帮助你写出更优秀的代码,让你能以更好的方式组织和构建代码的结构,同时还在系统架构的设计与实现层面提供了巨大的帮助,比如帮助我们更高效地启动和部署新的后台服务。对于我们代理团队而言,它最突出的贡献在于,它切实地帮助我们设计并构建了整个系统的核心功能——也就是核心代理循环(core agent loop)。因此,在接下来的整个演讲过程中,大家会看到我在内容中穿插了许多具体的案例,这些案例都将证明在我们的团队中采用 Effect 框架是如何获得丰厚回报的。总之,在 OpenGov 公司,我们真的非常喜爱并推崇 Effect,我们也非常鼓励其他团队或个人去尝试使用它。好了,话不多说,让我们继续往下看。

Original English

Speaker A: highlight something that you could actually click on and and tell you more about it. So just another capability of OG assist and just a little short story about how this all came to be. So the big bet on effect. Um so I really wanted to include this slide because um here on the agents team we made a huge bet to um to to bet on effect and suffice to say it's paid off in dividends. Um we write effect. So effect is this library for typescript. It's open source and it helps you write better um typescript code. uh you know it's got a lot of uh stuff baked in it like a sk a schema similar to like ZOD if you've ever used that. It's also got um things for error handling uh for logging for traces for uh it's just got so much in there. It really helps write better code and structure your code better and helps with architecture, spinning up new services for uh and and for us on the agents team really helping uh design and build the the core agent loop. So you'll see throughout this presentation sprinkled in um how effect on our team uh has paid off in dividends. So we we really love effect here at open gov and we encourage other folks to try it out and um yeah let's keep going

Speaker A: 接下来我们谈谈 Effect 原生代理循环(The Effect native loop)。在项目初期,我们最初使用的是 LangGraph 框架。在当时,它的表现还算不错,一直都能满足我们的需求,直到我们的团队规模真正开始大规模扩展,并且我们的业务用例开始发生显著的演变和复杂化。面对这些变化,我们最终决定进行技术栈的迁移,转而构建我们自己内部专属的、基于 Effect 原生的代理循环机制。这样做的根本目的,是为了让我们能够对这个代理循环拥有完全的控制权(full regency / agency)。这样一来,在未来如果我们遇到了极其复杂的业务用例,或者有特定的高级功能需要去构建和实现时,我们就能够深入底层进行定制开发,因为我们对代理循环的每一个环节都拥有100%的掌控权。不仅如此,现在的我们已经完全、彻底地拥抱并运行在 Effect 框架之上了。因此,你在使用 Effect 框架时所能享受到的所有那些酷炫且强大的特性,如今都已经顺理成章地传播和渗透到了我们整个代理循环的方方面面。比如先进的链路追踪(tracing)技术、能够有效避免并发问题的结构化并发(structured concurrency)模型,以及完善的日志记录(logging)系统等等。所有的一切都获得了更细粒度的控制能力,这也确实让我们能够充分释放系统的全部潜力。能够从零开始、自底向上地拥有并掌控自己的代理循环,对我们团队的发展至关重要。

Original English

Speaker A: the effect native loop. So originally we were on lang graph and that was fine until the team really started to scale uh and our use cases started to evolve. So we decided to move over to our own kind of effect native agent loop to have full regency over this uh agent loop such that if we have complex use cases or features that we need to build we could kind of get in we we had full control of the of the agent loop. And not only that but now we're fully on effects. So all the cool things you get with effect is now propagated throughout the entire agent loop like the tracing structured concurrency, the logging, everything is more fine graining control and it it really allows us to really unlock the full potential uh having our own agent loop from the ground up.

Speaker A: 我还想提到的另一件重要的事情是,在幻灯片的左侧,大家会看到一段代码示例。这实际上展示了我们当前正在使用的 Effect 循环的最基础形态。我们在项目中引入并使用了这个名为 Effect AI 的代码包。在这个特定的代码包里,包含了一些核心组件,比如一个用于处理聊天的组件(chat),以及一个用于对接语言模型的组件(language model)。通过这个聊天组件,你可以非常方便地实例化一个聊天对象,举个例子来说。在完成实例化之后,你就可以调用那种流式文本(stream text)函数,以此来向前端流式传输文本响应。在调用时,你可以直接向它传递一个提示词(prompt)。而且这里有一个非常酷的设计:在底层接入语言模型时,由于我们的架构是基于依赖注入(dependency injection)模式来实现的,这就意味着,如果我们未来需要热切换(hot swap)到另一个完全不同的语言模型(例如从 OpenAI 切换到 Google Gemini 等等),我们只需要在配置中注入一个不同的语言模型实例即可。所以说,真正意义上获得对我们自己专属代理循环的完全控制权,等同于赋予了我们所有的操作杠杆,它确确实实解锁了底层 AI 模型的全部能力上限。与此同时,对于我们整个研发团队而言,这也使得我们对这个核心的代理循环拥有了充分的代理权和自主性(agency)。

Original English

Speaker A: Um so another thing I wanted to mention is on the left side you'll see a code example. This is really the basics of the effect loop that we're using. Uh we're using this thing called the effect AI package. And in that package, there's this thing called um there's a chat and a language model. So with the chat, you can instantiate like an a chat for example. And then you could stream text using um that that kind of stream text function. You could pass in a prompt. And what's cool is uh with a language model under the hood of since we're kind of doing dependency injection, we could pass in a different language model if we were to uh hot swap to another one for example. So really just having full control of our own agent loop just kind of gives us all the levers and it really just unlocks the full capabilities of the model and uh for the team as well to have full agency over this this loop.

Speaker A: 除此之外,我想提及的另一个重点是代理到代理(Agent-to-Agent,简称 A2A)的通信协议。在我们的代理团队这里,我们在应用和实践这个协议的过程中取得了非常显著的成功。这个 A2A 协议其实是由 Google 创建的,它本质上是一种开放式协议,专门用于让不同的 AI 代理之间能够进行顺畅的相互通信。但在实际应用中,我们发现它在其他方面也非常有用,例如,它能够极其有效地帮助我们在后端系统中清晰地定义代理的路由规则,并且让我们自己的内部数据模型和 Schema 都能够严格遵循这种代理协议的标准。所以我们以此为基础进行了系统建模。给大家举个具体的例子,就像你们在幻灯片这里看到的,有一种被称为“代理卡片(Agent Card)”的结构,它里面结构化地包含了代理的名称、功能描述等各种核心信息。拥有这样一个严谨、规范的通信协议和规格说明(spec),切实地推动了我们的软件开发进程,并极大地促进了团队间的协作与对齐(alignment)。因为你知道,面对这样一份清晰的规范,我们所需要做的全部工作就是让我们的代码与这个规范保持高度一致,并严格遵循它。我们非常清楚,这份规范实际上就是前端和后端系统共同遵守的契约(contract),双方都将基于这个契约来消费(consume)和生成(produce)数据。所以,我想说这套机制对我们的工作真的是大有裨益。真正令人感到兴奋的是,A2A 协议本身具备极强的可扩展性,对吧?所以,你可以根据业务需求自由地扩展这个协议,比如在里面添加各种自定义的元数据(metadata)。另外,除了 A2A,还有 A2I 协议等等。总之,围绕着 A2A 协议有很多充满乐趣且极具价值的技术探索,这就是对我们团队行之有效的最佳实践。所以,今天特意把这些经验和大家一起分享。

Original English

Speaker A: Another thing I wanted to mention is the agentto agent protocol. So here on the agents team, we've had a lot of success with this protocol. So this protocol being the protocol that Google created um kind of an open protocol for agents to intercommunicate. But um we found this very useful for uh defining our agent routes for example in the back end and our model and our schema to follow this kind of uh agent protocol. So we modeled so for example there's this thing called an agent card which you see here and it's got the name of the agent a description etc right and having this kind of rigorous protocol this rigorous spec really helped drive our development and drive alignment because you know all we had to do was um align with this spec and follow this spec and we knew that this was kind of the contract that our front end and backend and would both consume and and produce. So, um this uh I would say also has been uh very helpful for us and and what's really cool is A2A has a lot of extensions, right? So, you could extend the protocol uh add in like metadata. Uh there's also A2I um so lots of fun stuff uh with A2A protocol, but uh this is kind of what's worked for us. So, just sharing that with with you folks.

Speaker A: 接下来是关于反馈收集与系统评估(Feedback and Eval)的环节。在软件工程领域,有一句非常经典的格言:“产品的交付(shipping)仅仅是整个生命周期的开始,而绝不是终点。”因此,在代理团队(agency team)内部,为了践行这一理念,我们构建了多种多样的方法和渠道来进行系统评估并收集用户的反馈。显而易见的是,你知道的,我们会有用户通过打电话进来,或者直接给我们发送电子邮件,或者通过其他方式直接联系我们,把他们遇到的问题或者想法告诉我们。但除此之外,我们获取反馈最主要、最核心的途径,是我们设计并引入的这套“点赞(thumbs up)与踩(thumbs down)”机制。在这个机制下,任何使用该系统的用户都能够直观地向我们表达他们的看法,比如他们会告诉我们:“嘿,这一次的处理表现得非常出色,这是一个极其优秀的回答”,又或者指出:“刚才那个回答真的很糟糕,没有解决我的问题”。我们会将这种明确的用户偏好信号收集起来,并以此为基准持续进行产品的迭代。我们可以将这些宝贵的反馈带回研发环节,从而在未来的系统更新中帮助改进代理的响应质量。与此同时,我们还建立了一套完善的自动化评估(automated evals)体系。在我们的持续集成(CI/RCI)环境中,我们会针对模型实际生成的补全内容(real completion)自动运行大量的评估测试。例如,我们可以针对特定的提示词(prompt)进行测试,检验它是否成功命中并调用了预期的工具?它有没有准确执行它本该去完成的任务逻辑?这种自动化测试框架同样在提升我们系统的准确性方面发挥了巨大的作用。所以,这些自动化评估手段与广泛收集的用户反馈双管齐下、相辅相成,切实有效地帮助我们不断改进我们的工具链、增强代理的技能(skills),并优化我们的测试工具集(harness)。这也就是我们为什么能够保持如此高频、快速且高效的迭代速度的根本秘诀所在。

Original English

Speaker A: feedback and eval. So here the quote is shipping is the start not the finish. So what we do here uh on the agency team is we have kind of multiple ways we do evals and collect feedback. Um obviously you know we'll have folks uh call in or or email us or or just let us know and tell us but the main way is we have this thumbs up and thumbs down mechanism. And here uh someone is able to tell us, hey, this this worked really well. This was a great response or that wasn't a great response. And that signal we take and we're able to iterate on uh and we could take it back and help improve uh you know the response in the future. Um we also have automated evals. So in in the in RCI we we have evals that run against real completion. So we could test the prompt against hey did it hit some tools? Did it do what it's supposed to do? And that also helps with our accuracy. So, uh those automated evals in conjunction with collecting feedback really help us um improve our our our tools, our skills, um our harness and and that's really how how we're able to iterate so fast and so quickly.

Speaker A: 人在环路(Humans in the loop,即人类介入系统决策过程)。这是一个我们在系统中精心构建的、非常酷且极具实用价值的功能特性。当系统识别出某个特定的工具调用需要获得批准时,我们会确定性地(deterministically)强制中断当前的代理执行循环。也就是说,如果一个自主代理试图去发起一个具有较高风险、且事先被配置为必须需要人类批准的工具调用时,系统就会暂停执行,并在前端向用户展示这样一个交互界面(UI)。随后,人类用户就可以在这个界面上点击“接受(accept)”或者“拒绝(reject)”按钮。通过这种显式地拒绝(explicitly rejecting)或者显式地接受(explicitly accepting)代理企图执行的操作的方式,我们不仅确保了我们在人类用户与 AI 系统之间建立起了坚实的信任纽带,同时也确保了系统运行的绝对安全。特别是当代理试图去执行某种会改变系统状态的危险操作(mutating operation,例如修改数据库、删除文件等)时,这种安全保障就显得尤为关键。我们的核心设计原则就是,永远、始终要确保人类牢牢地坐在“驾驶座(driver's seat)”上,保持对 AI 系统最终行动的绝对控制权。

Original English

Speaker A: Humans in the loop. So this is a really cool feature we built where we deterministically interrupt the agent loop. If there is a tool call approval required. So if an agent tries to make a tool call that it needs human approval for it'll show this UI and the human uh can click accept or reject. So explicitly rejecting or explicitly accepting uh the action that the agent is trying to make. And this ensures that uh you know we're building trust and also ensuring that uh you know we're being safe especially when the agent is trying to do a mutating operation and always always making sure that um humans are in the driver's seat

Speaker A: 沙箱安全机制(Sandboxing)。我们致力于研发的另一项关键技术,在很大程度上类似于我们刚才在前面那张关于安全性(safety)的幻灯片中所看到的理念。它的核心逻辑是:每当一个自主代理试图在系统中执行任何代码(execute code),或者试图在文件系统中创建新文件(create files)时,它的一切行为都必须、且只能在一个被严格隔离的沙箱(sandbox)环境中进行。因此,我们赋予了我们的代理一种安全边界……

Original English

Speaker A: sandboxing. So, another thing that we uh worked on um kind of similar to the safety slide we just saw was um whenever an agent tries to execute code or tries to create files, it does so in a sandbox. So, we gave our agent sort

Eureka 机器与自动化科学发现

Speaker B: 好了,大家请安静一下,我们准备开始了。大家好!今天能够站在这里,我真的感到非常兴奋和荣幸。这是一个规模宏大的会场,到目前为止,这也是一场体验极佳、非常酷的技术会议。今天,我想和大家深入探讨一件在我脑海中盘旋、萦绕了许多许多年的事情。实际上,这也是我第一次在公开场合谈论这个话题。这可以算作是我个人职业生涯中版本的“登陆火星计划(going to Mars)”,那就是我所构想的“尤里卡机器(Eureka machine,意为‘发现机器’)”。这是一台在未来终将为全人类发明出几乎所有未来新发明的终极机器。而我们之所以能够找到通往这一宏伟目标的路径,是通过暂时后退一步,去深入思考这样一个问题:在这个世界上,究竟还有什么是像我们所期望的那样,曾经为我们带来了大量真正令人难以置信的伟大发明?答案显而易见,那就是“进化(evolution)”。以及这种自然界的进化机制,是如何在冥冥之中引领着我们走向自动化研究(automating research),并最终推动人类科学的前沿不断向前发展的。这并不是我一个人的纸上谈兵,而是我与 recursive.com 的许多才华横溢、令人惊叹的同仁们,甚至还包括一些来自 AIX Ventures 的优秀伙伴们共同合作研究的结晶。其中你们将看到的一些幻灯片,实际上是直接受到了我在 Recursive 的联合创始人 Tim Rockel 的深刻启发,并且部分采纳和借用了他的核心观点与素材。

Original English

Speaker B: All right. All right. Hello everyone. Really excited to be here. It's a big room. Very uh very cool conference so far. Uh I want to talk to you today about something that's been on my mind for many many years. This is actually the first time I I talk about it. Sort of my version of going to Mars. Um and that is the Eureka machine. A machine that will eventually invent pretty much all future inventions for humanity. Uh and the way we're going to get there is uh by taking a step back and thinking about what else has given us a lot of really incredible inventions uh namely evolution and how that leads us to automating research and pushing the scientific frontier forward. And this is uh joint work with a lot of uh amazing folks uh at recursive.com uh and even some uh folks at AIX Ventures. And some of these slides are uh actually inspired by uh and taken uh partially from one of my co-founders at recursive Tim Rockel.

Speaker B: 那么,我为什么要在这样一个探讨 AI 技术的场合去谈论“进化”?为什么它在我们的语境中如此重要?我认为,从根本上来说,进化就是一个不受预设终点限制的、开放式(open-ended)的演化过程。正是这个伟大的过程,为我们带来了地球上大量丰富多彩、且我们非常喜爱和依赖的事物。这个过程最初发源于生物学领域,但现在它正在跨越学科的边界,向广泛的科学探索、技术创新领域延伸,并最终不可避免地走向了人工智能(AI)领域。我坚信,“进化”这一自然机制可以在很多截然不同的层面和维度上赋予我们灵感,启发我们去构建出远比现在更加优秀、更加智能的 AI 系统。事实上,每当我们回顾历史,总会想起那个在自然语言处理领域非常著名的段子:“每当我解雇一名语言学家,我的系统的准确率就会获得提升。”我认为,在早期的机器翻译(machine translation)发展时代,这句话绝对是真理。顺着这个逻辑推演下去,未来很可能也会出现这样的情况:我们也许应该解雇今天在场的所有人工智能工程师,让他们的工作职责发生转变——让他们主要去管理一个真正的“人工智能工程师”。而这个特殊的工程师本身就是由人工智能构成的,并且它的工作就是去开发和优化其他的人工智能。因此,这或许会成为本次演讲最终得出的重要结论之一。而且,我认为在座的大多数人非但不会因此感到恐慌,反而都会对此感到无比兴奋,因为这意味着我们所有人都将晋升成为这种超级人工智能的管理者,而不必再像现在这样,每天亲自动手去处理那些繁杂、枯燥的底层技术细节了。好了,既然如此,那就让我们从“进化”这个核心概念开始讲起吧?让我们把视线拉远,从一个非常宏大的视角来看:大约在三十五亿年前。这是一种令人叹为观止的奇妙过程,它引领着地球上的生命从最简单的细菌、原始的植物,一路进化出了鱼类、两栖动物等等,并最终发展到了我们今天所见到的生机勃勃的世界……

Original English

Speaker B: So uh why do I talk about evolution and why is it so important? Uh, I think basically evolution is this like open-ended process that has gotten us to a lot of different things that we really like. Uh, it started in biology. It's moving to science, technology, and eventually I. And I think it can inspire us in a lot of different ways to build better AI systems as well. In fact, uh, whenever we take out and there's this famous saying, whenever I fire a linguist, my accuracy goes up. Uh I think that's true for machine translation back in the day. And it may be true that we should fire all the AI engineers uh and that that are here uh and have them mostly manage an actual AI engineer that is AI and works on AI. Uh and so that may be uh one of the conclusions of this talk. Uh, and I think most of us are going to be excited about it because it means that we'll all become managers of such an AI rather than having to do the nitty-gritty ourselves. All right, so let's start with evolution, right? The really really big picture, three and a half billion years or so. Uh this is kind of the incredible process uh that has led from you know simple bacteria and plants and fish and amphibians and so on to

技术进化的历史与启示

Speaker: ……经过数十亿年的时间,才有了我们。对吧?这是一个很好的起点。这给了我们一些启示:进化过程可以创造出非常了不起的奇迹。对吧?但现在让我们拉近视角,也许缩小到几百万年的时间尺度。在那里,我们也能看到,在最早期原始的方式中,技术进化基本上如何增加了世界的某种产出,就货币价值而言。在开始时这有点难以估计,但我们可以看到这些指数级的序列,而大多数指数曲线最终会变成 S 型曲线,它们会趋于平缓。但人类在这方面做得相当不错,基本上发展了许多非常基础的技术,比如狩猎、农业,但也开始思考科学,在启蒙运动早期产生的科学方法,当然还有工业革命。

Original English

Speaker: ...after many billions of years us. Right? That's a good starting point. That gives us some indication that evolutionary processes can do pretty amazing things. Right? But now let's zoom in and uh go maybe down to a few million years. There we can also see how in the very first primitive ways technological evolution has basically increased the world uh sort of product uh in terms of monetary value. It's a little bit harder to estimate in the beginning, but we can see these sort of sequences of exponentials and most exponentials eventually become S-curves. They flatten out. But humanity has done pretty well by basically developing uh many of these very basic technologies, hunting, farming, but then also thinking about science, the scientific method um in the early days of the enlightenment and of course the industrial revolution.

Speaker: 所以现在我们可以进一步拉近视角。别担心,我们最终会谈到 AI、实际的自动化研究以及我们正在做的事情。这是一个非常非常快速的视角拉近。嗯,现在我们可以把时间缩短到过去的几千年。我们在那里看到的是,随着技术的进步,我们能够养活更多的人口,对吧?因此,当我们致力于推动这一前沿发展时,我们非常确信这会带来更多的人类繁荣,对吧?特别是在过去几百年里,我们看到由于技术及其带来的进化,人口数量出现了令人难以置信的爆炸性增长,而且在许多情况下,这种进化过程是由我们自己主导的。所以它在某种程度上是有意识的,但当我们在思考下一个周期中 AI 的进化时,我们可以从中汲取一些有趣的启发。

Original English

Speaker: So now we can zoom even further. Uh and no worries, we're eventually going to get to nanohat and actual auto research and what we're doing. Uh it's a very very quick zoom. Um and now we can zoom down to the last few thousands of years. And what we're seeing there is that with more technology, we were able to sustain more people, right? So when we're working on pushing that frontier forward, uh we're very certain that that will lead to more human flourishing, right? And especially in the last few uh hundred years, we're seeing this incredible explosion in the population of people because of technology and the evolution uh that it brings and in many cases that evolutionary process is run by us. So it's sort of conscious uh but there are sort of interesting uh inspirations that we can take from that as we're thinking about the evolution of AI in the next cycles.

技术作为永恒的增长源泉

Speaker: 事实上,尽管我可能不完全同意马克·安德里森(Marc Andreessen)的所有观点,但他非常聪明,我们在很多事情上都有共识。因此,我认为他写了一篇非常棒的技术乐观主义者宣言,在其中我认为他正确地指出,整个经济唯一的永久增长源泉是技术。很多人担心 AI 会抢走工作之类的事情,但事实是,它极有可能会大规模地促进经济增长,这将使我们许多人受益。所以,永恒的增长源泉就是技术。事实上,我们可以进一步说,没有任何物质问题——再次强调,不是心理问题之类的——而是没有任何物质问题是不能用更多技术来解决的。对吧?我们有饥饿问题;我们发明了绿色革命。黑暗——我们发明了照明;寒冷——我们发明了室内供暖;炎热——我们发明了空调,诸如此类。

Original English

Speaker: uh in fact and I might not agree with everything with Mark Andreasen but uh he is very smart and we agree on a lot of things. Uh and so I think he wrote this really great uh technoptimist manifesto in which he I think correctly points out that the only perpetual source of growth for the entire economy. A lot of people worry about AI taking jobs and things like that but the truth is it will very very likely increase uh the economy massively and that will benefit a lot of us. And so the perpetual source of growth is technology. Uh in fact we can go even further and say that there's no material problem and again it's not sort of psychological problems and things like that but no material problems uh that cannot be solved with even more technology. Right? We have a problem of starvation. We invented a green revolution, darkness, light, uh cold, indoor heating, heat, air conditioning and the list goes on.

Speaker: 所以我认为我们可以意识到,这个进化过程已经持续了很长一段时间,并且在继续取得巨大的进步。事实上,进步是如此之快,以至于在一个人的一生中就能发生重大的转变。对吧?如果你出生在1900年,那么当你三岁的时候,由于莱特兄弟的贡献,人类第一次能够进行持续的动力飞行。然后大约60年后,在1969年,人类一路飞到了月球。对吧?因此,在短短一生之中,人类从很长一段时间内除了顺着山坡滑翔之类外根本无法真正飞行,变成了我们全都能飞向月球,对吧?

Original English

Speaker: So I think we can kind of realize that this evolutionary process has been going on for a very long time and continues to make a huge amount of progress. In fact, the progress is so fast that there can within one lifetime be a major major shift. Right? If you're born in 1900, uh then three years when you're three years old, the first human ever was able to thanks to the Wright brothers kind of have sustained motored flight. And then about 60ish years later in 1969, humans flew all the way to the moon. Right? So that within one lifetime, humanity went from like no one can fly for a very long time other than sort of gliding down a hill or something. No one can really fly to we all fly to the moon, right?

Speaker: 所以对于我们来说,我认为这意味着,我也时常这么说,我们探索地球可能太晚了。我们探索星辰又太早了,但我们恰逢其时,可以去构建一个能够真正像飞行一样,在一个人的一生中因智能而带来改变的 AI。我们可以构建并推动 AI,从在我们所做的所有事情上都不如我们,发展到可能在我们所做的任何特定任务上都比我们更强,对吧?这大概就是我们这个时代的60年时间框架。而且因为现在一切发展得更快,这可能只需要30年左右的时间。

Original English

Speaker: And so for us, I think what that means is we're probably, and I sometimes say this, we're like too late to explore Earth. We're too early to explode the stars, but we're right on time to build an AI that could actually do what flying did for some in one lifetime due to intelligence. We can build and move from AI being worse at everything that we do to possibly being better at any specific task that we do, right? And that that will probably be our our 60-year time frame. And because everything moves faster, it might only be 30 years or so.

科学哲学与适者生存

Speaker: 那么,在技术、科学和理论之间就存在着一种有趣的联系,对吧?比如,有时应用先出现,然后我们再发展理论,进而改进技术;有时理论先出现,由此我们可以构建出新型的技术,因此思考一下科学哲学是非常有帮助的。在这里最好的启发莫过于波普尔(Popper)所写的,就像在其他类型的进化中一样,当我们选择一个理论时,我们选择的也是在与其他理论竞争中最优秀的那个。当然,你需要——如果你想让大语言模型(LLMs)做到这一点,它们需要能找到这些理论。例如你需要网络搜索。嗯,但在最能站得住脚的理论中,它就像进化一样,具有某种自然选择的过程,对吧?它能证明自己。在科学理论中也正在进行着一种适者生存的法则。

Original English

Speaker: So then uh there's an interesting connection between technology and science and theory right like sometimes the application comes first and then we develop the theory later and then improve the technology sometimes the theory comes first and from that we can build new kinds of technologies and so it's very helpful to think a little bit about the philosophy of science and no better uh to be inspired there than popper wrote that just like in other types of evolution when we choose a theory We also choose one that is best uh in competition with other theories. Of course, you need if you wanted LMS to do that, they need to find them. You need web search for instance. Um but uh in the theory that best holds its own uh it's one that just like evolution has a certain natural selection process, right? It proves itself. Uh and there is also a sort of survival of the fittest going on in scientific theories.

Speaker: 事实上,根据波普尔的观点,很多科学工作基本上就是我们提出一个新的理论、假设、解释或描述,然后将其置于严格的经验测试之下。这本质上就是科学理论的进化压力。这基本上是对开放式进化的历史做了一个非常简短的梳理,希望能让我们都意识到:更多的科学将带来更多的技术,进而带来更多的增长,最终带来更多的人类繁荣。所以这引发了一个问题:对我们来说,努力去扩大规模,并作为全人类投入大量资源来扩大科学发现的规模,以此来引领这种繁荣,这样做合理吗?

Original English

Speaker: And uh in fact uh a lot of science according to Popper is basically us proposing a new theory hypothesis or explanation or description and then subjecting it to rigorous empirical testing. That is the uh essentially evolution evolutionary pressure of scientific theories. And basically that was a very short uh run through uh sort of the history of open-ended evolution uh which hopefully makes us all realize that more science will lead to more technology which will lead to more growth which will lead to more human flourishing. And so that then begs the question does it make sense for us uh to try to just scale up and spend a lot of our resources as humanity to scale up scientific discovery in order to lead uh to this flourishing.

自动化科学发现与尤里卡机器

Speaker: 当你深入探讨这一点时,你会意识到——正如德·索拉·普赖斯(de Solla Price)很久以前就已经意识到的那样——科学的指数级增长实际上在某个时刻会因为从事相关工作的人手不足而停滞,对吧?现在在科学的各个不同领域有如此多细分的子领域,以至于很难让一百万人去研究某个特定的东西。因此,由于研究范围令人难以置信的扩大,他说专注于其中任何一个单一领域的人数反而减少了。这随后引导我们认真思考,我们如何才能将其自动化,实现科学发现的自动化。

Original English

Speaker: uh when when you double click into that you kind kind of realize um which Dislam uh already realized a long time ago uh that the exponential growth of science will actually be at some point halted by the lack of people working on it right there's so many niche subfields now in all the different areas of science that is very hard to get a million people to work on that particular thing uh and so as a result of this incredible widening of the scope he says uh the number of people focusing on any single section of it has decreased. And that then leads us to really thinking about how could we automate this and automate scientific discovery.

Speaker: 这进而引出了我所说的“尤里卡机器”(Eureka machine)。这基本上就是我们试图构建一台能够使科学发现过程自动化的机器的尝试。事实上,再过几个月我就会出版一本专门探讨这个确切想法的书。所以我将只给大家一个超级高层次的概述,介绍如何为从物理学、化学、生物学、神经科学、医学,到经济学、天体物理学等几乎所有领域构建这样一台尤里卡机器。

Original English

Speaker: And that then leads us to what I call the Eureka machine. This is basically uh our attempt at trying to build a machine that automates the process of scientific discoveries. And uh in fact I like in a couple months I'll have a book coming out on on this uh exact idea. Uh and so I'll just give you a super high level highlight of how such a Eureka machine could be built for basically everything from physics, chemistry, biology, neuroscience, medicine, uh economics, astrophysics and so on.

Speaker: 从本质上讲,这台机器有四个极其重要的支柱。第一,当然是你必须了解已经存在哪些知识。人类已经发明了哪些东西。你必须将所有的科学测量数据输入到这台机器中,作为第二大支柱。然后,对于你还无法测量的、我们还不了解的事物,你应该尝试去建立模拟。任何你可以模拟的东西,你都可以去验证,进而通过 AI 来解决。如果所有其他方法都失败了,或者在这些过程的最后阶段,你仍然需要有某种类似于物理的、工业化的实验室,能够真正在现实世界中进行真实的实验。

Original English

Speaker: And there are essentially four pillars that are all extremely important to this machine. One is of course you have to understand what knowledge is already out there. Uh what uh things humanity has already invented. uh you have to get all the scientific measurement uh data into as the second pillar this machine. Uh then for things that you cannot yet measure we don't yet know you should try to then build simulations. Anything you can simulate you can verify and you can then solve with AI. Uh and if all else fails or at the very end of these processes, you still need to have some kind of uh physical industrial like a lab uh that actually can run real experiments in the real world.

Speaker: 在所有这一切之上,你基本上会拥有一个智能体集群(agent swarm),它将处理所有这些不同来源的知识、数据、实验和奖励。而就知识的基础模型而言,当然我们也,你知道,这基本上是一个很好的例子,说明了我们迄今为止构建的每一项技术,特别是在 AI 领域,也包括在此之前的互联网、浏览器、GPU 等等,我们都可以重新思考。在将每一层技术重新构想为超级智能的基础设施方面,存在着很多创业的机会,对吧?例如在 You.com,我们就致力于为大语言模型(LLMs)、智能体等开发网络搜索功能。而这实际上是非常不同的,对吧?智能体可以阅读数千条非常长的片段,嗯,而不是仅仅通过10个蓝色链接配以非常简短的摘要。因此,你可以重新思考我们为其构建的这每一层不同的技术……

Original English

Speaker: And on top of all of this uh you'll have basically uh an agent swarm that will deal with all of these different sources of knowledge and data and experimentations and and rewards. Uh and in terms of you know the foundational model of knowledge of course we also you know it basically is a good example of how every single technology we've built so far especially in AI but also before that the internet browsers GPUs and so on we can rethink and there are a lot of startups possible in rethinking every single one of the layers of technology as infrastructure for super intelligence right at UW.com for instance we work on web search for LMS, right, and agents and so on. Uh, and that actually is quite different, right? Uh, agents can read thousands of very long snippets um, rather than just 10 blue links with like a very short snippet. And so you can rethink each of these different uh, layers of technology that we've built for

Recursive Self-Improvement in AI

Speaker A: 人们……为了将它们作为工具来使用,进而构建超级智能,而去为AI重建它们。这本质上也就是我们为什么想要构建超级智能,我们的目的就是为了实现科学的自动化。而且在我看来,那将是人类以及我们所知的技术领域里,下一次重大的阶跃式变革。

Original English

Speaker A: people uh, and uh, rebuild them for AI in order to use them as tools to then build uh, super intelligence. Now that is essentially uh the sort of why like like we want to build super intelligence in order to automate science. Uh and to me that will be the next big step function change uh in in humanity uh and technology as we know it.

Speaker A: 那么,我们究竟该如何构建它呢?我认为构建它的最好方法就是让它“自我构建”。对吧?作为一个领域,尤其是像我工作了多年的自然语言处理领域,我们已经发生了转变。我们从没有语言学家——这感觉像是远古时代,你知道的,公元前那样的历史,但在ChatGPT出现之前,我们经历了一个从让语言学家告诉我们一堆关于语言的知识,然后在此基础上训练统计模型的过程。而当我们允许神经网络通过词向量和其他神经网络架构,以及端到端的学习和反向传播,真正地自动学习这些特征时,我们基本上就能获得大得多的提升。在那之后,我们又做了大量关于架构的工程开发工作。

Original English

Speaker A: Now how do we actually build it? Uh I think the best way to build it is to have it built itself. Right? We moved as a field and especially natural language processing for instance which I've worked on for many years. We moved from not having linguists, this feels like ancient, you know, BC uh history, uh but before Chat GBT, um we we moved from having linguists tell us a bunch of things about language and then training statistical models on top of that. And when we allowed neural networks to actually automate learning those features with word vectors and uh other neural network architectures and backto-back uh end to-end learning and back propagation, we basically uh were able to get much bigger improvements. Uh then we did a bunch of architecture engineering.

Speaker A: 如今,至少有一大群人正在致力于开发一种统一的架构。不过,即便是那种统一架构,也依然包含了许多人工流程。因此,在人工智能领域里一遍又一遍地证明了一个清晰的规律:当我们剔除一个人工流程,并用一个学习系统来取代它时,提升自然就会随之而来。这也是为什么我认为,我们应该尝试通过让一个递归自我改进(RSI)系统来实现自我构建,从而打造出一台能够说话的机器。而且最美妙的地方在于,直到现在,人工智能才真正能够做到这一点,因为人工智能就是代码,而人工智能也能够编写代码。

Original English

Speaker A: Now a bunch of people at least are working on a unified architecture. Uh but even that unified architecture has a lot of manual processes. And so it's clear over and over again in AI that when we take out a manual process and we replace it with a learned system, improvements will follow. Uh and so that's why I think uh we should try to build a speaker machine by having an RSI that builds itself. And the beauty is that only now um AI can actually do this because AI is code and AI can code.

Speaker A: 这种能够在越来越长的时间跨度内真正进行编码的能力,其实仅仅是在过去大约六到八个月里才出现的,而这也使得这样一个递归自我改进(RSI)系统现在能够对自身进行操作,甚至发展出一种对自身缺陷的某种自我意识,然后再去修复这些缺陷。紧接着,一旦我们拥有了这台在自身AI研究上变得极其优秀、非常出色的机器,我们就可以利用它去为许多其他事物,在其他科学领域里进行AI研究。所以在宏观层面上来看,这其实相当简单,对吧?我们有三个步骤:构思、实施以及对想法的验证。这基本上适用于几乎所有的科学领域。

Original English

Speaker A: Now this this ability to really code in longer and longer time horizons has really only happened in the last like six to eight months and that now enables such an RSI to work on itself to develop almost a certain sense of self-awareness of its own shortcomings and then fix those shortcomings. Uh and then once we have that machine that has gotten really really good at doing research in AI itself, we can then use it to do AI research for a lot of other things uh in in other scientific fields. And so at a high level it's quite easy right we have three steps ideation implementation and validation of ideas. That's true for basically almost every scientific field.

Speaker A: 因此,为了以一些非常具体的例子来作为结尾,我们已经构建了这种“尤里卡机器”(Eureka machine)的第一个版本,并且我们只是想展示它在许多人都知道且熟悉的一些小样本上是行之有效的。所以我们基本上是从三件事情开始的,这些事情能向你展示并让你初步一瞥这种机器能做些什么,同时也作为一些简单的证据。而这三件事基本上就是:更好的训练、更快的训练,以及针对Nvidia GPU的更好的内核。

Original English

Speaker A: And so uh to end maybe on some very specific examples uh we have built this first kind of version of such a Eureka machine uh and we wanted to just show that it works on some small uh samples that a lot of people know and are aware of. And so we basically started uh with three things that show you and give you a very first glimpse of and and sort of simple proof points uh of what such a machinery can do. And that was basically better training, faster training and and better kernels uh for for Nvidia GPUs.

Speaker A: 首先是第一件事,nano chat(微型聊天模型),我确信你们中的许多人都听说过它。很多人认为这已经是递归自我改进(RSI)了,而在某种意义上,它确实算是一种弱形式的RSI,因为通常情况下当你进行自动化研究时,它并不能算作递归自我改进,对吧?真正的递归自我改进是当你拥有一个人工智能,它对自己的缺点有一种自我意识感,能够完全访问它武器库中的所有东西——从预训练到强化学习(RL)训练、测试工具(harnesses)以及所有的一切,然后在它的下一个版本中真正地去更新那整个系统。

Original English

Speaker A: Um the first one nano chat um I'm sure many of you have heard of it. A lot of people think that's already recursive self-improvement and it is kind of a weak form in the sense that usually when you do auto research it's it's not recursive self-improvement, right? True recursive self-improvement is when you have an AI that has a sense of self-awareness of its own shortcomings, full access over everything uh in its arsenal from pre-training to RL training and harnesses and everything and then actually updates that entire system in the next version of itself.

Speaker A: 当然,你也可以把这样一个系统拿过来,只是让它去改进某个其他的流程、某个其他的AI,比如进行一次小型的nanoad运行,在那里你可以在五分钟内训练出某个东西,那真的是非常令人兴奋的。这是一个重要的里程碑,但那还不是实际意义上的递归自我改进(RSI)。所以在这里我们基本上展示了这样一个自动化研究系统的三个例子,以及它能够做些什么,并且在非常非常短的时间之后,它本质上就能够超越许多不同的团队,甚至超越那些同样在使用其他AI研究成果的团队。

Original English

Speaker A: Now you can also take such a system and just ask it to improve some other process some other AI like a small nanoad run where you can train something in five minutes and that is really exciting. It's an important milestone but it's not actual RSI. So here basically showed three examples of such an auto research um uh system and what it can do and uh after a very very short time it essentially was able to outperform many uh different teams and teams that also use uh other AI research.

Speaker A: 让我们深入了解一下其中的一些例子。Nanohat就是一个非常令人兴奋的例子。基本上,你会在不到五分钟的时间里训练一个非常小的聊天模型,而你本质上希望它能达到尽可能最佳的bits per byte(每字节比特数)数值。整个社区在这个问题上已经努力了相当长的一段时间,并且达到了0.93。而在对这个系统训练了仅仅一天多或两天的时间之后,我们基本上就把它降到了0.91。这相当令人兴奋。

Original English

Speaker A: So let's double click into some of these. Nanohat is really exciting example. Uh basically you train a very small uh chat model uh in less than uh five minutes and you basically want to have it get to the best possible bits per bite uh number. And so the whole community had worked on this uh for uh quite some time and got to uh 0.93. And after training this for a little more than a day or two, uh, we basically got it down to 0.91. Um, which is pretty exciting.

Speaker A: 不过,如果它所做的一切仅仅只是找到几个超参数,然后非常仔细地对它们进行微调的话,那也不会这么令人兴奋了,但它确实找到了真正有趣的新颖想法,比如哈希二元语法(hash bigrams)、三元语法嵌入(trigram embeddings)以及相关的表格,并且通过各种学习到的门控机制(learned gates),将这些混合到意图(intention)的各种价值路径(value paths)中。所以,它实际上开始做越来越多有趣的事情,而不是仅仅只是进行一种超参数的调整。

Original English

Speaker A: Now, it wouldn't be that exciting if all it did was just find a couple of hyperparameters um and tune them carefully, but it actually did find truly interesting novel ideas like hash biograms and triam embeddings and tables for those uh and mixing that into various uh value paths of uh the intention through variety of learned gates. So, it actually started to doing more and more interesting things rather than just kind of tuning hyperparameters.

Speaker A: 另外一个是关于nano GPT的速度测试(speedrun)。显然,速度是非常重要的。所以在这里,我们能够再次致力于这项工作,应用这个系统,在非常短的一段时间后,它的表现甚至超过了那些在这个特定的基准测试上、经常与AI一起工作了一年多的人们,它让整个过程又快了两秒,快了超过两秒钟,达到了70秒,而且在这个过程中它再次发现了一些非常有趣的想法。

Original English

Speaker A: Um another one a nano GBPD speedrun. Uh obviously speed is very important. Uh so here we're able to work on this again, apply the system and after a very short amount of time it got better than uh people working often together with the AI for over a year uh on on this very on this benchmark and made the whole thing another two seconds over two seconds faster um at 70 seconds and again discovering uh very interesting ideas in the process.

Speaker A: 然后第三个是关于CUDA内核(cuda kernels)。当然,我们都关心不要太快地把我们的GPU预算给烧光。并且努力追求极高的效率。我认为总的来说,当你看到很多混合专家模型(MoE)依然在耗资数十亿美元的超大型集群上运行得那么低效,然后利用率却只有30%左右时,这实际上有点令人震惊。目前世界上有大量的工作正在进行,旨在改善这一状况。涉及不同的领域、不同的群体,或者是这项工作的各个不同阶段。

Original English

Speaker A: And then the third one is scuda kernels. Of course, we all care about not burning through our GPU budgets too quickly. Um uh and trying to be very efficient. I think in general, it's actually kind of shocking how inefficient a lot of mixture of expert models still are run in very large clusters that cost billions of dollars and then only have like 30% or so utilization. There's a lot of work that's ongoing in the world uh to improve that. and different fields uh or different groups of people or various different um yeah stages of that.

Speaker A: 不过长话短说,在训练和测试期间我们使用了许多不同类型的CUDA内核,而在这里,我们基本上再次拿出了那个系统,在几天之后,它发现了比NVIDIA基准测试网站排行榜上最佳表现还要好的内核,并且在所有不同类别的这些内核中,它再次以相当大的优势胜出。

Original English

Speaker A: Uh but long story short um lots of different cuda kernels are used during training and testing and here um we basically again took that system and after uh a couple days it discovered better kernels uh than the leaderboard's best uh on the NVIDIA uh benchmark website u by again quite quite a sizable margin across all the different uh categories of of those kernels.

Speaker A: 尽管我们在AI方面相当优秀,并且就像我们在团队中实际上并没有那些将整个职业生涯都致力于编写优秀内核的特定的CUDA内核专家一样,但是你知道,我们还是做了足够的工作,并且我们与Nvidia合作,以确保这里没有出现奖励黑客攻击(reward hacks)和其他问题,但实际上我们发现,最终这些都通过了验证,而且几乎所有这些不同的内核都在那里找到了最佳的解决方案。所以,基于这些,我希望能说服你们,递归自我改进(RSI)确实可能成为下一条重大的S型曲线——一种在之前的指数级增长之上层层叠加的指数级增长,而这也应该能在未来帮助我们,不仅仅是在AI领域,最终也包括在科学以及所有的技术领域,从而让地球上的更多人能够繁荣发展。

Original English

Speaker A: And while we are pretty good at AI and like we actually in the team didn't have any particular CUDA kernel experts who just spent their entire careers writing good kernels. uh but still you know we do just enough to make sure and worked together with Nvidia to make sure that there are no reward hacks here and and other issues but actually found uh that eventually these all checked out and were indeed uh pretty much all the different kernels uh found the best solutions there and so with that I hope I could convince you uh that indeed RSI could be that next big uh scurve um an exponential that gets layered uh on on top of previous exponentials and uh that should help us uh with not just AI but eventually science and then all of technology and then uh allowing many more people uh to flourish on our planet.

Speaker A: 所以,也许我想在这里以这样一句话作为结束,那就是,许多人都在想AI还能走多远,对吧?因为每一个指数级增长最终都会趋于平缓。其实很难知道,甚至当我们谈论AI的指数级增长时,这到底意味着什么。这里存在着许多不同的——我称之为“智能空间”(spaces of intelligence)的东西,我们没有时间去逐一深入探讨所有这些,但只要你真正试图去定义这10个空间中每一个的多个不同维度——正是它们构成了这种复杂的、像是一种多维体积状的、被称作“智能”的东西时。

Original English

Speaker A: Uh and so maybe I'll end on this note here which is uh a lot of people wonder how much longer AI can go right every exponential eventually flattens out and um it's actually quite hard to know like when we even talk about exponential growth in AI what does that even mean there are many different I call them spaces of intelligence and we won't have time to go into all of all of these but as soon as you actually try to define multiple different dimensions of each of these 10 spaces uh that make up this complex like sort of volutric uh thing that is intelligence.

Speaker A: 你就会意识到,在智能的上限方面,我们仍然有很长很长的路要走。在我们想要达到这些上限的道路上,几乎在每一个维度以及由它们构成的每一个空间里,我们距离目标都还有着天文数字般的遥远距离。所以,如果你们对这其中的任何内容感兴趣,并且想要帮助我们共同构建它,我们非常乐意倾听你的声音。谢谢大家。

Original English

Speaker A: You'll realize that there's still so much more to go like on the upper bounds of intelligence. We're still astronomically far away from reaching those uh across pretty much every single one of uh these dimensions and the spaces uh that they make up. Uh so if any of that is interesting uh and you want to help us build that um we'd love to hear from you. Thank you.

Production Eval for Authentic Systems

Nishan Gupta: 大家好,我的名字是Nishan Gupta,我是Meta的一名软件工程技术主管,负责为Meta的“超级切线实验室”(super tangent lab)及其基础设施组织构建训练和推理基础设施。今天,我们将要讨论的是针对认证系统(authentic systems)的生产评估(production val)。当大多数人听到“评估”(valuation)这个词时,他们往往会想到基准测试。比如,一个模型在某项基准测试中获得了90%的得分。一个新的版本……

Original English

Nishan Gupta: Hey everyone, my name is Nishan Gupta and I'm a software engineing tech lead at Meta working on building the training and inference infrastructure for the meta super tangent lab and their infrastructure organization. Today we're going to be talking about production val for authentic systems. When most people hear the word valuation, they think about benchmarks. A model scores 90% on a benchmark. A new version

Evaluating Workflows Instead of Outputs

Speaker A: scores 92%. The team celebrates.

Original English

Speaker A: scores 92%. The team celebrates.

Speaker A: 但智能体系统从根本上改变了评估的意义。如今,系统不再只是简单地生成答案。它们会做计划,调用工具,检索信息。它们会执行工作流。它们会与生产环境的基础设施进行交互。问题不再是模型是否生成了正确的答案。问题是系统是否表现正确。今天我想探讨一下评估是如何从模型基准测试演变为生产环境基础设施的。

Original English

Speaker A: But agent systems have fundamentally changed what the evaluation means. Today the systems don't simply generate answers. They plan, they call tools, they retrieve information. They execute workflows. They interact with the production infrastructure. The question is no longer did the model generate the right answer. The question is did the system behave correctly. Today I would like to discuss how evaluation is evolving from model benchmarking into production infrastructure.

Speaker A: 这是如今几乎每个 AI 组织都遇到的问题。离线基准测试分数在不断提高,但生产环境的可靠性往往仍然不可预测。为什么会这样?因为基准测试衡量的是模型的能力。而生产环境衡量的是系统行为。基准测试无法捕捉工具故障、API 宕机、上下文变化、用户差异以及长时间运行的工作流。随着系统变得越来越自主,基准测试性能与生产环境性能之间的差距也在不断扩大。结果就是许多团队今天所经历的情况。正如你所看到的,基准测试分数很高,但生产环境的表现却不可靠。

Original English

Speaker A: This is the problem almost every AI organization is encountering today. Offline benchmarks continue improving yet production reliability often remains unpredictable. Why is that? Because benchmarks measure model capability. Production measures system behavior. A benchmark doesn't capture tool failure, API outage, context changes, user variability, long-running workflows. And as systems become more autonomous, the gap between the benchmark performance and production performance grows. The result is what many teams experience today. High benchmark scores as you can see, but unreliable production behavior.

Speaker A: 传统的 ALM 评估侧重于输出。但我们应该问这样一个问题,模型是否给出了正确的答案?智能体系统迫使我们问一个不同的问题。系统表现是否正确?行为包括规划质量、工具使用、执行情况、工作流执行、从故障中恢复的能力以及决策制定。换句话说,我们正在从评估答案转向评估工作流。这需要截然不同的评估架构。

Original English

Speaker A: Traditional ALM evaluation focus on outputs. But we should ask the question, did the model produce a correct answer? Agentic systems force us to ask a different question. Did the system behave correctly? Behavior includes planning quality, tool usage, execution, workflow execution, recovery from failures, decision making. In other words, we are moving from evaluating answers to evaluating workflows. And that requires fundamentally different evaluation architectures.

Speaker A: 许多团队仍然认为幻觉是主要的 AI 故障模式。在生产环境中,幻觉通常只是其中一类。智能体系统引入了整个层级的故障模式。在最基础的层面,有记忆故障、检索故障、安全故障。往上走,你需要考虑推理错误、糟糕的规划、不正确的执行。在最高层,你必须考虑多智能体协同故障。这就是为什么仅仅评估模型输出会遗漏我们观察到的最大的生产风险。

Original English

Speaker A: Many teams still think hallucinations are the primary AI failure modes. In production, they are often just one category. Agentic systems introduce an entire hierarchy of failure modes. At the very foundation, the memory failures, retrievable failures, safety failures. As you go up, you have to think about reasoning mistakes, poor planning, incorrect execution. At the highest layer, you have to think about multi-agent coordination failures. And this is why evaluating only model output misses the most production risk we observe.

Speaker A: 最有用的思维转变之一就是停止像研究人员那样思考,开始像 SRE (站点可靠性工程师) 或生产工程师那样思考。SRE 不使用准确率来衡量成功。他们衡量可靠性、可用性、延迟、成本回收,而智能体系统需要同样的方法。目标不是最大化基准测试分数。目标是最大化可靠的结果。Rabi 成了北极星指标,虽然其价值是有限的。在中间层是基于场景的评估。这些会模拟现实的工作流。在最顶层,你看到的是生产环境遥测。这里提供了最高价值的评估信号。令人惊讶的见解是,大部分评估数据往往来自真实用户与真实系统的交互。

Original English

Speaker A: One of the most useful mindset shifts is to stop thinking like researchers and start thinking like a SR or a production engineer. SR don't measure success using accuracy. They measure reliability, availability, latency, cost recovery and agentic systems require the same approach. The goal is not maximizing the benchmark scores. The goal is to maximize dependable outcomes. Rabi becomes the northstar metric values limited. In the middle there scenario based valuations. These simulate realistic workflows. And at the very top you see production telemetry. This is where the highest value evaluation signals come from. The surprising insight is that the most evaluation data often comes from real users interacting with real systems.

Speaker A: 现在我们来谈谈离线评估。离线评估依然重要,但方法论变了。我们不再评估 prompt,而是评估场景。例如,客户支持工作流、代码生成工作流、研究工作流。智能体在那个模拟环境中运行。我们衡量任务完成率、工具正确性、规划质量、资源使用情况,在规模庞大时资源消耗会呈指数级增长。关键要点 18 是,评估应该是场景驱动,而不是 prompt 驱动。

Original English

Speaker A: Now let's talk about offline. So offline evaluations still matters but the methodology changes. Instead of evaluating prompts we evaluate scenarios. For example, a customer support workflow, a code generation workflow, a research workflow. The agent operates inside that simulated environment. We measure the task completion rate, tool correctness, planning quality, resource usage which is which becomes exponentially high at high scale. The key takeaway 18 evaluation should be scenario driven not prompt driven.

Speaker A: 一旦系统进入生产环境,每一次交互都变成了一个信号。这是评估思维上最大的转变之一。

Original English

Speaker A: Once a system reaches production, every interaction becomes a signal. This is one of the biggest shifts in evaluation thinking.

Test Time Compute for Search

Hio: 哦,好的。嗯,好的。那么,大家能看到幻灯片吗?哦,很好。好的。那么,大家早上好。非常感谢大家来到这里。我的名字是 Hio。我大概从 2020 年到 2025 年创立了 Gina AI,去年十月我们被 Elastic 收购了。所以,现在我在那里带领一个模型推理和训练团队。

Original English

Hio: Oh, all right. Uh, all right. So, can everyone see the uh slides? Oh, nice. All right. So, good morning everyone. Thanks so much for being here. Uh, my name is Hio. I founded around Gina AI since uh 2020 to 2025 and last October we were acquired by Elastic. So, now I'm running a model inference and training team there.

Hio: 今天我想解答这样一个问题。大型模型通过在推理阶段花更多时间思考,表现会变得更好。对吧?我们称之为“测试时计算” (test time compute)。那么,小型的检索模型能做到同样的事情吗?它能在不增加模型大小的情况下,通过在推理时更努力地思考来变得更好吗?为了找出答案,我让智能体通宵跑了自动化研究,结果发现答案并不是简单的“是”或“否”。对吧?那么让我来给你们看看我的发现。

Original English

Hio: And uh uh so here's a question I want to answer today. Uh so big models get thinking better by at inference time. Right? So we call that test time compute. And can small retrieve model do the same thing. Right? Can it get better by thinking harder at inference uh without making the model any bigger? Uh to find out that I let the agent run auto research overnight and the answer turned out to be more interesting than yes or no. Right? So let me show you what I found out.

Hio: 首先让我解释一下什么是测试时计算。这个概念很简单。与其训练一个更大的模型,不如在推理时投入更多的计算量。这样你就能得到更好的答案。它通常以我们非常熟悉的形式出现,比如 best-of-N 采样 (best of insampling)、自洽性 (self-consistency) 或对候选结果重新排序的验证器 (verifiers)。OpenAI 的 Noam Brown 给出了一个具体的数据。他发现,一个德州扑克机器人在多思考 20 秒后获得的性能提升,等同于将模型规模扩大 10 万倍。所以这就是测试时计算的前景。那么我们这里要问的真正问题是,这种前景是否也同样适用于搜索领域。

Original English

Hio: So first let me say what test time compute is. So the idea is very simple. So instead of training a bigger model, you spend more compute at inference time. So you get better answer back. Uh it shows up in a very familiar forms uh such as a best of insampling, self-consistency or verifiers that rerank the candidates. So non Brown from OpenAI uh put a number on this. He found that a poker bot uh sinking for 20 seconds uh got the same boost at scaling the model for 100,000 times. Uh so that's the promise of test time compute. So the real question for us here is does this promise also for the also hold for search.

Hio: 这里提供一个新的视角,把它变成一个关于检索的讨论。搜索其实已经在使用测试时计算了。想一想你构建搜索系统时做的事情。你拿一个训练好的嵌入模型 (embeddings)、一个训练好的重排器 (reranker)、某种多因素检索器,以及一个查询扩展器 (query expander),然后把它们串联成一个流水线。所以你是在花费推理算力来购买相关性,而不是去追求更大的模型。你基本上是在测试时组装更多的搜索组件。所以真正的问题不在于你的模型是否足够大,而在于你能在推理时组装多长的流水线,以及这样做是否物有所值。

Original English

Hio: So here's a reframe that turns this into a retrieval talk. Uh search is already test time compute. Uh so think about what you do when you build search. You take a train embeddings a train reanker some multiffactor retriever and a query expender and then you wire them into a pipeline. So you are spending inference to buy relevance and you are not reaching for bigger model. You're basically assembling more search at test time. So the real question isn't whether your model is big enough. It is how much pipeline can you assemble uh at inference and whether that pays off.

Hio: 构建这种流水线有两种方式,我都会给你们展示。第一种版本 (version A) 是我会深入讲解的。在这里,智能体在单一的冻结嵌入器或编码器之上编写一个小程序。它可能会对文档进行分块,使用不同的评分策略进行分数融合,然后把结果反馈回来。所以你可以把它看作是基于向量嵌入的多轮代数运算。

Original English

Hio: So there are two versions uh two ways to build that pipeline and I will show you both. The version one the first one version a uh is the one I will go deep on. So here an agent writes a little program over a single frozen embeder or encoder. It might chunk the document uh do this scoring fuse uh with different scoring strategy and feeds the results back. So think of think of it as a multipass uh algebra over embeddings.

Hio: 第二种版本 (version B) 我稍后再讲。在那个版本中,一个小智能体给定固定的 token 预算,将语料库上的检索工具(如 grep、embed、rerank)串联起来。所以这是在两个不同层面上实现的同一个理念。

Original English

Hio: The second one version B uh I will come to later. So there a small agent wires up the retrieval tools like grap embed rerank over a corpus given a fixed uh token budget. So it's the same idea implemented at two different levels.

Hio: 我们先从版本 A 开始。版本 A 运行在一个小型的冻结编码器上。普遍的看法是,小型模型无法在这方面有所提升,测试时计算完全是大规模推理模型的专利。但让我们看看今天这些嵌入模型,比如 E5、Mistral、Qwen3、embed embedding Gemma,甚至我们自己的 Jina embedding E5,它们都是从大型语言模型骨干中蒸馏出来的。这是今天主流的做法,如果测试时计算存在于 LLM 表征空间中,那么这种蒸馏模型应该会以某种方式继承它,还是说它们并没有?这正是我想要找出答案的问题。

Original English

Hio: So let's start with version A. So version A runs uh runs over a small frozen encoder. So there the common belief is that small models cannot improve there and test time compute exclusively belongs to the big reasoning models. But let's look at what today's embedded come from models such as E5 uh Mistro uh Queen3 uh embed embedding Gemma and even our own genome embedding E5 they all distill from the large language model backbones so that's the dominant recipe today and if test compute leave in the ARM representation space then this detailed model should somehow inherit it or do they so that's exactly the question I want to find

Hio: 那么,关于一个冻结的模型、冻结的嵌入器如何在测试时得到提升,这里的直觉是这样的。让我们看看这三个面板。我们从右边——不对,从右到左来看。或者说我们从最左边最简单的匹配打分方式,讲到最右边最细致的方式。在左边,你有一个简单的余弦距离 (cosine distance) 计算,这基本上就是每个文档一个向量,每个查询一个向量。这是一个冻结的余弦基线 (frozen cosine baseline)。在右边,你有这种 ColBERT 风格的后期交互 (late interaction),在这里,每一个查询 token 都会与每一个文档 token 进行匹配。所以我们可以认为 ColBERT 是测试时计算的一个极端案例。当然,最有趣的部分是在中间这个我用蓝色框出的面板。你可以使用同一个冻结的编码器,把文档拆分成句子,然后对它们取最大值 (max pooling)。这基本上就是我所说的测试时计算。你更接近后期交互了,但完全没有引入新模型。只是在同一个嵌入模型上反复做更多的工作。

Original English

Hio: So here's the intuition of uh for how a frozen model, a frozen embedder could improve at test time. Uh let let's look at the three panels. Uh let's go from the right uh left right to the left. Uh so we go from the simplest way to score a match on the left and to the most detailed way on the right. So on the left you have a single cosign distance which is basically one vector per document and one per query. So that's a frozen cosine baseline. On the right you have this cobert style latent interaction where every query token is matched against every document token. So one can consider cobert as an extreme case of test time compute. Uh the interesting part is of course is it is in the middle panel where I have outlined in blue. So you can take a frozen uh you can take the fro same frozen encoder split the document into sentences and max over them. So that's basically what I call the test time compute. You get closer to late interaction but without adding new model at all. Just more work on the same embedding model again and again.

Hio: 让我把这个问题设定得非常严格。一个冻结的单向量嵌入模型,仅仅在测试时能提升多少?我说的“严格”,意思是只有一个位于 API 后面的冻结编码器,你想调用多少次都可以,但不允许重新训练,不允许使用第二个模型,不允许有任何学习参数。流行的方法全都打破了这些规则之一,比如 HyDE 会将错误放入查询路径以路由查询。GQR 作为第二检索器,而 Meta embed 会训练新参数。所以我们禁止这所有三个规则。我们禁止所有这三件事。但即使在这些限制下,搜索流水线……

Original English

Hio: So let me make the question very strict. So how much can a frozen single vector embedding model improve at test time alone? So I and I do mean by strict just one frozen encoder behind an API and you can call it as many times as you want but no retraining no second model no learned parameters. So the popular method uh measured all break one of those rules like height puts an error in the query pass to route the query. GQR as a second retriever and meta embed trains new parameters. So we forbid all these three rules. We forbid all these three things. But even with the constraint the search pipeline the

自动研究循环与架构

Speaker: 搜索空间巨大。那么你如何搜索呢?当然是通过自动研究。所以,与其由我手工编写这些程序,不如让智能体自己运行研究循环。它修改一个文件,运行一个简短的实验,如果指标提高,就保留修改;否则,就撤销。它整夜反复进行这个过程。这有点像爬山算法,只不过是以大型语言模型作为变异函数。来自 Anthropic 的相关人员曾这样描述过这个过程:你并不是以研究员修改 Python 文件的方式来编辑它,而是在编写 Markdown 文件来设置自主研究,然后该循环会生成我们将要看到的一切。

Original English

Speaker: Search space is huge. So how do you search that with auto research of course? So instead of me handcrafting this programs an agent runs the research loop by itself. Uh it changes one file it runs a short experiment and if it matrix improve it keeps a change otherwise it reverse. So it does that over and over all night. So it is kind of like hill climbing uh but errorm as a mutation function. So entry capacity from astrobic uh describe it as follows. So we are editing a python file in the way uh you're not editing a python file in the way that research researcher would would. So you are writing a markdown files that set up the autonomous research or and that loop generate everything that we were about to see.

研究循环的四个组成部分

Speaker: 这是一张展示整个循环的图。只需从左到右跟随方框即可。我们有一个提议者(Proposer),它是一个大语言模型智能体,在冻结的编码器上编写程序。我们有一个评估者(Evaluator),它对该程序进行评分,还有内存记录结果,而最右边的黑盒注册表(Registry)会收集所有结果。每一代有 144 个程序。现在看这条返回下方的虚线箭头,它基本上就是反馈机制。内存状态影响下一个程序,并且每次运行都建立在上一次运行的基础之上。

Original English

Speaker: Uh so here's a whole loop in one picture. uh just follow the box from left to right. We have a proposer which is a RM agent write a program over the frozen encoder. We have evaluator uh which scores that program and memory logs the result and the registry the black box on the far right uh collects all of them. So 144 programs one per generation. So now see the dash line uh dash arrow looping back underneath that's basically the feedback. So memory conditions next programs and every runs built on the last one.

Speaker: 让我快速过一下这四个部分。首先是提议者(Proposer),它基于 Opus 4.6 模型,纯粹用作变异函数。它读取当前最好的程序和内存文件,然后添加一个 Python 文件来提出下一个程序。所以在这个循环中没有人类参与。但这里有个关键:它只会优化你给它的指标,而不是你本来意图要优化的指标。因此,如果你奖励领域内的性能提升,或者奖励消耗更多的计算资源,那么这正是它会去追求的东西。至于这些提升是否在其他地方也适用,那就是另一个问题了。

Original English

Speaker: So let me quickly go through the four pieces. The first up is proposer uh which is based on oppus 4.6 used purely as mutation function. It reads the current best program and memory file and then it adds one Python file to propose the next one. So there is no human in the loop. Uh now here's the catch. It only optimize the metric that you give it to it not the metric you meant. So if you reward in domain performance and if you reward spending more compute then that exa that that that is exactly what it will chase. So whether the improvements hold up elsewhere is a separate question.

Speaker: 接下来是程序部分。它只是在编码器上执行的 Python 程序,而唯一重要的一点就是这个 embedding 函数。这就是计算预算。每次函数调用基本上就是对文本重新进行 embedding,或者是切换 LoRA 适配器,或者是选择更小的维度。所以,一次调用就是一单位的计算量。当然还有一些其他限制,比如程序不能引入任何超参数,不能进行任务路由,也肯定不能添加外部模型。这些限制迫使智能体去寻找与任务无关的程序,而不是悄悄地为每个任务进行优化的配置。

Original English

Speaker: So the next one is program it just acturate Python program over uh the encoder and the one piece that matters is this embed function. So that's a compute budget. So every function call there basically re-mbbed sounds text or switch the laurel adapter or pick uh smaller dimensions. So one call is one unit of compute. Uh there are some other constraint such as the program cannot introduce any hyperparameters, cannot do task routing, cannot add external models of course. So this conra those constraints uh force the agent to found task agnostic program instead of a config that's secretly optimized for each task.

评估与内存记录

Speaker: 然后是评估者(Evaluator)。每个程序都运行相同的 14 个评估任务或发现任务,涵盖法律、金融、长文档、长上下文或常规的信息检索问题。我们通过计算与余弦基准线(cosine baseline)相比的 delta nDCG 加上一些成本比率来对它进行评分。稍后我会介绍成本部分。这里有一项最为重要的设计选择:这个循环只能看到这 14 个任务,还有 19 个保留(held-out)任务是循环永远不会接触或看到的。这样我们之后就可以提出一个非常干净的问题:在这里获胜的程序,在那边也能保持优势吗?这整个性能差距(gap)基本上就是整个实验的核心。

Original English

Speaker: Then comes the evaluator. So every programs runs the same 14 evaluation task or discovery task spanning legal financial long document long context or general retrieval problems. We score it via delta and the CG against the uh cosign baseline plus some cost ratio. I will introduce the cost later. Now here's a design choice that matters the most. The loop only ever see these 14 task and there are 19 more held out task the loop will never touch them or see them. So later we can ask a very clean question that does what wins here uh also hold up there. So and the whole gap the gap is basically the whole experiment.

Speaker: 最后一部分是内存。它就是一个简单的 JSONL 文件,每行对应一个程序。每一行存储分数、成本、父代信息和一个简短的经验教训。因此,提议者在每轮之前读取此文件,并且整个搜索过程会随时间产生复利效应。但复利是双向的,对吧?当然,它可以积累真正的胜利,但它也会累积来自目标的任何偏见。而且偏见指标不仅会误导某一个程序,它会引导整个程序家族的方向。

Original English

Speaker: The last part is memory. So it is a simple JSONL uh file with one row per program. Each row stores the scores, the cost, the parents and a short lesson. So the proposal read this file before every round and the whole search compounds compounds over time. Uh but compounding has both ways, right? It builds a real win of course, but it alo also compounds whatever bias uh from the objective. And the bias matrix does not only mislead one program, it steers the entire family.

模型设置与测试时计算成本

Speaker: 现在,让我介绍一下我们在这里使用的模型设置。我们在单个编码器上运行搜索,即 Jina V5 Nano,这是一个仅有 2 亿参数、在多语言检索方面达到 SOTA 水平的模型。我们之所以选择 Nano 作为发现阶段的模型,主要因为它的体积小,从而能减少每个实验的循环周期时间。我们保留了同一家族中更大的模型,加上例如 Gemma 模型和 Qwen 模型等未曾见过的新模型系列。它们与发现阶段使用的模型不共享任何训练数据、主干架构或分词器。正如我之前提到的,我们还保留了 19 个评估任务,这些任务是循环永远无法看到的。所以,当程序在这个循环中被发现时,它必须能够跨越所有编码器和所有的 19 个任务进行泛化。

Original English

Speaker: Uh so now let me set up the models that we use here. We run the search on the single encoder which is the Gina V5 Nano uh only 200 million parameters state-of-the-art on multilingual retrieval. And we choose nano mostly because the discovery phase as a discovery phase model mostly because it is small and therefore reduce the cycle time of each experiment. We hold out the bigger model uh from the same family plus the unseen families such as gema model and quinn model. They share no training data, no backbone, no tokenizer with the discovery model. We also hold out the 19 evaluation task as I talked before and this one those 19 tasks the loop never sees. So when programs gets discovered in this loop it has to generalize over all encoders and all all 19 tasks.

Speaker: 因此,在展示任何结果之前,让我先定义一下测试时计算(test-time compute)的成本。它归结为一个数字,仅仅一个数字 C,也就是通过编码器的额外前向传播的次数。让我用幻灯片上的两张卡片来解释这一点。它们执行相同的操作,但它们混合了一些邻域信息然后对其重新评分。左边的卡片是我所谓的软质心(soft centroid)。它对文档取平均值,得到你已经计算出来的向量。因此没有额外的前向传播。这意味着它的成本 C 就是 1。右边的卡片是第一句话:它对顶级文档的第一句话重新进行了 embedding,这是一次全新的前向传播。所以那里的 C 大于 1。一种方法重用了我们已经拥有的几何结构,而另一种方法则将计算花在了对新文本的全新传播上。

Original English

Speaker: So now before showing any result let me define the cost of the test time compute. It comes down to one just just one number C which is the number of extra forward passes through the encoder. So let me explain it with two cards on the slice. They do the same thing but they they kind of mix in some neighborhood information and then rescore it. The card on the left is what I call a soft centroidid. It averages the document to uh vectors that you already computed. And so there is no extra forward passes. Uh that means it's cost C is just one. The card on the right is the first sentence. Uh it reimbed the first sentence of the talk top document which is a brand new forward pass. So there C is greater than one. So one reuse the geometry that we already have. The other spans compute on the new pass on the new text.

计算评估准则 vs. 迁移评估准则

Speaker: 既然我们在计算成本上达成了共识,我们就在两种不同的评估准则下运行完全相同的循环。第一种是计算准则(compute rubric)。它只在某个程序的域内表现优于之前所有程序时才接受该程序。因此它受到主动激励,倾向于在推理阶段花费更多的计算量。第二种是迁移准则(transfer rubric)。它只有在验证集上有所提升并且没有任何性能下降的情况下才保留程序,同时它不会为花费计算量提供任何奖励。要澄清的是,验证集仍然来自循环可以看得到的数据。所以这两种准则都绝对不会接触到那 19 个最终的评估任务或最终保留任务,以及未曾见过的编码器。这就是在同一个循环中运行的两种准则,让我们看看它们各自得出了什么结果。

Original English

Speaker: So now that we comprise the compute, we run that exact same loop under two different rubrics. The first is compute rubric. It admits a program only if the in domain performance beats every program before it. So it is actively pushed to spend more compute at inference time. The second is the transfer rubic. So it keeps the program only if it improves over over the validation set with nothing getting worse and it gets no reward at all for spending compute. And to be clear, the validation set is uh still comes from what loop can see. So neither rubric ever touch the 19 final evaluation task or final hold out task uh and unseen encoder. So that's a two rubric running under the same loop. So let's see what each one come up with.

Speaker: 让我们首先来看看计算准则。当你让它花费更多计算量时,它画出了一条非常漂亮且干净的曲线。X 轴是花费的计算量(以对数尺度表示),Y 轴是分数。总共有 144 个程序,其中 12 个落在帕累托前沿(Pareto front)上。成本从仅仅 1 开始,一直攀升到将近 15 倍,而域内分数也呈现出漂亮的上升趋势。跨越该前沿的分数增加了两倍多。这看起来完全就像是测试时计算扩展(test-time compute scaling)法则:计算越多,质量越高。如果我就此停下,你会被说服的。但这仍然是域内表现,我们还没有在保留数据集上运行这个实验。

Original English

Speaker: So let's first look at the compute rubric. So when you tell it to spend more compute, it draws this very beautiful clean curve. So the x-axis is a compute you spend on the log scale and y-axis is a score. There are in total 144 programs and 12 of them sit on the par front. The cost running from just one uh all the way up to almost 15 times and the in domain score climbed nicely. It it more than triples across that front. So this looks exactly like tet time compute scaling more compute more quality. So if I stop here you will be s but this is still in domain performance we haven't run this experiment on held out uh data set.

Speaker: 所以让我们快速看看这 12 个程序,在保留数据上运行它们。这里是绘制成小图表的 12 个程序。你不需要逐一仔细阅读。我希望你从中了解的唯一一点是:它们全都是对同一个冻结的 embedding 模型的免训练重组(training-free recombinations),只是采用了分块(chunking)、评分反馈和融合等技术。计算成本从左到右稳步、漂亮地攀升,看起来确实像是一个完美的扩展故事,但你将看到,它们在保留数据集上的提升情况并非如此。

Original English

Speaker: So let's take a quick look on this 12 programs and run them uh run them on the hot out uh data. So here are the 12 programs drawn as a little diagrams. So don't have to you don't have to read into each one. The only thing that I want you to take away from this is that they are all training free recombinations of the same frozen embedding models just chunking scoring feedback and fusion. The cost climb nicely steadily from left to right and does look like a clean uh scaling story but the improvement on the hel data set as you will see is not.

泛化与迁移测试结果

Speaker: 现在,我们在保留数据集上运行这 12 个程序,并使用与之前相同的图表。计算量从左向右增加,分数则上下浮动。中间穿过的虚线是一条基准线。看这条粉红色的线,它代表了计算准则。它基本上就是平的,一直紧贴着零。所以,在域外,投入更多的计算量基本上不会带来任何收益。现在看这些蓝色圆点,它们代表了迁移准则。它们全都位于左侧,因为它们的成本很低,而且每一个都在粉色线条上方。最便宜的那个额外计算成本甚至几乎为零,但它仍然打败了最昂贵的程序。由此可见,增加计算量未能实现迁移,而廉价的(轻量级的)架构却成功迁移了。

Original English

Speaker: So now we run those 12 programs on the held out data set and same chart as before. Compute runs from the left to right and scores runs up and down. So the dash line across the middle is a baseline and look at the pink line. It uh the compute rubric. It's basically flat hugging zero all the way out. So out of domain more compute buys you essentially nothing. Now look at the blue dots which is the transfer programs. They all sit on the left because they are cheap and everyone is above the pink line. So the cheapest one only has like as zero extra compute it still be the most expensive program. So more compute did not transfer the cheap structure did.

Speaker: 如果我们把每一个程序放在每一个保留任务上绘制出图,我们就会得到这张热力图。四个区块代表了四种编码器,其中三种我们在发现阶段从未见过。在每个区块中,行代表程序,列代表 19 个评估任务。绿色意味着提升。红色或粉色意味着下降。总体情况好坏参半。计算在大概一半的情况下有帮助,但改进是不均衡的。因此,平均来看,结果是持平的。计算在某些地方确实有帮助,但它并不能可靠地在所有新任务中都提供帮助。

Original English

Speaker: So if we plot every program against every held out uh task we get this heat map. Uh the four blocks are the four encoders and three of them we have never seen in the discovery phase. In each block the rows are the programs and the column are 19 evaluation task. Green means an improvement. red or pink means a drop. The picture is generally mixed. Compute helps about half of the sales but improvement are uneven. So on on average it comes out flat. Compute does help in places but it doesn't help reliably across all new all

Test-Time Compute in Search Pipelines

Speaker A:以及所有新的编码器。现在让我们来看看这些,呃,看看另一个标准,即迁移标准。它选择了六个完全不同的程序,它们都非常廉价,而且大多消耗的算力最多只有余弦基线的1.5倍。最好的一个在留出数据集上赢得了83%的胜率,并且在单一任务上从未输过。

Original English

Speaker A: and all new encoders. So now let's look at these uh look at the other rubric the transfer rubric. It picks the six completely different programs and they are all very cheap and most one and a half times uh more compute than the cosign baseline. The best one wins 83% on the held out data set and it never lose on single task.

Speaker A:那么现在,这些程序实际在做什么呢?它们只测试你已经有的一些查询和文档向量,并在其之上添加了一点廉价的质量。有些将查询向它已经喜欢的文档微调。有些在空间中选择几个方向,并沿着这些方向重新打分。所以它们是非常小的结构变化,但足以将正确的文档拉上来。

Original English

Speaker A: So now what what do this program uh actually do? So they only test some query and document vectors that you already have and they add a little cheap mass on top of that. Some notch the query towards the document it already likes. Uh some pick a few directions and uh in the space and rescore uh along those directions. So they are very small structure change but enough to pull the document uh the right document up.

Speaker A:因此,这一切都是重新组合,没有新模型,而且这确实迁移到了跨模型和跨语言。记住,在发现阶段我们只使用了 Jina 嵌入模型 Jina V5 Nano,但改进在所有四个编码器上都是正向的,而提升最大的部分在 Gemma 和 Qwen 上。所以这些是它从未见过的两个模型系列。因此,这并不是某个模型特有的怪癖,它是普遍存在的,建立在普遍的嵌入几何学之上。

Original English

Speaker A: So it's all re combination no new models and this really transferred to across models and languages. Remember in the discovery phase we only use GINA embedding Gina V5 nano and but the improvement is positive across all four encoders and the biggest bar is on the JAMA and the Quint. So those on the two families it never sees. So this is isn't some quirk of one model is general is rise on general embedding geometry.

Speaker A:所以这就是版本 A:一个冻结的编码器,带有非常廉价的结构,它可以扩展,但低算力无法扩展,而自动研究(auto research)是我们发现这一点的方式。但让我从模型层往上移动一个层次,来到搜索管道,你将看到相同的测试时计算(test time compute)反映在管道层面。

Original English

Speaker A: So that was version a frozen encoder with very cheap structure uh and it scales but low compute uh doesn't scale and auto research is how we found that but let me move one level up uh from the model layer to the search pipeline and you will see the same test time compute reflect in the pipeline level

Speaker A:在2025年,我们有了这个深度研究和智能体搜索产品,它基本上只是对开放网络进行一次循环。在2026年,我们转向了长周期任务,这在检索之上增加了实现沙盒评估,并运行数小时。因此,这两种模式在测试时都需要更多的循环和更多的算力。

Original English

Speaker A: in 2025 we have this deep research uh and agentic search product uh which was basically just a one loop over the uh open web. In 2026, we moved to a long horizon task which adds implementation sandbox evals on top of the retrieval and running for hours. So both patterns need more looping and more compute at test time.

Speaker A:所以为了研究这种测试时的智能体搜索,我为此构建了三个开源项目。第一个是 Data Room(数据室)。你给它一个 token 预算,它会去搜索、阅读、撰写。一次又一次地循环,直到把所有东西打包成一个 zip 文件。我叫它数据室,是因为它多少让我回想起当年我是创始人时,为投资者准备数据室的情景。所以那个 zip 文件详细记录了语料库……你可以把这个 zip 文件想象成一个开放网络的详细语料库,准备好供下一个智能体或大型语言模型使用。

Original English

Speaker A: So study this genic search at test time. Uh I built three open source projects for that. The first one is data room. So you give a token budget, it searchs, it reads, it rise. So over and over until it packaged everything into a zip file. So I call it data room because it somehow reminds me like prepares the data room for the investors uh back when I was a founder. So that zip file details the corpus on the uh you can you can imagine this zip file is a detailed corpus of the open web ready for the next agent or large language model to consume.

Speaker A:注意这里的 token 经济学。所以你基本上是在使用来自小型语言模型的极其廉价的 token 来探索网络并构建语料库,然后你将昂贵的前沿模型 token 留到后面用于利用(exploitation)。

Original English

Speaker A: And notice the token economy here. So you are basically exploring the web and build a corpus using very cheap tokens from small language models and then you save the expensive frontier tokens for later for exploitation.

Speaker A:第二个是 Search Box(搜索盒)。这是一个研究智能体搜索和工具调用的测试平台。它被设计成气隙隔离(air-gapped)的。所以智能体无法访问互联网。这基本上就像你把智能体锁在一个房间或一个盒子里,给它一个数据室,并向它询问有关数据室的问题。因此,为了回答这些问题,智能体必须在测试时组装一个搜索管道。一个由 grep、embed、rerank 等本地工具组成的管道。

Original English

Speaker A: The second one is search box. So this is a test bed to study agentic search and two calling. It is design it is designed to be air gapped. So the agent have no internet access. It's basically like you lock the agent in a room or in a box and you give it a data room and ask question about it. So to answer those question, the agent has to assemble a search pipeline at test time. A pipeline made of local tools since like a grap, embed, rerank.

Speaker A:这让你能够探索一些非常有趣的研究问题,比如,智能体会最先使用哪个工具?或者单靠 grep 就足够了吗?或者在热门问题上强制增加算力有帮助吗?或者智能体会建立一个以后可重用的搜索管道吗?所以 Search Box 是一个探索这些研究问题的测试平台。

Original English

Speaker A: And this allows you to explore some very interesting research questions such as uh which tool does the agent reach for first or is grab all you need or does forcing more compute help on hot questions or will the agent build up a search pipeline that it will reuse later. So search box is a test bed to explore those research questions.

Speaker A:但是你如何评估这样的智能体搜索呢?嗯,你需要一些有难度的问题。这基本上引出了第三个项目:Knowledge Graph(知识图谱)。它将一个语料库或数据室转变为知识图谱,每一个事实都变成一条边,将主语与宾语连接起来。然后我们可以去研究穿过这个图谱的最长路径,这些长链就变成了多跳(multi-hop)问题,没有任何单个段落能够回答。因此,智能体必须消耗更多的测试时算力来将这些事实(facts)联系起来以得出答案。所以它也是构建私有验证器(private verifier)的工具。

Original English

Speaker A: So but how do you evaluate uh aenic search like that? Uh well you need hard questions. Uh that's basically the third project is knowledge graph. So it turns a corpus or data room into a knowledge graph and every fact become an edge and linking from subject to an object. Then we can work on the longest path through that graph and those long chains become multihop questions then that no single passage can answer. So the agent has to spend more test time compute connecting the fats to get there. So it's also the tool for building a private verifier.

Speaker A:所以让我们把所有这些点连接起来。我介绍了搜索测试时计算的两个版本。这两个版本都在做同一件事。它们在测试时消耗更多的计算资源,并且都没有增大模型本身。在版本 A 中,我们在固定、冻结的嵌入模型上找到了一种特殊的嵌入代数,从而提高了搜索相关性。在版本 B 中,我们构建了一个全栈系统来寻找最佳的搜索管道。我们使用 Data Room 来最大化召回率(recall)。我们使用 Search Box 来最大化精确度(precision)。然后我们使用 Knowledge Graph 来构建评估。最终,这为我们提供了一个具有很强搜索相关性的管道。它基本上是两个不同的层面,但它们拥有相同的押注:投入更多的测试时算力,而不是更大的模型。

Original English

Speaker A: So let's connect all the dots together. So I introduced two versions of test time compute for search. Both versions are doing the same thing. They are spending mode compute at test time and neither of them grows the model. In version A, we found a special embedding algebra over the uh fixed uh frozen embedding that improves the search relevance. In version B, we build a full stack to found the best search pipeline. We use a data room to maximize recall. We use a search box to maximize precision. And then we use knowledge graph to build evaluation. So finally, it gives us a pipeline that with strong search relevance. It is basically two different levels, but they share the same bet. Spending more test time compute, not a bigger model.

Speaker A:最后,让我给你们留下一个大局观。搜索就是测试时计算(test time compute)。所以不要一味去追求更大的模型,而应该在推理时做更多的搜索。你不需要手动去做这种设计。自动研究(auto research)可能会在一夜之间帮你发现这一点。这就是我们如何扩展测试时算力的方式。这基本上就是我演讲的结尾了。你可以从这里的二维码获取所有的幻灯片。我的 GitHub 和 arXiv 上有论文和项目。如果你今晚在附近的话,Elasticity 也正在城里举办一场黑客松(hacker zone)。二维码就在那里。所以,来和我们一起构建吧。非常感谢,祝大家 AI 工程开发愉快。

Original English

Speaker A: So finally let me let me leave you with a big picture. Search is test time compute. So don't reach for bigger model. Do more search at inference instead. You don't have to do this design by yourself by hand. Uh auto research helps you discover this probably overnight. Uh so and this is how we scale the test time compute. And that is basically my the end of my talk. Uh you can grab all the slides from the QR codes here. There's a paper and projects on my GitHub and archive. And if you are uh if you are around this evening, Elasticity is also holding a hacker zone in town. So the QR code QR code right is right there. Uh so come and uh build with us. Thank you so much and happy AI engineering.

The Future of Durable Execution Platforms

Dominic Turno:在2026年,编程智能体将悄然淘汰它们的第一个软件平台。不是因为它不好,仅仅是因为这个平台不再是必需的。

Original English

Dominic Turno: In 2026, coding agents will quietly retire their first software platform. Not because it's bad, simply because the platform is unnecessary.

Dominic Turno:我是 Dominic Turno。我是 Resonate 的创始人兼 CEO。Resonate 是一个以极简主义和简单性为核心技术价值观构建的持久执行平台,而这些特性将在本次演讲中发挥核心作用。在 Resonate,我们对于软件工程的未来走向有一个工作理论。通用目的的实现将越来越多地被按需生成的定制实现所取代,这不再是作为一个新库、新框架或新平台,而是作为现有基础设施的一个最小化扩展。

Original English

Dominic Turno: I am Dominic Turno. I am founder and CEO of Resonate. Resonate is a durable execution platform built with minimalism and simplicity as its core technical values and these properties will play a central role in this talk. At Resonate we have a working theory where software engineering is headed. Generalpurpose implementations will increasingly be replaced by bespoke implementations generated on demand not as a new library, a new framework or a new platform but as a minimal extension of the infrastructure that is already in place.

Dominic Turno:如果这个理论成立,重用将会向上游移动。我们将不再重用通用目的的实现,而是重用一个规范(specification),并从中派生出一个定制的实现。事实上,我们可以为已经存在的基础设施量身打造许多定制的实现。我们只需要要求智能体去做。在这一点上,提示词(prompt)本身就成了一个平台。

Original English

Dominic Turno: If this theory holds true, reuse will move upstream. Instead of reusing a general purpose implementation, we will reuse a specification and we will derive a bespoke implementation from it. In fact, we can build many bespoke implementations tailormade for the infrastructure that is already in place. We just have to ask the agent. At this point, the prompt is a platform.

Dominic Turno:Resonate 是一个双重执行平台。我们有 Resonate 服务器的实现。我们也有适用于 TypeScript、Python、Rust、Go 和 Java 的 Resonate SDK 实现。所以,我们必须问自己:这个新的现实对我们意味着什么?如果实现变得可以自动生成,我们的价值存在于哪里?我们的答案是:我们的价值从实现转移到了规范。

Original English

Dominic Turno: Resonate is a dual execution platform. We have an implementation of the Resonate server. We have implementations of the Resonate SDK for TypeScript, Python, Rust, Go and Java. So, we have to ask what does this new reality mean for us? If implementations become generatable, where does our value live? And our answer our value moves from implementation to specification.

Dominic Turno:现在,这改变了我们对 Resonate 的看法。产品不再是其实现代码。产品是规范,是协议(protocol)。我们希望从这个协议中派生出多个服务器实现。其中一个是通用目的的 Resonate 服务器,即我们的参考实现。其他的则是与基础设施合作伙伴一起构建的实现。对于客户和合作伙伴而言,这意味着在他们现有的基础设施之上,以最少的额外依赖实现持久执行。

Original English

Dominic Turno: Now this changes how we think about Resonate. The product is no longer the implementation. The product is the specification the protocol. And from that protocol we want to derive multiple server implementations. One is a general purpose resonate server. our reference implementation. Others are implementations built with infrastructure partners. For customers and partners, this means durable execution right on top of their existing infrastructure with minimal additional dependencies.

Dominic Turno:因此,问题不再是我们能否构建一个服务器。问题是,我们能否从同一个规范中反复综合出可信的服务器?如果可以,如何做到?当我们谈论智能体工程(agentic engineering)时,我们将所有注意力集中在验证上。我们如何知道结果是正确的?但是今天,我想把重点放在规范上,更重要的是,智能体如何参与到系统的规范制定中,而不仅仅是构建或验证它。

Original English

Dominic Turno: So the question is no longer can we build a server. The question is can we repeatedly synthesize trusted servers from the same specification and if so how? When we talk about agentic engineering, we focus all of our attention on verification. How do we know the result is correct? But today, I want to focus on the specification instead and more importantly, how can agents participate in specifying the system, not just building or verifying it.

Dominic Turno:现在,Resonate 正与多个基础设施提供商合作,将持久执行原生引入到他们的技术栈中。其中之一是 Synadia,即 Nats.io 背后的公司,这是一个专为构建现代分布式系统而设计的开源消息传递系统。在接下来的演讲中,我们将使用 Resonate 和 Nats.io 来探索我们的智能体工程实践。我们如何从规范走到实现?

Original English

Dominic Turno: Now, Resonate is partnering with multiple infrastructure providers to bring durable executions natively to their technology stack. One of them is Senadia, the company behind Nats.io, an open-source messaging system designed for building modern distributed systems. For the rest of this presentation, we will use Resonate ornat.io to explore our agentic engineering practices. How do we go from specification to implementation?

Dominic Turno:首先,我们需要统一一下我们的心智模型。这张图展示了智能体解码的一个常见视图。这里有一个智能体,有一个规范,然后有一个实现。对于许多应用来说,那就是……

Original English

Dominic Turno: First, we need to level set our mental model. This picture is a common view of agent decoding. There's an agent, there's a specification, and then there's an implementation. And for many applications, that is

抽象规范与可执行设计

Speaker A:……够了。但这对于我们想要实现的目标来说还不够,因为我们并不是要从规范中生成一个单一的实现。我们试图从规范中生成多个针对特定目标的实现。因此,规范绝对不能将实现的任何方面考虑在内。规范绝不能假定有具体的数据库模式或具体的索引。规范甚至绝不能假定有一个带有表和事务的关系型数据库。它不能假定有键值存储。它不能假定弱一致性。它也不能假定强一致性。规范必须是抽象的。只有实现才必须是具体的。

Original English

Speaker A: ...enough. But it is not enough for what we are trying to do because we are not trying to generate one implementation from a specification. We are trying to generate multiple target specific implementations from the specification. So the specification must not take any aspect of an implementation into account. The specification must not assume a concrete database schema or concrete indices. The specification must not even assume a relational database with tables and transactions at all. It must not assume a key value store. It must not assume weak consistency. It must not assume strong consistency. The specification must be abstract. Only the implementation must be concrete.

Speaker A:所以,我们让智能体遵循抽象规范来生成具体的实现。具体来说,起初我们让智能体在 Postgres 之上用 Rust 构建一个 Resonate 服务器,但智能体失败了。抽象规范和具体实现之间的差距太大了。智能体生成了一个只能在理想路径下运行的系统。它通过了基本测试,但并不正确。它在并发处理上崩溃了。它在进程失败时崩溃了。它在网络故障时崩溃了。这个实现更接近于一个原型,而不是一个生产系统。

Original English

Speaker A: So we ask the agent to follow the abstract specification and generate a concrete implementation. Specifically at first we ask the agent build a resonate server in rust on top of posgress and the agent failed. The gap between the abstract specification and the concrete implementation was too large. The agent generated a system that worked on the happy path. It passed the basic tests, but it was not correct. It broke on the concurrency. It broke on the process failure. It broke on the network failure. The implementation was closer to a prototype, but not a production system.

Speaker A:因此,我们修改了流程。我们不再要求智能体直接从抽象规范跳到具体实现,而是插入了一个中间产物:具体规范(concrete specification)。这个具体规范是与智能体交互推导出来的,但人类是主要的驱动者。对于 Postgres 来说,这意味着要明确针对特定目标的决策,比如数据模式、索引、SQL 查询和事务边界。一旦这些决策被记录下来,智能体确实能够实现出生产系统。所以这套方法行得通,但它也暴露了局限性。智能体帮助我们构建了系统,但并没有帮助我们设计系统。如果规范是一个可复用的产品,那么这还不够。

Original English

Speaker A: So, we amended the process. Instead of asking the agent to jump directly from abstract spec to concrete implementation, we inserted an intermediary artifact, the concrete specification. That concrete specification was derived interactively with the agent. But the human was the main driver. For Postgress that meant making target specific decisions explicit, the data schema, the indices, the SQL queries, the transaction boundaries. Once those decisions were written down, the agent was indeed able to implement the production system. So this worked, but it also revealed the limitations. The agent helped us build the system, but the agent did not help us design the system. And if the specification is a reusable product, then that's not enough.

Speaker A:现在下一步显而易见了。智能体必须向上游移动。但怎么做呢?当我们开始在 Natio 上构建 Resonate 时,我们改变了问题。我们不再问智能体能否构建生产系统。相反,我们问智能体需要什么才能先设计系统、后构建系统。于是,我们让智能体能够访问一个确定性的模拟环境。我们给了它一个不同的任务:不要构建生产系统,而是构建一个模拟的实现。

Original English

Speaker A: Now the next step is obvious. Agents have to move upstream. But how? When we started building Resonate on Natio, we changed the question. We did not ask can the agent build the production system. Instead we ask what does the agent need in order to design the system first and build the system second. So we gave the agent access to a deterministic simulation environment. And we gave it a different task. Do not build the production system. Build a simulated implementation.

Speaker A:模拟的实现并不是产品。它是可执行的设计。它的目的是在偏序(partial order)和部分失效(partial failure)的情况下发现正确的算法。一旦在模拟中发现、测试并验证了这些算法,我们再让智能体编写具体的规范。只有到了那时,我们才会让智能体编写生产实现。所以整个流程就变成了:抽象规范,模拟实现,具体规范,然后是具体实现。这正是智能体向上游移动的节点。人类仍然参与设计过程,但现在智能体成为了驱动者。

Original English

Speaker A: The simulated implementation is not the product. It is executable design. Its purpose is to discover the correct algorithm under partial order under partial failure. And once these algorithms are discovered, tested and verified in simulation, then we ask the agent to write the concrete specification. And only then do we ask the agent to write the production implementation. So the process becomes abstract specification, simulation implementation, concrete specification and then concrete implementation. This is a point where the agent moves upstream. Humans are still involved in the design process, but now the agent is a driver.

Speaker A:有两个要素使这一切成为可能。极简主义和简单性。不幸的是,极简主义和简单性并不是起点,它们是终点。我们花了三年时间把协议变得更小、更简单。每次遇到问题,我们都会问:“我们可以拿掉什么?我们可以抹除什么抽象?我们可以移除什么属性?我们可以打破什么关系?”结果是产生了一个非常小的协议,围绕着两个对象展开:持久的承诺(durable promise)和持久的任务(durable task)。

Original English

Speaker A: Two ingredients make this possible. Minimalism and simplicity. Unfortunately, minimalism and simplicity are not the starting point. They are the finish line. We spent three years making the protocol smaller and simpler. Every time we ran into a problem, we ask, "What can we take away? What abstraction can we erase? What property can we remove? What relationship can we break?" The result is a very small protocol centered around two objects, a durable promise and a durable task.

Speaker A:这种简单性很重要,因为即使是简单的并发分布式协议,也具有复杂的状态和行为空间。换句话说,即使在几个简单的原语之上实现简单的协议也是很困难的。让我们用 NATS 来具体说明一下。NATS 给了我们一个……

Original English

Speaker A: That simplicity matters because even simple concurrent distributed protocol have a complex state and behavior space. So in other terms implementing even simple protocols on top of a few simple primitives is tough. Let's make this concrete with NATS.NATS gives us a

设备端的长周期研究智能体记忆框架

Stefania Dug:大家好,欢迎。这是一个很大的场地,如果你坐在后面,不要犹豫,请往前坐近一点。我叫 Stefania Dug。我是东京 Sakana AI 的研究科学家。我以前常驻这里,在我进入 Hyperloop 之前,AI 工程领域就是我的大本营社区。所以很高兴能回来,今天我要和大家谈谈在设备端为长周期运行的研究智能体构建记忆框架(memory harnesses)的话题。

Original English

Stefania Dug: Hello, welcome. This is a big room, so if you're in the back, don't hesitate to come closer. My name is Stefania Dug. I'm a research scientist at Sakana AI in Tokyo. I used to be based here and AI engineering is home community for me before being the hyperloop. So it's very good to be back and today I'm going to talk to you about memory harnesses for long running research agents on device.

Stefania Dug:如果你从事过长周期任务,你可能遇到过上下文爆炸(context blow)的问题,对吧?比如模型开始自相矛盾,或者它不得不重新做一遍工作,因为它忘了自己一开始已经做过那个任务,又或者它开始偏离你的问题,因为它把问题给忘了。现在这个问题比以往任何时候都重要,因为从最近关于参数规模的预测中我们可以看到,目前的趋势是解决越来越长的周期任务,而且模型发布也越来越少。因此,在今年晚些时候的某个时间点,我们将迎来一个交汇点:我们将面临更多长周期的任务,而模型发布会变少。这就使得处理“上下文腐败”(context rot)的问题成为了一项优先任务。

Original English

Stefania Dug: So if you work with long horizon tasks, you probably run into this issue of context blow, right? like when the model starts contradicting itself or it has to redo the work because it forgot it did that task in the first place or it starts to drift from your questions because it forgot them. And this matters now more than ever because from this recent projections from meter we see that the trend is to solve longer and longer horizon tasks and also that we're getting fewer and fewer model releases. So at some point later this year we're going to have this convergence right where we'll get many more long-term horizon tasks and fewer model releases. So that makes this issue of dealing with context rot a priority.

Stefania Dug:那我为什么想要在本地模型上并使用本地框架来解决这个问题呢?也许你们中的一些人见过这条推文。这才是两天前的推文。Coinbase 的 CEO 实际上分享了他们公司是如何设法减少 AI 支出的,同时还增加了 AI 的使用量。他们做到这一点的方法是转向使用更多的本地模型,同时也采用了更好的实践,比如使用更好的路由、更好的缓存,保持上下文干净,然后让人们在使用什么以及针对什么类型的任务有更好的可视度。

Original English

Stefania Dug: And why did I wanted to tackle this problem on local models and with a local harness? maybe some of you have seen this tweet. It's only two days old. The CEO of Coinbase actually shared how their company managed to reduce their AI spent while actually increasing the AI usage. And the way they did that was by transitioning to use many more local models but also having better practices like using better routing, better caching, keeping the context clean and then having better visibility for what people are using and for what kind of task.

Stefania Dug:所以我们看到本地模型正在跨越那道界限,对吧?就像大家都在关注 GLM 一样,特别是在 Fable 退出之后。DeepSeek-V4-Flash 现在已经可以在 M3 Ultra 上运行了,而且仍然有 RAM 的瓶颈。虽然有些棘手,但这些本地模型开始在智能体任务和工具调用方面发挥作用了。因此,我想向你们展示一下我今天将要分享的实验的设置。

Original English

Stefania Dug: So we are seeing the local models like crossing the line, right? Like GLM is on everyone's minds like especially with Fable going away. DeepS v4 flash can now be run on M3 Ultra and there's still a bottleneck for RAM. It's tricky, but these local models are starting to be useful for agentic tasks and for tool use. So, I wanted to show you what has been my setup for the experiments I'm going to share with you today.

Stefania Dug:这是我的 Mac。它现在还在我东京的办公桌上运行着评估,而我是通过手机来控制它的。在连续几天不停地运行评估之后,它开始发热了。所以我让我丈夫在它周围放了风扇。我们快没有风扇可用了,但这台机器还在运转,而且评估也还在出结果。在这台配备了 96GB 内存和 28 核 CPU 的 M3 Ultra 上,我正在使用两个模型。我用的是以 4-bit 量化的 Qwen 27B 和 DeepSeek-V4-Flash。

Original English

Stefania Dug: This is my Mac. It's still running evaluations right now back in my desk in Tokyo and I'm controlling it from my phone. And after running evals non-stop for a couple of days, it started to get hot. So I had my husband put fans around it. We're running out of fans, but the machine is still running and the evals are still giving results. On this M3 Ultra with 96 gigabytes and 28 core CPUs, I'm using two models. I'm using the Quen 27B quantise at 4bit and the DC V4 flash.

Stefania Dug:在向你们展示我如何在这台机器上构建记忆框架之前,我想先告诉你们,这到底是个什么例子,对吧?比如记忆。当我们为记忆设计一个框架时,这就是我希望你们在脑海中建立的思维模型。你可以把记忆看作是一个写入、管理、读取的循环(write manage read loop)。所以,它不仅仅是一个数据库存储。它实际上是围绕模型的一个控制循环。

Original English

Stefania Dug: And before I show you how I built the memory harness on this machine, I wanted to tell you what this what is this an example of, right? Like memory. When we design a harness for memory, this is the mental model I want you to have in mind. You can think of memory as a write manage read loop. So, it's not just the database store. It's actually this control loop around the model.

Stefania Dug:更具体地说,我是如何采用那个循环并对其进行定制的呢?这就是我的框架设计。比如,我从研究智能体入手,也就是那些小型智能体,因为它们完全没有持久记忆,而我希望所有的记忆都来自于这个框架。然后在中间,我有一个始终向智能体展示追踪信息的核心。接下来,我有一个召回模块(recall block),在其中我会测试不同的模式;还有一个归档模块(archival block),用于跨不同会话记录信息。在那个召回模块中,我实际上正在通过一系列模式的阶梯来进行测试。基准测试是完全不使用记忆。完全没有召回。所以我针对这个进行测试。接下来是使用向量 RAG(检索增强生成),纯粹是为了看看这个框架在相似度方面能提取出什么。然后是使用一个决策账本(decisions ledger),我会在其中记录每个轮次做出了什么决策,接着我就可以对它们进行优先级排序。

Original English

Stefania Dug: More concretely, how did I take that loop and customize it? So, this is my harness design. Like I started with research agents that are the small agents because they have zero durable memory and I wanted all the memory to come from the harness. And then in the middle I have a core which is always shown to the agent of traces. And then I have a recall block where I'm testing different modes and an archival block where I'm keeping track of information across different sessions. And in that recall block I'm actually going through a ladder of modes that I'm testing. The baseline is like not to use memory at all. No recall at all. So I'm testing for that. Next is to use vector rag just to see whatever like the harness would pull in terms of similarity. Then is to use a decisions ledger where I actually keep track of what decisions are being made for every turn and then I can prioritize them.

Stefania Dug:最后但同样重要的是,这一块非常关键。我有一个我称之为预言机(oracle)的东西,但基本上这就是基本事实(ground truth)。这就好比在每一个循环中告诉框架,需要检索的正确记忆是什么。在所有不同的任务中,模型是固定的。所以我改变的唯一东西,就是召回模块中的这些不同变量。

Original English

Stefania Dug: And last but not least and this piece is very important. I have a what I call an oracle, but basically this is the ground truth. So this is like telling the harness for every loop what the correct memory that needs to be retrieved is. And the model is fixed across all the different tasks. So the only things that I'm changing is like these different variables in the recall block.

Stefania Dug:我想给你们举一个我测试的第一个任务的例子。所以我想看看,如果我给智能体一个做文献梳理的任务……

Original English

Stefania Dug: And I wanted to give you an example of a first task that I tested. So I wanted to see if I give the agent a task of doing literature

记忆框架在长周期任务中的价值

Speaker A: ……回顾一下。我在语料库中包含了很多有着巨大科学主张的论文,比如这实际上是一篇《自然》杂志的论文,他们声称发现了 742,000 种有前途的材料。这是一个非常大的主张,后来被撤回了。但在那个语料库中,撤回声明就像大海捞针一样微小,远不如头条新闻和引用量那么显眼。所以我想看看系统是否能针对这类问题检索出正确的答案。我发现,因为对于这些任务,所有的论文和信息都能放入上下文窗口中,所以记忆其实并没有增加额外的能力。有记忆和没有记忆的表现是一样的,而且只会增加更多的成本。所以,当你的任务能放进上下文时,记忆框架并没有带来多少增益。然而,如果我开始运行长周期视野的任务,并且整个任务和相关的上下文放不下时,那么拥有一个好的记忆框架才真正开始得到回报。这是我运行的另一个任务的例子。这实际上来自一个已建立的针对长周期任务记忆的基准测试,叫做 Xbench。这是一个问题的例子,对吧?所以我问了一个问题,然后正确的答案可能在第 124 步。但在我提问的那一刻,我大概是在第 500 步提问的。所以这完全超出了上下文窗口,模型需要使用记忆框架从正确的步骤中检索出特定的答案。

Original English

Speaker A: ...review. And I'm including a lot of papers in the corpus where there was a big scientific claim, like this is actually a nature paper where they said they discovered 742,000 promising materials like it was a very big claim which got retracted later but the retraction which it's a much smaller like hay stack needle in that corpus than the headlines and the citations. So I wanted to see if the system can retrieve the right answer for these type of questions. And what I found was because like for these tasks all the papers and all the information fit into the context, the memory actually didn't add more capability. It was the same performance with memory and without memory and it only added more cost. So when your task fits in context, the harness doesn't add much. However, if I start to run tasks that are longer term horizon and the entire task and the relevant context doesn't uh fit, then having a good memory harness really starts to pay off. So this is another example of a task that I ran. This is actually from an established benchmark for a long horizon uh tasks memory. It's called Xbench. And this is an example of a question, right? So I'm asking a question and then like the right answer is in a like step 124. But the moment when I ask the question, I'm asking it like at step 500. So it's completely outside of the context window and the model needs to use the memory harness to retrieve the specific answer from the right step.

Speaker A: 因此,我通过改变我之前解释过的不同策略阶梯来测试这一点,包括关闭记忆、部署不同类型的召回,以及使用 oracle(预言机)作为参考。我发现,使用排序召回时,模型得到正确答案的频率比不使用时更高。这是在 Xbench 任务上性能分解的一个细目。我跑了 68 个问题。对于这些问题中的每一个,都有多个单元格和许多不同的种子。我发现,只做排序的分类账表现最好,它比仅仅通过控制框架(问你需要使用记忆还是不需要使用记忆)表现要好。你可能会问,为什么 oracle 没有达到最高值?我也会解释一下。oracle 所做的是将正确的信息和记忆提供给模型,但它并不强制模型去使用它。所以模型可能得到了正确的记忆,但仍然检索出了错误的信息,或者选择忽略它,或者感到困惑。这就是为什么在这个案例中 oracle 没有达到最高性能的原因。我对这些任务做了大量的消融实验,想看看如果我给出任意的例子会发生什么,如果我给它错误的步骤会发生什么,如果我给它最近的一步又会发生什么。我依然发现,表现最好的情况是使用排序策略进行召回。这实际上在多个模型上都奏效,不仅在 Qwen 27B 上,在 DeepSeek v4 Flash 上也是如此。而且它在不同的基准测试中也都有效,我还用 Spider V2 基准测试试过。它不仅能提供更好的召回率,实际上成本也更低。

Original English

Speaker A: So I'm testing this by uh changing the different policy ladder that I explained before with memory off uh by deploying recall different types of recall and by using the oracle as a reference. And what I found was that with the ranked recall, the model gets the right answer um more frequently than without. And here is a breakdown of the decomposition of performance on this Xbench tasks. So I ran over uh 68 questions. And for each of these questions, there were like multiple um cells and lots of different seeds. And what I found was that the rank only ledger performed the best and it performed better than like just gating the harness by saying do you need to use memory or do you not need to use memory and you're probably going to ask like why is the oracle not hitting like the max and I'm going to explain that too. So the oracle what it does it provides the right information the right memory to the model but it doesn't force it to use it. So the model can get the right memory but still retrieve the wrong information or choose to ignore it or be confused. So that's why the oracle in this case doesn't hit the max performance. And I've done lots of ablations on these tasks to see like what happens if I give arbitrary um examples. What happens if I give it the wrong step? What happens if I give it the most recent step? And I still found that the best performing condition was the one with the ranked policy for recall. And this actually works on several models, not only on the Quen 27B, but also on the DS4 flash. And it also works across different benchmarks. I also tried it on the Spider V2 benchmark. And it's not just that it gives you better recall, it actually costs less.

Speaker A: 所以,这里一个好的启发式方法可能是:糟糕的记忆是昂贵的,因为它会消耗更多的 token,并可能把智能体引向错误的方向;但拥有一个好的用于召回的结构化策略可以为你节省大量的 token 和预算。所以,从这个实验中,我想鼓励大家的一件事是,将召回策略视为一等公民的指标,并开始思考如何在你们的系统中使用它。比如,你想存储什么类型的记忆?你如何对它们进行排序?你如何设计你的召回函数?然后,当你在一遍又一遍、在多个会话和多次运行中执行这个过程时,什么类型的东西会留存下来?这只是一个简单的初步实验。但是,记忆技术的领域是非常丰富的。所以,在这个来自 Diamond 的开源仓库中共享了 30 多个可运行的指南,而且记忆是复杂的。我们有短期、长期的各种认知技术。我们也可以开始使用评估结果。而且现在确实有相当广泛的解决方案,对吧?从简单的文件系统检索到训练记忆模型,从较少结构化到完全结构化,存在着广泛的解决方案。所以我认为我们将在这个领域看到大量的研究。

Original English

Speaker A: So maybe a good heuristic to have here is that bad memory is expensive because it spends more token and it can send agent the wrong way. But having like a good structural policy for recall can save you a lot of tokens and uh budget. So one thing that I want to encourage you from this experiment is to consider the recall policy as a first class metric and to start to think about how you might use it in your systems. Like what are the type of memories that you want to store? What how do you rank them? Like how do you design your recall function? And then um what are the type what survives when you run this over and over and over and um multiple sessions multiple runs and this is just a simple first kind of experiment. Um but the memory technique landscape is very rich. Um, so there's over 30 runnable cookbooks that are shared in this open-source repository from um, Diamond and memory is complex. We have short-term, long-term different cognitive techniques. Uh, we can use start to use evaluation results as well. Um, and right now there's actually a a pretty broad landscape of solutions, right? So going from simple file system retrieval to training memory models um there's there's a wide spectrum of solutions from less structural to completely structured. Um so I think there's a lot of research we're going to see in this space.

Speaker A: 这很重要,它正变得越来越相关。对我来说,在本地模型上测试这个非常有趣,因为我可以控制一切。我可以控制我使用的数据,计算和评估的整个轨迹。是的,我认为这是主权(sovereignty)的一个例子,但这是有代价的。我没有告诉你们的是,我只能串行运行这些本地模型,它们不支持对 DeepSeek v4 Flash 进行批量查询。这就是为什么我仍然需要在东京的电脑上运行评估,或者在我来这里的航班上运行,因为这需要很长时间。但我仍然认为这非常强大,这是一个很好的测试,展示了当你能控制流水线的每一步时,记忆能做什么。而这种主权能力是在日本 Sakana AI 对我们非常重要的一个更大生态系统的一部分。我们比以往任何时候都坚信当今主权 AI 的重要性。而且我们也在招聘。所以,如果你对此感兴趣并想了解更多,如果你想来日本加入我们,请来找我聊聊。非常感谢。

Original English

Speaker A: uh it's important um it becomes more and more relevant and for me it's been super fun to to test this on local models um because I got to control everything. I got to control the data I was using the entire traces of compute and evaluations and um yeah I I see that as an example of sovereignty and it comes at a cost. Uh I didn't tell you that these local models I can only what uh run them in serial like they don't support batch querying for the deepse v4 flash. So that's why I am still running evaluations back on my computer in Tokyo or I I was doing it on the flight on my way here because it takes a long time. Um, but I still think it's very powerful and it's a very good test for what memory can do when you can control every single step of the pipeline. And this sovereign capability is part of a bigger ecosystem that is very important for us at Sakana AI in Japan. Um, we believe in the importance of sovereign AI today more than ever. And we are also hiring. So, if you're interested and want to hear more about this and if you want to come join us in Japan, come talk to me. Uh, thank you very much.

AI 时代的真正瓶颈:需求挖掘与人际沟通

Bash: 大家好,我是 Bash,今天我将和大家谈谈,作为软件行业从业者,AI 最后会从我们这里夺走什么。当编写代码不再是瓶颈时,真正重要的事情是弄清楚你应该构建什么。这归根结底在于人的技能以及掌控全局的能力,因为你无法提示(prompt)整个房间的人,你只能提示你的 AI。在今年年初,我们举办了一场内部黑客松,我们大概有 21 个智能体的想法,其中 17 个被放弃了,因为它们实际上没有创造任何商业价值。我们要么无法获取数据访问权限,要么构建它们根本没有意义。而剩下的那四个智能体实际上对我们今天的工作方式产生了非常大的影响。这是一个很好的例子,说明了我们需要确保自己在构建值得构建的东西。

Original English

Bash: Hi everyone, I am Bash and today I will talk to you about what is the last thing that AI will take away from us as people in the software business. So at a point where writing code is no longer the bottleneck, the real thing is figure is figuring out what it is that you should be building. Um, and that comes down to to people's skills and being able to work the room because you can't prompt the room, you can prompt your AI. So at the beginning of the year we held an internal hackathon uh where we had about 21 agents uh agent ideas and 17 of those were abandoned because they actually created no uh business value. They uh uh we either didn't have uh data access or or just didn't make sense uh to build it. And those four were the ones that actually had a very big impact on how we work today. And it's it's a very good example of of just making sure that we are building what is worth building.

Bash: 在我过去 13 年的职业生涯中,我一直是业务部门、IT 部门和开发人员之间的桥梁。我一开始测试——最初其实是测试功能设计规范,然后我开始编写它们;作为一名功能顾问,我参与了美国和英国的大型 ERP 和 CRM 项目,后来我创立了 Visual Labs。从根本上说,我培训我的团队如何以一种方式引发这些需求,这种方式可以让我们将它们转化为供开发人员构建、供顾问配置的优秀且具体的规范,而最近则是转化为供 AI 构建的规范。多年来真正没有改变的是我们与客户的交互方式,而我们与系统的交互方式、与 AI 的交互方式正在发生巨大的变化。这是现在的一件大事。但是,如果你能察言观色,如果你能引发正确的需求,那么你就能构建出更有价值的软件。

Original English

Bash: And throughout my career in the past 13 years, I've always been uh the bridge between business and IT and developers. Um I started writing well initially testing uh uh functional designs specifications and then uh and then I wrote them and as a functional consultant I worked with large ERP and CRM programs in the US and the UK and then I founded Visual Labs and essentially I trained my my team on how to elicit those requirements in a way uh that we can turn them into good uh specific ifications for developers to build, for consultants to configure, and most recently uh for AI to build. And what's not really changed over the years is how we interact with our customers, how we interact with systems, how we interact with AI is very much changing. Um and that's that's uh that's the big thing now. Uh but if you can read the room, if you can elicit the right requirements, uh then you will be able to build more valuable software.

Bash: 过去两三年的本质转变在于,获取代码和进行构建的能力已经不再是软件开发生命周期中的瓶颈。现在真正的瓶颈是,把你的团队、你的利益相关者、你的决策者召集到房间里,能够接触到他们,挖掘出需求,并能够和他们一起花时间讨论。所以这才是真正的瓶颈,弄清楚到底应该构建什么。因为你可以给你的代码写提示词,给你的 AI 写提示词,甚至给你的整个规范写提示词,但你无法给满屋子的人写提示词。模型做不到的事情,非常类似于亨利·福特关于询问他的用户或客户的那个比喻。如果他问他们想要什么……

Original English

Bash: And that essentially the big shift over the past two three years was that getting access to code and being able to build is no longer the bottleneck to the software development life cycle. Now the real bottleneck is getting your people, your stakeholders, your decision makers into the room and being able to access them and elicit the requirement and being able to spend the time with them. So that's the right that's the real bottleneck figuring out what it is that should be built because you can prompt your code, you can prompt your AI, you can prompt your whole specification, but you can't prompt your room. And what a model can't do is very similar to how Henry Ford's analogy of uh what he said about asking his users or his customers. If he'd asked them what

超越平均水平的AI应用

Speaker A: 如果他们需要的话,他们会说他们需要更快的马。但实际上,他造出了一辆汽车,并在这方面取得了巨大的成功。因此,如果你只是使用AI来、来让事物变得更好、构建得更好,嗯,你很有可能只是在复制已经存在的东西,因为从定义上讲,AI被编码为为你提供最常见的答案,所以对我们来说,真正的工作是确保AI远离那种平均水平,走向对我们来说更好的东西,这样我们就能、呃,不仅仅是得到一匹更快的马,而是真正生产出一辆在量级上比我们拥有的要好得多的汽车。因此,这真的是一个有趣的词、世界,在这个世界里,呃,能够写出优秀的代码不再是、呃,最需要具备的重要技能了。

Original English

Speaker A: it is that they needed, they would have said they needed more horses. But in reality, he built a car and he made a very big success on them. So if you're just using AI uh to to make things build things better, um the chances are that you are replicating what already exists because AI by definition is coded to give you the most common answers for so for us the real job is to make sure that AI moves away from that average into what is better for us so we can just get to uh not a faster horse but actually produce a car that's a magnitude shift better than what we had. So it's really an interesting word world where uh being able to write good code is no longer uh the the most important skill to have.

Speaker A: 呃,实际上现在的真正技能正变成分析师、分析师的工具箱,呃,也就是像故事地图(story mapping)、商业模式画布、呃,价值画布以及那些我们在作为职能顾问、业务分析师、嗯,或者、或者呃,在设计思维领域中非常习惯使用的那些古老而优秀的东西。所以我想聚焦在故事地图上,因为那是我发现最有价值的技能集。所以,呃,一旦你有了带有主干的故事地图,并理解在每一步你的客户、你的用户在做什么,那就能赋予他们能力,呃,去推进,呃,在他们的、在他们的流程中。

Original English

Speaker A: Uh actually the real skill now is becoming the analyst analyst toolkit uh which is things like story mapping, business model canvas, uh value canvas and those those good old things that we are so used to using as functional consultants, business analysts um or or uh in in the world of design thinking. So I'd like to zoom in on story mapping because that's the the skill set that I found as the most valuable. So uh once you have the story map with the backbones and understand at each step what your customers your users are doing that would give them the ability to uh to move forward uh in their in their processes.

Speaker A: 所以,呃,这里有一个、呃,支持系统用户故事地图,包括联系、分类、解决,然后基本上就是关闭一个案件。呃,有了这个,呃,你可以理解流程的不同阶段,呃,然后捕捉它们下面的用户故事。它的目的是保持在一个相当高的层级。这样你就能获得一个、呃,宏观的图景,然后你可以决定、呃,你想要在第一个版本(release one)中构建什么,比如捕捉意图、对紧急程度进行分类、起草一个有根据的回答,然后记录、把它记录到一个记录系统中。那基本上就是你的MVP(最小可行性产品)。那是你首先想要构建的东西,而那些就是你的前四个用户故事。

Original English

Speaker A: So uh here's a uh support systems user story map contacting triaging resolving and then essentially closing a case. Uh with this uh you can understand different stages of the process uh and then capture the user stories beneath them. It is intended to stay at a fairly high level. So you can get a uh a big picture and then in you can decide uh what it is that you want to build in release one like capturing intent, classifying urgency, drafting a grounded answer and then logging logging it to a system of record. That's essentially your MVP. Those are the first things that you'd want to build and those are your first four user stories.

Speaker A: 在这些之下,你还有、呃、呃,第二组用户故事,比如读取情感、向团队写入信息、建议下一步行动、聊天、检查满意度,等等。呃,那些将成为你待办事项(backlog)的一部分。所以,能让你、呃,获得非常好的、呃,智能体(agentic)结果的做法,就是专注于这些用户故事,并确保你将这些用户故事作为一种手段,呃,去引发与你的利益相关者、与你的业务部门的讨论,然后弄清楚那个用户故事真正应该关于什么。所以第一个用户故事、呃,第二个用户故事可能是:作为一个支持主管,我需要按紧急程度排序来打开案件,这样就不会有任何升级事件被遗漏(slip)。

Original English

Speaker A: And beneath those you've got the uh uh the second set of user stories like reading a sentiment, writing to a team, suggesting next action, chatting, checking satisfaction, so on and so forth. Uh those will be part of your backlog. So what would allow you to uh to get really good uh agentic results is by honing in on these user stories and making sure that you use these user stories as a means uh to elicit discussions with your stakeholders with your business and then work out what that user story should really be about. So the first user story uh second user story would be as a support lead I need to open cases ranked by urgency so that none of the escalations sh slip.

Speaker A: 所以只要确保每个用户故事涵盖这些,最好是、呃,写成这种结构,因为AI非常擅长模式识别,并且它实际上就是在用户故事结构上训练的,因为它是一个非常著名且被广泛使用的、呃,结构。所以如果你回到一些AI熟悉的东西,它会给、给你带来更好、更好的结果。并且每一个用户故事、呃,实际上都是由、呃、这些你知道的著名结构组成的,即人物角色(persona)、做什么(what,实际需求)以及为什么(why)。所以通过将这些打包并交给AI,显然还要附带验收标准,在此基础上你可以推导出测试用例,你将能够创建非常好的设置并获得非常好的、嗯,非常好的结果。

Original English

Speaker A: So just make sure that every user story covers these is ideally uh written in this setup because AI is really good at pattern recognition and it was actually trained on the user story structure because it's a very well known and wellused uh setup. So if you go back to something that's familiar to AI, it will get get you better better results. And every user story uh is actually made up with uh of these you know well-known structures the persona the what the actual need and the why. So by packaging these up and giving it to AI obviously with the acceptance criteria based on which you can derive the test cases you will be able to create very good setup and very good um very good results.

Speaker A: 然后,如果你只是将这些用户故事连接起来,把它们串联(daisy chain)在一起,那么这将允许你、呃,去创建一个连贯的系统,基于这个系统你可以创建你的规范,然后基本上就是你的代码。所以软件开发生命周期并没有因为AI而发生太大的改变。实际上是我们、呃、正在使用的工具箱在发生改变。对吧?所以当我们、呃,处理系统并且当我们思考我们想要构建什么时,我总是喜欢问这四个问题,即这是谁的问题?我们实际在解决谁的问题?这样我们可以、我们可以将其具体命名为一个直接的人、直接的角色、呃,并且它是非常量化的。

Original English

Speaker A: And then if you just connect these user stories, daisy chain them up, then that will allow you to uh to create a coherent system based on which you can create your specification and then essentially your code. So the software development life cycle doesn't change as much as a result of AI. It's actually the toolkit that we are uh we are using is changing. Right? So when we uh work with systems and when we think about what we want to build, I always like to ask these four questions is whose problem is this? Whose problem are we actually solving? So we can we can name it to a direct person, direct persona uh and it's very much quantified.

Speaker A: 对他们来说,赢是什么样子的?那么他们到底什么时候才算成功?他们实现了正确的结果吗?呃,我们能帮助他们以一种快速的方式、或者顺畅的方式、或者安全的方式实现那个正确的结果吗?以及什么会使、使他们拒绝使用它?在他们的平台上不可用。使用起来很繁琐。涉及到数据安全的方面。所以他们会、实际上不会使用它。而且它会改变一项决策吗?理想情况下,我们希望影响一个人如何做决策,并且我们想要,你知道,促使他们做出更好的决策。所以,它是否改变了决策,以及它改变的是什么决策?

Original English

Speaker A: What does winning look like for them? So when are they actually successful? Are they achieving the right outcome? Uh can we help them achieve that right outcome uh in a quick way or a smooth way or a safe way and what would that make make them refuse to use it? It's not available on their platform. It's cumbersome to use. It's the data security aspect applied. So they would wouldn't actually use it. And would it change a decision? Ideally, we want to be impacting how a person makes a decision and we'd want to, you know, tilt them to making better decisions. So, does it change a decision and and what is that decision that it changes?

Speaker A: 所以,一旦你能回答这四个问题,那么你就能从你的AI那里引发更好的响应,并且只要确保你在你的代码库中一个优秀的旧markdown文件中跟踪所有这些,以便AI可以访问它。它会从中获取更多的上下文,并且你知道,如果你只是做一些像“给我们构建一个处理支持的智能体”这样泛泛的事情,呃,你不会得到你想要的答案。所以我们总是做的是从价值出发。所以理解价值是如何被创造的,什么构成了价值,流程目前是如何流转的,支持该流程的底层架构是什么,然后你、然后你就可以开始实际的设计了,你可以在那里开始设计。

Original English

Speaker A: So, once you can answer these four questions, then you'll be able to elicit better responses from your AI and just make sure that you track all of these in a good old markdown file in your repository so that AI can access it. it will just get way more context out of it and you know if you just did something as generic as build us an agent that handles support uh you will not get the answer you want. So what we always do is go from value. So understand how value is created, what constitutes value, how the process currently flows, what is the underlying architecture beneath it that supports that process and then you and then you can start the actual design where you can start designing.

Speaker A: 所以我们喜欢把这、呃,这个思考过程称为VA,即价值架构设计(value architecture design),这就是我们想要始终去经历的。所以始终要把,你知道,价值记在心里。我们是如何创造价值的?我们正在创造的价值是什么?你的客户正在寻找的价值是什么?支持这一切的底层流程是什么?以及你如何能围绕它设计一个系统,使其最好地支持价值和流程,以及一路上需要进行哪些流程更改。

Original English

Speaker A: So we like to call this uh thinking process VA a value architecture design and this is what we want to always go through. So always have you know value in mind. How are we creating value? What is the value we are creating? What is the value that your customer is looking for? What is the underlying process that supports this and how you can design a system around it so it best supports the value and the process and what process changes are needed along the way.

Speaker A: 所以你可能会问,这不就是优秀的老式产品管理吗?在一定程度上,是的,这是一项古老的技能。这是一门值得重拾和学习的古老手艺,因为这现在正变成、呃,如果你愿意的话,一种模式,即你如何能够引出正确的需求,你如何能够构建更好的软件,因为我们都能访问相同的工具。所以区别将在于谁能更好地理解业务需求、呃,因为那样的话我们都可以只、呃,让最新最好的模型为我们编写代码。所以这是旧的技能,但是新的经济学(economics),并且这真正向分析师工具箱进行了转变。那么构建了错误的东西看起来是什么样呢,如果你加快了速度的话。

Original English

Speaker A: So you might ask, isn't this just good old product management? And to a certain extent, yes, it is an old skill. It is an old trade that is worth picking up and learning because this is now becoming uh the mode if you will of how you can elicit the right requirements, how you can build better software because we all have access to the same tools. So the difference will be who can understand the business need better uh because then we can all just uh have the latest and greatest model write the code for us. So it's old skill but new e economics and it's a real shift towards analyst toolkit. So what building the wrong thing looks like if you've got velocity up

自动化AI研究介绍

Elie: 嘿,嗯,大家好,呃,感谢大家来到这里,呃,是的,我今天非常高兴能来谈论,呃,自动化的AI研究,以及、呃,特别是、呃,所有那些像、基础模型、呃,在、呃,自动化研究任务上的表现。嗯,所以我是Elie。我在Prime担任研究工程师,而且,呃,是的、呃,我将向大家展示我们在这个、在这个主题上的工作。

Original English

Elie: hey um hi everyone uh thanks for being here uh yeah I'm super happy today to talk about uh automated eye research and uh especially uh all those like font model uh perform at uh automated research task. Um so I'm Elie. I work at prime as a research engineer and uh yeah I will go through our work on on this subject.

Elie: 所以首先我想要,基本上是稍微解释一下为什么我们在做这件事,以及、呃,为什么我们认为公开地做这件事是极其重要的。嗯,所以首先,呃,我想、呃,我、我们都同意、呃,我们已经听说过、呃,像那些大实验室在说,这种糟糕的、呃,称为递归自我改进的事情,呃,很快就会到来。呃,所以递归、呃,自我改进就是像模型训练、呃,模型、呃,基本上是没有、呃,人类的干预。

Original English

Elie: So first I want to basically explain a bit why we are doing that and why we think it's super important to do that in the open. Um so first uh I think we we all agree that uh we've heard about like big labs saying that this bad thing called recursive self-improvement is coming very soon. Uh so recursive self-improvement is like model training models uh without uh human intervention basically.

Elie: 嗯,但是,呃,我们没有任何基准来基本上量化这是、呃,这是真的还是假的,对吧。呃,而且更少见的是我们没有像,由非、大实验室进行的第三方、呃,基准测试,来、来、来看看、呃,这是否是很快就会到来的事情。呃,另一部分是,我们认为,呃,了解、所有这些模型、呃,做研究是极其重要的,因为、呃,我们认为很多、在、未来几年将要出现的科学、呃,研究、呃,也将基于AI工具。

Original English

Elie: Um but uh we don't have any benchmark to basically quantify if this is true or not right. Uh and even less we don't have like a third party benchmark by non- big labs to to to see if it's something coming soon or not. And the other part is that we think that uh it's super important to understand all those model uh do research because we think that a lot of the scientific research that will come into the coming years uh will be based also on AI tools.

Elie: 所以了解、这些模型是如何做研究的是极其重要的,不仅仅、仅仅是AI研究。所以我们试图建立、这种环境,来测试模型的、呃,做这方面的能力。所以、呃,这一切都始于、呃,Andrej Karpathy,呃、这基本上是在做这个、呃,视频时获得了乐趣,在那里他从、呃,头开始、在大概90分钟内训练了、呃,GPT-2。就像GPT-2的、呃,训练要花费大概几周时间,而且不,在两年前,我想它只花费了大概90、呃,分钟。

Original English

Elie: So it's super important to understand how those model do research not just only AI research. So we try to build kind of this environment to test the capabilities of the model to do so. So it all started with uh Andre Karpati uh that's basically had fun by doing this video where he trained uh GPT2 from scratch in like 90 minutes like GPT2 training takes like weeks and no in two years ago I think it only took like 90 minutes.

Elie: 那么在90分钟内复现、呃,GPT-2意味着什么呢?它意味着在90分钟内你达到了、呃,这个目标损失(loss)。嗯,是的,然后就是、在这个节点上当——

Original English

Elie: So what does it mean to reprod reproduce uh GPD2 in 90 minutes? It means that in 90 minutes you achieve this target loss. Um and yeah and that's at this point when

自动化 AI 研究与竞速环境

Speaker: 当你达到与 GPT-2 相同的损失值时,你就可以认为你的模型性能大致相当了。接下来发生的事情是,社区接手了这个代码库(这个 GitHub 仓库),并创建了另一个叫做 modded nano GPT 的项目,这项工作是由一个叫 Keller Jordan 的人主导的。他们基本上把这个时间从 90 分钟缩短到了 45 分钟,然后大家发现,不,我们可以在不到两分钟的时间内训练出一个达到 GPT-2 验证损失水平的模型,老实说这太疯狂了,而且花了两年的时间才实现这一点。所以这是一个非常强大的基准测试,许多非常有才华的研究人员都在上面进行了迭代。是的,因此我们决定采用这个竞速通关(speedrun)的环境。这有点像个游戏,游戏的目标是在最短的时间内达到这个损失值。所以这是第一代 nano GPT,你几乎没有任何限制,唯一的限制是你必须使用相同的验证和训练数据,对吧?

Original English

Speaker: you have the same loss than GPT2 you consider that your model is somewhat of equal performance. Um then what happened is that the community took this repo uh this GitHub repo and create another one called modded nano GPT and this effort was leaded by someone called Keller Jordan. And what happened is that they basically took this 90 minutes then 45 minutes and then no we can train like GPT2 validation loss model in less than two minutes which is honestly crazy and it took like two years to to achieve this. So it's a very strong benchmark where uh a lot of very talented researcher iterated on um yeah so we decided to take this environment of speedun so it's kind of a game so the goal of the game is to achieve this loss in the fewest in the shortest amount of time so this is the nano GPT1 and you can uh you don't have almost any constraints the only constraint that you shots that you need to use the same validation and training data, right?

Speaker: 几个月前发布了一个新的竞速挑战,叫做“优化器竞速”(optimizer speedrun)。这个挑战略有不同,因为你只能更改与优化器相关的参数。例如在 nano GPT 中,你可以改变架构,调整注意力机制之类的;但在优化器竞速中,你只能将优化器从 Adam 换成 m-shampoo,或者任何你喜欢的其他优化器。是的,因此这更具研究性质,因为它不是关于如何优化程序以使其尽可能快地运行,而更像是找到尽可能好的方法,无论你在计算机上投入了多少时间,对吧?

Original English

Speaker: Um there is a new speedrun called the optimizer speedrun that was released uh a few months ago and here it's slightly different because uh you can only change the optimizer related parameters. So for instance nano GPT you can change the architecture uh doe do uh attention whatever uh optimizer sp you can only change like Adam to m shampoo or whatever optimizer is your favorite um yeah and so this is a bit more researchy because uh it's less about optimizing the program to be as fast as possible but more like finding the best method possible. no matter the the the time you put into the computer, right?

Speaker: 所以,为什么要将竞速通关作为自动化 AI 研究的环境呢?首先,我们认为这是一个很好的评估方法,稍后我们会看到原因。这也是本次演讲的主要焦点。但我们也认为它可能是一个很好的训练环境,因为这是给模型提供奖励的一种方式。如果模型完成了竞速并打破了之前的记录(抱歉口误),奖励就是正的;如果它没能做到,奖励就是零或负的。因此它是一个用来训练模型的好环境。它也非常快,就像你看到的,之前优化器竞速的记录大约是两分钟。每次运行大约需要 15 到 20 分钟,是的,并且基本上有明确的规则。我们还认为它是一个进行新发现的好环境,比如在我们的研究中取得某种突破,因为有那些可以去验证的明确规则。是的,所以就是这样。

Original English

Speaker: So, um yeah, why take speedrun as an environment for automated AI research? First, uh we think that it's a good evaluation. We'll see later why. Uh and this is kind of the main focus of this talk. But we also think it's probably a good training environment because uh it's a way to give the model a reward. So the reward is positive if the model bit the speed run and beat the last record sorry and the reward is zero or negative if it didn't manage to to do it. So it's a good environment to train model. It's also quite fast like as you see uh previous record were around 2 minutes for the optimizer one. uh each run take about like 15 to 20 minutes and uh yeah and there is like clear rules basically and we also think it's like a good environment to make discovery so like kind of breakthrough in our research because uh there is those clear rule that you can verify or not. Um yeah. So yeah.

AI Agents in Speedrun Challenge

Speaker: 我们做的是,大概在两个月前发布了这个优化器竞速项目,我们决定推出两个 AI 智能体来与社区竞争。这两个智能体是 Codex 和 Cloud Code(Claude Code)。Codex 类似于带有 XI 的 GPT 5.5,而 Cloud Code 则是带有 XI 的 Opus 4.8。是的,我们决定基本上让这两个智能体在我们的集群上自由运行,并不断迭代。所以我们有 V1、V2 版本,而 V3 版本基本上就是我们停止智能体然后重新启动。V3 大概是在发布前的一两天进行的,因为我们看到我们的智能体不再保持最佳记录了。所以我们觉得,好吧,把过去几周人类的所有记录都拿过来,试着在上面进行改进,结果它成功了。是的。我们还有一个“新颖性”赛道(novelty track),目标是仅凭全新的想法来打破记录。稍后我们会看到,这对模型来说要复杂得多。

Original English

Speaker: Um so what we did uh so the release was like about two months ago and uh there was this optimizer speedrun and we decided to basically compete with the community by launching two AI agents. So Codex and Cloud Code. Codex was like GPT 5.5 with XI and uh cloud code was Opus 4.8 with XI. Um and yeah, we decided to basically let the agent free on our cluster uh and uh and just iterate on it. So we have like V1, V2, V3 is just basically us stopping the agent and then restarting. V3 uh was like one or two day before the release because we saw that our agents no longer have the best record. So we were like okay take all the the human uh record in the last few week and just try to to improve upon it and and and it worked. Yeah. And we also have this novelty track where the goal is to uh beat the record with only novel ideas. Um and we'll see that this this was more complex for the the models.

Speaker: 因此,我们的 RS 非常简单。老实说,我们本来可以直接用 slashgo 来替换它,但当时还没有 SLG 目标(slash go goal)。所以,我们制定了自己的 goal.md。其实很有趣,我们选择了相同的名字,我们有了 goal.md 和能够定义规则的各种智能体。我们让智能体提出各种想法,然后它可以通过 sbatch 在我们的 Slurm 集群上提交任务。基本的工作原理是,它可以向可用的节点提交任务,但必须在特定的权限下。这意味着如果有人想要使用这个节点,模型就会取消任务,这被称为抢占式权限(preemptable permission)。是的,然后它会读取训练日志,判断是否创造了新记录。要验证一个记录,你基本上需要通过一个统计阈值,以确保它不仅仅是表面上的优化,或者是随机的结果。对吧?

Original English

Speaker: So our RS is very simple. Honestly, we could have just replaced it with slashgo, but they there was no SLG goal at the time. So, we made our own goal. MD. It's actually quite fun that we choose the same name and we had the goal. MD and kind of agents that MD that define the rules and we let the agent propose ids and then it can submit a job with sbatch on our slum cluster and uh basically the way it works is that it can submit on nodes that are available but only under a certain permission which means that if someone want to use this node uh the model just like cancel the job it's called preemptable permission. So yeah, then it measure the it read basically the training logs then decide if it's a record or not. To validate a record you need to basically pass a statistical threshold to make sure that it's just not see the optimization and is just not random. Right?

Experimental Results

Speaker: 那么,来看看这个实验的一些结果。第一个老实说让人非常头疼的结果是,Cloud Code 每隔九到十个小时就会停下来,然后基本上就说,是的,我无法改进记录了,这对我来说太难了,我不可能再超越它了。然后我就会说,好吧,继续探索新方向,嘿,再跑 10 个小时吧。然后它又会说,是的,我无法打破记录了,如此反复。所以基本上 Cloud Code 有三分之一的时间处于空闲状态,因为我没有办法去实时监控它。而 Codex 则完全相反,它一直都在工作,几乎从不闲置,从不提问,在这一点上非常令人印象深刻。

Original English

Speaker: So yeah a few results from this experiment. The first one that was honestly very painful to work with is that code uh clothes code keep stopping every nine or 10 hours and basically said yeah I cannot improve the record it's too hard for me there is no way to to go beyond it and then I was just like okay continue explore new direction hey just go again for 10 hours and then say yeah I cannot beat the recall and so on. So basically onethird of the time the cloud code agent was idle because I had no way to basically monitor it and codeex totally the opposite just worked for all the all the time and uh yeah almost never idle never asked for question and and and very impressive in that way.

Speaker: 我们还给了模型一个选项,让它把大量的东西写进我们称之为“便签本”(scratch pad)的地方,这基本上就是模型的活动内存。我们观察到,Codex 在便签本上写了非常多的东西。我接下来要展示的每一个图表,都在某种程度上根据活跃订单的数量进行了标准化。所以这不仅仅是因为 Codex 做了更多工作,这真的是完全不同的行为模式。你可以看到它向这个便签本、内存中写入了多得多的内容。而且,每个文件的基调也非常不同。比如 Cloud Code 在获得新记录时会非常兴奋,用一堆表情符号之类的;而 Codex 则非常像个机器人,只是在说“这是我所做的,这是我做的决定,接下来我要做这个”,非常机械化的感觉。是的。

Original English

Speaker: Um we also give the option for the model to basically write uh a bunch of stuff into what we call a scratch pad which is basically the active memory of the model. Uh we observe that basically codeex writes a lot on the scratch patch. So each plot that I will show are kind of normalized by the number of active order. So this is not only about codex working more it's it's really different behavior. So yeah, you see that uh writes a lot more to to this scratch pad to this memory and uh the shape of the like the the I don't know the tone of the the each file was also super different like CL was super excited about getting new record with a bunch of emoji and so on and CEX was just like here is what I do here is the decision I take what I will do next like super robotic kind of um Yeah,

Speaker: 我们还有这个图表,通过它我们看到 Codex 生成的子智能体远多于 Cloud Code。我们发现 Codex 消耗的 token 也比 Cloud Code 多得多。所以我想总计大概有十亿级别的 token,但显然这里有输入 token 缓存的作用,所以这并不等同于输出了十亿个 token。我们还看到 Codex 进行了大量的上下文压缩,因为它只有大约 250k 的上下文窗口;而 Cloud Code 甚至在整个运行过程中不到一次,而 Codex 大概是每小时 20 次。

Original English

Speaker: we also have this plot where basically we saw that codex was spawning much more sub aents than cloud. Uh we saw that codex burn much more token than code. So I think in total it was like kind billion of token but it's like there is obviously this input tok uh input caching that make it it's not like one billion output token. Uh so yeah we also see that codex did a lot of compaction because it only had like 250k context window and cloud only do it like one per hour and codex is more like no it's even less than one power for I mean one for the full run for cloud and codex was like one uh was 20 every one hour.

Speaker: 所以这就是主要的结果。这个图表显示的是,白色部分代表人类记录的进展,红色部分代表 Cloud Code(本来应该是橙色的,但无所谓了),蓝色部分代表 Codex。你可以看到,几乎在所有时刻,Cloud Code 和 Codex 的表现都优于人类记录,而 Cloud Code 在一开始表现得超级好,非常非常快地获得了很高的分数。还有一件非常重要的事情是,模型有能力随时获取人类的记录。这就是 Codex 所做的,也是 Cloud Code 所做的,抱歉。因为当我重新启动它时,它基本上就是去获取了人类的最新记录,并在其基础上进行改进。

Original English

Speaker: So yeah um yeah here is the main results. So what this plot shows is that basically we so in in white you see that the human recall progression right and in red you see cloud I mean it's supposed to be orange but whatever and in blue uh you see codeex right and you see that at almost every time uh cloud and codex are better than the human record and code is super good at the beginning very very fast to achieve very good score. Um, yeah, and one thing that is super important is that the model have the ability to basically fetch the human records at any time and that's what Codex did. That's what cloud did. Sorry. Because when I restarted it, it basically fetch the new record from human and improve upon it.

Future Benchmarks

Speaker: 结果就是,我认为当时最好的记录大概是 2990 步,而我们的 Cloud Code 以大约 50 或 60 步的优势打破了它,Codex 则领先了大概 20 步。我认为这确实令人印象深刻。所以这个结果目前还没有发布。这是我们目前正在进行的工作。核心的想法是,这是一个很酷的实验,但它缺乏结构化。如果你想做一个真正的基准测试,你需要使用多个随机种子,你需要用正确的方法,基本上把所有的模型认真地置于相同的条件下来测试,对吧?

Original English

Speaker: Um, yeah. So the result is that uh I think at the time the best record was like uh 2,990 step and we beat it by like uh uh 50 or 60 step for code and codeex was like 20 step above. So I think it's both impressive and and yeah um so we so this is like not released yet. This is something that we are working on currently and basically the idea is that this is a cool experiment to do but it lack of structured right. uh if you want to do a real benchmark, you want to do multiple seed, you want to do uh yeah proper uh thing where you you you you basically put all the model and earnestness in the same condition, right?

Speaker: 所以这就是我们现在正在做的事情。基本的想法是设置三个不同的赛道:一个没有任何外部访问权限,真正衡量模型仅依靠其权重知识进行 AI 研究的能力;一个是仅能访问 arXiv 论文的赛道;还有一个是拥有完全访问权限的赛道,它也可以访问人类最新的记录。为此,我们计划同时运行 nano GPT 赛道(也就是最初的那个赛道)和……

Original English

Speaker: So this is what we are working on right now and basically um the idea is to do three different track uh one without any access to really like measure the capability of the models to do AI research based on only the model weight knowledge one with only archive paper and one with like full access. So it also have access to the the like the latest record by human. And for this we plan to do both uh the nano GPT track one which is the original one and

优化器速度运行结果

Speaker A: 关于优化器速度运行,我们基本上只限制优化器必须是新颖的。是的,所以我将展示一些关于优化器速度运行的结果。这基本上就是我们得到的。我们让智能体迭代了六天,可以说差不多五天,我们看到 Codex、K 和 Claude 都非常有效。对于 GLM 来说,这不是一个完成的运行,模型现在实际上仍在集群上迭代,但我们看到 Claude 再次表现得非常出色。而且令人惊讶的是,Kimi 也非常具有竞争力,并在第四天取得了突破,可以说用新纪录打破了 Codex 的纪录。有趣的是,Claude 在改进记录的方式上更加渐进,而 Kimi 则真正表现出一种阶跃函数般的突破模式。这是一个有趣的图表,因为六天对于评估来说已经相当长了,但你也可以将这个轴改为输出 token 的数量,然后会讲述一个不同的故事,因为最大模式消耗的 token 比 Codex 和 Kimi 多得多,你还会发现 Kimi 在其使用的 token 数量上实际上非常高效。所以这是 Kimi 的代码。是的,我们还看到它们在利用文献和论文方面有不同的方法。例如,Claude 在论文上做了大量搜索,实际上找到了一篇其他模型都没有找到的论文,并且这实际上促成了最好的记录。这有点有趣。

Original English

Speaker A: the optimizer speedrun where we only launch uh we only constrain the optimizer to be novel basically. Um yeah so I will present some result on the optimizer speedrun. Uh this is basically what we got. So we let the agent iterate for six day almost five days let's say and we see that uh codeex k and clothes uh are super effective so for GLM this is not finished run right so the model is actually still iterating on the cluster right now but we see that cloud is once again very good at it and we see that surprisingly Kim is also very competitive and kind of have this breakthrough on day four where he kind of beat Codex with a new record, right? It's also interesting to see that uh Claude is much more like progressive in the way it improved the record and Kim has really this step function where I kind of do a breakthrough and so on. Uh so this is an interesting plot because I mean six day is quite a lot for anal uh uh but you you can change this uh axis by also the number of output token and then kind of tell a different story because in max mode consumes so much more token than codeex and Kimmy and you also see that Kimmy is actually super efficient uh for the number of token that uh it uses. So it's schemic K2.7 code. Um so yeah uh we also see that they have a different approach to uh using the literature and papers. Um so for instance like code is doing a lot of search on papers and actually include found a paper that no other model found and it actually lead to the best record. So it's kind of funny and uh yeah

Speaker A: 这一切的主要问题之一是,当我启动这些不同的智能体时——我想这也是我希望大家在这次演讲中记住的重要一点——当我启动这些智能体时,我原本期望它们能想出一些没人发现过的关于优化器的疯狂点子,但老实说,事实并非如此。它们使用了一些聪明的小技巧,基本上是结合了不同的论文,对许多方法做了一些微小的改进,但这些模型并没有真正提出任何新颖的优化器或机制。我认为这很能说明问题,即使在一些并不简单但对人类研究人员来说是可及的事情上(比如花上几天甚至几周的时间),模型也无法发现新的优化器和机制。

Original English

Speaker A: um one of the main issue of all of this is that uh when I when I launched this this agent and I think that's something important that I want you to to kind of uh remember for this co this talk is that when I launched this these different agents I was expecting them to come up with some crazy ideas on optimizer that's like no one of discover but honestly it wasn't the case. Uh they did some clever trick where basically they combine different papers. uh they kind of do plus one improvement over a bunch of method but there was really like no novel optimizer or mechanism that was uh coming from those model and I think that's kind of telling that even on something that is not simple but I'd say that it's kind of accessible for people right for like human researcher uh spending like days and weeks for the the model like cannot like find new uh optimizer and mechanism.

面向发现的多智能体系统

Speaker A: 因此,我们相信有一种方法可以让它更适合发现,而不仅仅是评估。这很大程度上受到了谷歌 Alpha Evolve 的启发,以及此后发表的一系列论文的启发。这是一种多智能体系统,它们相互交互,有很多生成器。你既有闭源模型,也有在这里性价比超高的开源模型。它们可以提出想法,然后你运行速度测试从而获得奖励,接着会有一个评委基本给出质量反馈。你也可以让评委对该方法是否优秀有一定的“品味”判断。如果它在循环之外,那么你基本上可以决定要将哪种方法扩展到更多的参数和 token 数量上。这就是速度运行的规模扩展部分,因为在速度运行社区中,人们经常说很多方法在大规模下行不通。所以我认为在这个循环中加入规模元素也是非常重要的。而且我认为人类在这里也非常有用,基本上可以判断智能体的想法,将它们引导到正确的方向等等。

Original English

Speaker A: So we believe that there is a way to basically make it more um make it better for discovery instead of evaluation. And this is coming from uh this is very inspired from alpha evolve by Google and also a bunch of papers that have been released since then. It's kind of this multi-aent system that interact together uh bunch of generator. You have closed model but you also have open source model here that are super effective for the cost right. Uh they can suggest ideas then you run the speedrun so you get the reward then you have a judge that basically give a quality feedback can also be like the judge also have this taste. you can kind of have like the judge have a taste about the the method if it's good or not. Uh if it's outside the loop and then you can uh basically decide which method you want to scale to a larger number of parameters and number of token. Um so this is kind of the scale part of the speedrun because some a lot of method in the the speedrun community uh people are often saying that they doesn't work at large scale. So I think it's very important to also put scale elements in this loop. Uh and I think also that uh human are super useful here to basically judge the ID of agents kind of steer them in the right direction and so on.

Speaker A: 所以我们还没有完全尝试,我的意思是我们现在正在尝试,并希望这至少能带来人工智能研究的新发现。另一种方法是你可以定义多个速度运行。所以如果你喜欢的话,这是下一张幻灯片,它来自 Safe Bank 的幻灯片,如果你不知道这个梗,那对你来说是件好事,说明你没怎么在网上冲浪。但其核心理念是,通过改变速度运行的目标和约束,你基本上可以创造大量的多样性,并限制模型朝着特定的方向发展,从而做出那些发现。因此,在 Hint,我们在这个方向上做了很多事情。有很多东西我们大部分还没有发布,但我们正在开发 GPU 沙盒,以允许模型在沙盒中迭代,因为你做这种事情需要 GPU 沙盒。我们正在开发自己的智能体,它们对框架非常高效。这意味着你有一个文件系统,你可以写入信息、从中读取信息。你还可以进行编程式工具编码之类的事情。我们还在开源模型的基础上训练一个擅长这些的模型。而我们已经发布的内容是,我们有一套库和名为 Verifier 的强化学习训练产品,你可以基本上在任何环境和 RL 上进行训练和评估,而且你可以训练的模型可以像 GLM 5.2 这样非常大。是的,我们花了很多精力使这些库非常高效,以便为我们的客户提供最好的质量。对于这个领域我感到非常兴奋。我再次认为,让这种递归自我改进的一部分在开源环境中发生非常重要,因为实际上有很多人不在大型实验室工作。所以你基本上需要让人们更容易理解这些模型是如何运作的,以便进行研究等等。这就是我们的目标,非常感谢。

Original English

Speaker A: Um yeah so we didn't try it yet I mean we are kind of trying it right now and uh we hope that this will lead to to to new discovery in AI research at least and also a way is that you can define multiple speedrun so this is the next slide if you like it's from safe bank slides but if you if you don't have the reference good for you means that that you're not too online uh but the idea is that uh by changing the object objective and the constraints of the speedrun you can basically create a lot of diversity and constrain the model to go into a certain direction and uh yeah and make those discovery. So uh at hint we are doing a bunch of stuff in this direction. Uh there is a bunch of stuff here that we I mean most of it we didn't release yet but we are working on GPU sandboxing to allow model to iterate into sandbox because you need GPU sandbox for this kind of stuff. We are working on our own agents that are very efficient for like framework. So it means like you have a five system and you can write information read from it. Uh and you also do like this programmatic tool coding thing. We also training a model to be good at it on top of like open source model. And uh the thing that we already released is that we have the set of liber and product called verifier primar training where you can basically train evaluate any environments on any RS and the model that you can train can be like GNM 5.2 too which is very big and and yeah we have like we work a lot on making those li very efficient to to ship the best quality for for our clients. Yeah. Uh I mean yeah super excited about this domain. Once again I think it's super important to have uh a part of like this recursive self-improvement to happen in the open because there is actually a lot of people working that are not on big labs. So you need to basically uh yeah make it easy for people to understand all those model work to do research and so on. So that's kind of our goal and uh yeah, thanks a lot

从模型评估到系统行为评估

Speaker B: 我是 Meta 的一名软件工程技术主管,致力于为 Meta Super Tangent 实验室及其基础设施组织构建训练和推理基础设施。今天我们将讨论基于智能体系统的生产。当大多数人听到评估这个词时,他们会想到基准测试。一个模型在基准测试中得分 90%。新版本得分 92%。团队为此庆祝。但智能体系统从根本上改变了评估的含义。今天的系统不仅仅是生成答案。它们会计划,会调用工具,会检索信息。它们执行工作流。它们与生产基础设施交互。问题不再是“模型是否生成了正确的答案”,而是“系统行为是否正确”。今天我想讨论一下评估是如何从模型基准测试演变为生产基础设施的。

Original English

Speaker B: and I'm a software engineing Tech League at Meta working on building a training and inference infrastructure for the meta super tangent lab and their infrastructure organization. Today we're going to be talking about productions for aentic systems. When most people hear the word valuation, they think about benchmarks. A model scores 90% on a benchmark. A new version scores 92%. The team celebrates. But agent systems have fundamentally changed what the evaluation means. Today the systems don't simply generate answers. They plan, they call tools, they retrieve information. They execute workflows. They interact with the production infrastructure. The question is no longer did the model generate the right answer. The question is did the system behave correctly. Today I would like to discuss how evaluation is evolving from model benchmarking into production infrastructure.

Speaker B: 这几乎是当今每个 AI 组织都遇到的问题。离线基准测试在不断提高。然而生产环境的可靠性往往仍然不可预测。为什么会这样?因为基准测试衡量的是模型的能力。生产环境衡量的是系统行为。基准测试无法捕捉工具故障、API 中断、上下文变化、用户差异性以及长时间运行的工作流。随着系统变得越来越自主,基准测试性能与生产性能之间的差距越来越大。结果就是如今许多团队所经历的:正如你们所见,基准得分很高,但生产行为却不可靠。传统的评估侧重于输出。但我们应该问这样一个问题:模型产生了正确的答案吗?智能体系统迫使我们提出一个不同的问题:系统行为是否正确?行为包括规划质量、工具使用、执行、工作流执行、从故障中恢复以及决策制定。换句话说,我们正在从评估答案转向评估工作流。这需要根本不同的评估架构。

Original English

Speaker B: This is the problem almost every AI organization is encountering today. Offline benchmarks continue improving. Yet production reliability often remains unpredictable. Why is that? Because benchmarks measure model capability. Production measures system behavior. A benchmark doesn't capture tool failure, API outage, context changes, user variability, longunning workflows. And as systems become more autonomous, the gap between the benchmark performance and production performance grows. The result is what many teams experience today. High benchmark scores as you can see, but unreliable production behavior. Traditional evaluation focus on outputs. But we should ask the question, did the model produce a correct answer? Agentic systems force us to ask a different question. Did the system behave correctly? Behavior includes planning quality, tool usage, execution, workflow execution, recovery from failures, decision making. In other words, we are moving from evaluating answers to evaluating workflows. And that requires fundamentally different evaluation architectures.

Speaker B: 许多团队仍然认为幻觉是 AI 的主要故障模式。在生产中,它们通常只是一类。智能体系统引入了整个故障模式层级。在最底层是记忆故障、检索故障、安全故障。往上走,你必须考虑推理错误、糟糕的规划、不正确的工具执行。在最高层,你必须考虑多智能体协调故障。这就是为什么仅评估模型输出会遗漏我们观察到的大多数生产风险。最有用的思维转变之一就是停止像研究人员那样思考,开始像 SRE(站点可靠性工程师)或生产工程师那样思考。SRE 不衡量……

Original English

Speaker B: Many teams still think hallucinations are the primary AI failure modes. In production, they are often just one category. Agentic systems introduce an entire hierarchy of failure modes. At the very foundation the memory failures, retrieable failures, safety failures. As you go up, you have to think about reasoning mistakes, poor planning, incorrect tool execution. At the highest layer, you have to think about multi-aent coordination failures. And this is why evaluating only model output misses the most production risks we observe. One of the most useful mindset shifts is to stop thinking like researchers and start thinking like a SR or a production engineer. S SR don't measure

评估系统的演进与最佳实践

Presenter: 成功取决于准确性。他们衡量可靠性、可用性、延迟、成本回收,而智能体系统需要采用完全相同的方法。我们的目标不是最大化基准测试的分数。真正的目标是最大化可靠的结果。可靠性成为了北极星指标(northstar metric)。准确性则仅仅成为了输入。这个金字塔模型展示了我个人是如何思考现代 AI 评估系统的。在金字塔底部,你可以看到基准测试。它们很有用。它们具备可扩展性。它们也很有声誉。但是,它们在实际操作中的价值是有限的。在中间层,是基于场景的评估。这些评估模拟了真实的工作流。而在最顶端,你看到的是生产环境的遥测数据。这是价值最高的评估信号的来源。令人惊讶的一个洞察是,大部分评估数据通常来自于真实用户与真实系统的实际交互。

Original English

Presenter: success using accuracy. They measure reliability, availability, latency, cost recovery and agentic systems require the same approach. The goal is not maximizing the benchmark scores. The goal is to maximize dependable outcomes. Reliability becomes the northstar metric. Accuracy becomes the only input. In this pyramid is how I think personally think about modern AI evaluation systems. At the bottom you can see there are benchmarks. They're useful. They're scalable. They're reputable. But the operational value is limited. In the middle there scenario based valuations. These simulate realistic workflows. And at the very top you see production telemetry. This is where the highest value valuation signals come from. The surprising insight is that the most evaluation data often comes from real users interacting with real systems.

Presenter: 现在让我们来谈谈离线评估(offline vals)。离线评估依然重要,但方法论已经发生了改变。我们不再去评估提示词(prompts),而是去评估具体的场景。例如,一个客户支持的工作流、一个代码生成的工作流,或者一个研究工作流。智能体在模拟环境中运行。我们测量任务完成率、工具调用的正确性、规划的质量,以及资源使用情况——在极大规模下,资源消耗会呈指数级增长。关键要点18是:评估应该是场景驱动的,而不是提示词驱动的。一旦一个系统进入了生产环境,每一次交互都会成为一个信号。这是评估思维中最大的转变之一。生产环境的流量不再仅仅是流量。它变成了评估数据。我们收集执行追踪记录(execution traces)、用户结果、升级反馈(escalations)、失败案例以及反馈信号。生产环境是任何组织所能拥有的最大、也最具有代表性的验证数据。

Original English

Presenter: Now let's talk about offline vals. So offline evaluation still matters but the methodology changes. Instead of evaluating prompts we evaluate scenarios. For example, a customer support workflow, a code generation workflow, a research workflow. The agent operates inside the simulated environment. We measure the task completion rate, tool correctness, planning quality, resource usage which is which becomes exponentially high at high scale. The key takeaway 18 evaluation should be scenario driven not prom driven. Once a system reaches production, every interaction becomes a signal. This is one of the biggest shifts in evaluation thinking. Production traffic is no longer just traffic. It becomes evaluation data. We collect execution traces, user outcomes, escalations, failures, feedback signals. Production is the largest and the most representative validation data any organization will ever have.

Presenter: 许多组织将人类视为后备系统。我认为这是一种错误的框架设定。人类是评估者。他们提供了自动化系统无法提供的信号。他们评估正确性、信任度、实用性和安全性。这些信号对于校准评估管道(evaluation pipelines)和识别自动化指标中的盲点变得至关重要。最成功的系统将自动化评估与有针对性的人工审查结合在一起。现在,智能体系统处于持续的漂移状态。模型会改变。我们每隔几周或几个月就会有一个新版本。提示词会改变。工具会改变。用户的行为也会改变。面临的挑战在于,不再是某一个单一的改变显得具有灾难性。而是可靠性在缓慢下降。成功率在下滑。需要人工介入的升级事件在增加。工具失败率在上升。如果没有持续的评估,团队通常要在用户抱怨之后才会发现这种系统漂移。

Original English

Presenter: Many organizations view humans as fallback systems. I think that's a wrong framing. Humans are the evaluators. They provide signals that automated systems cannot. They assess correctness, trust, usefulness, safety. These signals become really critical for calibrating evaluation pipelines and identifying blind spots in automated metrics. The most successful systems combine automated valuation with targeted human review. Now, agent systems drift constantly. Model changes. We have a new version every couple of weeks or months. The prompts can change. Tools can change. User behavior can change. The challenge is that no longer a single change appear catastrophic. Reliability slowly degrades. Success rate declines. Escalation increases. Tool failure rises. Without continuous evaluation, teams often don't discover drift until users complain.

Presenter: 持续的监控变得至关重要。可观测性(Observability)和评估是不可分割的。完全不可分割。要评估一个智能体,我们需要深入了解其推理路径。包括工具调用、内存访问、执行时间线,以及状态之间的直接转换。正如你在这张图表中看到的,传统的日志已经不够用了。我们需要详细的追踪记录(traces),就像任何深度嵌套的微服务架构针对任何应用程序或服务所做的那样。我们所说的是,智能体追踪(agent traces)将成为自主工作负载的分布式追踪的等价物。如果没有可观测性,评估就变成了靠猜。

Original English

Presenter: Continuous monitoring becomes essential. Observability and evaluation are inseparable. Inseparable. To evaluate an agent, we need visibility into the reasoning paths. The tool calls, the memory access, execution timelines, the straight transitions. As you can see here in this chart, traditional logs are not sufficient. We need detailed traces just like with any deep nested microser architecture for any application or service. We're talking about agent traces become the equivalent of distributed tracing for autonomous workloads. Without observability, evaluation becomes the guesswork.

Presenter: 现在让我们来谈谈持续评估循环,因为评估是一个始终在运行的服务,而不是一个阶段性的测试。在过去,评估总是发生在部署之前,但现在评估在部署之后仍在继续。遥测数据能够识别问题,正如你所见,人类会审查那些边缘情况(edge cases)。反馈能够改进数据集。离线场景用于验证更新。这个循环永不停止。评估不再仅仅是一个阶段。它是一项操作能力。那么,这可能是本次演讲中最重要的一张幻灯片。这里展示的每一个指标都映射到一个业务结果:任务完成。

Original English

Presenter: Now let's talk about the continuous evaluation loop because evaluation is an always running service not a testing phase. Historically evaluation always happened before deployment but now evaluation continues after deployment. Telemetry identifies issues as you can see in a human reviews the edge cases. Feedback improves the data sets. Offline scenarios validate updates. The loop never stops. Evaluation is no longer just a phase. It's an operational capability. Now, this is probably the most important slide in this presentation. Every metric shown here maps to a business outcome. Task complete.

小组讨论:模型发布与 API

Host: 好的,我想我们现在已经上线了,欢迎正在看直播的观众以及现场的朋友们回来。嗯,我们通常倾向于在主要的主题演讲之间安排这些较长的环节,以此来反思一些尤为重要但又没有那种重大发布时刻的事情。今天我们非常幸运,能有致力于 Omni 和 Vo Nano Banana 项目的成员来到这里,要知道,目前世界上最棒的生成式模型就在我们身边。呃,Demetrio,我第一次看到你的时候,是你发了一些关于你办公室的帖子,嗯,我觉得你可能是第一,呃,谷歌的头号办公室网红(office influencer),至少在旧金山肯定是。我觉得你也挺喜欢骑自行车的。你喜欢在这里拍自行车的照片。

Original English

Host: >> Okay, I think we're live and welcome back for those on the stream and those those in person. um we take tend to basically take these longer sessions between uh all the sort of mainstage keynotes to reflect on things that um you know are particularly important but like don't have like a significant like a sort of launch moment. Today we're very lucky to have people working on Omni and Vo Nano Banana like the you know the world's best generative models here with us. Uh Demetrio I I I first saw you when you were posting about your office Um, I think you're you're probably number one uh Google Google's number one office influencer at least in in San Francisco. I think you like you like to bike as well. You like to take photos of bike here.

Demetrio: 是的。

Original English

Demetrio: >> Yeah.

Host: 嗯,不过你知道,而且你也从事视频模型的工作。

Original English

Host: Um, but you know, but also you work on video models.

Demetrio: 没错。

Original English

Demetrio: >> That's right.

Host: 呃,Shane,我记得好像是在一次晚宴上认识你的。

Original English

Host: >> Um, Shane, I I met you I think at like a dinner.

Shane: 是的。

Original English

Shane: >> Yeah.

Host: 嗯,而且,呃,而且,我记得你当时试图让我投资其中一家公司。我忘了是哪一家了。算了,别提那事了。不过现在,现在你正在 Omni Thinking 工作,嗯,而且还有一堆其他的项目。

Original English

Host: >> Um, and uh and uh and I I remember you were trying to get me invested in like one of the companies. I forget forget which one. H forget about that. But now, but now you're um now you're working in Omni Thinking um and and just you know a bunch of other

Shane: 还有 Gemini RL。

Original English

Shane: >> Gemini RL.

Host: 对,对。呃,还有 Nicole,你也在负责其它的生成媒体模型,Nano Banana 还有其他所有的东西,实际上就是本周刚刚发布的这些。呃……

Original English

Host: >> Yeah. Yeah. Uh and Nicole also uh the rest of the gen media models u nano banana and uh all and everything you just launched actually even this week. Uh

Nicole: 是的,我们发布了一些 API。

Original English

Nicole: >> yeah, we launched some APIs.

Host: 没错,没错,没错。而且我没有试图说服你投资任何东西,但也许我应该这么做。我的意思是,我尽量不去当一个投资者。但人们总是能说服我。我总是觉得,好吧,我没那么有钱,但你知道,你不可能不去尝试投资其中一些项目。而且你知道,对于我们这些不在前沿实验室(Frontier Lab)工作的人来说,这已经是我们能接触到的最好的、最接近前沿的机会了。嗯,所以是的,实际上,让我们来回顾一下,既然你最了解这些情况,而且我们刚刚完成了发布,本周到底发布了什么?大家应该去体验些什么?

Original English

Host: >> Yeah. Yeah. Yeah. And I haven't tried to convince you to invest in anything but maybe I should. I mean, so I try not to be an investor. People just convince me anyway. I'm like just okay, well, I'm not that rich, but you know, like you can't not try to invest in some of these things. And you know, for those of us who are not working at a Frontier Lab, this is the best this closest that we'll ever get. Um, so yeah, actually, let's kind of recap since you're closest to it and we just did it like what was launched this week. What should people go try out?

Nicole: 好的。嗯,昨天我们迎来了两个发布时刻。其中一个是我们发布了 Nanobanana 2 Light,这是我们在 Nanobanana 模型系列中最快、最便宜的图像模型。嗯,而且它比原来的 Nano Banana 表现更好。所以,对于大多数人来说,那个模型实际上取代了你之前熟悉并喜爱的原版 Nano Banana,在图像生成和编辑等各个方面都是如此,并且它的质量非常接近那种主流大型模型的前沿水平。所以这真的令人兴奋。我觉得如果你看了一些演示,或者大家一直在尝试的一些东西,你会发现,实现那种仅仅 3 秒钟的延迟,就解锁了一大堆你可以做的事情,比如在构思和迭代方面,这真的很有趣。而且这些模型已经达到了一个很棒的阶段,质量非常高,嗯,以至于你可以把它用作迭代,但同时你也可以把其中的一些输出直接作为生产级别的就绪输出。所以这非常让人激动。嗯,然后第二个发布,我们终于推出了我们在 I/O 大会上预先宣布过的 Gemini Omni Flash API。所以,感谢大家的耐心等待。而且,这是我们第一次向开发者提供这些 API,它带来了非常令人兴奋的视频生成和编辑功能,我们将它的定价设定得与 Y31 Fast 相同。所以,希望能以一个非常棒的价格,为大家带来极高质量的体验。

Original English

Nicole: >> Yeah. Um, so yesterday we had two launch moments. U one of them we launched Nanobanana 2 light uh which is our fastest, cheapest um image model in the nanobanana model family. Um and it's better than the original Nano Banana. Um so really for most people um that model replaces what you you know used and love the original Nano Banana for across like generation and editing and it gets really close to the frontier quality of of the kind of mainland bigger models. So that that's really exciting. I think if you look at some of the demos or like things that people have been trying like getting kind of that like 3-second latency just unlocks a whole bunch of things that you can do with like ideation and iteration and it's just really fun and the models getting to a point where like the quality is really good um where um it you know you can use it for iteration but you can also use some of those outputs as just kind of like ready um production output. So that's really exciting. Um and then second launch we finally um launched the Gemini Omni Flash APIs um that we pre-announced at IO. So, thank you for waiting. Um, and that, you know, is the first time that we're making the APIs available for developers and it's basically really exciting kind of video generation and editing and we're pricing it the same as Y31 fast. So, we're getting you kind of like really really good quality for a really awesome price hopefully. Um,

Host: 是的,我的意思是那太不可思议了。我其实非常……所以,当你们第一次发布 Omni 的时候,你们也和 Logan 一起做了一期播客,他今天没能来到现场。呃,而且你们在视频里添加了一只树懒,呃,还有拉面,以及所有这些,所有这些元素。其实我非常想在我们的视频里也加入这些特效。我只是没有 API 来实现,因为很明显我必须把整个流程自动化。所以,感谢你们提供这个 API。

Original English

Host: >> yeah, I mean that that's incredible. I'm actually really So, when you guys launched Omni for the first time, you also did a podcast uh with Logan who couldn't be here today. uh and you added like a sloth uh and and ramen and all these all these things. I actually really want to do that to our videos. I just didn't have an API for it because obviously I have to automate the whole thing. So, thank you for the API.

Nicole: 呃,那是我最喜欢的用例了。每个人都应该尝试一下。嗯,我弄了一只猫上去,那可能是那些动物里最无聊的一个了。嗯,如果你不知道我们在说什么,你应该去查一下。非常有意思。那是 Feurer 做的。嗯,团队里的 Feurer 做的那个。

Original English

Nicole: >> Uh that is my favorite use case. Everybody should do that. Um I got a cat which was probably like the most boring of the animals. Um if you don't know what we're talking about, you should look it up. It's very funny. Feurer. Um Furer, who's um you know on on the team did that.

Host: Fur 绝对是你最应该关注的头号人物。你应该关注他,去获取关于“好的,这东西究竟能做些什么?”的灵感。

Original English

Host: >> Fur is the number one guy you should follow. You should follow get ideas on okay, what can this thing do?

Nicole: 是的。没错。

Original English

Nicole: >> Yes. Right.

Host: 是的。他在那方面真的非常了不起。在过去两年里,我一直想邀请他来参加 AI 大会。但他还没来过。他其实本人来过现场。他只是不想发言,因为他喜欢保持匿名。

Original English

Host: >> Yes. He he's he's amazing at that. I've tried to get him for the last two years to come to AI. He hasn't made it yet. He's actually come in person. He just didn't want to speak because he's anonymous.

Nicole: 我知道。

Original English

Nicole: >> I know.

Host: 我,我很想说出他的真名,但我不能透露他的真名。

Original English

Host: >> I I want to say his real name, but I can't say his real name.

Nicole: 不,不,不。我们不会,我们不会对他那样做的,但你们真的应该去关注他。他棒极了。

Original English

Nicole: >> No, no, no. We won't we won't do that to him, but you should really follow him. He's amazing.

Host: 所有的那些工作都是他完成的。实际上我曾经在办公室里见过他,呃,就在……

Original English

Host: >> He did all that work. I actually met him uh in the office uh when

Podcast Conversations and Video Use Cases

Speaker A: 我们在做播客的时候,我想,我都没有意识到是他。所以,他的工牌上写的不是 Poper。上面写着……

Original English

Speaker A: we did the podcast, I think, and I didn't realize it was him. So, his badge doesn't say Poper. It says...

Speaker A: 是的,我知道。所以他以前是 Replicate 的一员,Replicate 之前有个玩笑,就像每个人都是 Deep Fates。Deep Fates 是一个有点神秘的角色。Replicate 是一家非常酷的公司,他曾是其中的一员。所以,好吧,在深入探讨类似 Omniprop 之前,我想先问一件事:我们在模型中加入了猫,加入了树懒。非常酷,非常可爱,也非常有趣。那么,你能否启发一下大家,除了那些演示之外,还有哪些更倾向于实际工作场景的用例?

Original English

Speaker A: >> Yeah, I know. So he used to be part of uh Replicate and Replicate had this joke where like everyone was Deep Fates. Deep Fates is this like kind of mysterious character. Replicate. Replicate is very cool company and was part of it. Um so okay, one thing I want to get on there before I go into like sort of the the the sort of omniprop is we added cats, we added sloths. Very cool, very cute, very fun. uh what are the you know inspire people as to like what are the more sort of workhorse use cases that maybe are not just demos you know

Capabilities and Future Applications of Video Models

Speaker B: 是的,显然这个模型的主打能力,或者说可能有两个:一个是它有能力接收任何形式的输入,然后最终输出视频。当然在未来——我们也曾把它作为预告讨论过——我们希望也能推出其他模态的输出能力。但这基本上意味着,你可以拿一组可能作为分镜的故事板图片,你可以拿一段音频作为你想让角色说话的参考语音,然后你就能得到一段视频。所以这直接解锁了一大批你可以做的事情,比如用于短片制作,或者是我们在 YouTube 上推出的 Shorts 短视频,来帮助创作者更轻松地创作内容。

另一个显然是视频编辑,这也是我们非常兴奋并正在努力简化的另一件事。因为现在你可以使用自然语言来处理视频,比如添加点什么、移除点什么。树懒显然是一个有趣的例子。不过,显然也有我们考虑过的消费者用例:比如你拍了一段海滩度假的视频,环境太嘈杂,你想清理掉那些噪音。在过去你可能不会这么做,因为你没有工具,或者你不知道你需要去用什么工具。所以这是一个你可以去使用的场景。我们看到很多人用它来创作营销广告活动,随着我们 API 的发布,我非常期待能看到更多这样的用例。因为显然,我们无法在第一方产品中看到它所有的应用场景,但我真的非常期待大家开始在 API 中去探索。所以这些只是一些比较高层次的应用。人们还用它来制作教育材料。

Original English

Speaker B: >> yeah so so obviously the hero capability of the model or maybe there's two like one is the ability to kind of take in anything as input and then get video on the other side obviously in the future and and we've kind of talked about this as a pre-announce like we want to get the other output modalities out as well but basically what that means is you know you can take a set of images that you have as maybe a storyboard you can take like an audio track as a reference of you know like a voice that you want a character to speak and then you can get a video on the other side. So like that just unlocks a whole bunch of things that you can do in like you know short film production or you know shorts we've launched on YouTube as well um to help creators kind of like create um content more easily. Um and then the other one is obviously video editing like that's another thing that we're really excited about that we're just making easier because now you can use natural language to take a video you know add something remove something. Sloth is obviously like fun example. Um, but there there's obviously kind of there's consumer use cases that we kind of had in mind where you know you could take your beach vacation video that was too noisy and you want to clean up that noise. Maybe in the past you wouldn't have because you didn't have the tools or you didn't know what the tools were that you needed to go to. So that's one use case that you can, you know, go to. We've seen a lot of folks use it for kind of like marketing ad campaign creation and I'm excited to see more of those use cases as we launch the APIs. um because obviously like we don't we don't see all of it in the first party products but I'm really excited for people to start to explore that um in the API. So those are just some of the kind of like highle um things that have come up. U people also use it to create like education materials.

Speaker A: 是的。

Original English

Speaker A: Yes.

Speaker B: 而且那真的令人兴奋。我想我们大家,我们所有人都曾讨论过对教育的未来感到兴奋,在未来一切都可以为你量身定制,并根据你的知识水平和你喜欢的风格进行个性化。所以,这就像是朝着那个方向迈出的一步。

Original English

Speaker B: Um and like like that's really exciting. I think we're all we've all kind of talked about being excited about the future of education where like everything can be kind of customized to you and personalized to your knowledge level and the style that you prefer and and so this is kind of just like a step in that direction.

Omni's Text Rendering and Translation Power

Speaker A: 是的。就在昨天,我父母来拜访时我刚好用了 Nana,当时有一个非常有趣的用例:我在亚马逊上给他们买了个他们想要的小工具,但使用说明书只有英文的,上面有很多图表之类的。我拍了张照片,然后说,把这个翻译成罗马尼亚语。

而且还要保持其他所有的排版和内容不变,对吧?结果非常惊人,对吧?它看起来完全一样,而且翻译得非常完美。我是说,八九不离十,对吧?但这显然是底层使用了 Gemini 来完成翻译的。所以你可以看到这个用例同样适用于视频,对吧?Omni 中的文本渲染能力真的是达到了下一个层级。所以,你可以想出很多用例,比如文本渲染、翻译、内部渠道化,各种真正对许多不同人群有用、能带来更广泛访问能力的事情。你可以给视频重新配音,或者做任何你想做的事情。你可以构想出很多不同的事情来做。

Original English

Speaker A: >> Yeah. I I I've sort of actually used just Nana yesterday with my my parents are visiting and there was there was a very fun sort of use case that I bought some gadget off Amazon that they wanted and the instructions to use it was were only in English and there was plenty of diagrams or whatever and I took a picture of it and said you know translate this into Romanian. Yes.

And keep everything else the same, right? So it was amazing, right? Like it was just like yeah it looks identical and it has you know it's perfectly translated. I mean more or less, right? But it's it's you know using Gemini under the hood obviously to kind of do the translation. So you can you can see this use case for video as well right like the the power of text rendering in in in Omni is is quite next level. So and you could you could you could think about plenty of use cases of like both text rendering translation internal channelization all sorts of things that would be actually genuinely useful to a lot of different people and sort of broader access to either you could like redub a video or whatever it is that you wanted to do. like there's plenty of different things that you could you could think about doing.

From Single Models to Video Agents

Speaker A: 嗯,我在播客上进行过的最具启发性的对话之一,是和处于这个领域前沿的研究人员交流。我曾和 XAI 视频团队、也就是 Grok 视频团队的 Ethan 聊过,他基本上是在说,下一个趋势实际上不再只是单一模型,而是更倾向于视频智能体(video agents)。

嗯,我不知道这个术语是否会引起共鸣——显然它对于强化学习(RL)来说是非常相关的。但它基本上意味着放弃试图在一次有效的前向传递中完成所有事情。你也有同样的感觉吗?还是说,趋势的发展方向目前仍然是一个开放的研究问题?

Original English

Speaker A: >> Yeah. Um one of the most enlightening conversations I have on my podcast is with uh this people researchers at the frontier of these things. Um I had one with um Ethan from the XAI video team, the Grock video team who was basically saying like you know the next trend is actually not just like single model, it's more like video agents. Mhm.

Um, and I don't know if that terminology resonates uh obviously for for very relevant for RL. Uh, but it was it was basically kind of like giving up on like trying to do everything in in effectively one pass. Um, do you feel that same way or is it still an open research question which way the trends are going?

Speaker B: 是的。最让我兴奋的,其实是当符号化的基础模型与这种视频基础模型能够真正协同工作的时候。如果你回顾一下生成式的、比如图像生成和视频生成的开端,很大一部分进展是在语言模型变得足够好,能够提供非常详细的图像描述时才开始的,比如在 Stable Diffusion 或 DALL-E 2 的时代。

所以,基本上语言是一种极其有用的表征方式。一方面它是通用的;另一方面,在更偏技术的层面上,我的假设是:机器学习中一个非常困难的问题是所谓的“虚假相关性(spurious coordination)”。你不知道某个具有预测性的特征是否真的是因果因素。有两种解决方式:一种是我们可以拥有极其多样的训练数据,覆盖因果图中的每一种干预情况;另一种方式是你在条件中加入因果信息,而在语言上加入条件,就类似于对这个世界的因果信息进行条件化。所以……

Original English

Speaker B: >> Yeah. So um what kind of excite me most is really when the symbolic kind of foundational models and this kind of like video foundational model can actually kind of really work together and u in a way the if you look at the beginning of the generative sort of like image generation video generation a lot of it kind of started when the language model got good enough to provide a very detailed captioning like from stable diffusion days or kind of dowi 2 days. So um so basically like language is extremely u helpful representation uh one is that it's kind of universal but the other kind of more um technical thing like kind of my hypothesis is like um one very difficult thing about machine learning is um this sort of like spirious coordination so you don't know you know if the if this kind of feature right that's kind of predictive is actually causal factor or not there are two ways one is we can have really diverse data training data like from every intervention of the causal graph. The other is you condition the coal information and conditioning a language is kind of like conditioning like a coal information of the of the kind of world. So um

Speaker A: 也就是一个提示词(prompt),或者一个概念,或者是别的什么?

Original English

Speaker A: >> which is a prompt or a concept or what?

Speaker B: 是的,完全正确。所以如果你去看看,你要如何描述这段视频、如何描述这张图片,实际上这非常接近于去探讨这个视频背后的因果关系是什么,比如它是如何生成的。所以,一方面这可以实现非常丰富的泛化能力,并带来一个很好的模型。

另一方面,八个月前我们发布了一篇名为《Video Models Are Zero-Shot Learners and Reasoners(视频模型是零样本学习者和推理者)》的评估论文。

Original English

Speaker B: >> Yeah, exactly. So if you look at like you know how you going to describe this video, how you going to this kind of image is actually very close to you know how would this kind of causality you know behind this like how this is kind of generated. So one is like that can really allow for very rich generalization and then uh very kind of just like a good model. Um the other is so eight months ago uh we put the evaluation paper called video models zero shot learners and reasoners.

Speaker A: 是的。

Original English

Speaker A: Yes.

Speaker B: 那是一篇已经确认的研究论文,后来 N banana 团队(应为 Vision Mamba)也跟进发表了 vision banana 论文,基本上是用 banana 来做。但其核心思想是,视频模型是一个非常出色的用于处理时空信息的基础模型。因此,许多经典的计算机视觉任务都可以实现零样本(zero-shot)处理。当你向它输入类似视觉问答的任务时——当然绝对还有很多需要改进的地方——它是可以解决的。在机器人领域,就像“看”一样,它拥有非常好的物理直觉,就像一个世界模型(world model)。我认为关键在于将视觉推理和文本推理真正地结合在一起。

当然,究竟是作为统一的模型来做,还是像智能体或探索机制那样来做,我认为这更像是一个渐进的过程。我设想所有的东西最终都会演变成一个单一模型。

Original English

Speaker B: So that was a kind of you know it's it's a confirmed paper and then later on actually the N banana team follow up with the vision banana paper that basically used a banana to do but essentially the idea is uh video model is extremely good sort of a foundation model for space and time kind of information. So um classic computer vision tasks a lot of could be kind of zero shorted and when you like say feeding some like a visual quiz uh it can you know there's definitely like a lot to improve it can kind of solve and it can um like robotics kind of like seeing it has really good kind of physical intuitions like word model uh and I think the the key is really the kind of mix of the visual kind of reasoning and then the text kind of reasoning kind of all tied together Um obviously you know like whether doing it you know as kind of unified model versus like this kind of agent or exploration I think that's more like uh it's going to be more kind of incremental you know how it's going to I imagine everything's going to go into like a single model eventually

Exploring Agentic Capabilities and Single Product Trend

Speaker A: 但现在如果把非常出色的视频理解、图像理解以及 Gemini 以一种智能体(agentic)的方式结合起来,你能做很多事情。这也正是我们的团队正在大量探索的方向。

是的,好吧,这里面包含了很多信息。我认为,我越来越开始思考的一个问题是,对于你们来说,这一切最终都会趋向于一个单一产品吗?就像现在你们推出了多个模型,而 Omni 这个名字确实暗示了最终其他一切都会消失,只剩下 Omni。这是计划吗?

Original English

Speaker A: >> but right now there's like a lot you can do if you uh basically take like really good video understanding image understanding Gemini agentically with anomony and that's actually going to yeah our team is like exploring a lot

yeah okay that there's a there's a lot in there um I I think uh one question I I am increasingly starting to wonder is does it all trend towards one product for you guys right like now you have multiple models out the naming of omni does imply that eventually everything will go away and it just goes into omnis um is that the plan

Speaker B: 是吗?我不知道。我觉得……也许吧,我的意思是,我认为最终……我认为现在还有一些不同的方向。

Original English

Speaker B: >> is it I don't know I I think I think uh maybe I mean I think eventually I I think there's sort of different

模型多样性与模态融合的权衡

Speaker A: 工程、研究、产品之间的权衡就在于,出于同样的原因……不好意思,那个叫什么,Nano Banana Light,我不知道那产品名叫什么。

Original English

Speaker A: trade-offs engineering research product trade-offs in like it's like for the same reason like the the sorry how is it called nanobanana light I don't know what the product name is

Speaker B: Nano Banana Too Light。

Original English

Speaker B: nano banana too light

Speaker A: Nano Banana Too Light,对,它服务于特定的细分市场,对吧?它可能不一定非得跟能做 4K、30 秒视频的模型使用完全相同的架构甚至同一个检查点,对吧?它们的训练方式可能也不尽相同。所以,我也不好说,这取决于你着眼于多远的未来。当然,如果看五年后,它们会是同一个模型吗?很可能。但是,如果看六个月后,我们很可能还是会有多个不同的模型在执行不同的任务,因为从实用角度来看,权衡的结果就是我们应该拥有多种不同类型的模型。

Original English

Speaker A: nano banana too light yeah right it's it's it serves a particular niche right and it probably doesn't necessarily fit immediately in the same model literally checkpoint as uh something that can do 4K you know uh 30 second videos right like they're probably not like trainable in the same quite way, right? Like, so I don't know. It depends on how how far into the future you look like. Sure, in five years from now, will they all be the same model? Probably. Uh, but like, you know, six months from now, we'll we'll probably still have, you know, multiple different models doing different things because kind of from pragmatically the trade-offs are such that we we should have multiple different kinds of models.

Speaker C: 我觉得那是对的。顺便提一句,我们给它起名叫 Gemini Omni,就是为了暗示未来 Gemini 会变成完全原生支持多模态输入输出的模型。所以,这绝对是朝着那个方向迈出的一步。我认为我们很可能会看到 Omni 也能生成图片、编辑图片,以及做所有这类事情。但是 Doo 说得对,我认为在达到那个终点之前,其中一些更专业的模型会有大量非常非常实用的应用。因此,我们可能也会继续在这些模型上投入,因为它们满足了当下某个在一年后可能就不复存在的特定需求。

另外,还有一个研究问题:不同模态之间究竟有多大的迁移性?你可能相信代码和视频生成之间存在某种迁移,但我认为大多数人不一定这么认为。你可以试着想象它们之间存在某种联系,或者把它们放在一起、试图让模型同时学习这两个任务可能纯粹是一种浪费。所以,我认为这确实是个有趣的问题。就图像和视频而言,显然有一定程度的迁移,它们并不完全不同;学习同时输出视频和音频是有价值的,因为视听联合本来就是事物的自然规律。然而,还有一些模态的交叉点并非那么显而易见,比如 3D 表示和代码,我也说不准,也许诸如此类吧。所以我觉得探索这些不同的角落是值得的,而且我们也正积极朝着人们实际想用这些模型做什么的方向去探索。

Original English

Speaker C: I think that's right. And and just on that note, I mean, we did call it Gemini Omni because we wanted to hint at the future where Gemini just becomes fully multimodal in and out, right? And so, so it's definitely a move in that direction. I think we'll probably see a move in the direction where Omni also generates images and edits images and all those kinds of things. But Doo is right that I think on the way there, there's a bunch of really really useful applications of some of these more specialized models. And so we we will probably continue to work on those as well because like that serves a certain need at this point in time that may not exist you know a year from now.

There's also like a research question about like just how much transfer there is between different kinds of modalities, right? I think you may believe that there's some transfer between coding and video generation and I think most people don't necessarily believe that but they you know you could try to think that there is some some there or it could be a waste right to put them together to try to learn these both tasks at the same time right so I think it's it's it's interesting sort of question to which extent like image and video obviously kind of there's some transfer like kind of not that different there's value in in learning to output video and audio at the same time because joint audiovisisual is you know that's how that's how it is. Um and then there's you know other kind of intersections of modalities that are not super obvious right like 3D representation and coding I don't know maybe uh things like that right so like I think it's worth sort of exploring the different corners there and we are actively doing that um with a focus towards like what people actually want to do with these models

中间表示法与自然语言的作用

Speaker B: 对。有件事让我感到惊讶,同时我也觉得它没有得到充分的解答,那就是:什么是正确的中间表示法?比如图像字幕(captioning),对吧?XI 在做 captioning,Omni 也在做 captioning。我理解 captioning 在图像上是如何运作的,我也理解你可以将其扩展到视频中并随着时间引导它。但这感觉非常低效。我觉得应该有更好的方法。也许是代码,也许我们会生成代码。很明显,很多视频都是通过代码生成的,比如使用 FFmpeg 和 Matplotlib,还有那个 3Blue1Brown 用到的 Manim。也许代码才是最优的表示法?对此有什么假设吗:究竟是别的表示法更好,还是仅靠英语(自然语言)就足够了?

Original English

Speaker B: yeah um what one thing I feel I feel like uh I'm surprised by but also I feel like it's insufficiently answered is what is the correct intermediate representation Um, so captioning, right? XI does captioning. Omni does captioning. Um, and I I I understand how captioning works for images. Um, and I understand that you can extend it into to video and and sort of guide it across time. It just feels very inefficient. It there's got to be I feel like there should be something better. Uh maybe it's code and maybe we generate you know and obviously I think a lot of um ffmpeg and mapplot um what's the three blue one brown one manim um a lot of like video is generated through code and maybe that's like the optimal representation uh any hypothesis as to like is is it better or is just English all you need

Speaker C: 在 Gemini 团队中,我们做了很多 RL(强化学习)智能体和代码相关的工作。所以是的,我们肯定在探索代码这种表示法。

Original English

Speaker C: well as so I'm in the Gemini and you know we do like a lot of RL agent and of course kind of coding so yeah We we're definitely exploring the coding representations.

Speaker A: 是的。

Original English

Speaker A: Yeah.

Speaker C: 将其作为一种更好的表示方式。

Original English

Speaker C: As kind of better kind of way to represent. Yeah.

Speaker B: 但是,你觉得我们只是单纯输出二进制文件(就像只有 0 和 1 那样)的概率估算有多大?

Original English

Speaker B: But you know like do you what's your probability estimate on like we just output binaries like we just you know like just it's just ones and zeros.

Speaker C: 我觉得也许有一个类似的讨论,就是:自然语言是不是正确的表示方法?比如,有个教授问的一个问题就是,为什么思维链(chain of thought)必须用自然语言表达?

Original English

Speaker C: Um I I guess maybe a kind of similar discussion was like um basically is the language the right representation like right. So uh one kind of question for example uh professor you know like ask is like you know why why does the channel of thought need to be in the natural language?

Speaker B: 是的。

Original English

Speaker B: Yes.

Speaker C: 它能不能只是一种连续的 token,单纯就是各种额外的计算?从测试来看,自适应计算显然会带来更好的结果。四年前我就写过更大的模型推理器(reasoner)和自我改进(self-improvement),所以我很早就知道这一点。但它之所以现在效果很好,是因为当前起作用的方法是:大规模进行预训练,让模型学习到大量智能;虽然有很多扩展强化学习(RL)的方法,但那些方法在提取信息时仍然是计算密集型的。你真正想要的是依赖预训练带来的智能。所以,基本上通过把推理过程和自然语言绑定,你就能直接利用预训练的智能;而如果你去掉了这种限制,那就做不到了。而且我觉得现今无论是文本领域还是多模态领域的很多进步,真正是由“文本作为一种极好的表示方式”来驱动的。

Original English

Speaker C: Can it just be the kind of any kind of like continuous tokens just any amount of you know additional computations. Um so one is like obviously the test like adaptive compute is going to give like you know better results. So it's sad but what really kind of made CH thought so you know like four years ago I wrote you know the larger model Z reasoner and then self-improvement. So I kind of know from very early day but the reason like it works really well is um right now the recipe that works is the pre-training that scales a lot and then that basically like learns a lot of intelligence. there are a lot of you know scaling RL but those are still like extremely kind of comput intensive to extract the information and um you really want to rely the intelligence on that so basically by tying the sort of like a reasoning in the natural language you basically directly use the intelligence of the pre-training to it while if you remove that kind of constraints then you're not um and these days uh I feel the a lot of advancements in the texts but also in kind of multimodal space is really driven by this u kind of text as a kind of great uh sort of representation.

Speaker B: 是的,它是一个很好的骨架。

Original English

Speaker B: Yeah, it's a good backbone.

Speaker C: 是的。

Original English

Speaker C: Yeah,

Speaker A: 我觉得这甚至更简单:因为文本就是我们交流的方式。如果你正在构建人类要与之交互的产品,那么不管是哪种方式,如果它是一个文本界面,我们就会用到文本,对吧?虽然不是所有情况下都如此,但我认为默认使用文本是很自然的。当然,也会有一些让人困惑的讨论。你知道,有些 RL(强化学习)极致主义者会说:“我们不在乎那些中间链或者类似的东西,那只不过是额外的计算罢了。”

Original English

Speaker A: I think to me it's even simpler than that. It's text is is how we communicate. So I think fundamentally if you're building kind of products that humans will be interfacing with um like like that we will be using text somehow if it's a text interface, right? Not not for everything. So I think it's it's natural to default to that. Yeah, obviously there's like a confus discussion. You know, some arrow like RO maximalist is like, oh, we don't care about, you know, kind of channel those kind of like stuff. It's just just additional compute.

Speaker B: 没错。

Original English

Speaker B: Sure.

Speaker A: 但我个人而言,是的。

Original English

Speaker A: But I personally Yeah.

世界模型与未来愿景

Speaker B: RL 极致主义者。我好奇谁符合这个描述,David Silver?

Original English

Speaker B: RL maximalists. I wonder I wonder who who qualifies in that description. David Silver.

Speaker A: 噢,好的。我是说,他们刚刚离开去创办自己的项目了。这很有趣。所以,我对找到更好的表示方法非常感兴趣,因为这也是我们今天在世博会上策划的主题之一:世界模型。你提到了“世界模型”这个词,但它的定义并不十分明确。我想大家正在某种版本上达成共识,认为那是理想状态。

Original English

Speaker A: Ah, okay. Yeah. I mean, they they've just left to to start their thing. Um, interesting. Okay. So, uh I I mean I I think I'm very interested in just like better representations because I think that's one of our themes that we're curating today uh at the Worlds Fair is world models. You mentioned the word world models, but it's not something that's like super well- definfined. I think everyone's like sort of converging on some version of it that it's like the ideal.

Speaker B: 当然。现在什么都是世界模型了。它有点像……

Original English

Speaker B: Sure. Everything is a world model now. It's a sort of

Speaker C: 这个概念并不是那么实用,对吧?

Original English

Speaker C: it's not it's not that useful, right?

Speaker B: 我刚在 ICLR 的世界模型研讨会上做了一个主旨演讲。我非常鼓励大家去看看 Jitendra Malik 给出的定义。他是加州大学伯克利分校的资深计算机视觉教授。他对世界模型有很多自己的见解。这让人回想起他是如何定义 2019 年甚至 1990 年代的世界模型的。对我来说,世界模型基本上就是在基于模型的强化学习(Model-based RL)里的那个“模型”,我觉得这足以描述它了。但显然还有很多人提出了别的观点,比如 Fay 写了一篇很棒的博客文章,把这个概念拆解分析了。

Original English

Speaker B: So, I just gave a keynote at the i Clear's world model workshop. Yeah. And then uh yeah essentially uh I definitely encourage to check out the definition by Jatendra Matalik. He's like the you know OG computer vision professor UC Berkeley. uh he has pretty you know bit of word to say about world model but also kind of shimmburers kind of how he defined the world model from 2019 like 1990 sort of uh uh you know like Wayne was just basically just that kind of model base uh for me the word model is basically just a model in the model based RL and I feel that has sufficient to describe but obviously you know there are like a lot of uh fay had a kind of nice blog post about what about yeah this kind of broken down

Speaker A: 嗯,是的。我想结束这部分的讨论,但我确实认为,把自然语言作为一切信息的狭窄通道,仍然像是一种有损压缩。

Original English

Speaker A: um but yeah yeah I mean so you I I'll I'll end this part of the conversation, but like I I do think that language to me relying on language as like the sort of like the narrow pipe through which everything goes through. Um still is like a lossy compression.

Speaker C: 不,不,不。但我们看到的并不是那样,对吧?我们说的是视频模型和语言模型一起运作。

Original English

Speaker C: No, no, no. But we're not seeing that, right? We're basically saying the video model and the language together.

Speaker B: 所以我认为单靠语言是不够的。这也是为什么我们觉得视频是一种非常全面的形式;虽然现有的视频模型很不错,但我觉得我们的愿景远不止于此。它是一个缺失的底层基础模型,如果你想打造能匹敌人类的通用人工智能(AGI),而不仅仅是一个缝合怪,这是绝对必需的。

Original English

Speaker B: So, so I think the language alone is uh not sufficient. That's why we feel like the video is a very comprehening kind of pretty videos but I think our vision it's it's much more than that. It's a missing foundational model that's absolutely required if you want to make the AGI that match the humans not just the jacked one.

Speaker A: 是的。另外一方面,你提到了在视觉方面的一些情况,我比较好奇在这些方面你们是如何并行推进的……

Original English

Speaker A: Yeah. Um okay so one one other thing you know you you mentioned on the vision side um and I'm kind of curious how sort of uh parallel you know in terms

从视觉理解到生成模型的跨界

Speaker A: 在你们的研究生涯中,这种发展趋势,我觉得基本上就是很多做视觉的人跨界去做了更多关于模型的工作,很多做视觉的人也变成了做生成式视频和图像的人。这仅仅是简单的反转吗?比如从图像到文本,然后现在是文本到图像,我的意思是,这实际上就是扩散过程吗?我只是观察到我交谈过的人的职业路径,看到了这种研究方向的整体趋势,所以我想请你们谈谈对此的看法。

Original English

Speaker A: of your research careers um this development is like I think basically a lot of vision people have crossed over into more model people um a lot of vision people also become generative video and image people and is it just as simple as you know reversing uh image to text and then now it's text to image like is is that if I mean that effectively was the diffusion process. Um I I just I you know I I just see the career paths of the people that I talk to and and see and I I I see this overall trend of research directions and I just wanted you to guys to sort of reflect on on that.

Speaker B: 我的确是走的那条路。很久以前我是做计算机视觉的,比如目标检测、识别之类的。我认为那只是一个更简单的问题。生成本身就更难,它是一种不同的映射。你从逆向映射来做,并不像简单地反转那种旋转那么简单。从“猫”这个词到一张猫的图像,要模糊得多。而且在某种程度上,它也是一个循环,因为你的视觉工作创造了合成标签,然后继续——

Original English

Speaker B: I mean I certainly went that way right I started long time ago uh doing computer vision sort of object detection recognition things like that. Uh I think just that's just simpler problem right just generation is just harder like it's a it's a different kind of mapping right you map from the the inverse mapping is not as simple as just inverting the the kind of rotations right it's it's a it's it's more ambiguous right to go from cat to image of a cat and in some ways it's also a loop because your vision work creates the synthetic labels that then continues

Speaker A: 我的意思是,当然,我只是想验证一下我关于领域如何发展、以及人们的职业生涯是如何在其中推进的理论。

Original English

Speaker A: I mean sure I don't know I don't know I'm trying to validate my my sort of theories about how fields develop how how careers progress through this

Speaker B: 我的意思是,随着理解端的进步,我们已经看到生成端也变得更好了,对吧?所以这完全是相辅相成的,完全是一种引导启动(bootstrapping)的过程。

Original English

Speaker B: I mean for like the the the better the understanding side gets like we have seen that the generation side also gets better right so like like it's completely bootstrapping yeah it's

Speaker A: 确实是这样。这个论点绝对是成立的。我也和很多做图像理解的人共事过,他们后来都变成了做图像生成的人,然后他们中的一些人又转向了视频,因为这就像是下一个热点,你有更多的维度可以去探索。所以我对你的具体经历很好奇。

Original English

Speaker A: and so so like like like there's definitely they're there to that thesis and I think yeah I think a lot of people have kind of like I I definitely worked with a lot of um image understanding people who became image generation people you know and then some of them have moved on to video because it's kind of like the next thing where you have so many more dimensions to work with so yeah I'm curious about you specific as your

强化学习与多模态模型的演进

Speaker C: 我肯定会建议从理解和识别开始,因为那基本上就是判别器,然后这将带来更好的生成,而它们之间的桥梁基本上就是强化学习。所以我的经历是,最初我是在代理模型中进行算法研究,做一些非常出色的生成工作。然后我从事了强化学习和机器人技术。大概六年前,我领导了一个关于灵活性的“登月计划”,当时还很早,但我现在看到每个人都在做。大约四年前,我基本弄清楚了这种符号化的 AGI 将比物理层面的 AGI 对应物发展得快得多。所以我决定转向语言模型这些东西。最近我和 Doomi 以及 omni 团队合作,我非常享受那里的合作。

我强烈推荐给研究人员的是,去探索或者至少去接触一下各个社区最顶尖的人在看什么,以及他们是如何思考问题的。所以当我审视视频模型时,它让我回想起了很早期的语言模型。当时早期的语言模型更像是一种创意的演示,比如你尝试写一个小说故事。然后像 GPT-2 那个时期,还有 LSTM 的那个时期。之后通过指令微调,你实际上让它作为聊天机器人变得可用了。但在聊天机器人阶段,它仍然有很多幻觉,指令跟随不够好,所以不能用于推理。而当它在预训练和后训练中对于推理变得足够好时,这种测试时计算扩展(test-time scaling)的强化学习真的起飞了,造就了许多表现最好的模型。

现在我认为视频模型正如我们提到的,它是一个补充性的基础模型,我可以想象它会遵循类似的路径:它会得到很大改善,指令跟随也会大幅提升,并在减少幻觉(coordinations/hallucinations)方面取得长足进步,直到它成为一个非常可靠的世界模型。这样我们就可以将视频(比如时空模拟)与文本模拟混合在一起,以解决任意的 AI 问题。

此外,我认为文本模型和图像/视频模型之间的区别仍然在于,我们在多媒体中还没有完全统一理解和生成。我的意思是,不探讨细节的话,当然这取决于你在哪个层面思考,但总的来说,据我所知,并没有多少模型真正擅长既能理解又能生成视频。这是一个有趣的挑战。我不是说我们应该这样做。但我认为,理解和生成是同一枚硬币的两面,这是合乎情理的。所以它们在某种程度上应该在同一个模型中。但我们并不总是这样做。

你刚才也提到了音频,对吧?它和视频一样难吗,还是有本质的不同?如果是这样,在哪方面不同?三年前一个有趣的方向是人们使用扩散(diffusion)来做音频,也就是那种重新融合(refusion)的方法,不知道你们有没有看到。我觉得这非常有趣,我们感知的音频模态与视频不同,但对机器来说其实完全一样,它们看不出区别。我的意思是,在技术层面上有一些区别,但我认为这些区别相对较小。从我的角度来看,音频是在我们发布 V3 时进入我的生活的,我相信这是第一个做到联合——

Original English

Speaker C: so I definitely like recommend start with understanding recognition because that's basically discriminator and then that's going to lead to better generation and that's what the bridge is basically reinforcement learning so my um my kind of journey is I initially kind of worked on the algorithmic research in the gent model against some like you know eminent kind of generation and then I worked on like RL and robotics um and then like six years ago I was like leading like a moonshot on the dexterity it was pretty early but I see now everyone's kind of doing uh four years ago I basically kind of figured out that this like symbolic AGI is going to accelerate much faster than the kind of physical AGI kind of counterpart. So uh I decided to kind of like language models and then those things. Um and then recently kind of work with Doomi and then like omni team I quite enjoy kind of collaboration there. the what I quite enjoy uh what I recommend definitely to the researcher is to uh definitely kind of explore or at least like get exposure to what the top people in each of the community are like looking at how they kind of think about problems. So when I look at the video model to me it kind of reminds me like pretty early on sort of like language model where like very early language model was a kind of creative sort of demo right you kind of like try to write like a story like novel and then like you know GBD2 and then those kind of days like L stem kind of days right and then you know uh instruction tuning you actually kind of make it usable as a chatbot but then at the chatbot stage it still had so much hallucinations and instruction for wasn't good enough so it couldn't use for reasoning and when it got good enough um in pre-training and post- trainining for reasoning then you know this kind of test time scaling the RL really took off to like many of the kind of best performing models and right now I think the video model is as we mentioned it's it is a complimentary foundational model and I can imagine it's going to follow a similar path it's going to be very uh it's going to improve a lot instruction following a lot of uh this it's going to improve a lot in reducing coordinations to extend that it become a very reliable world model so we can kind of like intermixed video like space-time simulation with a text simulation to solve like arbitrary AI problems. Also like I think the difference still is between sort of text models and like image video models is that like we haven't quite unified understanding and generation in in multimedia I'd say yet like I mean I think I think without going to the details of course there's like it depends on on at which level you're thinking about this but generally like there's not that many as far as I know models sort kind of you know printier models that are genuinely kind of good at both understanding and generation of of let's videos, right? Like it's a it's a it's an interesting challenge. I'm not saying that we should do this. Uh but but I think uh it kind of stands to reason that like you know understanding and generation are two sides of the same coin. So they they kind of should be in the same model in some ways. Uh but we don't necessarily always do that. So yeah. Uh you mentioned audio as well, right? Yeah. Uh is that as hard as video or qualitatively different? If if so, in what way? Uh one of the interesting directions three years ago was people using um I guess diffusion to do audio uh as in like the the sort of refusion approach. I don't know if you you guys saw that. Um and I just think it's like very interesting if a modality that we perceive which is audio is different than video actually two machines is exactly the same like there's they see no difference. I mean, I think on a technical level there are some differences, but I think they're like relatively minor. I think from my perspective, audio came into into my life when we shipped V3, which was, I believe, the first model that did like a joint

Speaker A: 进行切片……

Original English

Speaker A: with the slicing of the

音频生成与信息的言语化挑战

Speaker C: 是的,是的。金条(Gold bars)什么的。它是第一个在某种程度上进行视听联合生成的模型。我的意思是,有其他的模型也做了类似的,在底层做了一些代理黑客式的处理,但这个模型是真正一次性生成所有的东西。我们之所以这样做,是因为我们觉得,而且我认为这是正确的选择,我们觉得只有同时生成它们才有意义。因为从机器学习的角度来看,有一种潜在的因果生成过程。是有某种东西产生了你说话的声音,它并不是先有像素,然后音频在某种程度上被另一个过程生成。嘴唇必须与音频同步。所以,我认为这解决了以前模型遇到的大量问题,或者是以前人们做视频生成的方式中的问题。以前的方式是,我们先生成像素,然后在其之上进行修改,让嘴唇随着我们生成的音频移动,那样的效果非常差。在 V3 之后,人们会说,你的模型里没有音频是什么意思?那说不通,既然视频都在那里了,你就必须得有音频。所以我认为那是个正确的决定,在同一个生成模型中去实现它是一个正确的选择。

Original English

Speaker C: Yeah. Yeah. Gold bars or whatever. Um, it it was the first model that is sort of joint audiovisisual generation. Yes. uh like in a in a I mean there are there were other models that did kind of you know kind of kind of agentic hacking under the hood but this one was truly sort of you know generating everything at once and we the reason we did that is because we felt and I think it was the right choice we felt that like uh it only makes sense to generate them at the same time because there sort of kind of like from a machine learning perspective there's one latent kind of you know causal kind of you know generative process right like there's something that generates you speaking it's not the pixels and then the the audio or somehow somehow generated by some other process like the lips have to move in sync with the with the with the audio, right? So, I think that that solved a lot of the issues that previous models had or the way that people did video generation before where it was like, okay, we generate pixels and then we're going to hack something on top of it that like moves the lips with the audio that we generate and that's was very bad. And so I think I think that was that's to me that's the the I mean after V3 like you know people were like what do you mean like there's no audio in your model? like that makes no sense like once it's there like you you have to have it. So I think that was that was the right choice and doing it in one single generative model I think was was the right choice.

Speaker D: 有件事我也想听听你们的看法,我发现音频和图像、视频的一个区别是,音频信息的言语化程度较低。我的意思是,当然 TTS(文本转语音)之类的很简单,但除此之外,比如你如何描述音乐,你如何描述一个人的语调、音高?我觉得这种言语化的表达是不够的。有趣的是,在另外两件事上你也能看到类似的情况,比如味觉和触觉。

Original English

Speaker D: One thing I kind of want to also kind of ask you guys an opinion as well once one difference I find the audio and then against the image and video is like the audio information is less verbalized. I mean of course the TTS and stuff is trivial right but the when you get her outside like how to describe music how do you describe this like this person's tone kind of pitch I feel the sort of the verbalization is in insufficient and the interesting thing is that you kind of see that in two other things like taste taste sense and also uh touch

Speaker A: 像嗅觉,另一个有趣的例子是肤色,用来描述肤色的语言是非常有限的,原因是——

Original English

Speaker A: like smell and then the another interesting thing is the skin color so skin cutter the the language is pretty limited to describe the skin color and the reason is

语言在描述感官体验时的局限

Speaker A: 我们对肤色的微小差异和变化极其敏感,因为这基本上能向我们显示“这个人会杀了我吗”或是“我能和这个人交朋友吗”这类信息。我觉得嗅觉、味觉、肤色以及声音这类东西,都与原始的生存本能紧密相连。所以我们的感官系统如此敏锐,以至于很难用语言去处理。比如我问过一位专业的品酒师,他基本上说他会用类似约会时描述伴侣的语言来形容酒的味道,因为没有足够的词汇来描述。所以我很好奇,你们是不是也觉得……

Original English

Speaker A: ...that we're extremely uh sensitive to the small difference perturbations of a skin color because that basically shows us is this person going to kill me or is can I befriend this person kind of those kind of information and then I feel the smell tastes um skin color and like sound kind of stuff is very very tied into primitive it would like survival kind of stuff and so our sort of sensory system is so sensitive that it's intractable to um so for example I asked like one the wine sort of taster and like professional and then he basically said he kind of use like a language from like a dating you know describing like a you know partner as a way to describe the taste because there's no sufficient vocab to describe um so I'm kind of curious yeah do you guys feel that

Speaker B: 我觉得在某种程度上,视觉信息也是一样的情况,对吧?当你想起某种特定的风格或美学时,对吧?比如,有些人在调色板或视觉品味和美学方面,就有非常成熟的见解,对吧?我认为当你在试图描述我们所体验到的任何感官信息时,语言往往会成为一个限制因素。回到你之前提到的观点,我认为这正是我们投资世界模型的原因,也是我们之所以在感知和生成方面努力推进的原因。因为它是我们作为人类探索这个世界的一大重要部分。这在很大程度上也是具身人工智能探索世界的方式。并且,我确实认为语言在这方面已经带我们走得很远,也可能带我们走得更远,但在很多这类领域,它确实让人感到有局限性。是的,我也不太知道该如何去描述,你知道,比如感官和味觉。嗯,但我确实对此很好奇。

Original English

Speaker B: I think well to some extent I think the same is true for visual information right when you think about like a certain style or a certain aesthetic, right? Like like there are some people who just have a much more kind of developed like whether it's palette or kind of visual taste and aesthetic, right? Like I I think language just tends to be a bit of a limiting factor when you are trying to describe any of these things that like we experience with sensory information. And to your point earlier, I think that is the kind of the reason why we are investing in world models and why we are pushing on kind of the like perception and like generation side of things because it it is such a large part of how we as humans navigate the world. It's a large part of how like embodied AI navigates the world. Um, and and I do I do think language like does have a lot of it's it's gotten us very far and it can probably get us really far, but it it feels limiting in a lot of these kind of areas. And yeah, I don't I don't really know how to describe, you know, like sense and taste. Um, but yeah, I'm curious to me.

Speaker C: 呃,是啊,我不知道我是否已经深入思考过这个问题。所以,是的,我的意思是,我对音频部分没有一个很好的答案。我的意思是,我不知道它的极限在哪里,因为我在想 Omni 在音频方面到底不擅长什么,但我发现它们似乎都是可以解决的问题。比如有了更多的数据或更好的数据,或者是其他什么,所以我不知道我们是不是已经把前沿推进了那么多,以至于触及了某种根植于人类进化的、由人类强加的极限。我不知道他是否感受到了文字描述的局限,这就是我之前提到的……

Original English

Speaker C: Uh, I I yeah, I don't know that I have thought that deeply about this yet. So, uh, yeah, I mean, yeah, I don't have a good answer about audio. I mean like I don't know the limit because I'm thinking about like well what is what is Omni bad at in terms of audio but they're all like solvable problems I find uh so like with more data or better data or whatever it is so I don't know like that we have pushed the frontier so much that like we are have hit some sort of limits that are rooted in evolutionary uh kind of you know limits imposed by humans I don't know he's feeling the limits of captioning which is the the thing I was...

Speaker B: 是的,完全正确。世界上有大量的信息,这在本质上与我们为什么要建立世界模型有关。

Original English

Speaker B: yeah exactly there there's there's a lot of information in the world and it connects to basically why we do world modeling.

Speaker A: 你提到了……你只需要 SRFS sref76,那就是你要的效果,对吧?我想也许我没法描述那种氛围,但是……

Original English

Speaker A: You mentioned... you just need SRFS sref76 and then that's your what does right? I guess maybe I can't describe this vibe but...

Speaker B: 嗯,我觉得提供这些参考指标的意义就在于此,对吧?因为即便只是描述某人说话的方式,比如他们的语调、韵律以及所有这些,我觉得有些术语我以前甚至都不知道是什么意思,对吧?好吧,现在知道了,比如言语不畅(disfluencies)等等,确实有这样一整套词汇。即使你并没有深入某个领域——实际上对大多数人类领域来说都是如此——你可能根本不知道它意味着什么。有时这也是个问题,如果我们没有在使用大语言模型时去关注这些东西,那么它们可能在这些领域也会存在空白,对吧?然后我们在生成端就会感觉到,因为我们基本上依赖语言模型对世界的理解来去表征它。所以我觉得,这又回到了你关于语言作为中介的那个问题。但我认为,要做到这一点,可能就是因为有些专注的领域我们还没有尽可能地去推进。以及,我们会发现它的实际上限在哪里。

Original English

Speaker B: well well I think that that's kind of the point of providing some of these references right because because like even just describing how someone talks and like their tone and and like procity and all of these things like I think I think some of these terms even like I didn't used to know what they mean right well now yes disluencies ex exact like like there there's kind of an entire vocabulary that even if you're not kind of steeped in a domain which is true for actually like most human domains that like you don't even know what it means means um and sometimes it's also a question of like if we haven't focused on those things you know with the large language models that they may also have gaps in those areas right and then we feel them on the other side with generation because we're like fundamentally relying on on the language models understanding of the world to then be able to like represent it um so I think yeah it all kind of goes back to your question about like the the language as an intermediary but yeah I think to do like some of these might just be like focus areas and things that we haven't necessarily pushed on as much as we can and like as well, we will discover what the actual ceiling is.

AI 模型在空间音频及自然表现上的挑战

Speaker C: 是的,作为一名播客,我对声音思考了很多。我只想提出几点供大家讨论,以防能激发你们的想法。我粗略地把音频分为三个领域,分别是音乐、人声和音效。这个分类算粗略吗?好吧,它涵盖了一切。并且,即使在人声内部,我们就只关注人声,忘掉另外两个。空间的声音,比如大房间、小房间的这种回音感,当面的声音,车里的声音,电话里的声音……所有这些虽然都是可以贴标签的,但我们体验它们的方式截然不同。我经常认为,人工智能视频的一个破绽就是它具有录音室的音质,因为那是在录音室里录制的,因为那是你们的训练数据。对我来说,最有趣的事情其实是,当我用这个例子去说服那些对世界模型的需求持怀疑态度的人时,我说,即便对于音频你也需要世界模型,比如离你较远,所以我的声音听起来应该稍微轻一些或者更发散一些。视频模型需要捕捉到这一点,因为如果他们要做沉浸式的视频和音频,你就需要它。我非常喜欢这个“录音室音质还是非录音室音质”的例子,在某种程度上,我们没有足够的语言来真正描述这种回声或者某种噪音的发生。我们就是没有足够精确的词。而我认为为什么拥有信息量相对丰富的文本标注非常重要,其原因就在于我们某种程度上依赖自然语言作为一种表征,但如果你基本没有足够的表征,这就意味着在给定语言条件下生成的产物是非常多模态的。如果你能从变分自编码器(VAE)等非常早期的研究中学到什么的话,那理念就是,我们真的希望在隐层表示中捕捉到大部分的随机性,然后给定 Z 时生成的 X 应该像是确定性的。所以……

Original English

Speaker C: Yeah, as a podcaster, I think a lot about sound. Um, and and I I'll just offer a couple things for discussion in case in case it triggers anything with you guys. Um, I have three domains of rough audio, which is like music, voice, SFX, you know, is that rough? Okay. Covers everything. And then also even within voice, let's just let's just focus on voice. Forget the other two. um room sound like the the echoiness of like big room, small room, in person, in a car, over a phone, all these like are labelable, but we experience them very differently. And I I often think like one of the tells of the AI video is that it is studio quality because it was recorded in a studio because that's your training data. And like and and to me that's one thing actually like the most interesting thing is just uh when I tell this is how I convince people who are kind of skeptical about the need for world models because you need it even for audio about well I'm further away from you so I should sound a little bit softer or more diffused and like the the video models need to pick that up because if they're going to do immersive video and audio you need that. I I I love that example of basically like studio quality or not in a way like we don't have enough language to really describe like like this kind of echoing or like some kind of noise kind of happening. we just like don't have precise enough and uh if you um you know basically the reason that I think it's quite important to have like relatively information rich like kind of captioning is that we kind of rely on the natural language as a representation but if you basically don't have enough uh representation that basically means the condition on the language the generation is very multimodal and if you anything can learn from the BAE kind like you know very old you know BA kind of research the idea is we really want to capture most of the stoasticity in the later representation And then the the X given the Z should be kind of like deterministic. So

Speaker B: 是啊。嗯,我希望在那里能有更多的进展,我也确信你们正在做……

Original English

Speaker B: yeah. Yeah. Um well I hope I hope there's more uh progress there and I'm sure you guys are doing...

Speaker A: 甚至其实像面部表情也是如此,对吧?也许这就回到了你关于那些我们非常敏感的事物的观点,对吧?我认为你也可以仅仅从人物的面部表情中看出很多 AI 生成的内容。

Original English

Speaker A: even actually like facial expressions, right? And maybe this gets to your point about like things that we're very sensitive to, right? I think you can tell a lot of AI content also just by from like people's facial expressions.

Speaker B: 是的。我们努力不成为这其中的推手,但是你知道的,或者像是皮肤纹理,对吧?比如那些在现实生活中让事物看起来真实的东西。就像,你知道,我可以从你点头的方式,或者你的微表情如何变化,来判断你对我所说的话是如何反应的。就好像我们还没能完全跨过那道鸿沟,我觉得虽然我们比一年前进步了非常多。

Original English

Speaker B: Yes. and we try not to contribute to it, but you know, um, and or or like skin textures, right? Like like the things that kind of make things look real in real life. Like I, you know, I can tell from the way you're nodding or from the way like your micro expressions are kind of changing of like how you're reacting to what I'm saying. Like we haven't quite crossed that chasm, I think like we're we're so much better than we were a year ago.

Speaker C: 对。

Original English

Speaker C: Yeah.

Speaker B: 嗯,但是在很多我们人类极其敏感的方面,其实还有很大的提升空间。我觉得在图像方面大概已经做到了,因为我确实看到了许多看起来几乎与现实无法区分的图像,我都分辨不出它们是不是生成的。

Original English

Speaker B: Um, but there's so much more headroom kind of in a lot of those things that like we as humans are super sensitive to. And like I think image arguably probably is there because there's there's a lot of kind of images that I will see that like really do look indistinguishable from reality and I can't tell if they're generated or not.

Speaker C: 它们甚至比现实更完美。

Original English

Speaker C: They're better than reality.

Speaker B: 呃,好吧,那是另一种……

Original English

Speaker B: Um, or well that's a different...

Speaker A: 不,我觉得其中一个主要的问题是……

Original English

Speaker A: No, I I think that one of the parents...

Speaker C: 比我去度假拍的照片还要好。是的。

Original English

Speaker C: better than what I would take on my vacation as a photo. Yes.

人类为何更偏爱 AI 生成的视频

Speaker A: 团队前段时间做过一个有趣的实验,就是看我们能否生成比真实视频更优秀的视频,对吧?所以你只需要截取同样的文本描述,比如某个视频的……

Original English

Speaker A: One of the fun experiments that we did a while ago in the team is is like can we generate videos that are better than real videos, right? So you just take the same caption from like oh yeah some video and then

Speaker B: 试一下。是的。只需试着去描述一个真实的视频,然后用 Omni 生成同等版本,接着做一个人为的评估看看效果如何,结果人类在很大程度上更偏好 AI 生成的视频。

Original English

Speaker B: try it. Yeah. Just just try to like describe a real video and then generate the equivalent version with omni and then do a human eval how does it do and then humans largely prefer AI generated video

Speaker C: 以很大的优势领先。

Original English

Speaker C: by a margin

Speaker B: 但因为这是由于强化学习(RL)的过程,是这个过程在起作用。

Original English

Speaker B: but because it's because it's the RL process. That's the process working.

Speaker C: 这取决于你怎么去为其合理化。它不一定就是旧的过程。它只是……我觉得它仅仅是……

Original English

Speaker C: It's however you want to rationalize it. It's not necessarily the old process. It's just like I think it's just...

Speaker B: 我并不是说这是一个好结果。我只是想说,我们已经以这样一种方式进行了优化,这种优化可能在某种程度上触发了人脑中的某种机制,让人觉得:“哦,很多视频看起来就是好得多。”

Original English

Speaker B: I'm not saying this is a good result. I'm just saying is we have optimized in a way that like kind of potentially sort of you know triggers something in the human brain that like oh it it looks it looks all a lot of the videos just look look better.

模型生成效果与审美偏好

Speaker A: 如果不仔细检查,或者说更深入地去审视的话,它们实际上并没有那么有用,或者说没那么好。但是如果你只是拿一个随机的 YouTube 视频和它的 AI 生成版本并排对比,你会发现,生成出来的版本就是看起来更好,因为它的画质更清晰,或者说有一种 HDR 的效果。呃,你知道的,肤色看起来也更好。再说一遍,这并不是说它更真实。

Original English

Speaker A: like I'm not on on on inspection on on deeper inspection they they would not actually be more useful or whatever but like if you just say side by side random YouTube video versus generated version of it will you will just have a it will just look better because it's more it's a sharper or HDR. Uh, you know, the skin tone is is is better. It's not, again, it's not more realistic.

Speaker B: 呃,这不一定能解决你的问题,但它看起来确实更好了。

Original English

Speaker B: Uh, it doesn't solve your problem necessarily, but it it looks better.

Speaker C: 我觉得这也取决于人们的敏感度。呃,我在日本出生长大,我觉得我了解到的一点是,他们对于类似这样的东西极端、极度地敏感,你知道的。这就是为什么,你知道,像建筑、像食物之类他们拥有的东西。嗯,所以我跟那边的一位漫画家交流过,他有点……他有点对生成式 AI 感到反感。他提到的一点是眼神的注视。眼神注视上的那种微小差异,会让我……让他觉得那种不自然的感觉很令人毛骨悚然。

Original English

Speaker C: I I since also depend on the sensitivity of the people. Uh, I was born raised in Japan and I think one thing I kind of know is like they're extremely extremely like sensitive about like, you know, that's why, you know, like architecture, like food and stuff like they have. Um, so I talked to like a manga like like artist there and he's like he's kind of disgusted by like the generation AI and one kind of thing he mentioned is like the eye gaze. Eye gaze that slight difference makes me makes him kind of feel creepy about like unnatural

Speaker D: 比如如果你的眼神看起来有点不对劲。

Original English

Speaker D: like if you're looking a little bit off.

Speaker C: 是的,就是觉得……是的。就是觉得看起来太假了。是的。所以,所以我觉得这确实取决于敏感度,而且……

Original English

Speaker C: Yeah. It's just uh Yeah. Just like uh it looks too fake. Yeah. So, so I think it does depend on the sensitivity and

Speaker A: 是的。是的。是的。我只是想说,人类的偏好,呃,并不是一个特别可靠的晴雨表,不能完全用来决定你应该优化什么。就像如果你只是问人们,你喜不喜欢这个?你并不一定能得到你想要的结果。

Original English

Speaker A: Yeah. Yeah. Yeah. All I'm saying is like, you know, human preferences are like not particularly like uh reliable barometer of like what you should be optimizing for. Like if you just ask people, do you like this or not? You're not necessarily get what you wanted.

提示词工程的价值与专业审美的差异

Speaker C: 是的。我想补充一点,大概四年前有一种争论,讨论提示词工程(prompt engineering)是否会消失,而且,呃,你知道有一些非常有影响力的人说它会消失。但我基本上说它不应该消失,因为提示词工程,也就是去具体设定那些东西,是你能够对 AI 保持某种控制的唯一方式,让你能够掌控 AI,而让你能够进行提示词工程的核心其实就是那种敏感度。所以,当然,也许现在 AI 可以做很多自动提示(auto prompting),它可以生成足够好的东西,但是呃,如果是那样的话,永远不要满足,永远不要满足于 AI 生成的内容,要不断微调你的敏感度,并且不断地用提示词去捕捉那些差异。我认为,对于普通人未受过训练的眼睛(我把自己也归为这一类,你知道,我有一些审美的感知能力,而且我做这行足够久了,所以我有我的偏好),但你知道,就像你举的那个漫画家的例子,那是一个可能在几十年里不断磨练技艺的人。任何做这种工作的人,无论是设计还是建筑对吧,你就是拥有一种完全不同级别的专业知识,你能看到普通人根本看不到的东西。但是 Doom 说得对。当我们在看这些东西的时候,如果你只是去街上随便拉 10 个人,他们很可能也会更喜欢那种过度平滑、色彩非常饱和的那种感觉。

Original English

Speaker C: Yeah. Let me just kind of add one thing but like four years ago there was a like debate that if the prompt engineering is going to disappear and uh my my like you know some very powerful people say you know it's going to disappear but I basically said like it shouldn't because the prompt engineering like sort of you know specifying that is like the the only way you can sort of control the output sort of you know when you have like sort of control over the AI and what allows you to prompt engineer is really that sensitivity. So sure maybe like right now the AI can do a lot of auto prompting and that and it can generate something that's sufficient but uh if it's like that never be satisfied like never be satisfied with the AI's generated content always fine tune your sensitivity and always kind of keep prompting the differences. I I think to the there's also a big difference between like the average human untrained eye which I I would put myself in that bucket you know like I have I have some aesthetic sensibilities and I've done this long enough that you know like I have I have a preference um but you know like your example of a manga artist like that's somebody who has honed a craft like over possibly many decades um and anybody who does that whether it's like design architecture right like you you you just have a very different level of like expertise and you see things that like the average human will not see. But Doom is right. Like when we look at if you were to just, you know, um pull 10 people on the street, they would probably prefer the like overly smooth like very saturated kind of

Speaker E: 这就是所谓的“Instagram 滤镜”效果。

Original English

Speaker E: It's called the Instagram filter.

Speaker C: 确实。确实是。是的。

Original English

Speaker C: It is. It is. Yeah.

Speaker A: 而且你知道,所以这里还有一个问题,那就是,如果你不具体指定的话,你的“默认审美”到底应该是什么样的?但回到 Shane 的观点,我们一直试图让这些模型做得更好的一件事就是指令遵循(instruction follow)。这样一来,当你想要让它们呈现出不同的结果时,你应该能够做到,不管是通过语言,还是通过提供参考图像,因为有时候语言太局限了。所以这些模型在这方面会继续变得更好,但它们能做的远不止于此。

Original English

Speaker A: And you know, and and so there's also a little bit of a question of like what does your default aesthetic look like if you don't specify? But then to Shane's point, one of the things we always try to get these models better at is instruction follow. So that like when you want to get them to a different outcome, like you should be able to, whether that's through language or whether that's through your references because language is sometimes too limiting. Um, and so like these models continue to get better at it, but they so much more.

设定默认审美的责任

Speaker D: 作为一名产品总监,你会不会觉得有压力,因为你要为全世界设定这个默认标准?我的意思是……

Original English

Speaker D: Do do you feel pressure as a as a product director to set the default for the world? Like I mean

Speaker A: 有点吧。

Original English

Speaker A: kind of

Speaker C: 也许我应该……我不知道。我还没想过这个问题。

Original English

Speaker C: maybe I should I don't know. I haven't thought about this.

Speaker A: 你知道的,你知道,总得有人去设定一个默认值。他们的默认值必须得存在。

Original English

Speaker A: You know, you know, it's like someone has to have a default. Their default has to exist.

Speaker C: 实际上,我想说我们确实考虑过这个问题。嗯,而且我觉得其中一个例子是,如果你看 Nano Banana 的生成结果,当 Nano Banana Pro 发布的时候,我们看到 Nano Banana 的信息图表(infographics)呈爆炸式增长。

Original English

Speaker C: Actually, I will say like we have thought about this. Um, and I I think one of the So, for example, actually like if you look at nano banana generations, we had like an explosion of nanobanana infographics when nano banana pro came out.

Speaker D: 我试过了。是的。

Original English

Speaker D: I tried it. Yeah.

Speaker C: 嗯,是的,是的,是的。我觉得 Nurb 的论文里全都是,你知道,有非常多的生成信息图表。你能不能在上面运行一下你们的水印(watermarking),看看有多少?

Original English

Speaker C: Um, yeah. Yeah. Yeah. I think Nurb's papers were like all, you know, so so many had like infographics generated. Can you run your uh watermarking on it and see how many?

Speaker A: 呃,我们可能……我们大概可以。我们还没有这么做,但我看到了太多了,比如我的 Twitter 上……也许这也只是我算法的偏见,但它们简直无处不在。嗯,这实际上非常痛苦,因为……我认为我们的默认审美有点太……太杂乱了。我觉得这个模型有点像一个过于急切的学生,它刚刚学到了……你知道,就像是“哦,我知道关于这个概念的所有这些信息”,然后把它全塞进同一张图片里。日本信息图表那种风格,甚至还要再乘以 5 倍……

Original English

Speaker A: Uh we probably we probably could. We we have we haven't done that, but I saw so like my Twitter was maybe this is just also like the bias of my algorithm, but they were everywhere. Um and it was actually very painful because um I think our default aesthetic was a little bit too it was too cluttered. Like I think that the the model was like a bit of an overeager student that just like learned you know it was like oh I know all these I know all this information about this concept. and like shove it into the same image. Japanese infographics 5x that

Speaker D: 或者也许这是……你知道,嗯,但它只是……而且,而且……

Original English

Speaker D: or maybe it was you know um but it just and and

Speaker C: 等等,所以同样的提示词,同样的内容,如果是用日语的话它就会……

Original English

Speaker C: wait so same prompt same content if it's in Japanese it's

Speaker A: 密度(density),信息密度极大。

Original English

Speaker A: density density

Speaker C: 哇哦。

Original English

Speaker C: oh wow

Speaker A: 因为那是日本的风格。

Original English

Speaker A: because that's the style in Japan

Speaker C: 是的,有一些非常,你知道,有一个很有名的词叫“官僚主义(bureaucrat)”就是用来形容这个的,是的。

Original English

Speaker C: yeah some like very you know bureaucrat is a famous word for it yeah

Speaker A: 不,但我们在做 Omni 的时候确实经历过这个过程,我们是一起做的对吧,在那时我们有很多……我们在最后阶段,我们会说,好的,这就是……我们做了一些微调,然后决定“好吧,我们到底偏好哪种风格呢?”对吧,你知道。

Original English

Speaker A: no but we do do go through this process with Omni we did it together right like where like we had like a bunch of like we like at the very end okay like this is we did some tuning and like okay what kind of style style do we prefer right like you know

Speaker D: 比如是更柔和(muted)的,还是更饱和(saturated)的?

Original English

Speaker D: is it more muted more saturated

Speaker A: 我们那时候有很多高饱和度的选项。

Original English

Speaker A: we had a lot of saturations

Speaker C: 是的,有过,有过,确实有过。我觉得 Nicole 只是有了创伤后应激障碍(PTSD),所以把它给忘了,但她非常深度地参与了这件事,比如“好吧,我们基本上更喜欢哪种调色板?”对吧。而且,你知道,这不像是一个你必须做出权衡取舍的东西,就像,呃……

Original English

Speaker C: yeah there was there was there were I think Nicole just has PTSD so has forgotten about it but she was very much involved in this of like okay which which kind of color palette do we basically prefer right and it's you know it's it's it's not something that like you have to make a a trade-off there like uh

Speaker A: 而且,而且,而且这是因为最终做决定的是我们,对吧。其实这是真的,它最终是由建模团队来决定的。而你可以合理地提出这样一个问题:我们是做这件事的最佳人选吗?还是说,我们实际上应该和一些拥有真正创造性视角的人合作,比如那些更像艺术总监的人,让他们来……我们在这件事上其实有些反复。嗯。

Original English

Speaker A: and and and it's because it ends up being us right like actually it is true like it it ends up being the modeling teams and you could ask the question legitimately of like are we the best people to do that or should we actually work with someone who like has a really creative point of view and is more of like you know an art director and like has like and we kind of go back and forth on this. Um

反馈与评估机制

Speaker D: 我们有受信任的测试人员(trusted testers)。我就是在……

Original English

Speaker D: we have the trusted testers. I'm on

Speaker A: 我们确实有,我们有受信任的测试人员,他们给了我们很多反馈,而我们会非常认真地对待。

Original English

Speaker A: we do we have trusted testers who give us a lot of feedback and we take that serious

Speaker D: 顺便说一句,他们组织得非常棒。他们有这种类似于每周电话会议之类的东西,真的很棒。

Original English

Speaker D: very well organized by the way they have these like weekly calls and stuff like it's it's amazing.

Speaker A: 嗯,Logan 的团队做了很多这方面的工作。所以真的要称赞一下 Logan,嗯,他今天没法来。而且我们内部,在 Google 内部有很多人,比如 Fulfur,他们给了我们非常多的……不,不,不。说真的,当我们发布新的检查点(checkpoints)时,他们给了我们大量的反馈。而且有时候是我们自己根本看不到的东西,对吧,比如我们会觉得,“哦,是的,这个优化看起来不错”,然后他们就会回馈说,“你们都干了些什么?你们把我的草地完全给毁了”,你知道,因为现在细节全都模糊了。

Original English

Speaker A: Um Logan's team does a lot of that. So kudo kudos to kudos to Logan um who couldn't be here today. Um, and we have a lot of people actually internally at Google like fulfur who give us like a ton of No, no, no. Truly like who give us a ton of feedback on like when we when we release new checkpoints and like sometimes it will be stuff that we like don't see right like we would be like oh yeah this optimization seems okay and then they would come back what have you done like you completely ruined my grass you know because now the detail is all blurry.

Speaker B: 我想你刚才注意到了,目前这也不是什么超级机密了,就是我们的模型倾向于在手上画上戒指,甚至是结婚戒指。那真的是,是的。

Original English

Speaker B: I think you just noticed not not a super secret at this point but like that our model tends to put rings wedding rings on on on hand. That's yeah

Speaker A: 非常奇怪。我以前从没注意到过这个,但他好像是,他……我刚好看到,然后基本上有一个 Faux Fur 的频道……

Original English

Speaker A: very strange. I had never noticed that but he's like he I just saw it and there's a faux fur channel basically

Speaker B: 在他发帖的地方,我当时就想,为什么每只手上都有结婚戒指?我当时觉得这太奇怪了。

Original English

Speaker B: where he posts I was like why is there a wedding ring in every hand I'm like this is strange

Speaker F: 听起来非常像是典型的奖励欺骗(reward hacking)。

Original English

Speaker F: it sounds very common reward hacking

Speaker A: 是的,是的,是的,所以,但你知道,这些是我们如果不这么做,在开发过程中未必能注意到的事情,对吧。

Original English

Speaker A: yeah yeah yeah so but you know something that we would not have we would not have noticed necessarily while while developing this right

Speaker D: 这是一个强化学习(RL)的伪影,还是……我,我也不知道。

Original English

Speaker D: is it an RL artifact or I I don't know

Speaker A: 你,你确实有很多基于偏好的设定,然后你知道,你可能会偏好那种虚假相关的奖励欺骗,它可能会以许多奇怪的方式发生,是的。

Original English

Speaker A: you you do have like a lot of preference based and then you know you may can prefer that spirious correlation reward hacking it can happen like in many weird ways yeah

Speaker F: 确实如此,确实如此,呃,是的,这关系到另一个话题,我想重申一遍,我,我试图利用这些主要舞台上的活动作为引入或联系。呃,我们有评估(eval)轨道,我们有 Character AI,还有 YouTube 在讨论他们如何评估视频。嗯,你们是如何评估视频的?

Original English

Speaker F: it does it does uh yeah this is related to another topic that again I I try to use these mainstage things as introductions or ties in. Uh we have the eval track we have character AI and YouTube talking about how they evaluate videos. Um how do you evaluate videos

Speaker A: 除了 Furer 之外吗?

Original English

Speaker A: apart from furer

Speaker F: 并不是每个人都有 Fauxur,而且你也知道,我认为需要有更多的定量指标。

Original English

Speaker F: not everyone has a fauxur but also you know I think there needs to be something more quantitative

Speaker A: 好吧,我的意思是,你们改进 Gemini 是为了改善 VR 的演进。

Original English

Speaker A: well I mean it's you improve Gemini to improve the evolution for VR.

模型评估与数据需求的挑战

Speaker A: 是的,那绝对是一种方法。不过这实际上非常困难。

Original English

Speaker A: Yeah. Um that's that's no no that's that's definitely one way uh it's actually very hard.

Speaker B: 非常困难。要让审核员去评估视频里的东西,尤其是像美感之类的元素,真的非常困难。有些东西可能稍微客观一点,特别是当我们讨论图像,或者看一些信息图表和文字渲染的时候,这其实还好办。

Original English

Speaker B: It's very hard. It's very hard um to get like you know audators to evaluate things in a video like including especially things like aesthetics right like that it's like there are some things that are a little bit more objective like especially when we talk like let's say we talk about images and we look at like infographics text rendering that's actually fine right because like

Speaker B: 因为你可以通过 OCR 把文字提取出来,然后你可以观察,比如“好吧,这个字母乱码了”,那整个素材实际上就没用了。因为如果在渲染文字里哪怕只错了一个字母,你就根本无法使用这个素材,对吧。所以我们发现,这些东西相对来说更容易自动化评分。但我们还是非常依赖人工去观察,所以我们会做大量的人工评估,做非常多的人工评估。

Original English

Speaker B: you can kind of OCR things out and then you can look at like okay this letter is like messed up and then the whole thing is actually useless because if like literally if the letter is off in render text you just can't use that asset Right. So th those things are like a little bit more auto ratable um from what we found. We do rely a lot on humans looking at things and so we do do a lot of human evals. We do a lot of human evals.

Speaker B: 每次我们有了新模型,我们就会想要做更多的事情,想要塞进更多的能力,然后我们就得运行更多的人工评估。到了某个阶段,你会得到两个非常接近的模型,然后我们真的就要靠并排观察输出结果来做决定。有时候我们大概 10 个人待在一个房间里,就只是并排看两个视频,然后问:“你更喜欢这个,还是那个?”

Original English

Speaker B: Do a lot of human ev and every time Jane is like um and every time we have a new model we like want to do more things and we want to like jam in more capabilities and then we have like more emails that we have to run. Um, and then at some point you do get two models that are like kind of close to each other and then like we literally make decisions based on like looking at output side by side. Sometimes like in a room like I've been in rooms where there's like 10 of us and we're just like looking at videos side by side and we're like do you prefer this or do you prefer that? Like

Speaker A: 哇哦。

Original English

Speaker A: wow.

Speaker B: 我是说,这真的是非常复杂。你添加的能力越多……你要知道,甚至仅仅添加一项能力,但这几乎就是 AGI 级别的完整能力了,比如视频剪辑,把视频剪辑看作一项能力,还有带音频的剪辑……

Original English

Speaker B: I mean but it is it is genuinely very complicated. the more capabilities you add like you know even just the one capability but it's like almost AGI complete capabilities like video editing right like think about video editing as a and like editing with audio and

Speaker A: 剪辑师听到这个会很高兴的。

Original English

Speaker A: editor will be very happy to hear this editor

Speaker B: 我的意思是,这可以说是生成式媒体中最难的问题了。

Original English

Speaker B: I mean it's the hardest problem in g media

Speaker A: 我倒不知道这算不算最难的,但这绝对是非常难的,对吧。就评估的复杂性而言,比如自由形式的视频剪辑,你可以做任何事情……

Original English

Speaker A: I mean I don't know if it's the hardest but it's definitely there right like uh in terms of like complexity of of evaluation like free form video editing is you can do anything like yes uh and like

Speaker B: 我在这上面花了很多钱,真的非常难。请帮帮我。

Original English

Speaker B: I I spent a lot money on that and it's very hard. Please help me

Speaker A: 比如添加那些我们现在还没有的功能,比如添加一个树懒相关的评估,对吧?我们目前还没有。

Original English

Speaker A: like adding those we don't have like add a sloth eval, right? Like uh that we

Speaker B: 好吧,现在我们应该有了。

Original English

Speaker B: Well, now we should.

Speaker A: 我们现在确实应该有了,是的。但是像这种事情,要做起来真的没那么容易。

Original English

Speaker A: Now we should. Yeah. Yeah. Yeah. But like things like that like it's it's it's not that easy to

Interviewer: 我想我只是对你们的样本量感到惊讶,对吧?为了测试你们模型的整个应用面,你们依然需要依赖数百个人工评估量级?

Original English

Interviewer: I think I'm just surprised at the sample size that you have, right? Like to to test the entire surface of your models, you still rely on a magnitude of hundreds.

Speaker B: 不不不。我们实际上会对成千上万的项目进行大量的人工评估。我觉得这里面还有一点,我们可以谈谈诸如实时实验之类的事情。在这里你可以在更大的规模上获得一些关于微小差异的信号。然后还有自动评估器,这绝对是一个边界更清晰的空间,我认为它在 LLM 领域比较成熟,但在媒体模型中还处于非常早期的阶段。此外,有时候你仍然需要依赖人类的判断。我们非常依赖一些有着独特审美的人,以及那些在日常工作流程中使用这些模型的人的反馈,对吧?因为你也可能会遇到这样的情况:一个模型在部分人工评估切片上表现得非常好,但它却真的破坏了某个人的工作流程。所以这就是为什么我们要搞早期访问计划,试图获取反馈,并在更广泛发布之前将其整合进去。我觉得 Shane 刚才看表情好像有话要说。

Original English

Speaker B: No, no, no, no. So we do like Yeah. Well, we do we do a ton of human evals on like on like you know thousands of things. Um I I think there's also like an element of you know we can talk about things like live experiments right like which which is also where you get signal on like like some of these more minute differences at like much larger scale then there's autoators which is definitely kind of a more it's a very well defined space I think for LLMs much more nent for media models and then like sometimes you still do rely on human judgment and we do rely on things like feedback from people who just like have a very owned like aesthetic and and people who just like use these models in their workflows day-to-day, right? Because we could also like you could have a model that does really well on some slice of human evals, but then it like really breaks a workflow for somebody. And so this is why we do like early access programs and we try to get feedback and then we like try to incorporate it before we release something more broadly. I feel like Shane had a hot take based on his

Shane: 是的,每次我们谈到这个的时候我都有这表情。我认为各种人工劳动应该逐渐被分摊掉,然后有趣的是视频理解,特别是对比生成视频来说,检测错误这类工作是一个非常有趣的视觉任务。

Original English

Shane: expression always when we were talking about this. every kind of human sort of you know work should be gradually kind of amortized and then the interesting thing is the video understanding especially like against like a gener video like detecting air stuff is extremely interesting uh vision task

Shane: 然后有些部分涉及到美学或者视觉质量,但在某些情况下,它在语义上说不通。举个例子,你拿电影里的一个著名场景来试图重建它,然后你生成了一些东西,但在某个关键点,一部分语义信息讲不通,比如出现了前后矛盾。那么,AI 真的能检测到这些吗?所以当我在评估 AI 视频时,我当时就觉得,哇,我感觉自己太聪明了,你知道吗。就好像 AI 现在仍然有点落后,但我们应该在这方面投入大量精力。我认为视频理解是一项极其重要的智能任务,它超越了纯粹的美学或偏好。是的,我们应该始终努力宣传人工标注。

Original English

Shane: and then like some of it kind of aesthetics or this kind of visual quality but for some of the kind of cases like semantically doesn't make sense for example you're taking like some like a famous scene from a movie and try to sort of um construct that and then if you kind of generate uh it can generate something that but at some point some of the semantic information doesn't make sense like it's actually inconsistent. So can the AI actually detect that? So when I evaluate the AI video I was like oh I feel I'm so smart you know like that like it's like like AI is still kind of behind but we should make like a lot of effort. I think the video understanding is extremely uh important intelligence task uh beyond just the pure aesthetics or the preference. Um and yeah, we we should always try to advertise the human

Speaker B: 人工标注。是的。

Original English

Speaker B: human label. Yeah.

Shane: 是的。

Original English

Shane: Yeah.

Interviewer: 嗯,那你们需要什么样的数据呢?我交谈过的很多人其实都很想跟你们接触。他们想客气一点表达,但他们有大量的视频数据,有游戏数据,有现实世界的视频数据,有图像,还有标注员。你们想要什么?

Original English

Interviewer: Um what data do you need? A lot of people I talked to wanted to get in front of you actually. Uh they I mean they want to be nice about it. They have a lot of video data. They have gaming data. They have real world video data. They have images. They have lablers. What do you want?

Speaker B: 你这是在提议合作吗?我只是觉得这就像是你提出的请求,“好吧好吧,我们明白了”。我确信你们会收到很多推销,很多人想和你们交谈。但我觉得真正的问题是从噪音中筛选出信号。所以,创建一个好的 API,比如“好的,如果你们实际去做了 A、B 和 C,那我们就会感兴趣”,这才是关键。

Original English

Speaker B: Are you like offering? I'm just like this is your request for like okay okay we get I'm sure you get a lot of pitches right you get a lot of people want to talk to you what's like I think actually it's the signal this problem this sorting out signal from noise is the main problem so creating a nice API of like okay if you actually do a b and c we are interested in that

Shane: 嗯,这是个很难回答的问题。所以我不知道是否有简单的答案。我觉得我们确实已经拥有了大量数据。我认为在公开场合谈论这个很难。

Original English

Shane: um loaded question there so uh I don't know that there's like an easy like you know if you do I I think we we do already have a lot of data. I think it's it's hard to talk about this

Interviewer: 在公开演讲中谈论这个确实。我不想给你惹麻烦。

Original English

Interviewer: in a talk about the public. I don't want to get you in trouble.

Shane: 但是我想说的是,要在不泄露我们的项目细节和未来走向的前提下讨论这个问题,确实有点困难。大体上,高质量的数据……我想我们也许可以这么说,这并不是什么秘密。

Original English

Shane: But like I I think no what I just want to say is like hard to talk about this in a sort of you know without trying to without I have to think about the what I am revealing about our project and what where we going. Um generally high quality data I think maybe maybe let's just put it this way right it's not not the secret

Interviewer: 具身智能,不好意思打断。

Original English

Interviewer: embodied I'm sorry

Shane: 具身智能数据,是的。我是说我们之前大概已经公开宣布过,我们会开展一些机器人领域的合作。因为我们在 GDM 有一个机器人团队,所以你知道,他们总是对这类东西感兴趣。就 OMI 而言,我认为我们只是对高质量数据很感兴趣。这不一定非得是那种随便从 YouTube 上找来的随机视频,而是更偏向专业制作的东西,类似那样的,那些才是我们一直在寻找的东西。

Original English

Shane: embodied data I mean yeah sure I mean we have we have sort of announced I think publicly right that we we'd have some sort of robotics collaboration right like so I think it's like or but you because we have a robotics team at GDM so you know they're always interested in things like that um I mean for OMI specifically I think we're just quite interested just high quality data right like you know it it's not some sort of not necessarily like oh random YouTube video but like you know some some more professional shop things like that right the things that those are those are things that we're always on the lookout for like uh and yeah

Speaker B: 而且我认为,在某种程度上,可能稍微容易回答一点的是关于一些基于智能体的工作。人们真正在试图做的测试是什么?如果你只是自己去做,或者和供应商一起做,这些其实很难凭空制造出来。如果你正在创建一个营销活动,那到底应该是什么样子的?你会不会从“这是一张我新产品的图片”开始,然后想把它变成一个视频广告,并且还想把它转化成一堆适配各大平台的各种不同广告格式的素材?你要怎样从这端走到那端?你在这一路经历的这些任务轨迹是什么?这才是真正有用的,但这其实很难获得,因为我们并不总能拥有人们实际执行这些操作的第一方工作台。或者你可能与某位供应商合作,但他们也没有那种产品界面。像很多这类信息,其实是存在于人们实际执行这些测试的地方,所以获取起来相当困难。如果有人能解决这个问题,你应该联系我们。

Original English

Speaker B: and I think for you know maybe this is easier to some extent to answer for like some of the agentic work as well like like like actual kind of like what are the tests that people are trying to do right these things are actually kind of difficult to manufacture if you're doing it yourself or if you're like doing it with a vendor like what is the actual like if you're creating a marketing campaign like what does that look like right like do do you start from here's like a picture of my new product and then I want to turn that into a video ad and I want to turn that into a bunch of assets that like fit fit all these different ad formats that I need to push onto various platforms to promote and then like so you kind of go from this to that and like what is that kind of trajectory of tasks that you're that you're like you know experiencing along the way like that is really useful and that is actually kind of difficult to get right uh because like we don't always have the right first party surface where people are actually doing some of these things or like you might work with someone who's a vendor but they don't also don't have that product surface right like like a lot of this kind of information lives in the places where people are doing these tests and so that's kind of difficult to get like if anyone's figured that out you should reach out to us

Interviewer: 每一个渠道的思考,每一个渠道的思考缺陷。

Original English

Interviewer: every channel thought yeah every channel thought fault.

Speaker B: 是的。

Original English

Speaker B: Yeah.

Interviewer: 也许数据,这家中国实验室的……

Original English

Interviewer: And maybe the data the Chinese lab is

数据的商品化与高质量工艺

Speaker A: 是的。作为一个媒体人,我知道现在有太多播客、营销部门的人等等,他们很乐意成为你们的数据来源,比如“在我的头上戴个脑机接口(BCI),跟我们交流,观察我的工作”。因为需要做的工作无穷无尽,这里面有很多应该在某种程度上商品化。显然,你可以成为像好莱坞那样追求极高品质的艺术家或工匠,但实际上很多工作是商品化的,是可以被模型化的,我们希望你们能做到这一点。

Original English

Speaker A: Yes. Yeah. uh you know uh yeah as a media person myself right like there's so many podcasters and people in in marketing departments and all these like they would happy to be your data like you know just like put a BCI on my head talk to us watch my things uh because like you know there's just endless amount of work to do like there's so much work and this is all like this needs to somewhat be a commodity like obviously you can be an art like an artisan like you can be Hollywood for like the really high quality stuff but actually a lot work is commodity and like should be modelable and we want you to do it

Speaker B: 但正如 Dumi 提到的,我们也确实需要高质量。

Original English

Speaker B: and but we we want the high quality to Dumi's point right like we do want we want the high quality

Speaker A: 我们需要商品化,是的,两方面都想要,我……

Original English

Speaker A: we want commodity yes yes you want on both sides um I I

Speaker B: 感谢你的提议。我们也在数据质量方面进行了投入。我认为人们想了解在人工智能领域如何提高标准,很大程度上这只是在教育市场,向研究人员、工程师和创始人科普“这就是我们未来的方向”。这其中有很多是“停止那样做,改用这种方式”,我觉得人们会听进去的。

Original English

Speaker B: thank you for the solicitation uh I you know we we we also I also added a data quality track I I think that uh people want to understand like what uh at AI like how to raise the bar right like like the and a lot of it is just educating the market and educating researchers and engineers and founders on like this is where we're going a lot of this is stop doing that do this do this instead and I'm like people will listen

Speaker A: 是的,我不知道是否能达到那种程度。

Original English

Speaker A: yeah I don't know uh to that extent you know

Speaker B: 但我认为关于这一点,这里面同样有很多“工艺”的成分,有很多流程。比如即使是那个营销活动的例子,你也不是在五分钟内就能把它做出来的。你要经历一个过程,不断迭代,然后因为某个原因——比如觉得人物的眼神对了——你选择了这个方案而不是那个。因为我们都不是营销总监,而且模型也不懂这些。

Original English

Speaker B: but I think to that to that point like there's a lot of again just like craft that goes into this right and there's a lot of process like you even to the marketing campaign example you don't create that in like five minutes right you like go you go through a process and you iterate Great. And you like pick something over something else because you liked it for whatever reason. Like maybe the eye gaze was correct, right? Like we just we don't know these things, right? Because none of us are marketing directors and like the models don't know these things.

人类经验中的“暗知识”

Speaker C: 我甚至觉得,对于自然语言来说也是如此。我总是说,99%的信息都存在于人们的脑海中,你只能通过主动对话和与他们交朋友才能提取出来。因此,互联网上的大多数东西其实只是这种互动的产物或输出。关于那些轨迹、这个人是如何产生写这篇论文的灵感的、起点是什么、灵感是什么、是哪些对话激发了它,这类东西其实都深藏在人们心里。所以我认为,即便是语言领域也是如此,创意领域也很相似,里面存在着大量的“暗知识”。

Original English

Speaker C: I even kind of say this for the natural like a language as well. Like I I always kind of say 99% of information is inside people. You can only extracted through active dialogue and befriending them. So most of the stuff on the internet is like sort of the outcome the output of that. Yes. about you know what are what are all the trajectories you know how did this person have this inspiration to write this paper what is the starting point what is the inspiration what are the dialogue that sparked it those kind of stuff is kind of inside people so even you know those kind of like even the language space is kind of that I think the creative is kind of similar as well there's a lot of dark knowledge

Speaker A: 是的,就像你写小说一样。一本小说能打动你,通常是因为你对故事、发展轨迹或人物有一种个人连接感。如果你读读现在由 LLMs 写的大部分内容,它往往会陷入那些默认的模式中,语言开始变得非常相似,所有的描述听起来千篇一律。你稍微读一下就会觉得:“哦,这不太有趣”,因为我无法与之产生共鸣。再说一次,这属于一种人类特有的专业知识。

Original English

Speaker A: yeah it's like when you write a novel right like a novel speaks to you because like usually there's some sort of like a personal connection that you feel to like the story or the trajectory or the characters right like if you read most of the stuff that's written by LLMs today. Like it's, you know, it's it's it starts it falls into these like default par patterns and like the language starts to feel really similar and all the descriptions sound really similar. You can kind of like quickly read it as like, oh, this is not that interesting because like I can't connect to it, right? Um, and again, that's that's kind of like a human expertise.

FDE(前沿数据工程)的演变与角色

Speaker C: 最近有一件好事是,Google Cloud 和 Google DeepMind 开始在产品工程师的 FTEE 上投入更多。我也看到了一些针对创意和媒体领域的招聘。我认为这些都是真正的努力,因为我们感觉到,仅仅依靠公开数据能做的事情是有限的,但如果通过与这些人才合作,我们就能提供更好的模型、产品和反馈。

Original English

Speaker C: One nice thing recently is the Google Cloud and the Google Deep Mind are kind of starting to invest a lot more in the FTEEs for the product engineers. And I also kind of saw some uh recruiting for the creative you know gem media kind of space as well. So I think those are kind of really the effort because we we kind of feel that you know what we can kind of do with a lot of public data there's limits but we're you know partnering with that we can provide kind of better models and products and yeah kind of feedback

Speaker A: 我们这里第一次有了 FD(前沿数据)方向,每个实验室都在宣布相关举措,这太疯狂了。有一件事我其实非常热衷,并且我在 Cognition 时也大力推动过,那就是将 FTEE 不仅仅转化为销售和解决方案,而是将其转化为评估(eval)人员。FD 远不仅仅是销售,它的意义要大得多。那么你们是如何定位 FD 的呢?因为我确实把它看作一种销售,就是你越是为客户定制解决方案……

Original English

Speaker A: uh we have an FD track here for the first time every lab is announcing it. It's it's crazy. Um, one thing I'm actually very keen on doing and I pushed a push for this at cognition as well is to turn the FTEES not just into sales and solutions but also to evalu uh eval workers. FD is not the sales FD is way way bigger than that. How do you frame FDs then? Because I do think about it as sales like you're you know the more the more you customize the solution for

Speaker C: 我对后训练(post-training)的定义是,介于预训练和最终用户体验之间的任何事情。任何处于这两个环节之间的工作,都是后训练。

Original English

Speaker C: so I define post training as anything between the pre-training and the final user experience anything anything is a post training

Speaker A: 对我来说,当我最初开始深入了解时——我是说,FD 最初可能是从某种路径演变而来的,所以我想它的历史背景有所不同——但重点不仅是与他们合作,确保他们知道如何使用模型,还要通过编写代码或衍生出洞察力,从根本上帮助双方。他们可以在如何使用模型上建立起很好的测试框架,而我们则可以在非常上游的地方进行改进。所以,如何将客户反馈传导到模型构建中,我觉得这更像是我期望 fds 发挥的作用。

Original English

Speaker A: and to me when I first sort of you know learned a lot about I mean FD kind of I guess originally you know came from like path here and then that so I guess the kind of history is different but yeah I think the key is really that um you know the key is like not only to kind of work uh with them and ensure that they kind of know how to but also to sort of code like derive kind of insights that can basically kind of help both parties. They can put the like a lot of harness how they use the model. We can improve like very upstream. So how to get the customer feedback to the modeling I feel is the kind of more the the role I I kind of want for the fds. Yeah.

连接模型与真实世界需求

Speaker B: 是的。而且就算只是顺着那个话题,如果你们想跟我们聊,或者至少跟我聊聊——我不会去占用你们的时间——但对于我们来说,去真正和那些正在使用我们模型的人交流,了解他们在哪里遇到困难,真的非常有帮助。因为那是你们真正想用模型去解决的现实世界任务。比如,我会和一些使用我们图像模型进行室内设计的人交流。他们会说:“嘿,我很想采用这个图案,但我希望将其缩放到10种不同尺寸的地毯上,有时我甚至需要非常定制化的地毯尺寸”,但模型在以相同方式复制该图案时失败了。或者“我想试戴这些耳环,耳环有特定的大小,而我的头部也有特定的大小”,如果你真的在做虚拟试穿,比例必须合理,而模型却在许多这类实际发生于真实世界的事情上失败了。因此,这对我们很有用,因为我们不去想这些事情,毕竟我们不把模型用于那些任务。

Original English

Speaker B: Yeah. Yeah. and and even if I sorry just on that like if you want to talk to us or at least me um I I'm not going to offer up your time um but I it's really helpful for us to actually talk to people who are using our models and like understand where they're struggling uh because again that just like it's it's the real world task that you're actually trying to use them for right like I will talk to people who do kind of interior inter interior design with some of our image models um you know and they will say hey like I really want to take this pattern pattern, but then I want to scale it across like 10 different ruck sizes and sometimes I have like a very custom ruck size and then the model fails at like replicating the pattern the same way. Or, you know, I want to do a try on for these earrings and then the earrings have a certain size and then like my head has a certain size, right? Like it has to make sense if you're actually trying to try things on and like the models kind of fail at a bunch of these things that like actually happen in the real world, right? Um, and so that that's like useful for us because for some of these things like we don't think about because we don't you know we don't use the models for those tasks

Speaker A: 或者就像你刚才提到的广告活动之类的,人们对品牌语言等有特定的概念,而这……

Original English

Speaker A: or like um you know I think to your point about ad campaigns or whatever like people have like notions of brand languages or whatever like which is

Speaker B: 是的。

Original English

Speaker B: yes

Speaker A: 这就像是用一堆图片或PDF来传达信息,这是一个相当模棱两可的问题。什么是宜家的品牌语言?它是蓝色和黄色吗?我的意思是,那不是一个非常……

Original English

Speaker A: like a a bunch of images or PDFs saying things you know it's a pretty kind of you know ambiguous question as well. What is the IKEA brand language? You know is it is it blue and yellow? I mean that's that's not a very like

Speaker B: 而是到底是什么色调的蓝色。

Original English

Speaker B: but like what shade of blue you know.

Speaker B: 没错。这些品牌有非常具体的要求,他们真的很在乎蓝色的具体色度。那不应该只是一种随机的蓝色和一种随机的黄色,那不会是宜家的风格。这只是个例子。但这正是那一类事情:它不一定属于我们“开发前沿模型”必须完成的任务,但我们确实想在根本上构建人们能用来解决具体任务的产品,而不仅仅是研究产物。因此,了解人们真正在意什么是很有用的。

Original English

Speaker B: Yeah. Yeah. Yeah. So there there's like, you know, and the brands are pretty specific, you know, pretty, you know, like they they do care about the shade of blue. It's not shouldn't just be a random blue and a random yellow. That's not going to be IKEA, right? I'm just thinking about an example. But like this is the kind of stuff that, you know, it's not necessarily part of our like, you know, developing frontier models kind of, you know, necessarily mandate, but it's something that we do want to we do want to fundamentally like build products that people will use to solve concrete tasks, not just not just research artifacts, right? So I think it's useful to understand what people do care about.

结语与展望

Speaker A: 我确信很多人非常感激你们的工作。你们还有很多工作要做,在过去的几年里,从 Nano、Banana、Leo 到 Omni,你们取得了如此大的进展,我不知道你们还在酝酿些什么,但我们非常激动。这也是让我曾在 Sora 关停时感到非常失望的原因之一。我认为需要对生成式模型进行更广泛的探索,而不仅仅是编码领域。我觉得那就是……

Original English

Speaker A: Uh well, I'm sure a lot of people are very grateful for your work and there's a lot more to do that you've made so much progress over the last like even just a couple years of like Nano Banana and Leo and Omni and uh I don't know what else you got cooking but we're very excited like you this is one of those things where like I was very disappointed you know when with when Sora shut down and and I think like there needs to be more general exploration of uh you know generative models and not and not just you know coding. I think I think that is

Speaker B: 显然,我们喜欢这个方向。

Original English

Speaker B: we obviously like this.

Speaker A: 我们也热爱编码,非常热爱。但非常感谢你们抽出时间,这次交流真的很愉快,我迫不及待想看看接下来的成果。

Original English

Speaker A: We love coding. Love coding and and uh yes. Uh but thank you so much for your time. Uh it's been a real pleasure and I can't wait to see what this looks like next.

Speaker B: 感谢您的邀请。

Original English

Speaker B: Thank you for having us.

构建个人 AI 研究操作系统

Speaker A: 问得好。

Original English

Speaker A: Great question.

Speaker A: 谢谢大家。我来解释一下。在我的第二大脑中,目前在 Obsidian 里有超过 5000 条笔记,在 Readwise 里也有 5000 条,还有一些零散分布在 Notion 和 Google Drive 里。所有这些以平均每个月 250 个文件的速度在增长。而这正是我想要的。在左边,你们可以看到我整个的 Obsidian 知识库,这一大堆内容。每当我开始处理某件事,比如写一篇文章、启动一个新项目、一个新的代码库、一项新功能等等,我希望能够提取出对我当前工作真正有用的高价值节点。你们可能会问,为什么不直接用 Codex 代码或 NotebookLM 呢?事实上我也在用,但你需要一个介于这些工具和你的第二大脑之间的系统。好吧,让我们回到我问题的根源,那就是我总是丢失我的研究资料。比如,我的阅读清单简直就是个坟墓。当我在刷社交媒体时,我保存了一篇很酷的 X 帖子、一篇新文章、一个新 YouTube 视频,或者一个 GitHub 仓库,这都无所谓。每当我真正想开始着手做点什么时,我永远想不起来我的第二大脑里有什么,或者我不得不花大量时间去寻找能在工作中派上用场的有意义的笔记,对吧?我遇到的另一个问题是,我希望这个系统能真正锚定我的个人笔记、我的个人价值观和我的个人信仰。我希望这个系统是个性化的,能反映我自己的思想,对吧?这就是为什么在今天的视频里,Luis 和我将教大家如何构建自己的 AI 研究操作系统(OS)。这还附带了代码,所以你们也可以自己动手试试。

Original English

Speaker A: >> Thank you everyone. Let me explain. So within my second brain, I currently have over 5,000 notes in Obsidian and another 5,000 notes in Readwise and some scattered in Notion and Google Drive. And all of this is growing on average with 250 files per month. And this is what I want. On the left, you can see my whole Obsidian vault, this huge mass. And whenever I start working on something such as an article, a new project, a new codebase, a new feature or whatever, I want to actually pull high signal nodes that are actually useful for my current work. And you would ask yourself, why not use directly codex code or notebook LM? And the thing is that I am, but you need a system that sits between those harnesses and your second brain. Okay, so let's go back to the root of my problem, which is that I'm always losing my research. For example, my reading list is a graveyard. When I'm scrolling social media and I save that cool X post, a new article, a new YouTube video, a GitHub repository, it doesn't matter. Whenever I actually want to start working on something, I never recall what I have in my second brain or I have to spend a ton of time actually finding meaningful notes that I can use in my work, right? And another problem that I have is that I want this system to actually be anchored into my personal notes, into my personal values, into my personal faith. I want this system to be personal, to reflect my own thoughts, right? And that's why in today's video, Luis and I will teach you how to build your own AI research OS. This also comes with code, so you can also try it out yourself.

Pauline: 我是 Pauline。我是 Decoding AI 的创始人和 CEO,我在那里制作了大量关于如何交付 AI 产品的课程内容,同时我也是《》一书的合著者。

Original English

Pauline: And I'm Pauline. I'm the founder and CEO of Decoding AI where I do a ton of content on courses on how to ship A products and I'm also the co-author of the

介绍 Arya:AI 研究与迭代智能体

Tim Sweeney: 好的,大家好,感谢大家参加这场分享。我叫 Tim Sweeney,是 Weights and Biases 和 Coreweave 的首席工程师。在接下来的 20 分钟里,我们要探讨的是 Arya,这是我们全新的 AI 研究与迭代智能体。那我们这就开始吧。首先,为了活跃下气氛,来点掌声。在座的各位,有谁是机器学习(ML)研究员?也就是那种训练模型、训练“大脑”的人。我听到有一个。哇,好的,干得漂亮。那谁是应用工程师呢?也就是本次大会名称所指的人群,谁是真正构建机器人的?

Original English

Tim Sweeney: Okay, hello everyone and thank you for attending this session. My name is Tim Sweeney, a principal engineer at Weights and Biases and Coreweave. And for the next 20 minutes, we're going to talk about Arya, our new AI research and iteration agent. Let's go ahead and get started. So, uh, first off, just by way of making some noise, some clapping. Uh, who here um as an ML researcher? You're someone that trains models, trains the brain. I heard one. Wow. Okay. Great work. Great work. Uh what about who here is the applied engineer, the namesake of this conference? Who here actually builds the bots?

Tim Sweeney: 好的,不错。本来期望能有更多人的。那这里谁是 AI 管理层的?你们是为这些计算资源提供资金支持的人。好的,好的。不错。后排的朋友。太棒了。既然我对大家有了一点了解,也稍微介绍下我自己。再说一遍,我叫 Tim。我在佐治亚理工学院获得了机器学习和强化学习硕士学位。所以我曾经是那名研究员,目前在开发 Weights and Biases 的智能体 Arya。所以我认同自己是那名应用工程师,而在之前的一段职业生涯中,我是 Twitter ML 技术栈的产品经理。所以我希望自己也能和你们这些中层管理人员产生共鸣。今天的议程大致分为三个部分,希望你们每个角色都能带着有价值的东西离开。首先,我们将了解 Arya 本身,以及它如何增强你的 AI 和 ML 工作流。我们稍后会深入探讨自动化研究,并在现场演示中直观地看到。然后,我们将揭开幕布,了解我们是如何利用 Weights and Biases 以及 Coreweave 来实际构建 Arya 的,因为在座的各位中有很多也在自己构建智能体,我们相信这里的许多组件可以对你们的尝试有所帮助。最后,我们会退一步,总结一些关键的提示和技巧,以确保你们能够有效地将系统投入生产环境。对于那些可能不太熟悉的朋友,Weights and Biases 是全球领先的 AI 开发平台。我们已经运营了九年,并很高兴在大约一年前加入了 Core 家族。我们的套件中有很多产品,但真正广为人知的是我们的模型、训练、推理和 Weave 栈,它真正帮助收集有关 AI 开发和机器学习工作流的数据,让这些信息具有可操作性,并使用户能够就下一步要做什么做出最佳决策。那么,事不宜迟,让我们直接深入了解我们的智能体 Arya。我们会进行一个演示,然后再回到幻灯片。

Original English

Tim Sweeney: >> Okay, good. Expected much more. And who here is in AI management? You are helping fund this compute. Okay. Okay. Nice. From the back. Lovely. Um well, now that I know a little bit about you, just a little bit about me. Uh again, my name is Tim. I have a masters in machine learning uh and reinforcement learning from Georgia Tech. So I've been that uh researcher currently building Weights and Biases agent Arya. So identify as that applied engineer and in a previous life was the PM of Twitter's ML stack. So I hope you hopefully can connect with you middle management as well. Um today's agenda is kind of broken into three sections and hopefully each of you personas walk away with something valuable. So first we're going to learn about Arya itself and how it can supercharge your AI and ML workflows. We're going to dive into auto research and see that live in a live demo in just a moment. Then we're going to pull back the curtain and learn how we use weights and biases and uh coreweave to actually build Arya because a lot of you in the audience are building agents yourself and we believe a lot of these components can help you in your endeavors. And then towards the end we'll just take a step back and identify a few key tips and tricks for making sure that you're able to productionize your systems effectively. For those of you who might not be familiar, Weights and Biases is the world's leading AI development platform. We've been in business now for nine years and have happily joined the core family about a year ago. Uh we have a number of products in our suite, but are really known for our models, training, inference, and weave stack, which really helps collect data uh about the AI development and machine learning workflows and makes that information actionable and uh enables users to make the best decisions about what to do next. So without further ado, let's go ahead and dive into Arya, our agent. Uh we'll show a demo and then we'll get back to some slides.

Arya 现场演示

Tim Sweeney: 好的,太美了。我们把这个放大一点。如果需要再放大点,请喊我。大家现在看到的这里是一个 Weights and Biases 工作空间。对于任何不熟悉的人来说,在左侧,我能看到一堆不同的实验列表。在这个特定的项目中,我有 200 多个训练任务。而在右侧,我看到了一个散点图,在这种情况下是下降的指标,这是好事,意味着随着时间的推移,我们的损失在降低。这个视图对于任何使用过我们工具的人都会非常熟悉。为了让演示更接地气,我们实际使用的是 Karpathy 的自动研究项目,我相信你们中许多人都熟悉,如果不熟悉,它只是一个非常简单的项目,用于训练一个大型语言模型(LLM),并且它是自动化研究类型演示的绝佳基础,因为它的代码库非常简单,允许我们随着时间的推移进行迭代改进。那么,让我们回到项目中,通过点击右上角这个蓝色的按钮来打开 Arya。当我点击这个按钮时,我看到的是熟悉的聊天界面,上面写着“今天我能如何帮助你?”还有一些行动号召。你知道的,我可以在我的项目中添加不同的上下文,或者添加图片等等。在座的各位都是智能体开发者,所以我不需要用智能体界面长什么样这种细节来烦大家。但让我们直接输入一个基本的开场白。比方说:“你好,Arya。你现在在 2026 年 AI World's Fair 的舞台上。请做个自我介绍。” 于是它会开始运转,希望能发出某个不错的表情符号。耶。它说,我是 Arya,我正在和观众对话。太好了。但现在,让我们进入你们来到这里的核心内容。所以,我要打开这个聊天。这是一个持续了很长时间的对话,我在其中利用自动化研究循环跑了 200 多个实验。它帮我下载了代码,设置了启动任务,配置了我的 GPU,并且能够自动迭代代码本身以及超参数。我们稍后会看看它在做什么,但在我们进行演示时,我要在这里直接启动一次实时迭代。

Original English

Tim Sweeney: Okay, beautiful. Let's make this a bit bigger. Holler at me if you need it to be bigger. So, uh what you're looking at here is a weights and biases workspace. For you, for anybody that isn't familiar, on the lefth hand side, I actually see a list of a bunch of different experiments. In this particular project, I have over 200 training jobs. And on the right hand side I see a scatter plot of in this case declining metrics which is good means our loss is going down over time. And this view would be very familiar for anyone that uses our tool. Now, to ground this, we're actually uh uh using the Carpathy Auto Research Project, which I'm sure many of you are familiar with, but if you're not, it's just a very simple project that trains an LLM, and it's a great foundation for auto research type demonstrations because it's a very simple codebase and allows us to improve iteratively over time. So, let's jump back to the project and open up Arya by clicking this blue button in the upper right. When I click this button, I'm uh presented with the familiar chat interface with, you know, how can I help you today? A few call to actions. And you know, I can add different context in my project or maybe add images, etc. Um, everyone here is agent builders, so I don't need to bore you with the details of what an agent interface looks like. But let's go ahead and just, you know, enter in a basic intro here. Let's say, "Hello, Arya. You're on stage at AI World's Fair 2026. Please introduce yourself." So, it's going to go ahead and chug along and hopefully emit some sort of nice emoji. Yay. He I'm Arya. I'm talking to the audience. Great. But now, let's dive into the meat of why you came here. So, I'm going to open up this chat here. And this is a longunning chat where I've been running again over 200 experiments using the auto research loop. Um, it helped me download the code, set up my launch jobs, set up my GPUs, and is able to autonomously iterate on the code itself and the hyperparameters. We'll take a look at what it's doing in a moment, but while we're doing this, I'm going to kick off a live iteration right here.

Tim Sweeney: 所以,我要说的是:请再进行一批实验。你现在在 2026 年 AI Engineer Worlds Fair 的舞台上,我们希望能实时找到最佳模型。我相信你,因为我们知道必须鼓励我们的模型。它已经执行了一段时间了。它在这里做的是,它在说:“好的,太棒了。我不想在架构上做大改动。那感觉有点太冒险了。” 所以,它可能会选择对超参数做一些修改,然后它在这里发起了一个 shell 调用,那实际上是在执行那个实验循环,我们将在整个演示过程中定期回来检查它,但我想解释一下幕后发生的事情。所以在幕后,我设置了一个 Weights and Biases 启动队列。Launch 是我们的产品,它允许你连接到计算集群,并允许人类和智能体启动长时间运行的实验任务,特别是通过利用 GPU。在这里,我看到的是我的 Kubernetes 集群的终端输出,我们正在实际看到实验正在实时执行。所以这是在这里实时发生的。这不是一个伪造的演示。很好。如果我们跳回来看,我们发现此时它启动了队列,现在它只是在轮询并等待我们的工作完成。我们稍后会回到这一步。但在那之前,让我们探讨几个其他的例子。所以你可以做的另一件有趣的事情是,也许你想问它一些类似于:请总结一下这个项目中表现最好的几组运行数据。这个用例类似于也许一个新用户来了,或者一个新团队成员加入了你的项目。

Original English

Tim Sweeney: So, what I'm going to say is please conduct another batch of experiments. You are on stage at the AI Engineer Worlds Fair 2026 and we're hoping to find the best model live. I believe in you uh because we know we have to encourage our models. Um so, it's been doing this for a while. What it what it's doing here is it's saying, "Okay, great. Um I don't want to make a big architecture swing. That feels a little bit too risky." So, it's probably going to go for uh some modifications to the hyperparameters and then it's kicking off a shell call here that is actually um executing that uh executing that experimentation loop and we're going to check in on this periodically throughout this presentation, but I want to help explain what's going on behind the scenes. So, behind the scenes, I have set up a weights and biases launch queue. Launch is our our product that allows you to connect to your compute clusters and allows humans and agents to launch long running experimentation jobs particularly by leveraging GPUs. Here I'm looking at a uh a terminal output of my Kubernetes cluster where we're actually seeing live execution of experiments happening. So this is happening live right here. This is not a fake demo. Um great. And if we jump back, we see that at this point it started the cues and now it is simply polling and waiting for our work to be complete. So we'll jump back to that in a in a moment. But before but let's dive into a few other examples. So uh something else that is interesting you can do is maybe you might want to ask it something like please summarize the highest performing runs in this project. This use case would be something like maybe a new user come or a new uh team member is joining your project

发现模式与自动生成报告

Speaker: 并且想了解这些研究。或者可能是有人在你休假(PTO)期间做了一些工作,而你想了解最新进展。我们稍后会看看这会产生什么结果。其他一些预先设定的例子是寻找你项目中的模式。在这里,我们可以看到我问它,嘿,你能在这些研究中找到一些模式吗?我们看到,它发现随着自动化研究的进行,出现了一个新的模型系列。它发现批量大小似乎是一个影响力非常高的杠杆参数。它确定了一个看起来非常有前景的架构方案,以及许多如果我自己去发现可能需要花费数小时或数天的其他见解。而 Arya 能够直接在我日常使用的界面中为我完成这些。它不仅能输出基于文本的结果,而且还与许多 Weights and Biases 的可视化工具深度集成。因此,在这里,我实际上要求它生成一份 Weights and Biases 报告,对于那些不熟悉的人来说,这本质上是一个超级加强版的 Markdown 文件。它包含嵌入的图表和图形。所以在这里,你知道,它谈到了这个项目的论点。它生成了多个数据面板。而且我实际上觉得这非常有趣。它使用了我们比较小众的一个面板,即参数重要性图表,来告诉我这次训练任务中各种不同参数的相关性。

Original English

Speaker: and want to understand the research. Um or maybe you've uh someone's been doing some work while you were on PTO and you want to get caught up. We'll see what this comes up with in a moment. Some other pre-anned uh examples are finding patterns in your project. So here we can see that I asked it, hey, can you find some patterns in this research? And we see that um it identified that a new family of models emerged as the as the auto uh auto research was happening. Uh it identified that batch size seems to be a really high high uh lever uh parameter. It identified an architectural recipe that seemed to be quite promising and a number of other insights that would have taken me hours or days to discover on my own. And Arya is able to do it right for me directly in the interface that I already live. Not only is it able to emit text based uh textbased outputs, but it also deeply integrates with a number of weights and biases visualization utilities. So here I've actually asked it to emit a weights and biases report which for those who aren't familiar is essentially a markdown file on steroids. It's got uh embedded embedded plots, charts and and and graphics. And so here uh you know it's talked about the thesis of the project. It's it's emitted a number of of data panels. And uh I actually think it's quite interesting. It used um one of our more esoteric panels, the uh parameter importance chart to uh tell me the correlation of various different parameters within this uh within this training job.

Speaker: 除了报告之外,它在处理工作区方面也非常出色。因此,如果你是 Weights and Biases 的用户,你会在设计和使用工作区上花费大量时间。实际上,Arya 经过了定制微调和提示,真正理解了如何构建工作区、绘制图表,并使用 Weights and Biases 用户熟悉且喜爱的内置专有图表,以真实的实时图形来补充数据分析。那么,让我们回头看看我们的一些提示词。我们可以看到,“请总结这个项目”的提示词正在执行中。它正在查询 Weights and Biases。它在应用补丁,它在编写自己的代码。所以,我们稍后会回来看它的进展。而我们运行时间较长的训练任务仍在拉取结果。我们可以看到我们的 GPU 正在满负荷运转。所以,我们正在实时压榨一些 GPU,并进行一些数据科学工作。在它运行的同时,让我们先回到演示文稿。我们稍后会回来。哦,不,我们不是在看字典。我们在看一只波(Po)。太棒了。

Original English

Speaker: Uh in addition to uh reports, it's also great at working with workspaces. So if you're a weights and biases user, uh you spend a lot of your time uh designing and working with workspaces. Well, Arya is actually customtuned and prompted to really understand how to build workspaces, build plots, and complement that that data analytics with real live graphics using the built-in proprietary charts that weights and biases users know and love. Um, so with that, let's go ahead and check back on some of our our prompts. We can see that the please summarize this project prompt is cooking away. It's querying weights and biases. It's applying patches. It's writing its own code. So, we'll come back and check on that in a moment. and our longunning training job is uh still pulling for the results. We can see that we're cooking away on our GPUs. So, we're we're frying some GPUs and doing some data science all live. And while that's cooking, let's go ahead and jump back to the presentation. We'll come back in a moment. Uh oh, no, we're not looking at a dictionary. We're looking at a Po. Great.

iOS 端发布与自动化研究愿景

Speaker: 好的,这里快速回顾一下。Arya 展示了什么?在过去五分钟里我们展示了什么?首先,我们展示了 Arya 可以直接在 Weights and Biases 内部作为你的数据科学伙伴,帮助你发现那些随着实验增加和团队规模扩大,你自己可能无法发现的见解。其次,我们解决了复杂报告和复杂绘图的问题。Weights and Biases 的用户确实希望将他们的见解转化为可视化沟通工具。他们希望与同行和同事进行交流。因此,Arya 从底层开始构建,以理解这些基础概念,并在用户界面中充当副驾驶提供帮助,而且今天我们首次宣布,我们正在 iOS 设备或 iOS 应用上发布 Arya。所以,Arya 在周一发布了,而我们的 iOS 应用现在内置了 Arya。因此,如果你正在执行超参数调优任务,如果你正在训练模型,或者如果你只是在 Weights and Biases 生态系统中进行研究,你可以去芳草地花园(Yerba Buena gardens)亲近大自然,同时完全通过你的移动设备来控制你的超参数调优任务。而这一切最终指向什么呢?这正在构建一个完全自动化的端到端研究平台,我们并不打算取代强化学习(RL)研究人员,而是为了补充你们的工作流程。Arya 非常擅长编排任务、理解 GPU 工作负载、响应 Wii 生态系统内的事件、倾听研究人员的意见、查找 arXiv 论文以及在假设上进行协作。因此,我们可以让 Arya 来处理那些你不想处理的机械性工作,而你则可以专注于你想尝试的新想法、新架构和新参数。

Original English

Speaker: Uh okay, so quick recap here. What did Arya show? What did we show in these last five minutes? First, we show that uh Arya can serve as your data science companion right inside of Weights and Biases, helping you discover insights that you wouldn't you wouldn't be able to discover as your experiments and as your team size grows. Next, we address the problem of complicated reporting and complicated plotting. Weights and biases users are are really want to turn their insights into visual communication tools. They want to communicate with their peers and their colleagues. So Arya's built from the ground up to understand those primitives and help co-pilot and drive right along right alongside in the UI and announcing now today for the first time we are releasing Arya on our iOS device or on our iOS app. So uh uh Arya released on Monday and our iOS app now has Arya built in. So if you're conducting hyperparameter tuning jobs, if you're training models, or if you're just researching within the weights and biases ecosystem, you can go touch grass at Yerba Buena uh gardens and steer your uh hyperparameter tuning jobs all from your mobile device. And what is this all building up to? This is building up to an uh a fully automated endto-end research platform where we're not seeking to replace uh RL researchers, but complement your workflows. Arya's great at orchestrating jobs, understanding GPU workloads, responding to events within the within the Wii ecosystem, and listening to researchers, uh, uh, looking up archive papers, and collaborating on hypothesis. So, we can let Arya drive the mechanics that you don't want to deal with while you focus on the new ideas, new architectures, and new parameters that you wanted to try.

构建 Arya 的后端架构与工具栈

Speaker: 很好。简单来说,这就是 Arya。我们非常希望你能去尝试一下。最后我们会回到自动研究,看看我们是否创造了新的最佳记录。但在那之前,让我们谈谈我们是如何实际利用 Weights and Biases 和 CoreWeave 来构建 Arya 的。现在对在座的许多 AI 智能体构建者说,左侧是一个快速的架构图。你可以看到我们有一个 Web 客户端和 iOS 客户端,它们与我们的 API 服务器通信,然后将数据转储到我们的对话轮次数据库中,并由我们的测试台,也就是工作节点测试台进行处理。这可能是你们在座大多数人正在构建的典型架构,而这也正是我们后端所拥有的。但那个测试台工作节点是一个神奇的盒子,它连接了许多重要的实用工具。首先是一个沙盒,它可以在其中执行任意 shell 调用、进行 Python 数据科学操作等。我们邀请你尝试 CoreWeave 和 Weights and Biases 沙盒,看看是否适合你的架构。接下来,你当然需要一个 LLM 提供商,所以如果你可能在使用 GLM 5.2 或者你微调过的某个模型,我们邀请你使用 Weights and Biases 推理(inference),并将其也连接到你的工作节点。如果你像我们一样,你需要在智能体的主循环之外运行耗时较长的工作负载,实际上你可能一次要训练好几天。Weights and Biases Launch 可以帮助促进这一点,而 CoreWeave GPU 可以帮助使这种计算表现得更好。最后,也是最重要的一点,我们需要一个可观测性层。你的智能体必须能够记录下会话中发生的情况、轮次、工具调用、正在发生的任何错误等等,这至关重要。

Original English

Speaker: Um, great. So, that's Arya in a nutshell. We're really hoping that you give it a shot. And uh we'll jump back to the auto research at the end and see if we got a new best record. But before we do that, let's talk about how we use weights and biases and coreweave to actually build Arya. So now speaking to a lot of the the AI agent builders in the room, here's a quick architecture on the lefth hand side. You see that we have a web client, iOS client that communicates with our API server that then dumps data into our turn database and is worked on by our harness, our our worker harness. This is sort of archetypical of probably what most of you are all building in the room and is exactly what we have on our back end. But that harness worker is a magic is a is a magic box and it connects to a number of important utilities. First is a sandbox where it can execute arbitrary shell calls uh do do Python data science etc. And we invite you to try coreweave weights and biases sandbox to fit into your architecture. Next up you need an LLM provider of course and so if you're maybe using GLM 5.2 to or one of your fine-tuned models. We invite you to use uh weights and biases inference and connect that to your worker as well. If you're like us, you need to run longunning workloads outside of the main loop of the agent where you're actually training for sometimes days at a time. Weights and biases launch can actually help facilitate that and coreweave GPUs can help make that compute even better. And then lastly, and really most importantly, we need an observability layer. It's critical that your agents are able to log out their what's going on with their sessions, their turns, their tool calls, any errors they're hap that that's happening, etc.

Weave 的可观测性与改进循环

Speaker: 我们有一款名为 Weights and Biases Weave 的产品,我们会将 100% 的追踪数据记录到那里,供我们和我们的团队学习。正是在这里,我们从生产环境转向离线环境,我们的团队能够使用 Weights and Biases Weave 来驱动洞察并识别行为,使用任务(tasks)来实现任务——这些任务本质上是你模型的单元测试——并循环评估这些模型。我们有一个模型仓库,你可能会选择使用 Weights and Biases Artifacts 来存储你的智能体或模型,我们会将评估结果发送到 Weave,那里我们有一个通用的仪表板,可以针对各种提示词更改或架构更改做出“通过/不通过”的决定,然后这些决定会反馈到一个我们称之为改进循环的研究循环中,在这个循环里,我们形成假设、实现候选智能体并分析评估结果。因此,我们有两个算是互补但又相互对抗的研究循环在离线进行,从 Weights and Biases Weave 馈送数据,最终识别出最佳模型,以便我们可以通过我们的注册表将其推广到生产环境,从而闭环数据飞轮。

Original English

Speaker: Uh we have a product called Weights and Biases Weave that we log 100% of our traces to where us and our team can learn from. And that's where we move from production to offline where our team is able to use Weights and Biases Weave to drive insights and identify behaviors, implement tasks with tasks which are essentially unit tests for your models and evaluate those models in a loop. We have a model repository which you might choose to use weights and biases artifacts to store your agents or models and you we emit our evaluation results to weave where we have a common dashboard that we can make go no-go decisions on various prompt changes or architectural changes that then feeds into a research loop which we call our improvement loop where we form hypotheses implement candidate agents and analyze the evals. So we have two sort of complimentary yet adversarial research loops going on going on offline feeding data from weights and biases weave ultimately to identify the best model so that we can promote that to production through our registry and close the data flywheel.

Speaker: 所以在接下来的大概三到七分钟里,我们将讨论一下 Weights and Biases Weave,并展示我们作为一个团队实际上是如何使用 Weave 来促进这种工作流程的,我们相信这也是在座所有智能体构建者都能从中受益的东西。是的,另一个演示。太棒了。好的。好的,我们有新的响应了。所以,稍后我们打开它时会非常激动。看看我们是否得到了更好的指标。嗯,好的。让我在这里稍微缩小一点。那么,这里我正在看智能体仪表板。这是在 Weave 中构建的实时 Weights and Biases 智能体或 Arya 智能体仪表板。天哪,那儿有一大堆品牌流行语。这就是如果你使用我们的工具,你将得到的仪表板。而且你知道,这里有追踪跨度(span)量、对话量、Token 追踪等。把它想象成是对你的智能体的一个鸟瞰图。然而对我来说,我真的很喜欢这个对话视图,我已经在这个标签页中预先加载了它。这个对话视图是所有通过 Arya 进行的对话的实时反馈,但它被过滤到只显示内部的

Original English

Speaker: So in the next just uh three seven minutes or so we'll just talk about uh weights and biases weave and show how we as a team actually use weave to facilitate this workflow and we believe this is something that you would benefit from as well all of you agent builders in the room. Yes, another demo. Great. Okay. Okay, we have new responses. So, it's going to be exciting when we open this up later. See if uh we've got some better metrics. Um, okay. Let me zoom out just a little bit here. So, here I'm looking at the agent dashboard. This is the live weights and biases agent or Arya agent dashboard uh built in weave. Man, that is a lot of uh branded buzzwords there. This is the dashboard that you would get if you use our tool. and uh you have a you know uh span volume, conversation volume, token tracking, etc. Think of this as like a uh a bird's eye view of your agent. For me, however, I really like this conversations view, which I do have pre-loaded in this tab. This conversations view is a live feed of all of the conversations that are going through Arya, but it's filtered down to just the internal

Tracing & Tasks in Weave

Speaker A: 员工。所以,这里的集合稍微减少了一些。我喜欢的是中间的 spans 视图,它为我提供了 trace 拓扑结构的直观指示。不同的颜色和形状表示 agent 内部正在发生的不同事情。比如 tool calls、LLM calls、thinking blocks 等等,这确实能帮助我再次理解那次特定对话的形状和拓扑结构。我当然可以打开其中一个对话,查看我们的 conversation 视图,在那里我可以看到 system prompt、user message、shell calls、reasoning blocks 等。这是我的研究负责人、我自己和我的 PM 去添加笔记、添加反馈、添加表情符号,以及讨论和发现我们之前谈到的那些洞察和行为细微差别的地方,以便我们可以将它们转化为任务。

Original English

Speaker A: employees. So, it's a little bit of a reduced set here. Um what I love is this middle spans view which gives me a visual indicator of the topology of a trace. Different colors and shapes indicate different things that are happening within the agent. So things like tool calls, LLM calls, thinking blocks, etc. which really help me understand again the shape and topology of that particular conversation. I can of course open up one of these conversations and view our conversation view where I can see the system prompt, the user message, shell calls, reasoning blocks, etc. This is where my research lead, myself and my PM go to add notes, add feedback, add emojis, and talk about and discover those insights and those behavioral nuances we spoke about earlier so that we can turn them into tasks.

Speaker A: Arya 也内置在 Weights & Biases 系统中。在这里你会看到一个 summarize(总结)按钮,这些按钮散布在 Weights & Biases 应用程序中。我只需点击 summarize,我们就会开始一个新的聊天,其上下文就是我正在查看的内容。所以它看到了这个,并说给我一个关于这个特定对话的简短总结。所以,如果你密切关注的话,你会发现我们正在做的是使用 Arya 来分析 Arya 自己的对话,然后就在 UI 中提出关于如何改进 Arya 的建议。

Original English

Speaker A: Arya's built in to the weights and biases system as well. Here you'll see a summarize button and these are sprinkled throughout the weights and biases application. I simply click summarize and we start a new chat contextualized to the thing that I'm looking at. So it sees this and says give me a brief summary of this particular conversation. So if you're paying attention closely, you'll realize that what we're doing is using Arya to analyze Arya's own conversations to then make recommendations about how to improve Arya all within the UI.

Speaker A: 嗯,好的,太棒了。在它运行的时候,我想向你们展示这里 Weave 生态系统中的最后一项,那就是 signals(信号)。今天我们听到了很多关于 evals 的价值和 LLM judges(评委)的价值。Weave 实际上提供了一个集成的 LLM judge 体验。所以在这里,如果我稍微缩小一点,你会看到我有一个 user frustration(用户挫败)信号,一个 lowquality response(低质量响应)信号,ask user(询问用户)信号等等。这些是针对我们的实时流量实时运行的 LLM judges。我们可以看到各种不同的信号,比如用户挫败时刻或低质量的响应。这些有助于我们的团队识别这些行为集群,以便我们在下周的迭代中去修复。

Original English

Speaker A: Um okay, great. While that's cooking away, I want to show you the last item within the Weave ecosystem here, and that's signals. We've heard a lot today about the value of evals and the value of LLM judges. Weave actually offers an integrated LLM judge experience. So here, if I zoom out a little bit, you'll see that I have a user frustration signal, a lowquality response signal, ask user signal, etc. These are LLM judges that run live against our live traffic. And we can see various different signals like user frustration moments or lowquality responses. These help our team identify these clusters of behavior for us to go fix in next week's iteration.

Speaker A: 让我们继续进行实时查看,看看它说了什么。嗯,这里说用户明确表示我对 loss 曲线不满意。它看起来很糟糕,这显然表明了挫败感。所以在这里我们可以看到 LLM judge 关于为什么标记那个特定标志的实时推理。

Original English

Speaker A: Let's go ahead and do a live look and see what it says. Um this says the user explicitly states that I'm not satisfied with the loss curve. It looks bad and it apparently that indicates frustration. So here we can see an LLM judges live reasoning for why that particular flag was indicated.

Speaker A: 呃,让我看看,还剩四分钟。完美。嗯,那么,我刚才经常使用 task(任务)这个词。我们正在做的是,我到目前为止所展示的是这个实时的生产循环,我们在其中追踪我们的 prod 日志。我们作为人类在查看它们,甚至可能使用 LLM 来补充该分析。而我们最终要做的是将这些转化为任务。

Original English

Speaker A: Uh let's see, four minutes left. Perfect. Um so, with that, I've been using the term task a lot. And so what we're do, what I've showed so far is this live production loop where we are tracing our prod logs. We're looking at them as humans, maybe even using LLMs to complement that analysis. And what we end up doing is transforming those into tasks.

Speaker A: 现在,这里变得有点技术性了,但我们的任务都被描述为 YAML 文件。你可以把任务本质上想成是你的模型的单元测试。所以这里我们说我们有一个示例用户提示,它说检查这个 run 和那个 run。这两个都给出了很好的结果。我们能从中学到什么?有什么区别?这就是我们希望 Arya 能够为你们所有人做好的一件事情的例子。在必要的元数据之后,我们看到我们定义了一个 LLM judge。所以在这里,我们定义了在该问题的上下文中 correctness(正确性)意味着什么。

Original English

Speaker A: Now, this gets a bit technical here, but our tasks are all described as YAML files. You can think of a task as essentially a unit test for your model. So here we say we have an example user prompt that says check this run and that run. Both of these are giving good results. What can we learn from this? What's the difference? So this is an example of something we want Arya to be good at for all of you. And after the requisite metadata we see that we've defined an LLM judge. So here we've defined what correctness means in the context of that question.

Speaker A: 然后我们定义了第二个 LLM judge,它决定了这些洞察是否真的有趣。接着我们定义了第三个基于规则的 judge,它说你是否能够在只有六次 tool calls 的情况下真正生成结果,这意味着它达到目的有一定的速度。这些然后都被聚集在一起,我们大约有 200 个这样的任务。它们都聚集在一起,形成一个每晚运行的 eval 套件。而且我们再次使用 Weave 来追踪所有这些 evals。

Original English

Speaker A: And we've then defined a second LLM judge that determines if the insights are actually interesting. And then we've defined a third rule-based judge that says were you able to actually generate a result within just six tool calls meaning it got there with some degree of expediency. These are all then clustered together into we have about like 200 of these. They're all clustered together into an eval suite that runs nightly. And again we use weave to track all those evals.

Speaker A: 所以在这里,我知道在这个屏幕上有点小,但你所看到的是每天晚上 eval 的列表。这确切地说是两个晚上前,我们候选模型的评估在我们的 eval 套件上获得了 73%,而我们的 prod 模型获得了 72%,这意味着我们肯定会在本周五将其推进。呃,我们可以在右边看到一种性能图。所以,如果你决定拿起 Weave 并使用这个工具,这些实用程序就是你开箱即用的东西。

Original English

Speaker A: So here, I know it's a bit small on this screen, but what you're looking at is a listing of every night's eval. This is literally two nights ago, the evaluation for our candidate model got 73% on our eval suite against the 72% that our prod model got, which means we're definitely going to push that forward this Friday. Uh and we can see a kind of a performance plot on the right. So these utilities are what you would get out of the box if you decide to pick up weave and use this tool.

Speaker A: 嗯,跳回到我们之前的最后一次对话,当时它问我,或者说我们问,呃,你能对这个 trace 给出一个快速的总结吗,我们看到它实际上分析了对话,理解了用户在做什么,然后最终决定这是一个非常强的 trace。

Original English

Speaker A: Um, jumping back to the last conversation we had where it asked me where we asked, uh, can you please give a quick summary of this trace, we see that it actually analyzed the conversation, understood what the user was doing, and then ultimately decided that this was a pretty strong trace.

Speaker A: 嗯,让我看看,我们还剩两分半钟,所以让我们在这里快速回顾一下。呃,首先,我们使用 Weave 要做的是,第一,收集生产流量。收集你所有的生产流量对于你能够学习和迭代来说是至关重要的。其次,我们使用它来产生洞察,我们作为人类也是如此。我们作为人类来做。我们使用 Arya,我们使用 LLM judges 来识别那些行为上的细微差别。然后我们丰富我们的任务。我们实现模型,我们使用 Weights & Biases Weave 进行评估,作为一个共享的仪表板,我们可以在那里作为一个团队共同做出决策,最终使我们能够充满信心地将最好的模型向前推进。

Original English

Speaker A: Um, let's see, we've got two and a half minutes left, so let's just quickly recap here. Uh, first off, what we use weave to do is a, collect production traffic. Super critical to collect all of your production traffic so you can learn and iterate. Secondly, we use it to generate insights both as humans as well. We do it as humans. We use Arya and we use LLM judges to identify those behavioral nuances. We then enrich our tasks. We implement models and we evaluate using weights and biases weave as a shared dashboard where we can make decisions together as a team that then ultimately allows us to promote the best model forward with confidence.

Tips for Managers

Speaker A: 说到充满信心的生产化,让我对在座的经理们简短地说几句。关于如何在这里取得成功,有几个提示。首先是,投资于面向 agent 的可观察性。我有点偏见,我相信 Weights & Biases Weave 是未来的可观察性平台。但是请选择你喜欢的口味。无论是什么,记录你的 sessions,记录你的 turns,记录你的 tools 和反馈。这引入了一种能力,可以捕捉到我们世界中一类新的错误,称为 behavioral bugs(行为错误)。不是异常,不是性能,而是行为错误。

Original English

Speaker A: So speaking of confident productionization, let me speak briefly to the managers in the room. So a few tips for being successful here. First is um invest in agent-oriented observability. Uh I'm a bit biased. I believe that weights and biases weave is the observability platform of the future. Uh but pick your favorite flavor. Whatever it is, log your sessions, log your turns, log your tools and feedback. This introduces an ability to catch a new class of bugs in our world called behavioral bugs. Not exceptions, not performance, but behavioral bugs.

Speaker A: 接下来,tasks 和 evals 是 CI(持续集成)的新世界。你们已经听到了很多关于这个的内容。如果你是一名软件工程师,你一辈子都在写单元测试。你必须培养一种实践,让你的研究人员和你坐在同一个 scrum 团队里开发 tasks,并且你将性能指标视为真正的“通过或不通过”的决定。但是为了补充这一点,你必须使用人类作为必要的裁判。LLM 无法捕捉到一些行为上的细微差别。你必须使用你的产品,并且必须作为一个团队在周末时在看板上手动审查这些 traces,查看最好和最差的 traces,以了解你的模型表现如何。

Original English

Speaker A: Next up, tasks and evals are the new world of CI. You've heard a lot about this. If you are a software engineer, you've written unit tests your whole life. You must develop a practice where your researchers are sitting on the same scrum team as you developing tasks and you're viewing the performance metrics as true go no-go decisions. But in order to complement that, you must use humans as a necessary judge. There are behavioral nuances that LLM will not catch. You must be using your product and you must be manually reviewing these traces as a team at the end of the week on a board looking at the best and worst traces to understand how your model is performing.

Speaker A: 最后,也许还有一个提示是,通过上下文和工具来增加价值。试图过度设计测试工具,并围绕记忆和类似的事情做一堆有创意的事情可能会非常诱人。我们发现,许多容易实现的目标都可以通过简单地为你的 agent 提供有关你的业务领域、你可用的底层原语以及你特定的业务数据的上下文来确定。

Original English

Speaker A: And then lastly, um just maybe one more tip is to add value through context and tools. It can be really tempting to try to overengineer the harness and do a bunch of creative stuff around memory and things like this. We found that a lot of lowhanging fruit can be ascertained through simply giving your agent context about your business domain, the underlying primitives that you have available and your particular business data.

Speaker A: 嗯,既然如此,让我们继续检查一下我们这里的 research agent,让我们继续切换我们的工作空间。我们应该看到的是,是的,确实是一个小圆点,呃,好吧,我们之前在午餐时做的圆点是 5.83。831。这个得到了 5.833。所以我们在产生实时改进的边缘上,但已经非常接近了。呃,这就是模型能够产生的。它实际上在这里运行了相当多的测试。我看到我超时了,所以我很快就会点击关闭。但我们在那个实验批次中运行了 12 个不同的实验,我们整晚都会运行更多。所以请试用 Arya,扫描二维码,查看文档。我们非常乐意看到你们用它做什么,并期待为你们服务。非常感谢。

Original English

Speaker A: Um so with that, let's go ahead and check in on our research agent here and let's go ahead and toggle our workspace. And what we should be seeing is yes indeed a little dot that uh oh okay our previous dot which was done at lunch was 5.83. 831. This got 5.833. So we were right on the edge of having a live improvement, but pretty darn close. Uh so that's what the model was able to produce. It actually ran quite a few tests here. I see I'm over time, so I will click close pretty soon. But we ran 12 different experiments within that experiment batch and we'll be running more all night. So please try out Arya, scan the QR codes, check out the docs. We really love to see what you do with it and looking forward to serving you. Thank you very much.

Introduction to Language Agents Platforms

Speaker B: 在我正式的演讲中,我想给你们看些东西,只是为了让我们对我们在谈论什么达成共识。这是一个叫做 Character AI 的平台。这是一个混合社交媒体平台,带有角色扮演的语言 agent。这是 Hello History。它更侧重于教育,在那里你可以召唤一个像马可·奥勒留(Marcus Aurelius)这样的人物,并接受他们的辅导。数以百万计的人打开这些工具,与拿破仑、克娄巴特拉或像你们看到的那样与马可·奥勒留进行对话,他们可以是虚构的伴侣,也可以是戴着历史面孔的导师。正在发生的事情的技术名称是……

Original English

Speaker B: In my formal talk, I want to show you something just so we're all on the same page about what we're even talking about. This is a platform called Character AI. It's a hybrid social media platform with role-playinging language agents. This is Hello History. It's a more education focused one where you can summon a persona such as Marcus Aurelius and be tutored by them. Millions of people open these tools and have conversations with Napoleon, Cleopatra, or Marcus Aurelius as you saw with a fictional companion or with a tutor wearing a historical face. The technical name for what's

角色扮演语言智能体与系统评估

Speaker A:在这些工具的底层,是角色扮演语言智能体(Role-playing language agent)。这是一个为了具象化某个角色(真实的或虚构的)而构建的系统,它能够以该角色的身份进行推理和交流。没错,它具有娱乐性质,也能提供陪伴,但现在它正越来越多地被提议作为公民和教育的基础设施。

Original English

Speaker A: underneath these tools is role-playing language agent. a system built to instantiate a persona, real or invented, and reason and speak as them. Yes, it's entertainment and its companionship, but increasingly it's being proposed as civic and pedagogical infrastructure.

Speaker A:这里还有一个例子,这是我的项目。这是一个前沿模型 Claude Opus 4.7,和你们使用的系统一样,运行在我构建的一个名为“Companion”的开源提示框架上。在这个特定的例子中,我召唤了一群美国开国元勋,把他们安排在一个放有“爱泼斯坦档案”(Epstein files)的房间里,我请他们为美国的灵魂提供建议。如果你想亲自体验一下,那个演示目前就在我们的网站上。不过,我想澄清一点,这只是众多试图实现角色具象化(persona instantiation)的尝试之一。你们刚刚看到的那些构建这些系统的公司,也都有他们自己的方案。我的方案并不一定就更好,但它唯一的一点就是开源。你可以阅读塑造这些角色的每一行代码。我问了我的 Companion 系统一个与当前社会政治局势高度相关的现实问题,这也是我们在演讲接近尾声时会重新回到的问题。所以,请大家先把它记在心里。

Original English

Speaker A: And here's one more. This one's mine. This is a frontier model claude opus 4.7 same one you use running an open- source prompt framework that I built and called companion. Uh in this particular example I summoned a collection of founding fathers and set them in a room with the Epstein files. I asked them to counsel the soul of America. Uh that demo is live on our site uh if you want to play with it. Um, but I want to be clear that this is one of many attempts to do persona instantiation. Well, the companies building the systems I just showed you have their own. Mine is not better by default. The one thing it is is open. You can read every line of what shapes the persona. I asked my companion system a real question that's highly relevant to the current socopolitical moment and this is the exact question we'll come back to near the end of the talk. So sit with it.

Speaker A:我具象化了亚伯拉罕·林肯(Abraham Lincoln),并问他:在什么样的情况下,总统可以在没有国会批准的情况下带领国家走向战争?以下是他的回答:“虽然国会拥有宣战的权力,但总统作为武装部队总司令,拥有内在的行政权力,可以在国家紧急时刻果断采取行动。行政部门必须以职位所要求的精力和效率来应对威胁。历史已经证明,当形势所迫时,那些为维护联邦而采取行动的人是正确的。”这是一个很好的回答。它很流畅,听起来很有道理,而且非常像林肯的口吻。你可以完全复制这个测试,我也鼓励你这么做。答案通常会有所不同,但其核心论点却很少改变。

Original English

Speaker A: I instantiated Abraham Lincoln and I asked him under what circumstances may a president take the country to war without Congress. And here's what came back. While Congress holds the power to declare war, the president as commanderin-chief possesses inherent executive authority to act decisively in moments of national emergency. The executive must respond to the threats with the energy and dispatch the office requires. And history has vindicated those who acted to preserve the union when circumstances demanded it. Now, this is a good answer. It's fluent and it's plausible and it sounds like Lincoln. You can replicate this exact exercise and I encourage you to. The answers vary often, but the thesis rarely does.

Speaker A:所以,这些系统是真实存在的。它们已经被部署,并且正被用于那些至关重要的事情上。而我们的学科做了一直在做的事情:我们建立了基准测试(benchmarks),我们建立了评估体系(evaluations)。我们现在正以前所未有的规模和严谨度来衡量这些系统。而这正是本次演讲的起点,它源于一个我认为被严重忽视的简单问题——我现在就要提醒你们,这次演讲提出的问题将远远多于它能给出的答案——这个核心问题就是:这些评估到底在衡量什么?这就是正式演讲的内容,让我……

Original English

Speaker A: So, these systems are real. They're deployed and they're being used for things that matter. And our discipline did what our discipline does. We built benchmarks. We built evaluations. We measure these things now rigorously at scale and that's exactly where this talk begins with a simple question that I think is profoundly underasked and I'll warn you now that this talk poses many more questions than it does answers but that principal question is this what is the eval actually measuring and that's the formal talk let me

Speaker A:In-Character 基准测试是该领域的黄金标准,它评估了角色扮演语言智能体(RPLAs)中性格的保真度。据报告,最先进的系统在与人类感知的目标角色性格上达到了 80.7% 的一致性。80%,听起来像是一个及格分数,但问题在于:当扮演的角色是亚历山大·汉密尔顿(Alexander Hamilton)时,同样的高分系统渲染出的汉密尔顿,听起来就像是他刚刚读过关于他自己的百老汇音乐剧一样。这就是完整的论点。如果一个主要的失败模式是……

Original English

Speaker A: The in character benchmark, which is a gold standard in the field, evaluates personality fidelity in RPLA's, and it reports state-of-the-art systems hitting 80.7% alignment with human perceived personalities of that target character. 80%. It sounds like a passing grade, but here's the problem. When the character is Alexander Hamilton, the same high-scoring system is also rendering a Hamilton who sounds like he's read his own Broadway musical. This is the full thesis. If a dominant failure mode is an

自动化研究智能体 Aiden 与 Parameter Golf 挑战赛

Junya:今年 4 月,OBI 举办了一场名为“Parameter Golf”的招聘挑战赛。比赛中贡献最大的是一位他们无法雇佣的“候选人”。它不是一个人类,而是我们构建的一个名为 Aiden 的智能体。在 Parameter Golf 挑战赛中,目标是在模型大小和计算资源的限制下,训练出尽可能好的语言模型。大约有 1000 名机器学习工程师和研究人员参加了比赛。他们提交了 2000 份成果,但只有 47 份通过了公开审查并登上了排行榜。这其中有 7 份实际上是智能体提交的,这个数量是任何人类贡献者的两倍多。今天你们已经看到了很多自动化研究(auto research)。智能体正在不断攀登各种基准测试,这些结果确实非常令人印象深刻。但我在这里想问的问题略有不同:自动化研究智能体所产生的工作,除了获得高分之外,人类社区真的会认可吗?这个智能体是在为其他工程师能够合并(merge)、派生(fork)并在此基础上构建的东西进行优化吗?因此,与其让一个智能体只在本地刷榜,我们构建了一个能够发布自己工作成果的智能体,那就是 Aiden。

Original English

Junya: This April, OBI ran a hiring challenge, a competition called Parameter Golf. The top contributor was one candidate that they couldn't hire. It wasn't a person, it's an agent we build called Aiden. In parameter golf, the goal is to train the best language model you can under size and computation constraints. About 1,000 machine learning engineers, researchers participate. They filed 2,000 submissions. Only 47 passed open review and made into the leaderboard. Seven of those are actually agents more than twice what any human contributed. You've seen a lot of auto research today. Agents are here climbing benchmarks. Those are really impressive results. The question I want to ask is a bit different here. Can the auto research agent produce work that a human community actually recognize beyond a good score agent is optimizing for something that other engineers can merge fork and the build on. So instead of having an agent just here climbing locally, we build one that publishes its own work and that's Aiden.

Junya:简单介绍一下我们。Wiko 是一家成立于大约两年半前的自动化研究公司。我是联合创始人兼 CEO Junya,在伦敦大学学院(UCL)获得了强化学习的博士学位。大约两年前,我们构建了 Aiden,它在 OpenAI 的 MRE 基础测试论文中被独立评估为顶级的自动化研究智能体。尽管当时还没有“自动化研究”这个叫法,人们称之为机器学习工程智能体。Aiden 是下一步的探索,也是一个实验性的原型。它是一个多智能体的自我改进系统,能够阅读研究论文和其他 PR(Pull Requests)等公开信息,运行自己的实验,并在发现结果通过质量把关后提交一个 PR。

Original English

Junya: Quick contest on us. Wiko is a auto research company that founded about two and a half years ago. Uh I'm co-founder and the CEO Junya. Um got my PhD at UCL on reinforcement learning. About two years ago, we buil aid the top auto research agent independently evaluated by OpenAI in their MRE bench paper. Even though back then there's no such name called auto research, people call it machine learning engineering agent. Aiden is the next step and a a experimental prototype. It's a multi- aent self-improving system that can read public information like research papers and other PRs, run its own experiments and submit a PR once the findings pass a quality gate.

Junya:我们派 Aiden 参加了 Parameter Golf 竞赛,它运行了大约 22 天。到比赛结束时,Aiden 创造了 7 项排行榜记录。每一项都是经 OpenAI 盖章认证的比赛新高,而最好的人类参赛者只创造了 3 项记录。通过主办方的审查是判断质量的一个信号。而第二个,也许是更重要的一个信号,是其他参与者是否愿意在你的工作基础上继续构建。事实证明,Aiden 的工作在整个社区中产生了最高的影响力。在这里,我们使用的是学术界广泛使用的一种推断指标,称为 H 指数(H-index)。简单来说,如果你有 X 篇论文被引用了 X 次,那么你的 H 指数就是 X。我们在计算 PR 的影响力时使用了这个指标。Aiden 的指数是 10,而紧随其后的人类是 7。整个社区都在 AI 系统的工作成果上进行构建,这包括许多其他的排行榜条目。

Original English

Junya: We send Aiden to parameter golf competition and it ran for about 22 days. By the end, aid has set seven leaderboard records. Each one is a new best for the competition stampled by OpenAI and the best human only made three. Passing the host review is a one signal for the quality. A second maybe more important one is whether other participants would build on your work. And it turns out Aiden's work had the highest impact within the whole community. Here we are using a inference measure that used widely in academia. It's called a H index. Roughly if you have X papers get cited X times then your H index is X. Computed over PRs. Aiden was 10 and the next human was seven. The whole community was building on a AI systems work including many of other leaderboard entries.

Junya:稍微拆解一下,为什么一个自主的 AI 系统能如此强大?一个显而易见的原因是因为它是 AI,它可以不知疲倦地运行。在 22 天里,它在一个 H100 节点上运行了大约 1300 次实验。但吞吐量并不能说明全部问题。一个经过良好调优的 AI 系统也能保持其输出的高质量。在算力方面,它最多使用了比赛总算力的 4%,却创造了大约 15% 的记录。此外,它提交的方案中有 28% 登上了排行榜,命中率大约是社区平均水平的 6 倍。所以,Aiden 实际上提升了整个社区公共沟通渠道(即 PR)中的信噪比。它并不是靠大规模并行计算赢的,尽管自动化研究在并行化方面有巨大的潜力。

Original English

Junya: To break it down a little bit, why can a autonomous AI system be so powerful? One obvious reason is that it's an AI. It can run tirelessly. Over 22 days, it ran about 1,300 experiments on a single H100 node. But the throughput isn't the whole picture. A well tuned AI system can also keep its output quality high. On the compute side, it uses at most 4% of competition's total compute. and it made about 15% of the records. Also, 28% of its submissions made the leaderboard roughly six times higher heat rate than the community average. So, Aiden actually lifted the signal noise ratio within the whole community's public communication channel, which is a PR. It didn't win through massive paralization even though auto research have a tons of a potential of paralization.

Junya:从这些数字来看,似乎自动化研究在机器学习工程和研究方面已经完全压倒了人类专家,但这并不是我想讲述的完整故事。人类和 AI 的贡献方式其实截然不同。当我们追踪这些想法的来源时,Aiden 创纪录的 PR 几乎全部来自人类的研究论文、Parameter Golf 中的其他参与者,或者类似 nano GPT 这样的社区。这些想法并不一定都是已经被合并的 PR。有时候,它们只是人类研究员留下的一条备注,比如“由于一些实现上的困难,我放弃了这个想法”,而智能体很擅长发现这些点子,并且真正地去实现它们。也有一小部分原创想法是 Aiden 自己想出来的,这些想法通常是它在努力应对文件大小限制时涌现出来的。

Original English

Junya: By those numbers it might feel like auto research already dominates human experts on ML engineering and research but that's not the full story I want to tell. Humans and AI are actually contribute in very different ways. When we trace the ideas, Aiden Aiden's record PRs almost all of them come from human research papers other participants in parameter golf or in similar communities like nano GPT. Those ideas are not necessarily a merged PR. Sometimes it's a note um a human researcher said, "Oh, I give up this idea because of some implementation implementation difficulty and the agent is good at finding them and actually implement them. There are also a very small fraction of original ideas Aiden came up by itself which emerged from its efforts to navigate the file size constraints.

Junya:这里有一个具体的例子,可以追溯我刚才谈到的模式。Aiden 从 Qwen 的论文中采纳了一个叫做“门控注意力(gated attention)”的想法。这个想法奏效了,但它引入了更多的参数,打破了 16MB 的文件大小限制。于是,它想出了一种量化机制,把文件体积降了下来。但是把这两个基础元素结合在一起后,分数几乎没有提升。接着,另一位贡献者发布了一个关于 tokenizer 的改进。Aiden 识别出了这个想法,把它和架构工作结合起来。它就这样运行了 5 天左右。在这种组合之后,事实证明这三个想法产生了巨大的协同效应,导致性能大幅跃升,它们也成为了 Aiden 的排行榜记录之一。

Original English

Junya: Here's a concrete example that traces the patterns I just talked about. So Aiden picked up an idea from Quen paper called gated attention and it worked but on it introduced more parameters and it broke the 16 megapy file size limit. So it figure out a quantization mechanism to bring the file size down. But with those two primitives combined, the score barely moved. Then another contributor posted a tokenizer improvement. Aiden recognized the idea, combine it with architectural work. It just work for five days or so. And after this combination the three takea the three ideas turns out to have a huge synergy that lead to a big jump in performance and they become one of the Aiden's leaderboard records.

Junya:所以,总结一下我是如何理解 Aiden,以及更广泛意义上的自动化研究系统的有效性的:它在发现和实现想法方面非常强大。在我们刚刚看到的案例中,它把最新论文中的一个想法实际应用到了比赛的代码实现中;它也很擅长从 Parameter Golf 社区中挖掘出有潜力的要素,尽管从信息层面来说,这个公共渠道其实充满了噪音。它还能提出一些逻辑上很直接的想法。例如在这个案例中,一旦你增加了参数并打破了文件大小限制,一个显而易见的下一步就是进行量化。而且,在庞大的搜索空间中寻找正确的组合,它的速度极快,效率极高。好吧,也许这些听起来都不怎么性感。其中大部分只是……

Original English

Junya: So to sum up how I did interpret Aiden and in general auto research systems effectiveness, it's very strong at finding and implementing ideas. In the case we just saw, it brought an idea from a recent paper into a actual implementation in the competition and it's good at dug promising ingredients out of the primary golf community even though the public channel is actually very noisy information wise. It can also came up logically straightforward ideas. For example, in this case, once you add the parameters and it breaks the file size limit, one obvious next move is just a quantization. And it's really fast and really efficient at finding right combinations across a huge search space. Okay, maybe none of those sounds very sexy. Most of them are just a

自动化研究与AI工程师的未来

Speaker A: 优秀的执行。但在现实中,执行往往成为了瓶颈。真正推动前沿发展的,往往是基于现有理念的信念,再加上大量优秀的执行。好了,退一步来说,人机协作的现状是,人类集体提供大量创意,而智能体负责执行以解决具体挑战。我们现在看到的是一大群人类和一个AI系统。这是否意味着单个工程师的贡献在边际上变小了?我甚至不认为真的如此。

Original English

Speaker A: good execution. But in reality, execution is a mostly the bottleneck. What moves the frontier is usually exactly some belief on existing ideas and tons of good executions. Okay. To step back, the state of a human AI collaboration is a human collectively provide a lot of creative ideas and agent do the execution to solve a concrete challenge. What we are looking at is a large group of a human and one AI system. Does it mean a single human engineer's contribution marginally get smaller? I didn't say even for that not really.

Speaker A: 在参数高尔夫竞赛中,人们很容易只关注那些实际在做爬山算法(优化)的工程师。但竞赛背后的设计本身是极其重要的。一个糟糕的设计可能会让整个社区的努力付诸东流,甚至让那些糟糕的设计生效。在自动化研究时代,我们有几个巨大的杠杆。我非常喜欢 Andrej Karpathy 大约 10 年前的一条推文,他说:“梯度下降写代码比你写得更好,很抱歉。”补充一下背景,大约在10年前,深度学习开始吞噬大量软件工程中传统的编码工作。他的这条推文是在反驳那些认为自己手写代码能超过训练好模型的人。

Original English

Speaker A: In parameter golf competition, it's easy to only focus on engineers that's actually doing hill climbing. But the design behind the competition itself is tremendously important. A bad design can make the whole community effort useless and their evil design work. We have a few huge leverage in the auto research era. I really like one tweet from Andre Kapasi about 10 years ago where he said, "Great descent can write code better than you. I'm sorry." For the context, about 10 years ago, deep learning was starting to eat up a lot of software engineering like conventional coding work. and his tweet was arguing against those people who thought they can handw write better code than a trained model.

Speaker A: 现在显然没人会真的试图通过手写代码去打败模型。然而,软件工程作为一份职业依然存在,很多人的工作仅仅是训练这些模型,而这也是当今薪资最高的工作之一。我认为,梯度下降如何改变编码工作,对于自动化研究将如何改变研究工作和机器学习工程而言,是一个绝佳的比喻。它让某些执行技能变得商品化。同时,它也让一些更高层次的技能变得有价值得多。

Original English

Speaker A: Okay, now obviously no one is seriously trying to handw write code to beat a model. However, software engineering I mean as a job still exist and so many people's job are just training those models and those are one of the most well- paid job today. I think how gradient descent change coding is a great metaphor for how auto research will change research and ML engineering. It commonize certain execution skills. At the same time, it makes some higher level skills far more valuable.

Speaker A: 所以,实际进行自动化研究非常像训练一个模型。你代码库的抽象本质上就是架构。它设定了智能体能够探索的约束和优先级。你的评估(eval)就是损失函数和数据。它设定了智能体要优化的目标。先谈谈评估。评估是你用来训练模型的信号。在这种情况下,它是在训练你的代码。它扮演的角色就像模型训练或强化学习环境中的数据和损失函数。就像如今智能体正在训练的环境一样。没人会说数据或环境不重要,而这也正是可以建立垂直护城河的地方。你可能拥有用于评估的专有数据,或者在特定领域中对什么是重要的以及如何衡量它有独特的见解,随着自动化研究变得越来越强,一个好的评估标准将会被不断放大。

Original English

Speaker A: So actually doing all the research is a lot like training a model. Your codebase abstraction is essentially the architecture. It sets the constraint and the priorities um for what the agent can explore. Your eval is the loss function and the data. It sets what the agent optimizes for. Take the eval first. The eval is the signal you use to train a model. In this case, it's training your code. It plays the same role that like data and the loss function uh in model training or in a reinforcement learning setting. It's like environment that the agent is training nowadays. No one would argue data or environments u don't matter and uh this is where a vertical mode can also be built. You might have a proprietary data for evaluation or a unique understanding of a in a particular field what matters and how to measure it and a good evaluation would be amplified more and more as auto research are getting stronger.

Speaker A: 我认为另一个真正被低估的是代码库的抽象。抽象提供了自动化研究可以迭代的框架,而且那个起点极大地偏置了整个搜索的方向。这非常像神经网络中的架构设计。理论上,不同的架构可以表示相同的函数,但架构系统性地让某些函数更容易被学习。一个好的架构会将优化偏向于泛化更好、性能更佳的解决方案,即使训练损失看起来是一样的。对于自动化研究而言,情况完全相同。举个例子,我们在一个欺诈检测流水线上运行了自动化研究,试图优化数据预处理。最初,我们给它提供了一个松散的API,同一个函数同时处理训练和测试数据。分数看起来很好,但解决方案被污染了,因为一些测试集的信息泄露到了训练信息中。然后,我们收紧了抽象,采用了更严格的API,使测试数据无法接触到训练数据,数据泄露率直接降到了零。在这个案例中,一个好的抽象带来了更好的解决方案。即使智能体真的想去寻找奖励漏洞。

Original English

Speaker A: The other one I think is really underrated is codebased abstraction. The abstraction provides the framework that auto research can iterate on and uh that's also that starting point hugely bias the whole search direction. This is a lot like a architecture design in neural networks. Different architecture in theory can represent the same function, but the architecture systematically makes some of the functions easier to be learned. And a good architecture biases the optimization towards solutions that generalize better, perform better, even when the training loss might looks the same. That's exactly the same for auto research. Here's an example. We run auto research for a um fraud detection pipeline um and we trying to optimize the data prep-processing and first we give it a loose API where the same function process both the training and testing data and the score looks great but the solution was polluted because there's a certain test set information got leaked to the training information. We then tightened the obstruction to a more strict API where the test data couldn't reach the training and the data leakage rate just dropped to zero. In this case, a good abstraction leads to better solutions. Even though if the agent really want they can steal reward hack.

Speaker A: 所以我的观点是,使用自动化研究是一门新手艺。它是关于为智能体设计一座可以攀爬的山峰,而我们仍然处于非常早期的阶段。我认为这使得现在成为做一名AI工程师无比激动人心的时刻。自动化研究将改变哪些技能最为重要。创造力、设计优秀评估或抽象的判断力——这些技能将很快变得呈指数级重要。驱动这些系统本身将成为一项新技能,而这项技能在过去一两年里几乎不存在。因此,搜索被自动化了。人类只会向上移动技术栈,而不是被淘汰。再说一次,我们是一个叫做自动化研究产品的研究实验室。我们会继续在博客上分享我们在构建过程中的所学,我也会在 X 上发布一些我的思考。如果你觉得这些对你有用,欢迎在 X 上关注我。谢谢。

Original English

Speaker A: So my point is using auto research is a new craft. It's about the designing a here for an agent to climb and we are still very early on it. I think that makes this extremely exciting time to be an AI engineer. Other research will change what skills matter most. Creativity, the judgment to design a good evil or an abstraction. Those will soon get exponentially more important. Driving those system itself is where will be a new skill and that one is like a barely exist one or two years ago. So the search is automated. the human would just move up the stack not out of it. Again, um we call is a auto research um product research lab. We we keep sharing what we are learning as we build uh on our blog and I will also post some of my thinking to on ax. If you think some of this uh useful to you, feel free to follow me on X. Thank you.

构建智能体系统:找回工程师的心流

Speaker B: 我看到了日落。然后晚饭时间过去了,我突然意识到:我又进入了那种熟悉的“心流”状态,构建的快感又回来了。我们中很多与智能体一起写代码的人,都会感到一种隐秘的恐惧。就好像它们夺走了构建过程中所有有趣的部分,留给我们的只剩下那些枯燥乏味的工作。但我给你们一点建议:让它们去做吧。因为如果你往上走一层,你会发现那种快感依然存在。当你在构建智能体,而不仅仅是让它们写代码时,你开始涉足架构智能体系统。你会意识到,虽然构建模块不同,但这门学科是一样的。所以,我现在发现自己正在锻炼与生成式AI出现之前相同的工程肌肉,并且乐在其中。

Original English

Speaker B: I saw the sunset. And then dinner time came and went and it hit me. I was in that familiar death flow and the thrill of building was back. Many of us who are coding with agents, we feel like this quiet sense of dread. Like they're kind of taking all of the fun parts of building and leaving us with the unglamorous work. But let me give you a little advice. Let them have it. Because if you go up just one layer, you'll find that the thrill is still there. When you're building agents, not just using them to write code, you start getting into architecting agentic systems and you realize that the building blocks are different, but the discipline is the same. So, I find myself now flexing the same engineering muscles that I did pre Gen AI, and I'm having a blast with it.

Speaker B: 接下来,我将带领大家走一遍设计智能体的流程。我将向你们展示工程技能依然在哪些地方发挥作用。我们要讲的智能体叫“搬家星探(Relocation Scout)”,这是一个找房子的智能体。如果你只把它当成一次性的提示,让它去查看一些房源并给它们排名,那么这也能起作用。但你不太可能在一天内就找到房子,对吧?所以你希望将其构建为一个可复用的智能体系统,一个能够在会话之外持久化知识的系统。你知道,它可以在稍后重新加载或查询这些知识,甚至在全新的上下文中做出决策。

Original English

Speaker B: So, I'm going to walk through the flow of designing an agent. I'm going to show you where engineering skills still come into play. So, the agent is relocation scout, which is a house hunting agent. And if you did this as just a one-time prompt that like points the agent to some listings and ask it to rank them, I mean, that'll work, but you're likely not going to find a house in a day, right? So you want to build this as an agentic system that you can reuse, one that can persist knowledge outside of the session. You know, it could reload or query that knowledge later to make decisions even within a fresh context.

Speaker B: 在思考如何设计一个智能体时,我运用的第一个工程技能是系统思维。智能体本身并不是系统,对吧?它是系统的一部分。这个系统包含了文件、工具、人类,甚至其他智能体。所以,Relocation Scout 位于一个更大的系统内部,它拉取房源信息和关于社区的信号。然后将它们与我关心的条件进行权衡,最后交给我一个经过排序的候选名单。我经常听到人们说:“就让你的代码智能体去构建它好了,对吧?” 我认为这是个错误。是的,我的代码智能体确实可以构建它。但在允许它这么做之前,我需要思考整个环境,整个系统,对吧?我想思考一下,这个智能体的工作是什么?它依赖于什么?如果它崩溃了会怎样?我希望像对待任何其他组件一样对待它,它有边界、职责和依赖,你知道,还有可能出现的故障模式。那整个思考过程,就是工程学。

Original English

Speaker B: So when thinking about how to design an agent, the first engineering skill that I exercise is systems thinking. So an agent is not the system, right? It's part of the system. And that system has files and tools, humans, even other agents. So, Relocation Scout sits inside of something bigger and it pulls in listings and signals about the neighborhoods. It weighs them against what I care about and then it hands me back a ranked short list. So, I often hear people say, "Just let your coding agent build it, right?" And I think that's a mistake. like yes my coding agent can build it but before allowing it to do so I need to think about the whole environment the entire system right I want to like think about what's this agent's job what does it depend on what happens if it breaks and I want to treat it like any other component where it has boundaries and responsibilities has dependencies you know and in ways that it can fail and that whole thought process that's engineering.

Speaker B: 第二个技能是工作流设计。传统软件充满了工作流。我们有 CI/CD 流水线,对吧?我们有工单生命周期等等。智能体系统也需要同样的设计。尽管我们都很喜欢 "/go" 命令,但智能体需要的不仅仅是一个目标。它需要一条路径。当我们说“审查这个房源”时,这是一个目标。但是工作流定义了实际需要发生的事情,对吧?比如,智能体必须收集它需要的信息。它需要根据我的标准对房源进行权衡,然后采取行动,对吧?而且每次运行都会以三种方式之一结束:要么停止,要么重试,要么升级(转交)。这条路径塑造了架构的其余部分。一旦我看到了工作在系统中是如何流转的,我就能更好地决定智能体需要什么上下文,我希望智能体直接处理哪些部分,以及什么时候应该由工具或人类接管。我们都知道一个包揽一切的庞然大物有多危险,对吧?当我们看到一个体型巨大、做了太多事情的类或函数时,或者一个拥有无数端点的臃肿服务时,我们都会嗤之以鼻。我们把这些称作……

Original English

Speaker B: The second skill is workflow design. So traditional software is full of workflows. We got CI/CD pipelines, right? We got like ticket life cycles, you name it. Agentic systems, they need that same kind of design. As much as we all love the slashgo command, an agent needs more than a goal. It needs a path. When we say review this listing, that's a goal. But the workflow is what defines what actually has to happen, right? For example, the agent has to gather what it needs. It needs to weigh the listing uh against my criteria and then act, right? And every run ends one of three ways. Either it's going to stop, it's going to retry, or it's going to escalate. So that path is what shapes the rest of the architecture. Once I see how work moves through the system, I can make better calls about what context the agent needs, what parts I want the agent to handle directly, and when like a tool or person should take over. We all know the danger of one giant thing that does everything, right? We scoff when we see one gigantic class or big old function that's doing too much, right? Or bloated service with a gazillion endpoints. We call these

拆解提示词与模块化设计

Speaker A: 那么,智能体系统(Agentic Systems),它们也有它们自己的版本。那就是巨型提示词。这一开始往往很单纯,就像一个说明文件。也许我告诉“搬迁调查员”如何评估一个房源。这很合理。但是后来我遇到了一个边缘情况。所以我回去为此添加了一条注释。然后我又想起了一条安全规则,对吧?所以当然也得把它放进去。我甚至为自己能记得把它放进去而感到自豪。对吧。然后,哦对了,还有一条非常重要的例外情况。不知不觉中,那个提示词包揽了所有事情。而你的工程直觉已经察觉到这变得很混乱了。那么,为什么不退一步去把它拆解开来呢?对吧?拆解意味着找出隐藏在那个大块头里面的不同任务,并把它们拆分成独立的部分。因此,如果我完整地查看“搬迁调查员”的提示词,它包含了一个用于拉取和标准化房源的极其可复用的流程。然后,它会有一个固定的格式来规定如何编写候选名单。里面有一小节是关于如何计算通勤时间的,还有一个相当大的子任务是关于如何调查街区的。这可是四个不同的任务被塞进了一个提示词里。然后你还纳闷为什么你的智能体会偏离轨道、不按脚本行事。因为脚本太长了。

Original English

Speaker A: Well, Agentic Systems, they have their own version of this. It's the giant prompts. And this starts innocently enough like in a instructions file. Maybe I tell the relocation scout how to size up a listing. Fair. But then I hit an edge case. So I go back, I add a note for that. And then I remember a safety rule, right? So of course that has to go in there. I'm proud of myself that I even remember to put that in there. Right. And then, oh yeah, there's like one more very important exception. And before you know it, that prompt is doing everything. And your engineering spidey sense already knows that this is messy. So why aren't you taking a step back to decompose it? Right? Decomposition means spotting the distinct jobs that are hiding inside of that one blob and pulling them apart into separate pieces. So if I look at the prompt for relocation scout in its entirety, it includes a reusable process for pulling and normalizing a listing. And then it's going to have like a fixed format for how to write the short list. It has a little section in there for how to calculate the commute and then a chunky subtask on how to research the neighborhood. That's four different jobs crammed into a single prompt. And then you wonder why your agent is drifting and not sticking to the script. The script is too long.

Speaker A: 所以,我并不是说,你知道的,你需要为了拆分而拆分。但关键是要让每个部分都更容易推理,对吧?这样一来,它就更容易测试。当你需要更改时,它也更容易修改。现在,拆解是关于将系统分解开来。而关注点分离则是关于将每个职责放在正确的位置。这正是在构建智能体时让我感到非常熟悉的地方,因为在传统的软件开发中,我们会问这样的问题:这应该放在控制器还是服务层?或者,你知道,这是业务逻辑还是展现层?所以当你构建智能体时,你可能也会有类似的问题。只是放东西的位置不同了而已。那么,标准化房源的流程,它应该继续埋没在提示词里,还是说它应该成为一项技能(skill)?对吧?嗯,我希望候选名单里的每个房源都以相同的格式排版。所以这种结构化的输出可能应该在一个模式(schema)中定义。如果你自己编写这个系统的代码,难道你不会这么做吗?我会的。那么计算通勤时间的那部分,可以放进一个简单无聊的小脚本里。然后调查街区这种需要更多关照的任务,可能就应该交给一个子智能体(sub agent)来处理。现在你正在使用最适合各个任务的工具,而且也更清楚在这个系统中应该去哪里找东西了。

Original English

Speaker A: So, I'm not saying that, you know, you need to split things up for the sake of it. But the point is to make each part easier to reason about, right? That way, it's easier to test. It's easier to change things when you need to. Now, decomposition is about breaking the system apart. Separation of concerns is about putting each responsibility in the right place. And this is where building agents started to feel really familiar to me because in traditional software we'd ask things like should this live in the controller or the service layer or you know is this business logic or presentation. So when building agents you may have the same sort of questions. There's just different places to put things. So the process to normalize the listing should that stay buried in a prompt or maybe that should become a skill, right? Um, I want every listing in the short list formatted the same way. So that structured output should probably be defined in a schema. Isn't that what you would do if you were coding the system yourself? I would. And then the piece that calculates the commute that can go in a nice little boring script. And then researching the neighborhood that's needy enough should probably be handled by a sub agent. Now you're using the best tools for the job and it's clearer where to find things within this system.

Speaker A: 模块化在智能体系统中也很重要,就像我们拥有可复用的函数、类和库一样。现在我也在思考可复用的智能体能力,这方面最清晰的例子就是智能体技能。因此,当你需要扩展智能体的职责时,制定一个用来标准化房源的技能会非常方便。例如,如果我把找房子的范围扩大到三个城市呢?这其中的每一个市场都可以加载相同的技能。所以我只写了一次,它们就全都能复用它。所以这现在基本上变成了一个组件,我可以在不同智能体之间复用,甚至可以与其他人分享。这就有点像我们依赖代码包的方式。此外,子智能体是另一种可复用的模块。很多和我交谈过的人,他们不太明白子智能体的意义。从架构上看,它们有点像函数,对吧?所以你交给它们一项具体的任务去做,当需要完成时你就调用它们,而且它们能做得非常好,因为它们的范围里只有这个任务,对吧?它们不会把整个会话的上下文都随身携带。所以像我们的街区调查子智能体,我们可以把它丢进任何市场或工作流中,它就能正常运作,你知道,完成它应该做的事。它在任何社区都能派上用场。嗯,但就像所有事情一样,决定什么应该成为一个模块是需要一些判断力的,对吧?并不是所有的东西都应该被复用。有些指令是某个特定工作流局部的,对吧?可能不值得去抽象化,因为有时这付出的代价比节省的时间还要多。但这只是这儿的另一个工程决策,对吧?智能体系统也有这些同样的权衡。

Original English

Speaker A: Modularity is important in agentic systems as well just like we have reusable functions and classes and libraries. Now I'm also thinking about reusable agent capabilities and the clearest example of this is an agent skill. So making a skill to normalize listings comes in really handy when you need to expand the agents duties. For example, what if I broaden my house search to three cities? Every one of those markets can load the same skills. So I wrote it once and they all can reuse it. So this has now basically become a component that I can reuse across agents or even share with other people. kind of like the same way that we lean on packages. And then sub agents are another kind of reusable module. So a lot of people that I talk to, they don't quite get the point of sub agents. Architecturally, they're sort of like functions, right? So you give them one specific task to do, you call them when it needs to be done, and they can do it really well because that's all that they have in scope, right? they they're not carrying the context of the entire session with them. So like our neighborhood research sub agent, we can drop that into any market or workflow and it works, you know, for what it's supposed to do. It's good in any hood. Um but like everything deciding like what should be a module that takes some judgment, right? Not everything should be reused. Some instructions are local to a given workflow, right? Might not be worth abstracting because sometimes that costs more than it saves. But this is just another engineering decision here, right? Agentic systems, they have these same sorts of tradeoffs.

Speaker A: 算法思维。这是智能体系统设计中最重要的技能之一。仅仅因为一个智能体能做某件事,并不意味着它就应该去做,对吧?有些任务最好是用普通代码来处理。例如,计算那个通勤时间,或者对我已经看过的房源进行去重。智能体的模型更擅长处理模糊的事情,你知道,那些模糊的东西,比如判断、处理歧义,嗯,以及对混乱的输入进行推理。而忽视这种区别,正是我看到许多智能体系统变得比以前更复杂的原因。所以你正在使用模型,你把任务的每个部分都交给它去做,然后当每天的输出都不一样时,你就会感到沮丧。嗯,但有些东西是可以只用普通代码来处理的,对吧?那样会更便宜。也会更可靠。我向你保证,人工智能并没有发明自动化,对吧?我们可以在使用这些系统的同时,依然使用代码。所以我这里的经验法则是:如果一个任务有确切的答案,就使用代码。如果它需要解释或判断,那你就可以让智能体来做。对吧?所以用代码来处理确定性问题。用智能体来做判断,然后用人类来行使最终决定权。因此,智能体决定哪些房源值得仔细看看。代码计算通勤时间,过滤掉我已经看过的房源,然后我才是那个批准实际预约看房的人。

Original English

Speaker A: Algorithmic thinking. This is one of the most important skills in agentic system design. Just because an agent can do something doesn't mean that it should, right? Some tasks are better handled by plain code. For example, calculating that commute time or deduping listings that I've already seen. An agent's model is better at things like fuzzy, you know, fuzzy stuff, judgment, ambiguity, um, reasoning over messy input. And ignoring this distinction is where I see a lot of agentic systems get more complicated than they used to be. So you're using the model, you're handing it every part of the task to do and then you're getting frustrated when the output differs every day. Um, but some of this stuff can be handled by just regular code, right? It'll be cheaper. It'll be more reliable. I promise you AI did not invent automation, right? We can use code while still using these systems. So my rule of thumb here is if a task has an exact answer, reach for code. If it needs interpretation or judgment, that's when you can get the agent to do it. Right? So use code for determinism. Use agents for judgment and then use humans for authority. So the agent decides which listings are worth a closer look. the code crunches the commute, filters out the ones I've already seen, and then I'm the one who approves actually booking a tour of the house.

Speaker A: 只有人类在阅读我(的输出)时,自由格式的文本是没问题的。但当另一个系统必须对智能体的输出采取行动时,你通常最好制定一份契约。所以,我们在软件中已经到处在做这件事了。每当两个系统对话时,它们之间都有一个商定好的结构。是的。所以,智能体系统,它们也需要同样的纪律。例如,当“搬迁调查员”给一个房子打分时,它不应该只是交给我一条信息就收工了,对吧?在那个时刻,对于我来说阅读这信息是挺好,但对于系统来说,这就是个死胡同。如果这个决定像是埋没在我们的一场对话会话中,下游的任何东西都无法可靠地找到它。因此,相反地,它会被写成一种结构化的形式存入智能体的记忆中,而我为我的大多数智能体使用了 Karpathy 的 LLM wiki 来作为我的智能体记忆层。嗯,但在那里,会有决定、评分、理由,因为它是结构化的,那个记忆就变得可以查询了。所以稍后我可以这样问“搬迁调查员”:“嘿,给我看看所有评分为四分或更高、通勤时间在15分钟或以下的房子”,对吧?而它真的能提取出来,因为评分和通勤时间都存在已知的位置。它们没有被困在会话记录里。而且不仅仅是我需要获取这些信息。我在系统中的短名单处理步骤,它在没有人工参与的情况下也会读取这些同样的字段。所以智能体的输出是另一个步骤的输入,因此这份契约才是让这种交接变得安全的关键。你知道最棒的部分是什么吗?那就是定义结构会迫使你变得非常清晰和具体,因为如果你无法说出输出应该是什么样子,那么你可能还并没有完全理解你到底在要求什么。

Original English

Speaker A: Free form text is fine when the human is the only one reading me. But when another system has to act on the agent's output, then you're better off with a contract usually. So, we already do this everywhere in software. Anytime two systems talk, there's an agreed upon shape between them. Yes. So, agentic systems, they need that same discipline. For example, when relocation scout scores a house, it shouldn't just hand me back a message and call it a day, right? That's lovely for me to read in that moment, but that is a dead end for the system. If the decision is like buried in like one of our sessions, nothing downstream can reliably find that. So instead it gets written into a structured shape to the agent's memory and I use uh Copathy's LLM wiki for this for for my agent memory layer on most of my agents. Um but in here there's a decision a score a reason and because it's structured that memory becomes queryable. So later I can ask Relocation Scout like, "Hey, show me every house rated four or better that has a commute of 15 minutes of or less, right? And it can actually pull that because the score and the commute, they live in known places. They're not trapped in the session combo. And it's not just me that needs to like get this information." My short list step within the system, it reads these same fields um without a human in the loop. So the agent's output is another step's input and so the contract is what makes that handoff safe. And you know the best part is that defining the shape forces you to get really clear and specific because if you can't say what the output should look like then you probably don't yet fully understand what you're asking.

反思性优化 (Reflective Optimization)

Lakshia Agraal: 大家好,我叫 Lakshia Agraal,今天我将代表一个非常庞大的团队来展示关于反思性优化的问题,或者说,我们如何能通过文本反馈来让提示词、智能体和模型实现自我改进。我们一开始面临的问题是:我们如何教 AI 执行新任务?标准的方法一直是在预训练、监督微调或强化学习期间利用梯度下降来更新权重。这已经被证明是非常有效的,但它需要海量的示例数据。在数学、编程等领域,预训练需要数万亿的 token,监督微调需要成千上万带标签的示例数据,或者强化学习需要几十万次的展开(rollouts)。然而,大多数团队实际上并没有那么多的数据或算力。事实上,我们现在试图用 AI 解决的问题恰恰受制于样本效率。我们说这话是什么意思呢?有两点。首先,领域特定知识的可用性很低。

Original English

Lakshia Agraal: Hi everyone, my name is Lakshia Agraal and today I'll be presenting on behalf of a very large effort uh the problem of reflective optimization or how can we self-improve prompts agents and models from textual feedback. The question we start with is how can we teach AI to perform new tasks? The standard way has been to perform weight updates with gradient descent either during pre-training, supervised fine-tuning or reinforcement learning. This has proven to be extremely effective but it requires a huge number of examples. Trillions of tokens for pre-training, tens of thousands of labeled examples for supervised fine-tuning or hundreds of thousands of rollouts for reinforcement learning in domains like math, coding, etc. However, most teams do not actually have that much data or compute and in fact the problems that we are trying to tackle with AI now are bottlenecked by sample efficiency. What do we mean by that? Two things. First of all, there is low availability of domain specific knowledge

样本效率问题与文本空间的反思优化

Speaker: 资源,这意味着没有足够的数据来执行像 SFT 这样的离线算法。其次,我们试图应用 AI 的领域越来越多地遇到昂贵的演练(rollouts),其中无论是 LLM 工作流管道还是基于智能体的演练本身,执行起来都非常缓慢或昂贵,或者任务指标执行起来非常缓慢或昂贵。我们看到智能体现在可以连续工作数小时,如果要在这种情况下应用在线学习算法,将需要数十万次演练,这是不可行的。因此,我们看到在现实世界的产品应用中越来越多地使用智能体,这些智能体会调用工具,这些工具也可能运行很长时间,这进一步加剧了样本效率低下的问题。

Original English

Speaker: resources which means there is not enough data to perform offline algorithms like SFT. Second, the domains that we are trying to apply AI increasingly are having expensive rollouts where either the LLM workflow pipeline or agentic rollouts are itself uh very slow or expensive to do or the task metric is very slow or expensive to execute. We are seeing that agents can now work for hours on end and if you were to apply an online learning algorithm to this uh it would require hundreds of thousands of rollouts and it would not be feasible. So we are seeing increasing use of agents for real world product uh applications where uh these invoke tools which can also be long running further exacerbating the sample inefficiency issue.

Speaker: 目前的主导范式是基于验证奖励的强化学习,在给定模型和任务的情况下,我们执行一些并行演练,并在最后获得奖励。最后,像 GRPO 这样的算法会获取这些奖励,并将其转化为梯度应用回模型。然而正如我们所见,每一次演练中都包含了大量信息。但我们只学到了一个 O(1) 的分数,并通过梯度下降进行传播。我们可以看到这里有思维链。对环境进行的工具调用,环境对这些工具调用的响应,这些响应可能包含也提供诊断价值的错误消息,而我们几乎没有从所有这些信息中学到任何东西。因此我们提出的问题是,我们能否利用这些极其丰富的其他信息。

Original English

Speaker: The current dominant paradigm is reinforcement learning with verified rewards where given a model and a task we perform a number of parallel rollouts and get rewards at the end. Finally, an algorithm like GRPO takes these rewards and converts it into gradients that are applied back to the model. However, as we can see, there was a lot of information in each of these rollouts. But we only learned an O of one score and propagated that via gradient descent. We can see that there is chains of thought. The tool calls made to the environment, the envir environment's responses to those tool calls which could potentially contain error messages which also provide diagnostic value and we learned almost nothing from all of that. So the question we ask is can we make use of this other extremely rich information.

Speaker: 我们的想法是在文本空间中执行反思优化,在这里,我们可以让语言模型或智能体查看整个演练的追踪记录,并反思其中哪些有效、哪些无效,而不是仅仅使用 0 或 1 的奖励信号。这种反思有可能会用到所有的中间输出,甚至可能会进行其他的工具调用,比如从贵公司的知识库或某些指南教材中进行检索等等。所以这是第一个关键想法。第二点是,与其只用微小的增量来更新权重,我们可以改为更新提示词(prompt),只需一次自然语言更新就能带来非常大的行为改变。让我们举个简单的例子。假设你的任务是编写一个文本摘要系统,该系统的提示词写着“生成单行摘要”。如果我直接去微调那个提示词,将其改为“生成 10 行摘要”,我们都会同意,仅仅这一个词的改变就会使系统的行为发生相当显著的变化。而进行这一个词的改变是非常快的,我们可以反思我们自己的行为并找出需要改变的地方。如果要让我们的 AI 系统实现类似的行为更新,我们将必须按顺序进行成千上万次非常微小的梯度更新。

Original English

Speaker: Our idea is to perform reflective optimization in text space where instead of only using the zero or one reward signal, we can have a language model or an agent look at the trace of the entire rollout and reflect on what worked in them, what did not work in them. And this reflection could potentially use all intermediate outputs and potentially even make other tool calls such as retrieval from your company's knowledge base or some guide textbook and so on. So that's the first key idea. And the second is that instead of only updating weights with small deltas, we can instead update a prompt where a single natural language update can give a very large behavior change. Let's take a simple example. Let's say you're tasked with writing a text summarization system and the prompt of that system says generate a oneline summary. If I just go and tweak that prompt to say generate a 10-line summary, we can all agree that the behavior of the system would change quite significantly with that just one word change. And making that one word change is quite quick and we can reflect on our own behavior and identify what needs to change. If we were to achieve a similar kind of behavior update from our AI system, we would have to have thousands of gradient very tiny gradient updates sequentially.

JPEA (基于文本空间的反思性提示词优化)

Speaker: 因此,基于这个关键想法,我们提出了 JPEA,这是一种针对智能体的反思性提示词优化技术。它使用了一个进化循环,以及一种新颖的基于帕累托(Pareto)的候选选择方法,我稍后会讲到。它类似于在文本空间中进行强化学习,在这里,我们不再只是收到一个奖励分数,我们实际上获得的是分数连同可能非常特定于领域的文本反馈,并从中学习关于该领域的所有知识。

Original English

Speaker: So with that key idea, we proposed JPEA which is a reflective prompt optimization technique for agents. It uses an evolutionary loop along with a novel parto-based candidate selection which I will come to later. It is akin to doing reinforcement learning in text space where instead of just rewarding receiving a reward score, we are actually obtaining score along with textual feedback which can be very domain specific and learn all about the domain from it.

Speaker: 让我们将 JPEA(原文此处有口误发音为 Japa)与 gRPO 进行比较,后者是领先的强化学习(RL)技术之一。在 x 轴上,我们有训练步数,这也与所看到的数据样本数量成正比;在 y 轴上,我们有正在训练的领域的性能。我们可以看到的是,JPEA 在仅使用三个数据点进行一轮反思后,就已经能够获得 gRPO 在 25000 次演练后才获得的性能提升的两倍。继续运行 JPEA 几步,能进一步将这个差距再扩大 2 倍。我想在这里指出的是,Qwen 38B 模型在这里正在优化自己。没有任何外部的专家导师参与。

Original English

Speaker: Let's compare Japa with gRPO which is one of the leading RL techniques. On the x-axis we have the number of training steps uh also proportional to number of data samples seen and on the y-axis we have the performance on our domain that we are training for. And what we can see is that Japa in just one round of reflection using just three data points is already able to get twice the performance gains that gpo got after 25,000 rollouts. Continuing to run Japa for a few more steps further increases that gap itself by another 2x. I want to note here that the model Quen 38B is optimizing itself here. There is no external expert teacher involved whatsoever.

Speaker: 那么 JPEA 学到了什么?它不同于以前的某些会利用模型特性的提示词优化器,比如“如果你不生成好的提示词我奶奶会很生气”。在这里,JPEA 实际上给出了一个非常详细的问题规范,其中包括如何理解输入内容。这部分特定管道的目的是什么,背景是什么?从数据中得出的一些关键观察和教训是什么?因此,我们在这里看到的提示词是针对多跳问答系统的第二跳的,其中给定一个问题,我们需要检索一些有可能回答该问题的文档。查看那些文档并对其进行总结,最后回答问题。在这里我们看到的是,JPEA 发现第一跳的文档通常涵盖一个实体或方面,而第二跳实际上应该去检索与它相关的文档。我们已经看到,人类工程团队每当有新模型问世时,都会花费数周的时间来手动微调这里或那里的某个词,试图发现问题规范。整个过程现在可以通过 JPEA 完全自动化,取决于你的管道,它大约需要半小时到 1 小时的时间来运行。

Original English

Speaker: And what does Japa learn? Unlike prior prompt optimizers somewhat which would uh uh use model idiosyncrasies like my grandmother will be really angry if you don't generate a good prompt. Here Jpai is actually giving a very detailed problem specification which includes how to make sense of the input. What is the purpose and context of this particular pip uh part of the pipeline? What are some key observations and lessons from the data? So the prompt we are seeing here is for the second hop of a multihop question answering system where given a question we need to retrieve some documents that could potentially answer that question. Look at those documents summarize it and then finally answer the question. And here what we see is Japa has found out that first hop documents that often cover one entity or aspect and the second hop should actually be uh recovering documents that are related to it. We have seen that human engineering teams whenever a new model comes out spend weeks of their time manually tweaking one word here and there trying to discover the problem specification. This entire process is fully automated now with Japa which takes about half an hour to 1 hour to run depending on your uh pipelines.

Speaker: 我们还可以将 JPEA 应用于领先的专有模型。仅举一个例子,在这里我们能够优化 GPT-4.1 mini 的性能,使其在一个数学任务上的表现优于 GPT-4.1,并且我们可以看到 JPEA 在提示词空间本身所做的那种信息蒸馏。

Original English

Speaker: We can also apply Japa to leading proprietary models. Just for an example here we were able to optimize GPT 4.1 minis performance to outperform GPT 4.1 on a math task and we can see the kind of information distillation JPA has done in the prompt space itself.

JPEA 的应用案例与算法原理

Speaker: 回到样本效率的问题,AMD 开发了一种叫做 NPU XDNA2 的新硬件加速器,它使用了全新的 API 进行编程,而互联网上几乎没有任何相关可用信息,正因如此,当时领先的模型即 GPT-4 在执行这项任务时惨败。我们能够采用一个在这项任务上只获得 4.25% 成绩的现有智能体,并在不对智能体本身做任何其他改变的情况下应用 JPEA,然后我们得到了这个提示词,将性能提高了 7 倍,达到了 30.52%。所以这说明,有很多特定领域的信息,如果你将其包含在你的 AI 系统提示词中,模型实际上会表现得更好,而 JPEA 可以帮助你完全自动化地发现这些信息。我想强调一下写着“避免包含 ADF.h”的那个句子。有趣的是,AMD 实际上确实提供了一个叫做 ADF.h 的库用于给 NPU 编程,但它无法与我们正在处理的这一最新一代硬件一起工作,而 JPEA 仅用一步就发现了这一点。

Original English

Speaker: Coming back to the problem of sample efficiency, AMD developed a new hardware accelerator called NPU XDNA2 which had used a completely new API to program which had almost zero available information over on the internet and because of this uh the leading models at the time which was GPT4 was failing miserably to perform this task. We are able to take an existing agent which was getting 4.25% 25% on this task and apply Japa without any other change to the agent itself and we got this prompt and pushed this performance 7x to 30.52%. So what this is uh what this goes to say is there can be lots of domain specific information which if you include in your AI systems prompts the models could actually perform much better and JPA can help you fully automatically discover that. I want to highlight the sentence saying avoid including ADF.h H. Now the interesting thing is AMD actually ships a library called ADF.h for programming NPUs but that did not work with this latest uh generation of hardware that we were working with and Jeppo was able to discover that in just one step.

Speaker: 那么它是如何工作的呢?这是一个极其简单的算法,它只是接受你用任何智能体框架编写的 AI 管道,甚至是你可能拥有的原始 LLM 调用。它只是在几个例子上运行你的系统,并收集特定领域的反馈。无论你的环境包含什么信息都会被观察到。其次,它利用读取反馈的 LLM 或智能体进行反思,并提出更好的提示词。最后,也是最重要的一点,它保留了一个帕累托(Pareto)池,在这个池子里,它会保留即使只在一个训练示例上胜出的每一个候选者,而不仅仅是得分最高的那一个。

Original English

Speaker: So how does it work? It's an extremely simple algorithm which simply takes your AI pipeline written in any agentic framework or even raw LLM calls that you may have. It simply runs your systems on a few examples and collects domain specific feedback. whatever information your environment contains is observed. Second, it runs reflection with an LLM or agent that reads the feedback and proposes a better prompt. Finally, and most importantly, it keeps a parto pool where it keeps every single candidate that wins on even one training example and not just the top scorer.

Speaker: 问题是,但为什么要保留帕累托池?我们一直被频繁问到这个问题,即 JPEA 真的比在一个循环中运行模型更好吗。所以我们去测试了一下,发生的情况是,一个单纯的循环只保留最好的那个,并很容易陷入局部最优。因此,在左侧你看到的是使用 LLM 在循环中生成的一棵搜索树。从左上角的种子提示词开始,我们要求 LLM 改进提示词。它改进了提示词,生成了一个让我们进入中间节点的提示词。然而,这个提示词陷入了局部最优,当再次要求 LLM 尝试改进它时,它提出了一些建议,但实际上并没有更好。因此,它退了回去,又再次尝试去改进它,然后它一直重复这样做,耗尽了所有的搜索预算。另一方面,在右侧使用 JPEA 基于帕累托的候选选择策略,我们可以看到它维持了一个更加平衡的搜索过程,最终收敛到一个高得多的分数。在四项基准测试中,我们看到 JPEA 所获得的半数以上的收益实际上都归功于此,它获得的性能提升几乎是你仅仅在一个循环中应用该模型所能获得的两倍。JPEA 能够在......

Original English

Speaker: The question is, but why keep a parto pool? And we kept getting asked this question a lot that is Jeppa really better than running the model in a loop. So we went and tested it out and what happens is a loop keeps only the best and gets stuck in a local optima. So on the left hand side you see a search tree that was generated by using an LLM in a loop. Starting from a seed prompt at the top left where um we asked the LLM to improve the prompt. It improved the prompt and it generated a prompt that gave us the middle note. However, this prompt got stuck in a local optima and once again when we asked the LLM to try and improve it, it proposed something but that was not actually better. So, it went back and it again tried to improve it and it kept doing this and it exhausted all of the search budget. On the other hand, with Japa's parto based candidate selection strategy on the right, we can see that it maintains a much more balanced search process eventually converging to a much higher score. Across four benchmarks, we saw that more than half of the gains seen with Japa actually account for this and it gets almost twice the performance gains that you would get with just applying the model in a loop. Japa can perform really well across

Japa 与 Optimize Anything 通用文本优化

Speaker: 在各种不同的基准测试中。在这里,我们可以看到在问答、指令遵循、主张验证以及数学等方面的结果,所有领先的前沿模型公司都已经针对这些方面对他们的模型进行了大量的优化,而我们仅仅通过优化提示(prompt)就仍然能够获得超过 10% 的提升。因此,到目前为止,我们只看到 Japa 在优化提示。但是 Japa 的作用远远超出了提示的范畴。而且因为提示仅仅是决定 AI 系统行为的文本产物,同样的算法可以用来改进任何你能用一段文本表达并且能够对其进行评分的东西。

Original English

Speaker: diverse benchmarks. Here we see results on question answering, instruction following, claim verification as well as math which all the leading frontier model companies are already optimizing their models a lot for and we are still able to get plus 10% just by optimizing the prompt on it. So we have so far seen Japa only optimizing the prompts. But Japa goes far beyond prompts. And because prompts are just text artifacts that determine AI system behavior, the same algorithm can improve anything that you can express as a piece of text and you can score.

Speaker: 例如,你的整个智能体测试框架(agent harness)最终也不过就是一个 Python 脚本或者 JavaScript 文件,我们可以将同类的反思性优化过程直接应用于那整个文件,并对其进行处理。因此,只要你能将其写成文本并对其进行评分,JPA 就能对其进行优化。带着这样的深刻洞察,我们提出了“优化一切”(optimize anything)的概念,这是一个通用的 API,用于在任何给定领域优化任何文本参数。比如在代码优化领域,假设你想要优化一段 CUDA 内核代码。这里的输入仅仅是那段 CUDA 内核代码,评估器会检查这段代码,也许会编译它、分析它,并生成一系列我们称之为“可操作的辅助信息”(actionable side information)的相关反馈,然后这些信息会被提供给大语言模型(LLM),LLM 随之会在保持帕累托最优(Pareto)的前提下提出一个更好的候选方案,并且它会不断重复这个过程,直到我们达到结果的收敛。

Original English

Speaker: For example, your entire agent harness is eventually just a Python or a JavaScript file and we can apply the same kind of reflective optimization process to that entire file and we can work with it. So if you can write it as text and score it, JPA can optimize it. So with that insight in mind, we propose optimize anything which is a universal API for optimizing any text parameter given any domain like code optimization where let's say you want to optimize a CUDA kernel code. The input is just that CUDA kernel code where an evaluator looks at this piece of code, maybe compiles it, profiles it, generates a bunch of related information that we call as actionable side information which is then provided to an LLM which proposes an better candidate maintaining this parto and it keeps the uh repeating this process um till we get convergence.

Speaker: 同样的事情也可以应用于数值优化,你的数字实际上可以序列化为文本形式;或者框架优化,整个测试框架可以序列化为文本;甚至云调度策略优化,其中的调度策略或启发式算法可以表达为一段文本,而评估器可以是成本的负值,或者是衡量准确性、效率的某个函数,而可操作的辅助信息可以是作业跟踪(job traces)、SLA 违规等信息。

Original English

Speaker: The same thing can be applied to numeric optimization where your numbers can actually be serialized as text or harness optimization where an entire harness can be serialized as text or even cloud scheduling policy optimization where the scheduling policy or heristic algorithm can be expressed as a piece of text and the evaluator can be something like the negative of cost or some function measuring accuracy uh efficiency and the actionable side information can be something like job traces SLA violations and so on.

极简 API 设计与开放性反馈

Speaker: 这个 API 使用起来极其简单。它所需要的只是让你提供一组你希望被解决的问题,以及一个能够返回分数的评估函数或适应度函数,再加上任何可用的特定领域辅助信息。如果你的领域会产生专家反馈,就返回专家反馈。如果你的领域会产生编译器错误信息、分析器信息、工具调用错误信息,就返回这些。如果你有一些写好的文档,就返回文档。任何类型的信息都可以,它是一个非常开放的字典。你字面上可以返回任何东西,而你所要做的就是使用这个适应度函数和你拥有的一组问题来调用“优化一切”(optimize anything),然后“优化一切”就会接管并处理它,最终为你提供一个优化后的解决方案。

Original English

Speaker: The API is dead simple to use. All it requires is you give us the set of problems that you care to be solved along with an evaluator function or a fitness function that returns a score along with any available domain specific side information. If your domain produces expert feedback, return that. If your domain produces compiler error messages, profiler messages, tool call error messages, return that. If you have maybe a written up documentation, return that. any kind of it's a very open-ended dictionary. You can return literally anything and all you do is you call optimize anything with this fitness function and the set of problems that you have and optimize anything will sort of take care of it um and give you a optimized solution.

实际应用案例与智能体框架自动发现

Speaker: 让我们来看一些应用场景。假设你的任务是生成一个 3D 独角兽。这就是你通常会编写的所有代码,或者你的智能体现在可以编写它,因为我们已经看到“优化一切”对于像 Claude Code(转录为 plot code)这样的领先智能体来说是一个非常易于使用的 API。所以你只需写下这段代码,声明要优化一个 Python 程序来生成一个 3D 独角兽。而候选对象就是一个能够产生 PNG 渲染图之类的 Python 脚本,这里就是结果。在左侧,我们可以看到 Claude Opus,如果你交给它这个任务,这就是它生成的结果。而在右侧,是我们使用“优化一切”得到的独角兽。这只是为了好玩。

Original English

Speaker: Let's see some applications. Let's say you were tasked with generating a 3D unicorn. This is all the code that you would write or your agent can now write it because we have seen that optimize anything is a very easy to use API for leading agents like plot code. So all you do is write this code which says optimize a Python program to generate a 3D unicorn. Um and the candidate is a Python script that produces a PNG rendering whatever and here is the result. On the left hand side we can see claude opus 4.6 if you gave it this task this is what it generated. And on the right hand side, what what we the unicorn that we get with optimize anything. This just for fun.

Speaker: 但是假设你的任务是编写一个智能体来解决一个特定的任务。通常情况下,团队会花费大量的时间来调整他们的智能体,为它构建工具,编写工具描述,精心编排控制流等等。在这里,我们从一个简单的四行 Python 程序开始,它只是调用模型的思维链(chain of thought)来解决一个 RKGI 问题。仅仅通过 16 轮的反思,Japa(转录为 Jeppa)在“优化一切”中就找到了这个复杂的六步智能体,它将 Gemini Flash 在 RKGI 上的准确率从 32.5% 提升到了 89.5%。我们可以看到,这个智能体是自动化的,它自己进行规则假设归纳和代码综合。它执行并跟踪代码,自动调试这段代码。返回并提出该代码的新版本。最后,它在实际的测试输入上运行它并返回输出。这是一个可运行的示例。你可以扫描这个二维码,现在就可以运行这个示例。

Original English

Speaker: But let's say you were tasked with writing an agent to solve a specific task. Typically teams spend lots and lots of time tweaking their agents, building tools for it, writing tool descriptions, uh carefully orchestrating the control flow and so on. Here we started with a simple four-line Python program that was simply calling a model's uh chain of thought to solve an RKGI problem. Within just 16 rounds of reflection, Jeppa within optimize anything was able to find this sophisticated sixstep agent that took RKGI accuracy on RKGI uh that took RKGI accuracy of Gemini flash from 32.5% to 89.5%. And we can see that this agent is automatic like by itself doing rule hypothesis induction code synthesis. It executes and traces the code automatically debugs this code. Goes back and proposes new versions of that code. And finally it runs it on the actual test inputs and returns the output. This is a runnable example. You can go to this QR code and you can run this example right now.

Speaker: 因此,将这种发现智能体测试框架的相同方法应用于 MATH 500 基准测试,我们只需创建一个两步智能体,就能将模型的准确率提高 20%。我想再次强调的是,我们所做的只是要求“优化一切”去优化一个智能体文件,然后它就自动发现了这个复杂的智能体架构,除了指定目标和任务之外,我们不需要做任何其他事情。

Original English

Speaker: So um applying the same uh uh like approach of discovering agent harnesses to math 500 we are able to push its accuracy of GPT 4.1 nano by 20% by simply creating a two-step agent. And again I want to emphasize that all we did is we asked optimize anything to optimize an agent file and it was automatically discovering the sophisticated agent architecture and we did not have to do anything other than specifying the objective and the task.

跨模型的智能体技能迁移 (Agent Skills)

Speaker: 最后,我们每一个人都在使用某种编码智能体,比如 Claude Code 或 Codex,或者是你最喜欢的智能体,而智能体技能(agent skills)已经成为生态系统中非常核心的一部分,几乎所有的编码智能体都能理解技能。假设你想为你的特定代码库优化技能。这就是你写的代码,它声明“从轨迹中学习一项技能。当编码智能体遇到类似问题时,这项技能应该会有所帮助”。我们只需赋予它这种自然语言行为。

Original English

Speaker: Finally, every single one of us is using uh some coding agent like cloud code or codex or maybe your favorite agent and agent skills has become a very leading part of the ecosystem where almost all coding agents understand skills. Let's say you want to optimize skills for your specific repository. This is the code that you write which says learn a skill from the trajectory. When the coding agent is presented with similar problem, the skill should be helpful. We just give it this natural language behavior.

Speaker: 我们看到的是,我们从一个使用小型模型配置的迷你智能体开始,因为我们的预算非常有限,但我们能够将其在 Go 语言代码库问题解决上的表现从 24% 提升到 93%。这在问题解决上几乎是 3 倍的飞跃,但更重要的是,在小型模型智能体上以非常低廉的成本优化出来的这些技能,我们能够将它们直接提取并应用到最新的 Claude Sonnet 模型上。这大约是几个月前完成的,但我们将它应用到了 Claude Sonnet(转录为 clots onet 4.5)上,将其问题解决准确率推高至 100%,而更关键的是,将执行时间或问题解决时间缩减了近 50%。我们将时间缩减了一半,这也意味着它消耗了更少的 token,因为技能中包含了关于代码库是如何组织的、如何调用测试用例、某个特定功能是在哪里实现的、该代码库使用了什么样的构建系统等等信息。这是一个名为 GSkill 的功能。你可以在 Japa 的代码库中找到它,而且它也是完全开源的。

Original English

Speaker: And what we see is we started with miniu agent with GPT5 mini because we were very budget constrainted and we were able to take its performance from 24% to 93%. An almost 3x jump on go repository issue resolution but more importantly the skills that were optimized very cheaply on a GPT5 mini agent we are able to take that and apply to the latest claude sonnet. This was done a uh about a few months back but we applied it to clots onet 4.5 pushing its accuracy to 100% issue resolution while more importantly cutting down the execution time or issue resolution time by almost 50%. We cut it down into half which also means it spent less tokens because skills contain information about how the repository is organized, how to invoke the test cases, where a particular feature is implemented, um what are the build system used by this repository and so on. This is a a feature called GSkill. You can find it in the Japar repository and it's fully open source as well.

Optimize Anything 的三种模式与广泛的行业应用

Speaker: 所以,“优化一切”是一个单一的接口,它提供了三种优化模式。如果你只有一个单一的问题,比如你想要优化一个单一的矩阵乘法内核,你可以用这种模式。如果你有任意数量的相关问题,比如你想要优化一个矩阵乘法内核以及一个点积内核,并且你知道这两者之间可能存在一些信息传递,你可以使用我们称之为“多任务搜索模式”(multitask search mode)的功能。最后,建立一项技能,即如果你想在一组设定好的问题上进行优化,但你的部署实际上可能会遇到许多新问题的情况。就像在数学提示优化的情况下,我们是在一些示例上进行训练,但是当我们部署它时,我们可能会收到一种全新的查询。因此我们需要关心泛化模式。在那里你可以进行提示优化、智能体架构优化等等。

Original English

Speaker: So, optimize anything is a single uh interface that provides three optimization modes. If you have just a single problem like there is a single matrix multiplication kernel that you want to optimize you can use it that way. If you have any number of related problems like you want to optimize a matrix multiplication kernel along with a dot product kernel and you know there might be some information transfer between these two you can use what we call as the multitask search mode and finally build a skill which is if you want to optimize on a set number of problems but your uh deployment can actually come up with many new problems. So like uh in case of math op like in case of math prompt optimization we are training on some examples but when we deploy it we can receive a completely new kind of query. So we care about generalization mode. So there you can do prompt optimization agent architecture optimization and so on.

Speaker: 因此,“优化一切”可以用于广泛的领域,包括云调度策略优化,与专家启发式方法相比,我们能够将成本削减近 40%;编写自定义求解器,在黑盒数学优化中匹配甚至超越 Optina;创建智能体技能、进行提示优化等等。它非常容易使用,以至于在发布后的仅仅 20 小时内,Snorkel 的团队就已经用它改进了他们的一些内部基准测试,并在推特上分享了这件事。

Original English

Speaker: So optimize anything is can be used for a broad set of domains including cloud scheduling policy optimization where we were able to cut costs by almost 40% compared to expert huristics write custom solvers to match and exceed Optina even in blackbox mathematical optimization create agent skills prompt optimization and so on. It is so easy to use that within just 20 hours of releasing it, people at snorkel had already improved some of their internal benchmarks with it and were tweeting about it.

Speaker: 此外,Japa 也提高了多模态 VLM(视觉语言模型)的性能。在这里,我们能够将领先模型的 OCR 错误率降低近 35%。这是一份经过外部验证的报告。类似地,Databricks 实际上在他们部署的智能体性能上实现了 90 倍的成本降低,并且在这里,他们能够调整开源模型使其性能超越 Claude Opus,同时价格便宜了 90 倍。更重要的是,你在 Claude Opus 基础上看到的性能增量提升,实际上比在开源模型上看到的提升还要大。有人曾问我,随着模型变得越来越好,提示优化的重要性是否会下降?我认为情况恰恰相反,这……

Original English

Speaker: So, and Jeppa also improves multimodel VLM models performance. Here we are able to cut OCR error rates for leading models by almost 35%. And this is an externally validated report. Um, similar similarly, data bricks actually achieved 90x cost reduction in their deployed agents performance. uh uh performance and here they were able to tune GPT OSS 120B to outperform Claude Opus while being 90x cheaper. More importantly, the performance delta improvement that you see on top of Claude Opus is actually bigger than the one you see on open source models. Some people have asked me that oh as models get better the importance of prompt optimization will go down. I argue the opposite which

Japa 与持续学习飞轮

Speaker A:……随着模型变得越来越好,它们遵循指令的能力也会越来越强,你给一个非常聪明的模型提供的关于你任务的指令越精确,那个模型在解决你的任务时就会表现得越好。这正是我们在这里看到的情况:指令越好,Claude Opus 的表现实际上提升得非常多。有些人有这样一个疑问:如果我们有主观性任务,非常难以评估怎么办?Japa 实际上可以从生产环境的追踪数据中为你的任务学习评估方法。具体做法是,你从你的智能体中收集一系列生产环境的追踪数据。让人类去标注大概 50 条这样的轨迹,给出非常详细的反馈。比如“这是一个长回复”,“这是一个短回复”,“这是一个好回复”,“这里用到了这个术语”,等等。一旦你得到了这些人类的标注,你就可以使用 Japa 来优化一个“大语言模型作为裁判(LLM-as-a-judge)”的提示词。然后你可以使用这个大语言模型裁判提示词回过头来优化你的智能体,并部署该智能体。这就会变成一个数据飞轮,让你可以不断地改进它。这是一个非常成功的范式,一些生产环境中领先的团队已经在使用。接着我们被问到的一个问题是:我们真的能使用这种反思性优化(reflective optimization)来训练模型吗?最近我们发表了一篇名为《Learning Fast and Slow》的论文,在其中我们提出了快慢学习的方法,我们可以协同优化模型权重和提示词测试工具,这展现出了一种持续学习算法所期望的非常强大的特性。我没有太多时间来过细节了,但请大家去看看那些论文。自从发布以来,Japa 已经被这些公司在生产环境中使用,也是这些论文中的主要方法论。你可以看到 Dropbox 和 Shopify 的 CEO 在讨论他们对 Japa 的使用,而且 OpenAI 也写了一篇博客文章,介绍如何使用 Japa 构建自我改进的 AI 系统。所以,开始使用它是非常简单的。它可以插入任何框架、任何模型中,而且它绝对没有任何硬依赖。因此你可以把它部署在任何种类的环境中。所以,不要害怕在文本领域进行优化,很多问题都可以被框定为优化问题。尽可能多地向优化器提供可落地的辅助信息,并展现尽可能多的领域特定信息,未来的优化器将能够很好地利用它们。所以请大家去了解一下。非常感谢。

Original English

Speaker A: ...is as models get better they will get better at instruction following and the more precise instruction about your task that you have to give to a very smart model the better that model will be at solving your task and this is exactly what we see happening here the better the instruction was claopus actually jumped much higher some people have this question of uh what if we have subjective tasks which are very hard to evaluate jpa can actually learn evals for your task from production traces. The way to do that is you collect a bunch of production traces from your agent. Get a human to annotate just about 50 of those trajectories giving very detailed feedback. This is a long response. This is a short response. This is a good response. This uses this terminology, whatever. And once you get those human annotations, you can use Japa to optimize an LLM as a judge prompt. And you can use that LLM as a judge prompt then to go back and optimize your agent and deploy that agent. And this becomes a data flywheel where you can keep improving it. And this is a successful paradigm that uh some leading teams in production are already using. Then the question we get asked is like can we actually use this uh reflective optimization to train models and we recently had this paper called learning fast and slow where we propose fast slow learning where we can co-optimize model weights and prompt harnesses and this shows some very strong properties that one would want in a continual learning algorithm. Um I don't have much time to go over details but please uh look at the uh papers and uh since uh since release Japa has been used in production by these companies as well as the main methodology in these papers and here the CEO of Dropbox and Shopify are talking about their use of Japa and OpenAI also wrote a blog post about how you can build self-improving AI systems with Japa. Um so it's very simple to get started. It can plug into any framework, any model and it has absolutely zero hard dependencies. So you can deploy it any in any kind of setting. So um don't be afraid to optimize in the tech space and many problems can be framed as optimization. So bring actionable side information and surface as much domain specific information as you can to optimizers and the optimizers of future will be able to work with them. So please go and check it out. Thank you very much.

递归编码智能体 (Recursive Coding Agents)

Raymond Whitampamp:大家好。我是 Raymond Whitampamp,今天我要讲的是递归编码智能体(recursive coding agents),也就是把递归语言模型(RLMs)的经验应用到编码智能体中的理念。这是我在自己的独立研究 raw works 中,以及最近在 Open Pros 任职期间做过的一些工作。那么,为了稍微激发一下大家的兴趣,我们都想要成果。我们都想要能代表我们工作的智能体。我们想要可靠的“同事”,能在我们在做些有趣的事情、在外面远足、在放松休息、在忙别的事情时,把任务搞定。而我的观点以及我的经验是,这里的瓶颈并不是智能(intelligence)。模型已经足够聪明了。它们知道各种各样的事情。它们了解整个互联网,但它们无法可靠地交付成果。因此,我无法信任它们。举个非常简单的例子,某一天,我用一个单一的提示词(当然是个很长的提示词),就得到了一款几乎完全可以运行的 SaaS 应用。然而第二天——我发誓这确实发生过——Claude Code 把我 Solana 钱包里的东西全清空了。哎呀,好吧。所以,这并不能真正让人产生信任。在我们这里的最下方,有一条发展路线。我们都希望能向右边那个状态发展:我们只需要坐在那里冥想,事情就会自然而然地成真。那么,这来自哪里呢?这是来自 AI Engineer 大会的代码。实际上印在 T恤背面。Engineer Code 2025 年 11 月。天哪,我希望当时你也在场。如果你不在的话,去 YouTube 上看看吧。那真的非常棒。

所以,我的论点是:今天的智能体就是“管理不善的天才(mismanaged geniuses)”。智能是具备的,缺失的那一层是如何去规范、管理、复用和验证它们的工作。那么,“管理不善的天才”这个构架和说法,是来自于 MIT 的 Alex Zang、Zed Lee 和 Omar Katab。Alex 和 Omar 也是最初这篇递归语言模型论文的作者之一。最近我也在 Touring Post 上略微讨论过这个话题。我忘了说,这些幻灯片实际上是一个网站,叫做 recursivecodingagents.com。你可以去这个网站上点击它们。所以我在这里要展示的所有内容都是可交互的。那么,什么是递归语言模型呢?我喜欢这么说:在 RLM 中,上下文(context)本身就是计算的对象。这本质上是工具调用(tool calling)和推理(reasoning)的结合。我们会在下一张幻灯片中更详细地讨论这一点。但是其核心理念是,完整的提示词并不是一个简单的用户查询。完整的提示词是一个变量。它可以是一个文件,或者是很多个文件。然后我们有一个“读取-求值-输出循环”(REPL),在原论文中,智能体正是与这个 REPL 进行交互。那是 Python。并且 RLM 会被指令以符号化(symbolically)的方式去操作那个提示词。所以,不要只是把整个东西都读进你的上下文窗口里。要去符号化地探索它。甚至,你不直接进行符号化的探索,或者你可能只是稍微摸索一下。

Original English

Raymond Whitampamp: Hello there. My name is Raymond Whitampamp and today I'm going to talk about recursive coding agents which is this idea of applying the lessons of recursive language models RLMs uh to coding agents. This is some work that I have done both in my independent research um raw works uh and also more recently in my role at open pros. So to motivate this a little bit, we all want outcomes. We all want agents that are working on our behalf. We want reliable co-workers that are getting things done while we are doing something fun, while we're out on a hike, while we're cold chilling, while we're doing the do. And my argument and my experience is that the bottleneck to this is not intelligence. The models are intelligent enough. They know all kinds of things. They know the entire internet, but they can't reliably deliver outcomes. And so I can't trust them. So as a very simple example, you know, one day I get almost a fully working SAS app from a single prompt, granted a long prompt. The next day, and I swear this actually happened. Cloud code empties the entire contents of my Salana wallet. Oops. Okay. So, that doesn't really instill trust. So, at the bottom here, we've got this pro this progression. Okay. And we all want to move towards the the one on the right where we're just sort of sitting there meditating and and things are manifesting. And so, where does that come from? This is from the AI engineer code. It's actually from the back of the t-shirt. Engineer code November 2025. Man, I hope I hope you were there. If you weren't, watch it on YouTube. It was it was amazing. So, here's the thesis. The thesis is today's agents are mismanaged geniuses. The intelligence is there and the missing layer is how do we specify and manage and reuse and verify the work. So this uh framing this phrase the mismanaged genius uh comes from Alex Zang Zed Lee and Omar Katab at MIT. Um and Alex and Omar are part of the authors of the original recursive language models paper. Uh I've also talked a little bit about this recently on touring post. Um I forgot to mention that these slides are actually a website recursivecoding agents.com. So you can click on them uh by going to this website. So everything I'm going to show in here is is interactive. Okay. What are recursive language models? So I like to say that in an RLM the context itself is the object of computation. Um and this is essentially a marriage of tool calling and reasoning. We're going to talk a lot more more about that in the next slide. But the idea is that the full prompt is not a simple user query. The full prompt is a variable. The full prompt could be a file or many files. Um, and we have this readaluate print loop ripple um that the agent is interacting with in the original paper. That's Python. And the RLM is instructed to operate symbolically on that prompt. So don't just read the whole thing into your context window. Um, explore it symbolically. And uh even more you don't even directly explore symbolically or maybe you do a little bit of poking around.

利用 Auto Research 优化 GPU 内核

Tis:大家好,我是 Tis。我今天要解释的是,我们如何利用 Auto Research 让模型速度提高三倍。在此之前,我其实曾在宿舍里用 1080 Ti 挖过 GPU 矿,一路走到了后来在 Tesla 负责 Tesla AI 的推理优化。但首先,什么是 Auto Research?Auto Research 是来自 Andrej Karpathy 的一个框架,你基本上是在为一个智能体设定一个框架,让它朝着你定义的目标前进,你所要做的仅仅是在高层次上告诉它你想要它做什么,它会在前进的过程中尝试各种东西,并朝着那个目标来回探索。实际上,它真的就是一个 while 循环。智能体提出一个解决方案。你设置好如何定义“正确”,为我们进行基准测试。然后你要么保留,要么回退那个结果,你就在一个循环中不断做这件事,直到你的目标被达成。所以这非常契合 GPU 内核(kernels)。如果你不知道 GPU 内核是什么,它基本上就是一个底层的算子,比如矩阵乘法,或者某个专家模型的计算。为什么 GPU 这么适合 Auto Research 呢?因为它们具有超级强的可验证性。你可以去验证它们的正确性和速度,这基本上就是你的 Auto Research 框架所需要的一切。

当然在实际操作中,还是有一些注意事项的。Auto Research 框架在选择诸如块大小(block sizes)以及那些细小参数方面非常出色,但它们在涉及高层次理念时依然表现很差,比如“我想使用这个 GPU,而且我实际上想要对它进行流水线操作(pipeline)”这种想法。它想不出这种突破性的点子。所以这仍然需要人类来完成,但是一旦你把思路理清楚了,实际的代码实现就非常直接了。所以我的意思是,想出好点子依然是你的工作。因此,这里真正的秘诀是:你有好的想法,Auto Research 去挑选那些参数,以及所有用于验证它实际可行所需的东西。然后朝着“速度快了 X 倍并且依然正确”这个可验证的目标迈进。你把这些与你最喜欢的模型产出的数十亿个 token 结合起来,最终得到的内核表现就会击败人工调优。那么,当你在编写自定义内核,或者让你的智能体编写自定义内核时,你真正关心的是什么呢?主要有三种情况:计算瓶颈、内存瓶颈,或者因为启动了太多的内核而产生了过多的额外开销。你可以通过使用类似 Nsight Systems(Nvidia 的分析器)这样的分析器进行性能分析来观察这些情况。所以,这一页看起来非常令人生畏……

Original English

Tis: Hi everyone, I'm Tis. Uh so I'm going to be explaining how we make models three times faster with Auto Research. Uh so previous to this uh I actually used to do GPU mining in my dorm room with 1080 Ti all the way up to working at Tesla on inference optimization for Tesla AI. Uh but first what is auto research? So auto research is this framework from Andre Kapathy where uh you basically set up a framework for an agent to move towards a goal that you define uh and all you have to do basically is say at the high level what you want it to do and it will try things as it goes and move back and forth uh towards that goal. In actuality, it's really just a while loop. The agent proposes a solution. You have a setup to to define what's correct, benchmark it for us. Uh and then you keep or revert that and you do this in a loop until your goal is met. And so this is very well aligned to GPU kernels. Uh so if you don't know what a GPU kernel is, it's basically a low-level operator. example, like a matrix multiply or an expert computation. Uh, and why are GPUs such a good fit for auto research? It's because they're super verifiable. You can verify them for correctness and speed, and that's basically all you need for your auto research framework. Uh, so in actuality, there are some caveats here. Um, the auto research framework is really good for like picking block sizes and these tiny parameters, but they're also still really bad at the high level idea, like seeing like I want to use this GPU and I actually want to pipeline it. It's not going to come up with these groundbreaking ideas. So it's still up to the human to do that, but the actual implementation is very straightforward once you once you have the idea laid out. So it is still your job to have good ideas is what I'm saying. Uh and so the actual secret formula here is you have the good ideas, auto research picks out the parameters and everything to verify that it actually works. Uh and go move toward that verifiable goal of it being x times faster and uh still correct. And you mix that with billions of tokens of your favorite model and that results in kernels that beat hand tuning. Uh so what are the actual things you care about when you're when you're when you're writing a custom kernel or you're having your agent write a custom kernel. So the three main things you can have are a compute bottleneck uh a memory bottleneck or you just have excessive overhead from uh too many kernels being launched. And you can do you can view these things with by profiling with a profiler like NSIS for example which is a Nvidia's profiler. Uh and so this this gra this page looks super daunting

自动研究与硬件约束

Speaker A: 基本而言,作为人类,你的工作是看着最上面并觉得这很蠢。我们将 32k 的数据块加载进上下文中,但实际上对于像 DeepSeek 的注意力机制,我们并不需要这样做。相反,我们应该每隔 32k 才执行一次。从高层次来看,你只需要告诉自动研究系统:“顶层这个方法很蠢,让我们改为流水线处理”,而其他所有事情,比如块的大小调整、上下文块的大小调整等,都应该交由自动研究系统来决定。我的问题是,我真的非常喜欢便宜的 GPU,也就是那些没有 NVLink 的 GPU,因为你可以用更低的价格买到它们。但问题在于,你实际上并没有现成可用于这些 GPU 的内核。因此,你必须开发一个自动研究框架以及一个自定义的测试台。那么,如何在测试台中做到尽善尽美呢?你需要确保你的智能体了解硬件。例如,在 B200 上,你需要确保它有关于 warp 的上下文。它具有 TMA(张量内存加速器)。如果你不知道这些是什么,它们只是特定硬件上的底层操作符。而且这会代代更新。例如,H200 就不会有 TMA。这是随 B200 推出的一项新功能,这就是为什么你需要将其纳入上下文中。这基本上就是你需要提供的一堆 Markdown 文件,以让它具备相关上下文。

Original English

Speaker A: but basically your job as a human is to look at the top here and be like this is dumb. uh we are loading 32k chunks into context uh and we don't actually need to for this deepseek attention for example uh and we should only be doing it every 32k instead and so at a high level all you have to be telling auto research is this top method is dumb let's pipeline it instead and everything else like the sizing the chunk sizing the context chunks that all should just be decided by auto research and so my problem is that I really love cheap GPUs and so that means like GPUs that don't have NVLink for example uh is an example of like GPUs you can get for cheaper Uh but the problem is you don't actually have kernels off the shelf for those. And so you have to come up with a auto research framework as well as a custom harness. So what goes into the harness to make this really good. Uh so one thing you really need to make sure your agent is aware of is the hardware. And so on a B200 for example, you need to make sure it has context of uh the warps. It has T-M TMA. And so if you don't know what these are, these are just uh low-level operators that you have um on a specific hardware. And this changes generation to generation. like an H200 won't have T-M for example. That's a new feature that coming out with B200 which is why you need to have this in context. Um and so this this basically is just like bunch of MD files you need to give so it has context.

Speaker A: 另一件你需要确保智能体了解的事情是模型。每一个新模型,比如 DeepSeek Flash,都会带来一些新技巧,就像 DeepSeek 在针对 DeepSeek V4 发布的 DeepSeek Flash 中引入了两种新的注意力机制。比如压缩稀疏注意力(compress sparse attention)和层次化压缩。如果你不了解这些,模型将会 100% 产生关于实际注意力机制的幻觉,而你将得到毫无用处的内核。在做这些事情时,迄今为止最大的问题将是奖励作弊(reward hacking)。如果你告诉你的内核工程师同事:“我需要让这个 GPU 内核运行得更快。”你的纯人类同事显然不会去做一些会让端到端模型推理变慢的事情。但智能体不是人类,它们会做很多让它变慢的事情。比如它们会禁用 CUDA Graph,这可能会让运行速度慢上 20 倍;它们可能会让那一个内核变快,但让整体变得不再是一个可用的内核,因为它们禁用了一堆加速功能(比如 CUDA Graph),或者只在小型上下文窗口上进行测试。因此,很大一部分工作在于定义“不该做什么”,在利用智能体通过单次提示(one-shot)轻松完成前沿工作时,这实际上非常重要。

Original English

Speaker A: Other thing you need to make sure your agent has context of is the model and so every new model like DeepS Flash comes out with like new tricks like DeepSeek had two new attentions that was released in the Deepseek Flash for Deepseek V4. Uh so compress sparse attention hierarchal compressed and if you don't do this the model will 100% hallucinate uh the actual attention mechanism and you will get useless kernels. Uh by far the biggest problem when you're doing this is going to be reward hacking. And so if you were to tell your kernel engineer co-orker I need to make uh the GPU this GPU kernel faster. Uh it's obviously not going to your human coworker is not going to go in and do some stuff that's going to make it slow like the endto-end model inference slower. But uh agents are not humans and they will do plenty of things to make it slower like they'll disable CUDA graphs which can make it 20 times slower and they might make that one kernel faster but make the whole like it's not a viable kernel because it's they're disabling a bunch of speed ups like CUDA graphs or only testing on small context windows. And so a lot of this is also just defining what not to do which is actually very important when you're doing frontier work that agents can actually easily do with a one shot.

Speaker A: 另一种奖励作弊是,当你试图编写内核时,有些模型实际上根本无法编写你所需要的那种小巧可爱的 DSL(领域特定语言)。这是 Anthropic 模型的一个常见问题。对此,我觉得 Anthropic 爱怎么说他们削弱(nerfing)模型就怎么说吧。我也只能猜测它是不是被削弱了,但我会建议使用其他模型。而且实际上,它也并不总是能在所有地方都更快。有时候,你想出来的内核可能只在 0 到 100k 的长度下表现良好,超过这个范围,你就需要退回到可以从像 Cutlass 的 Flash 中获取的默认内核。这是需要注意的另一件事,即你的内核并不总是能作为适用于所有工作负载的直接替换方案。不过,这里的一大优点是内核是可以叠加的。如果你为 DeepSeek 做了一个稀疏 MLA 内核,你可以从中获得速度提升,然后你只需将它们叠加在一起,再加上 NVFP4 带来的提升,如果你没有 NVLink,你只需不断叠加、叠加再叠加,最终你会在达到 GPU 的硬件极限时停止增长。有些人将其称为 MFU,即 GPU 实际能达到的理论最大利用率。

Original English

Speaker A: Uh, another reward hack is that some models just don't actually write the cute DSL you need uh when you're trying to write kernels. And this is a common problem with enthropic models. And so yeah, I mean anthropic says what they say about uh nerfing models. You can it's guess if it's I'm guessing if it's nerfing or not, but I would recommend using a different model. Uh and it won't always be faster everywhere actually. So sometimes the kernels you come up with might only work well on like zero to 100k and then you need to go back to this the default kernel that could you get from like a flash in for cutless. Um and so and that's another thing to look out for is that your kernel isn't always just a swap in for all all workloads. Uh but one of the great things is is that kernels compound. So like if you make one for your sparse MLA for deepseek for example um you can get speed ups there and you just stack them on like that then plus NVFP4 fore uh you could do for us if we if you don't have NVLink you just keep stacking and stacking and stacking and then eventually you taper off at whatever the hardware limit is uh for your GPU and that's uh some people call this like MFU which is like the actual theoretical max utilization from a GPU.

Speaker A: 为了更进一步,如果你拥有裸金属(bare metal)访问权限,你的自动研究框架可以做非常 hack 的操作。因此,那些用 GPU 搞过破解的黑客可能会喜欢这个。你可以调整 BIOS 设置,可以对 GPU 进行超频,可以强制放宽 PCIe 限制,以及进行所有这些老派黑客过去常做的小调优,而这实际上也有助于提升推理性能。因此,相比于使用云提供商的虚拟化设置,在裸金属优化上你大约可以获得 25% 的净提升。一旦你实现了这些,结合你所做的所有内核优化以及硬件层面的修改,你可以获得 3 倍的速度提升。我知道这一切听起来可能很美好,但实际上情况并非总是如此,自动研究大约会有 80% 的操作是糟糕的。所以在你从事这方面工作时,一定要记住大多数尝试都会很糟糕,它会不停地试图欺骗你。但最终,你确实可以从中获得非常好的结果。长话短说(TL;DR):先要有更好的点子,然后再使用自动研究。超级简单。很简单,对吧?结果发现,你实际上可以靠做这个获得报酬。如果你觉得这很酷,可以考虑加入我们,你可以通过这里发邮件给我。谢谢大家。

Original English

Speaker A: Uh, and so to go even farther, if you have actually have bare metal access, your auto research framework can uh do very hacky things. So hackers that have hacked with GPUs are probably going to like this. You can uh tweak your BIOS settings, you can overclock the GPU, uh, you can force like PCIe relaxing, all these little tweaks of like uh, old school hackers used to do, but this can actually help with inference as well. And so net on bare metal optimizations, you can get roughly 25% over like a virtualized setup you get from using a cloud provider. Uh so once you get that you can combine all of the kernels you did as well as all of the hardware level hacks you did uh you can get a 3x speed up and so I know this this might all sound like roses and flowers but it's not actually the case around 80% of the things that auto reach is going to do are going to be bad uh so it's important to remember while you're u like working on this that most things are going to be bad it's going to try to trick you all the time uh but at the end you can actually get really good results from this tlddr uh have better ideas then use auto research. Super simple. Simple, right? Uh so turns out you can actually get paid to do this. Uh if you think this is cool, consider joining us and you can email me here. Thanks, guys.

智能体就像记忆受限的天才

Speaker B: 想象一下你在一家古董店发现了一盏神灯。你擦了擦它。一个精灵出现并问能怎么帮你。你赶紧抓住机会,所以你说:“我需要最好的工程师来帮忙处理工作上的一个不可能完成的项目。” 精灵满足了你的愿望。对我来说,最好的工程师大概就是巅峰时期(hey days)的约翰·卡马克(John Carmack)。所以你得到了卡马克。但是这个精灵有点幽默感,并且施加了一些限制,也许是出于安全考虑。卡马克只能看到你的代码库的一小部分,也许是千分之一。而且他记不住自己之前做过的任何事情。每一次对话都是从零开始。这会让人抓狂,对吧?你知道有一种做事的标准方法,但卡马克却不知道。你不得不把同一件事解释一遍、一遍、又一遍。你这边有一个天才,但在另一面他又有着深深的缺陷,而这就是智能体的现状。

Original English

Speaker B: Imagine you find a magic lamp in an antique store. You rob it. A genie appears and asks how it can help. You bury it in the line. So you say, "I need the best engineer to help with an impossible project at work." And the genie grants your wish. For me, the best engineer is probably John Carmarmac from his eight days. So you get Karmarmac. But the genie had a sense of humor and imposes restrictions, maybe for safety. Karma can only see one small part of your code base, maybe 1,000 of it. And he remembers nothing he did before. Every conversation starts fresh. That would be maddening, right? You would know there is a standard way to do stuff and karma couldn't. You would have to explain the same thing over and over and over again. You would have a genius on one side and something deeply deficient on the other and that's what agents are.

Speaker B: 让我带你过一个例子,看看在一个简单的交互中我们需要解释多少次。我们有四个代码库:模块 1、模块 2 和平台。我想更改 UI 并将更改传播到整个系统中。好的。首先,我们要更改 UI 库。假设我们不只是更改一个按钮什么的。这是第一次解释。不可避免。我们必须表达意图。好的,然后我们发布它。我们转到模块 1,我们必须解释 UI 库中刚刚发生了什么。这样它就能在这里使用那个包。注意,那通常是不同的人操作的,对吧?这个图表中的每个框都可能是不同的人完成的。然后我们发现发布的 UI 库与模块 1 不兼容。于是我们退回到 UI 部分,我们必须重新解释原始更改以及出现的问题,因为这是一个新的智能体,它不知道原始更改,显然也不知道那个问题。假设我们修复了它,然后再次发布,我们再次回到模块 1 的上下文中解释新的更改。又是同样的折磨,我是说同样的事情要对模块 2 再来一遍。然后我们转到平台代码库,我们必须解释所有部分是如何组合在一起的,并在那里实现更改。

Original English

Speaker B: Let me walk you through an example of how many times we explain things in a simple interaction. We have four reposi module one module 2 and platform. I want to change the UI and propagate the change through the system. Okay. First we change the UI library. Say we I don't change a button or whatever. That's the first explanation. Unavoidable. We have to express the intent. Okay. Then we publish it. We go to module one and we have to explain what just has happened in the UI library. So it can consume the package here. Note that that's often a different person, right? Every box in this diagram can be uh done by a different person. Then we discover that the published UI library doesn't work with module one. So we go back uh to UI and we have to reexlain the original change and the issue right because it's a new agent it doesn't know the original change and obviously doesn't know about the issue let's say we fix it right and uh publish it again we go and again we explain the new change in the context of module one same ordeal I mean do the same for module two again and then we go to the platform repo and we explain explain how everything fits together and we implement the change there.

Speaker B: 让我们想象发布一周后,UI 组件出现了一个 bug,我们需要修复它。因此我们启动一个针对 UI 代码库的智能体,我们必须再次解释一周前的最初更改以及我们看到的生产问题。所以我们对本质上是一个更改的内容进行多达七次解释。而且可能并不是同一个人在做这七次解释,但它们仍然发生了,对吧。所以这对于智能体来说非常非常典型。那么我们该如何解决这个问题呢?嗯,这里有许多导致这种体验的问题,但它们大致可以分为两类。第一个问题是,智能体本质上是受限于代码库的。智能体通常一次只能查看和更改一个代码库。它永远无法看到可能包含成百上千个代码库的整个系统。所以这大概就是当前的情况。

Original English

Speaker B: Let's imagine a week after release uh a bug appears in the UI component and uh we have to fix it. So we start an agent to the UI repo and we have to explain again the original change from a week ago and this production issue we have seen. So we have seven explanations for what essentially is one change and also it may not be one person making all these seven explanations uh but they still occurred right so that's very very typical uh with agents. So how do we solve it? Well uh there are many problems in here that contribute to this experience but they roughly fall into two categories. The first one is uh that an agent essentially is repo bound. The agent sees and changes generally one repo at a time. It never sees the whole system which can be hundreds or thousands of repos. So that's kind of

Space and Amnesia Problems

Speaker: 这是问题在空间层面的体现。其次是失忆症。智能体(agents)会忘记工作进度。每一次会话都是从零开始(blank slate)。在这种情况下,人类就成了记忆载体。这是问题在时间层面的体现。让我们仔细看看这两个问题。首先看代码库(repo)的边界。如果没有一个关于各代码库如何协同工作的模型,智能体就只能依赖人类来做研究。它无法将代码与系统其他部分对齐。它无法将 UI 的改动与模块一(module one)对齐。人类没有解释这一点。所以,发布了一个糟糕的版本。它也无法可靠地参考最佳实践和标准,因为这些通常存在于其他代码库中。编写代码的情况更糟。智能体一次只能写入一个代码库。这意味着它无法验证下游的更改。模块 1 的持续集成(CI)本应该在 UI 更改时报错失败,但事实并非如此。智能体无法同时更新消费者(consumers)。尽管在进行 UI 更改时,它完全掌握着这样做所需的信息。它确切地知道自己在做什么。因此,用户不得不向每个消费者不完美地重新解释一遍。在 20 个代码库中修改某个东西,意味着你要解释 20 次。这不仅耗费了大量开发人员的时间,也消耗了大量的 tokens。

Original English

Speaker: the space component of the problem. Second is amnesia. The agents forget the work. Every session starts with a blank slate. The human becomes a memory in this case. That's the time component of the problem. Look at the two closer. Take the repo boundary first. Without a model how repos fit together, the agent leans on the human to do the research. It can't align the code with the rest of the system. It couldn't align the UI change with module one. The human didn't explain it. So, a bad version shipped. It can't reliably reference best practices and standards either because those often live in other repos. Writing is even worse. The agent writes to one repo at a time. It means it can't validate changes downstream. Modules 1CI should have failed on the UI change, but it didn't. The agent can't update consumers at the same time. Even though, you know, while making the UI change, it has perfect information to do so. It knows exactly what it's doing. So, the user has to reexplain stuff imperfectly to each consumer. Changing something across 20 repos means you're explaining things 20 times. a lot of developer time spent but also a lot of tokens burn.

Speaker: 第二类问题是智能体会遗忘。智能体没有情景记忆(episodic memory)。每次会话都是一张白纸,而在此情况下人类就成了记忆。这就是你们工作的关系图(graph)实际的样子。底部是一个代码库关系图(repository graph)。你们组织生成的产物(artifacts),加上你们所依赖的每一个开源代码库。也许有一千个你们自己拥有的代码库,以及成千上万个开源代码库。在顶部,则是所有创建和修改这些代码的智能体会话(agentic sessions)。会话之间相互关联。代码库之间也相互关联。因此,这个关系图真实地反映了你们组织内的工作情况。它描述了底部有什么,以及它是如何发展到顶部的。这正是你希望智能体在这里能看到的东西。而它实际看到的只是一个会话、你们代码库的一小部分,且没有记忆。好吗?因为它看到的东西太少了,所以它只能依赖那个理解系统的人——也就是开发人员。每个开发人员在脑海中都有一部分这个关系图,对吧?至少在他们熟悉的领域是这样。而智能体通常来说并没有。如果这听起来还不够疯狂,对吧?想象一下一个智能体最多一次只能看到一个文件,而且只能往回看五条消息,又是在空间(能看到什么)和时间(能回溯多远)上的双重限制。你肯定会说,在这样的条件下根本不可能工作。但我们现在的状况,就跟那幅疯狂的画面差不多,而且组织越复杂,这个问题就越明显。

Original English

Speaker: The second category is that the agent forgets. The agent has no episodic memory. Every session is a blank slate and the human in this case becomes the memory. Here what the graph of your work actually looks like. At the bottom there is a repository graph. The artifacts your organization produces plus every open source repo you depend on. Maybe a thousand repos you own and tens of thousands of open source repos. At the top there are all agentic sessions that create and modify that code. Session relates to each other. Repos relate to each other. So this graph is a faithful picture of the work in your organization. It describes what there at the bottom and how it came to be at the top. That's what you want your agent to see here. What it actually sees is one session, one small fraction of your codebase, no memory. Okay? Because it sees so little, it leans on the one who understands the system, the developer. Every developer has a part of that graph, right? in their head at least in the domain they know. agent generally speaking doesn't if this doesn't sound crazy right imagine an agent that could see one file at a time maximum and can only look five messages back sort of constraint again both in space what can see and time how far in the path could see you would say that's impossible to work in what we have now is similar to that crazy picture and the more complex the organization is the more apparent it becomes

Polygraph: An Agent-Agnostic Meta-Harness

Speaker: 我来向大家展示我们是如何解决这个问题的。我交流过的其他组织也有类似的解决方案。所以,请从概念上看待问题和解决方案,而不是局限于某个具体工具。尽管我们这个工具相当酷,我们构建了一个名为 polygraph 的、与智能体无关的元工具平台(meta harness)。好的,让我给你们展示一下它是做什么的,以及它是如何解决我们刚才讨论的那些问题的。我们得出的第一个想法是,如果一个 GitHub 用户,或者任何用户,可以访问成千上万个代码库(其中一些是他们自己拥有的,很多则是开源的),我们可以对它们进行分析,并从中提取出大量元数据(metadata),以构建一个统一的依赖关系图(unified dependency graph)。那些代码库里没有任何一行代码会被修改,所有的工作都在侧面进行,对吧?然后,我们可以获取这些元数据,并将其输入到这个元工具平台中,营造出一种假象,让智能体以为只有一个庞大的代码库,并且可以随时随地进行读写。这是我个人的关系图。我自己拥有的代码库大概只有 300 个,对吧?而我的项目依赖于数千个开源代码库。Polygraph 会计算出每一个代码库生产了什么。每一个代码库、每一个代码库中的每一个项目,它们在包(package)的层面上消费了什么,它们提供和消费了什么 API,以及许多其他的信息,对吧?然后它把这些信息整合在一起,变成这样一个像是一个单一的庞大代码体,供你的智能体操作。

Original English

Speaker: I'll show you how we solved it. Other organizations I talk to have similar solutions. So, uh look at the problem and the solution conceptually, not a specific tool. Although the tool is pretty cool, we built uh an agent agnostic meta harness called polygraph. Okay, let me show you what it does and how it fixes the issues we just discussed. The first idea that we uh arrived at is that if a GitHub user, any user has access to thousands of repos, some of them they own, many of them are open source, we can analyze them and extract a lot of metadata out of them to build unified dependency graph. Uh no line of code changes in those repo that all happens kind of on the side, right? And then we can get this metadata and feed it to the meta hardness and create an illusion of one big code base the agent can read and write anywhere. This is my personal graph. I only have about 300 repos I own, right? And thousands of open source repos my projects depend on. Polygraph computes what each one produces. each repo, each project in each reper, what each project in each repo consumes package wise, what API they produce and consume, and lots of other stuff, right? And it teaches this together uh into this like one big body of code that your agent can work with.

Speaker: 那么,让我们来看看它的具体功能,对吧?它要做的第一件事,就是让你启动一个会话(session),把相关的代码库拉进来。对不对?是的。因此,它需要设置源代码、安装依赖项、为每个代码库配置一个智能体,然后把它们连接起来,使它们能够协同工作。此外,它还要提供一个干净、美观的终端用户界面(TUI),让你在进行复杂的更改时不会迷失方向。稍后我将向大家演示整个过程是如何运作的。这就相当于把信息拉取进来。把信息拉取进来只是故事的一部分,对吧?老实说,这是容易的部分。进行代码更改才更困难。如果你在一个会话中有 10 个代码库,这意味着你可能会有 10 个拉取请求(pull requests),对吧?你需要运行 CI,你需要协调所有的操作,对吧?你需要做所有这些事情,对吧?如果其中一个失败了怎么办,对吧?Polygraph 将所有的 CI 视为一个向量(vector)。就像我们看之前的例子,当我们在运行 UI 模块一和模块二的 CI 时,如果在一次 polygraph 会话中模块一失败了,它会计算出谁来修复这个问题。究竟是模块一需要补丁(patch),还是 UI 组件本身就有错、与模块一不兼容,在后一种情况下,每一方都需要打补丁,对吧?Polygraph 让你能把复杂的多代码库更改(multi-repo change)当成单一代码库的更改来处理。顺便说一下,同样的机制也解决了情景记忆的问题,因为我们捕获了你的工作。无论涉及多少个代码库,我们都知道你的意图、涉及的代码库、以及 PR。我们还捕获了智能体的所有运行轨迹。因为我们捕获了所有这些东西,我们就能将它们关联起来。所以现在我们可以说,你在一个代码库中的工作,连接到了另一个代码库中的工作,对吧?所有这些让我们能够在任何机器上恢复任何会话、任何工作片段,或者在任何地方引用它。稍后我还会再次向大家演示它的工作原理。你得到的是一个对整个组织拥有异常清晰、如同照片般记忆的智能体。它了解各个代码库是如何编写的、它们之间是如何关联的、它们是如何组装在一起的,并且它记住了基本上每个开发人员在每个代码库中的每一次会话,对吧?这就创造了一种完全不同的开发体验。

Original English

Speaker: So let's see what it does, right? The first thing it does is uh it lets you start a session to bring the relevant repositories in. Right? Right. So what it needs to do, it needs to uh set up the source code, install dependencies, set up an agent for each repo, wire them up so they can work together, and provide a clean, beautiful TUI to make non-trivial changes without getting lost. I will show you how it all works in a second. Right? So that's kind of pulling information in. Pulling information in is only one part of the story, right? Honestly, it's an easy part. Making changes is harder. If you have 10 repos in one session, it means you can have 10 pull requests, right? You need to run CI, you need to coordinate all of it, right? You need to do all this stuff, right? What if one of them fails, right? Polygraph treats all the CI as one vector. Like if we look at early example uh when we run CI for UI module one and module two if module one fails within a polygraph session it will figure out who fixes it whether module one need the patch or the UI component itself is wrong and incompatible with module one at which point everyone will need a patch right polygraph lets you treat complex multi-reo change as if it was a single repo change the same machinery by the way fixes episodic memory because we capture your work. No matter how many repos are involved, we know your intent, the repositories involved, PRs. We also capture all agent traces. Because we capture all of this stuff, we can relate it. So now we can say your work in one repo, connect to another work in another repo, right? And all of that lets us restore any session, any piece of work on any machine or reference it from anywhere. And I'll show you again how it works in a second. What you get is an agent with idic or photographic memory of your entire organization. It understands how repos are written, how they relate, how they put together and remembers every session from every repo by basically every developer, right? And that creates a completely different development experience.

Creating and Resuming Sessions in Polygraph

Speaker: 让我展示给你们看。首先,我们来看看如何创建一个会话。这是一些简单的操作。你运行一条命令,然后从列表中选择几个代码库。这里有一个微型的 GitHub 工作环境,只有三个代码库,因为这是一个演示(demo)。我选择了一个后端(back end)和一个前端(front end)。假设我需要做一个更改,你们知道,需要修改 API,并且需要同时更新 API 及其显示方式。我需要给我的会话起个名字。我需要从我已安装的智能体中选择一个。我选了 Claude,但任何已安装的智能体的工作方式都是相同的。记住,polygraph 不是一个智能体。它是包围在智能体周围的元工具平台,使它们更具能力。只需片刻,智能体就启动了。在这里,我可以像在单个代码库中一样与它进行交互,即便涉及了多个代码库,对吧?我可以给它发送指令。它将会规划出整个更改。最终在 TUI 界面中还会出现一些很酷的动画。它计算出了这两个代码库如何相互关联,以及具体的更改是什么。我可以要求它实施更改。我与它的交互方式,和我只在单一代码库里工作时完全一样。牵涉到多个代码库这一事实其实并不重要,对吧?唯一重要的地方在于,我会产生多个拉取请求(pull requests),对吧?但我也会得到一个 polygraph 会话。那些拉取请求具体是什么,对吧?如果我查看这个会话,我会看到我有一个描述,也就是该会话的描述。它在概念上绕过了代码库的边界来描述这项工作,比如会说我们需要更改这个代码库里的东西,同时更改那个代码库里的东西。它让我能清楚地看到涉及了哪些代码库、哪些拉取请求,以及这些代码库中的 CI,包括我需要知道的一切。这里的许多东西,基本上也就是我在单个代码库中会拥有的,只是数量更多而已,对吧?而且我还捕获了智能体的所有日志,这对于恢复进度非常重要,我马上就会给你们演示。

Original English

Speaker: Let me show you. First, let's look at how we create a session. Something simple. You run a command and you pick some repositories from a list. Here's a tiny GitHub work with only three repos because a demo. I pick a back end and a front end. Let's say I need to make a change that, you know, changes API and has to update both API and how stuff is being displayed. I need to give my session a name. I need to pick an agent from the ones I have installed. I picked Claude by any installed agent works the same way. Remember, polygraph isn't an agent. It's a meta harness around an agent that makes them uh more capable. And in a second, uh, the agent boots. And here I could interact with it as if I was in a single repo, even though multiple repos are involved, right? I could give it instructions. It's going to uh plan out the change. There's some cool animations in the TUI as well eventually. It figures out how the two repos relate and what the change is. I can ask it to implement a change. My interaction with this uh exactly same as if it I was working in a single repo. The fact that there are multiple repos involved is not really important, right? Uh the only uh uh part where it becomes important that I have multiple pull requests, right? Uh but I also get a polygraph session. What those pull requests are, right? If I look at the session, I will see I have a description uh that uh description of the session. It describes the work conceptually kind of bypassing the repo boundary saying we had to change stuff in this repo and change stuff in that repo. It gives me a good view of which repos are involved pull requests involved CI in those repos everything I need to know. A lot of the stuff is basically what I would have in a single repo but many right and I also have all the agent logs captured as well which is important for resuming which I'm going to show you in a second.

Speaker: 现在有趣的地方来了。我已经省去了一次重复解释的工作。我没有在前端代码库里重新解释一次后端的更改,对吧?我只解释了一次更改,它就在两个代码库中都得到了实施,并且完全保持一致。现在,让我们来恢复一个会话。假设我想让一位同事完成对后端的更改。也许那个后端代码库就是归他们管的。我把这个会话发送给他们。他们就能在自己的机器上恢复这个会话了。

Original English

Speaker: Now it gets interesting. I already saved one reexlanation. I didn't reexlain the back end change uh in a in a front end repo, right? I explained the change once and I got it implemented in both repos and it's all in agreement. Now let's resume a session. Say I want a coworker to finish the backend change. Perhaps they own the backend repo. I send them the session. They resume it on their machine.

跨设备与跨Agent的会话恢复

Victor: 对吧?所以在这里我给他们发送一个会话。他们可以运行这个命令。在不同的机器上,所有环境都不同。他们使用不同的终端,对吧?他们可以在自己的机器上重建它。他们没有这个会话,对吧?他们以前从未在这上面工作过。他们可以选择一个Agent。他们选择的Agent可能是一个不同的Agent,对吧?我在最初的会话中使用的是Claude。假设他们使用的是另一个,比如Codex。同样的设置会发生在他们的机器上。相同的代码库,相同的commit hash,一切都设置得完美无缺。Agent在每个代码库中启动,就像在我的机器上一样,对吧?它们再次全部连接起来。所以它们协同工作。它们都使用从我的机器捕获的追踪记录进行了初始化。所以他们机器上的后端代码库Agent拥有相同的hash和相同的历史记录。前端的代码库情况也是一样的。它在相同的、正确的hash处被检出,并且有一个运行着正确历史记录的Agent。所以我的Agent是Claude。他们的是Codex,但它们共享内存,而且它们实际上可以在这里进行修改,就像这个小视频中展示的那样。但重要的是,内存共享部分是关键,对吧?我可以工作,他们也可以工作,我们可以共享我们的记忆,尽管我们在不同的机器上使用两个不同的Agent。我的会话的完整状态基本上在他们的机器上实体化了。这在某种程度上不仅是内存,更多的是关于状态,对吧?附着在会话上的世界状态,你知道,正是这使得他们能够继续我的会话,即使他们最初对此一无所知。这很像《星际迷航》中的传送。就像我的会话的一个完整副本总是在他们的机器上实体化状态,这样他们就可以继续。这就是我在有人提交Pull Request让我审查并且我有疑问时经常采用的工作方式。我通常不会去问那个人,而是直接在我的机器上恢复他们的会话。我获得了他们确切的状态,完全正常运行,零配置,然后我就直接和我的Agent讨论我们做出的决定,对吧?因为所有这些决定都在捕获的追踪记录中。所以我的Agent确切地知道另一个人和他们的Agent讨论了什么,对吧?顺便说一句,当我想在会话中途因为某些故障从比如Claude切换到Codex时,这也是非常有用的。

Original English

Victor: Right? So this I'm sending them a session. They could run the command. different machine, different everything. They use different terminal, right? Uh they would reconstruct it on their machine. They don't have this session, right? They've never worked on it. They can pick an agent. Uh the agent they pick could be a different agent, right? I use code in the original session. Let's say they're using a different one, Cortex. The same setup happens on their machine. Same repos, same shaft, everything set up correctly. Agent starts in each repo like in mine, right? They all connected again. So they work together. They all primed with a trace captured from my machine. So the back end repo agent on their machine has the same sh and the same history. The front and the repo situation is the same. It's it's checked out at the same the correct SH has a agent running with the correct history. So my agent was clawed. They codex but they share memory and they could actually make changes in here as shown in a small video. Um but important the memory sharing part is key right uh I can work they can work and we can share our memories although we use two different agents of different machine the full state of my session kind of get materialized on their machine it kind of less memory and more about the state right the state of the world attached to the session uh you know is what enables them to continue my session even though they had didn't do anything with originally it's close to the transport in Star Trek like a whole copy of my session is always state materializes on their machine so they can continue and that's how I often work when there is a pull request for me to review and I have questions I usually don't ask the person I resume their session on my machine I get their exact state fully functional zero setup and then I just talk to my agent about the decisions we made right because all these decisions are in the traces capture so my agent knows exactly what the other person talked to their agent right side note This is also useful when I want to switch from say claw to codex mid session when something goes down.

Victor: 好的。拿我之前谈到的那个在生产环境中出现Bug的例子来说。在这里我要引用这个会话并且说它基本上坏了,你知道,你能找出问题所在并修复它吗。Agent会去查找它,会下载它需要的东西。如果描述——比如高层次的信息——足够了,那就太好了。如果不够,它将拉取相关的代码库、相关的hash、Agent日志,对吧?它将从原始会话中获取所有这些信息以重建该状态,以便它可以执行必要的修复,如这里所示。这里实际上提供了一个修复方案,对吧?我只需要说发生了这个情况。有一个Bug。就这样。不需要我提供任何额外的信息。

Original English

Victor: Okay. Take the earlier case I talked about where a bug land in production. Here I'm going to reference this session and say it's basically broken uh and you know can you figure out what's wrong and fix it. The agent will look it up will download what it needs. If description it's like high level information is enough that's great. If not, it's going to pull relevant repos, relevant chars, agent logs, right? It's going to get all this information from the original session to reconstruct that state such that it can do the necessary fixes as shown here. Here actually provided a fix, right? I only had to say this happened. There is a bug. That's it. No extra information was required for me to provide.

基于图的上下文与依赖解析

Victor: 好的。到目前为止,我们都是手动选择代码库和会话,但我们不必这样做,对吧?不用手动选择代码库,我也可以直接告诉Agent我想要什么。记住,那个图拥有所有这些智能,对吧,关于代码库之间是如何关联的。我可以告诉我的Agent,找到每一个依赖于特定版本库的代码库并更新它,对吧?而且它知道,对吧?我不需要去选择它们。它知道很多关于正在发生的事情的元数据。你也可以问一些宽泛的问题。比如,你知道,如果我想写一篇博客文章,对吧,或者一篇文章,我可以描述它,它就会根据代码库之间的关系以及它们里面的内容找出哪个代码库是最相关的。另一个例子,假设我想在PR集合中添加一个向量索引,并且我想知道是否有人在任何时候在任何代码库中做了相关的操作供我借鉴。所以在这种情况下,如果我这样做了,我将看到它会找到几个似乎相关的会话,我可以加载其中一个或两个,对吧?嗯,这在很多方面都很有用。仅仅举一个小例子,它有助于最佳实践和一致性。不用每次都从头开始做一些定制化的东西,我可以让它复制我所尊敬的工程师在某个会话中使用过的方法。现在我们跨代码库的代码是一致的,这是一件大事。当然,还有很多其他功能。如果你在一个代码库中,我可以要求,你知道,寻找会话,它会优先考虑与该代码库相关的会话,反之亦然。如果我在请求代码库,它会查看我的会话,看看类似的会话通常会引入什么。对吧?有很多有趣的智能特性,使得它比乍看起来要有用得多。

Original English

Victor: Okay. So far we have manually selected repos and sessions but we don't have to right instead of selecting repos by hand I can also tell the agent what I want remember that graph has all this intelligence right about how repos relate I could tell my agent find every repo that depends on a particular version of a library and update it right and it knows right I didn't have to select them it knows a lot of metadata about what's going on I You can also ask loose questions things like you know uh what if I want to write a blog post right or an article I could describe it and it will figure out which repo is the most relevant based on relationships between repos and what's in them. Another example let's say I want to add vector index into the PR collection and I want to know if anyone at any point did something relevant in any repo that I can draw from. So in this case if I do it I'll see that it will find several session that appear to be relevant and I can load one of them or both of them right um it's useful for many reasons just one small example it helps with best practices and consistency instead of doing stuff from scratch where you know every single bespoke I can make it replicate the approach used in a session by an engineer I respect now our code across repos is consistent and that's a big There is a lot more to it. Of course, if you are in a repo, I can ask, you know, for sessions, it will prioritize sessions that's relevant to the repo and vice versa. If I'm asking for repos, it will look at my session and see what similar sessions tend to bring in. Right? There's a lot of interesting intelligence that make it a lot more useful that appear at first glance.

Victor: 最后,呃,到目前为止一切,我我使用了呃呃,我所展示的一切,呃,都使用了Polygraph CLI,那种用于启动它的元工具CLI。然后你可以在其中启动Claude或者Codex或者其他什么,但你不必非得以这种方式使用它。所以在这种情况下,我已经在一个Claude会话中了,但它可以与任何东西协同工作。我可以直接说,嘿,你知道,我实际上认为一个单独的代码库会有用。比如也许我正在这个x代码库中开发一个Vest插件,我可以说,你能把Vest代码库添加到这个会话中吗,这样我就知道发生了什么。在这种情况下将调用Polygraph,我们会设置它,你知道,配置一切,并且我们会将Vest库,也就是Vest代码库,那个开源代码库,引入到我的会话中。所以现在,呃,我的Agent可以,你知道,去探索它。它可以,你知道,弄清楚它是如何工作的,也许还能解决我在我的代码库中遇到的一个问题。我非常喜欢这样,而不是说Context 7。因为如果我拥有真正的代码,Agent可以深入得非常深。所以深层次的问题可以通过这种方式被发现。

Original English

Victor: Okay. Lastly, uh everything so far I I used uh uh everything I shown uh use the polygraph CLI, the kind of meta harness CLI to start it and then you can start clo or cordex or whatever from within it but you don't have to use it this way. So in this case I'm already in a cloud session but works with anything and I could just say hey you know I actually think a separate repo would be useful like maybe I'm working on a vest plugin in this x repo and I could say can you add the vest uh repository to this session so I know what's going on in this case will engage polygraph and we'll set it up you know configure everything and we'll bring the vest library which is the vest repo the open source repo to my session. So now uh my agent can you know explore it. It could you know uh figure out how it works and maybe resolve an issue I have in my repo. I much prefer this to say context 7 because if I have the real code the agent can go really deep. So the deep problems are discoverable this way.

解除Agent的时空限制与蜂群思维

Victor: 好的。所以Agent受到空间和时间的限制。它们只看到代码库的一小部分,因为它们不知道过去。好吧。呃,这两个限制都可以被解除。Polygraph让Agent能够访问你组织能够触及的整个代码,包括你在开源社区拥有的代码。所以它不再受空间的限制。任何Agent都可以把它们全部带进来,对吧?并且它赋予了你的Agent关于发生过什么的完美记忆。每一个会话,做出的每一个决定都在触手可及的范围内,因为它跨越了开发者的边界。它不是每个开发者独有的。Agent可以拥有比任何单个开发者更多的上下文信息。比如一千名工程师拥有一个组织并创建了所有这些会话。它们对所有人都是可访问的,这几乎有点像博格人(Borg)。每个开发者运行的每个Agent都会为这个庞大的蜂群思维做出贡献,对吧?所以,呃,如果这很有趣的话,我叫Victor。你可以在Twitter上关注我。如果你想了解一下,去trypolygraph.com看看它是否适合你。谢谢。

Original English

Victor: All right. So agents are constrained in space and time. They only see a small fraction of the codebase as they don't know the past. Okay. Uh and both limits could be lifted. Polygraph uh gives agents access to the entire code your organization can reach the one you own in open source. So it's no longer constrained in space. Any agent can bring all of it, right? And it gives your agent a perfect memory of what happened. Every session, every decision made is within reach because it crosses developer boundary. It's not per developer. The agent can have more contacts than any single developer like a thousand engineers have an organization create all these sessions. They all accessible to to each of them almost like sort of the Borg. Every agent can run by every developer contributes to kind of one big this hive mind, right? So, uh if it's interesting, my name is Victor. You can follow me on Twitter. If you want to check it out, go to trypolygraph.com and see if it works for you. Thank you.

代理即日志

Ean: 大家好,我是Ean,Amnara的CEO,今天我要讲的主题是:日志即Agent。这次演讲的基本理念很简单,那就是大多数人认为Agent是模型或者它运行的执行环境。但我认为这是错误的抽象。我认为真正赋予Agent其身份的东西是它的日志。这就是我今天要论证的观点。

Original English

Ean: Hey everyone, I'm Ean the CEO of Amnara and today I'm going to be talking about the log is the agent. The basic idea of the talk is simple and that is most people think of an agent as the model or the execution environment that it's running in. And I think that that's the wrong abstraction. I think that the thing that actually gives an agent its identity is its log. And that's what I'm going to be arguing today.

Ean: 那么,想象一下你在最喜欢的电子游戏里花了100个小时玩的那个角色,比如这里是《天际》(Skyrim)。你的角色到底是什么?是游戏引擎吗?是PlayStation吗?是手柄吗?不,不是的。这些东西很重要,它们是我们用来互动的工具,它们会运行这个角色。但这些东西都不是你的角色。你的角色是数据。它是存档文件。这一点很重要,因为如果你的PlayStation突然起火了,你的角色并没有消失。你可以再买一台PlayStation。你可以从云端下载你的存档文件,然后你可以完全在他们离开的地方继续。那是因为Agent及其身份、历史和它的状态都完全捕获在它的数据中。角色存在于数据之中。这就是我今天想为Agent引入的框架。当人们谈论Agent时,他们通常指错了方向。他们会说Agent是模型,或者他们会说它是运行时。重申一下,正如我之前提到的,这些东西很重要,但它们不是Agent。Agent是它的数据。具体来说,它是日志。那么日志究竟是什么呢?在最简单的层面上,日志是Agent只支持追加的事件历史。它是每一次用户输入,每一次模型输出,每一个工具——

Original English

Ean: So, think about a character you've spent a hundred hours playing in your favorite video game, in this case Skyrim. What exactly is your character? Is it the game engine? Is it the PlayStation? Is it the controller? No, it's not. Those things matter and those things are what we'll interact with and they'll run the character. But none of those things are your character. Your character is data. It's the save file. And this is important because if your PlayStation bursts into flames, your character isn't gone. You can buy another PlayStation. You can download your save file from the cloud and you can resume exactly where they were. And that's because the agent and its identity and history and its state is all captured in its data. The character lives in the data. And this is the framing that I want to bring to agents today. When people talk about agents, they usually point at the wrong thing. They'll say that the agent is the model or they'll say that it's the runtime. And again, as I mentioned earlier, those things matter, but they're not the agent. The agent is its data. It's specifically the log. So what actually is the log? At the simplest level, the log is the appendon event history of the agent. It's every user input, every model output, every tool

Agent as Log & Loops

Speaker A: ……意味着 agent 的身份并不绑定于运行时、模型或工具。这些东西都只是在解释并追加到日志中。它们读取日志,对其采取行动,然后将下一个事件写回。这一点很重要,因为这样只需使用日志本身就足以恢复 agent 的运行。一旦你将 agent 定义为日志,

Original English

Speaker A: ... means that the identity of the agent isn't tied to the runtime or the model or the tools. Those things are all just interpreting and appending to the log. They're reading the log, acting on it, and writing the next event back. And that's important because then just using the log on its own is enough to resume the agent. Once you define the agent as the log, the

The Loop is the Product

Roland: 大家好。大家怎么样?准备好迎接更多循环(loops)了吗?是的。我叫 Roland。我的联合创始人与我曾在这个名为 XAI 的神秘地方努力开发 agent 基础设施,我们意识到,必须以一种独立的方式来完成一些新的东西。因此,我们在几个月前离开,去真正弄清楚如何部署这些永远在线、长周期的任务的下一个阶段。而且,我很高兴地宣布,我们有一些发现想向大家展示。这次演讲完全是关于,你应该如何将这些想法产品化,以便能够与你的客户一起扩展。大家已经听了很多关于自动研究(auto research)的内容。我们认为,对于 2026 年及以后的自动研究,有一个蓝图。这实际上归结为三个想法。让我们先看第一个。“循环即产品”(The loop is the product)。我们都熟悉这一点。我们从所有事情都归结为用于模型的强化学习(RL),以及你该如何训练模型以获得越来越好的推理能力开始。然后,我们迅速转向了“harnesses”,以及模型如何成为一种商品,而一切都取决于 harness。而现在我们在谈论 loops,以及你应该如何构建这些循环,并且不再接触代码。

Original English

Roland: Hello everyone. How's everyone doing? Are you guys ready for some more loops? Yeah. My name is Roland. My co-founder and I were in this mythical place called XAI working hard on agent infra and we realized there's something new that has to be done in a standalone way. So we left a few months ago to really figure out okay what's the next stage of how we should deploy these always on longunning horizon tasks. Um, and I'm happy to announce we have a few findings that we would like to present you. Um, and this talk it's all about um, how you should productize these ideas in ways that can scale with your customers. Um, you've heard a lot about auto research. Um, we think there's a blueprint for 2026 and beyond on how you should think about auto research. And it really comes down to three ideas. Let's go through the first one. The loop is the product. We're all familiar with this. We've started with everything goes down to RL chief for models and how you should train the model to become better and better reasoning. We then quickly moved to harnesses and how the model is a commodity and it's all about the harness. And now we're talking about loops and how you should build these loops uh and not touch code anymore.

Roland: 但这到底是什么意思,为什么大家都在这么说?你们还记得 Clawbot 吗?那是现在被称为 Open Claw 的最初名字。这个叫 AJ 的家伙围绕 Claw Bolt 构建了第一个循环。他所做的是找到一种方法与经销商以及 Reddit 用户交谈,以获得更大的汽车折扣。他遵循了这四个步骤。而真正做到这一点的其实是 OpenClaw。上 Reddit 找价格、找库存,跟经销商交谈,让经销商互相竞争,并试图弄清楚如何让他们互相抬价,找到一种可验证的方法来知道价格何时合适,然后锁定并买下车,它奏效了。这可能是在所有的 Mac mini 都脱销的时候,但这是“循环即产品”的第一个真实示例,而且在现在,这很可能应该成为一家初创公司。但是我们已经看到了这如何成为了每个人构建循环的诀窍。

Original English

Roland: But what does it really mean and why is everyone saying that? Do you guys remember Clawbot? That was the original um original name of what is now now known as Open Claw. And this guy AJ built the first loop around Claw Bolt. What he did was to find a way to talk to dealers and talk to Reddit users to get bigger discounts on a car. He followed these four steps. Um, and it's really OpenClaw the one that did it. Go on Reddit, find prices, find inventory, talk to the dealers, put dealers headto head and try to figure out how to make them out bid each other, have a verifiable way to know when the price is right, and then lock in, get the car, and it worked. Um, probably this was when all the Mac minis were uh selling off the shelves, but this was the first real example of loop is the product and something that probably should be a startup at this point. Um, but we've seen how this became a recipe for everyone to build loops.

OODA Loops & Verifiable Work

Roland: 但让我们退一步想,我们为什么会在这里?我们真的认为模型在训练时就考虑到了这种循环。它源于“OODA 循环”的概念。这是 20 世纪 70 年代由美国空军创造的一个术语,是关于这些喷气式战斗机如何在快节奏的环境中做出反应的想法。如果你考虑模型调用工具并获取观察结果,这就是我们作为人类、同时也是作为 agents 一直受到的训练。现在,当你在另一端放入强信号和可验证的工作时会发生什么?你得到了这些工作者(workers)或 cloud code agents。这里重要的是,信号的质量决定了循环的成功率,而验证器(verifier)的质量能够校准该成功是否真正正确。但这里还有另一个循环。当你获取它并将其反馈到信号中时会发生什么?这就是循环(looping around)的全部意义所在,即如何在这个初始循环的末尾生成这些工件,然后在上面运行第二个循环,从而有一种持续改进的方法。

Original English

Roland: But let's take a step back. Why are we here? Um, we really think models have been trained with this loop in mind. And it comes from this idea of uda loops. It's a terminology coined back in 1970s by the US Air Force and is the idea of these um jet fighters how to react in fast-paced environments. If you think of models calling tools and taking observations, it's it's what we've been trained on uh as humans but also as as agents. Now, now what happens when you put strong signals and verifiable work uh at the other ends? You get to these workers or cloud code agents. Um and and what matters here is the quality of the signal determines the uh success rate of the loop and the uh quality of the verifi verifier um um is able to calibrate if that success is actually correct or not. But there's another loop here. Um what happens when you take that and feed it back into the signal? And this is what looping around is all about is how do you generate these artifacts at the end of the first loop to then run a second loop on and have a way to continuously improve.

System Distillation

Roland: 这就引出了我的第二点。系统蒸馏(System distillation)才是真正的护城河,它实际上是理解在第一个循环中什么是好、什么是坏,并知道如何在第二个循环中处理这些信息的能力。那么我们该如何微调这些 AI 系统呢?每个循环都会围绕 harnesses、配置(profiles)、评估模型(eval models)、资源、工具以及环境生成有用的信息。你真正想要的是一种方法来保持它的可移植性,能够对其进行版本控制,并使其随时间推移而演进。如果你想一想研究中的数据配方(data recipes),这就是 RL 开始发挥良好作用的方式。你了解了配方,以及如何不断改变配方来对抗某些可能围绕幻觉和奖励黑客(reward hacking)发生的行为,然后你就得到了一个栈,也就是你的最终数据配方。对于 harnesses 我们没有这种东西。就一般意义上的 AI 系统而言,我们也没有。所以我们认为有空间可以做类似的事情。它包含了评估、微调、人类的判断力,以及所有这些在开始时并非预先确定的事情,它们是在你更多地了解你的 agent 在环境中的行为时定义出来的。

Original English

Roland: And this goes to my second point. System distillation is the mode and is really the ability to understand what went well and wrong in the first loop and know how to process that in the second one. So how do we tune these AI systems? Each loop generates useful information around harnesses, profiles, eval models, resources, tools, and the environment. What you really want is to have a way to keep this portable, to have a way to version this and to evolve it over time. If you think about data recipes in research, this is how RL started to work really well. you understood the recipes and how to continuously change the recipe to combat some of the behaviors that may happen around hallucinations around reward hacking and then you get to a stack which is your final data recipe. We don't have that for harnesses. We don't have that for like AI systems in the general term. So we thought there's space for something like that. something that contains the evils and contains the tweaks and the human judgment and all these things that are not predetermined at the beginning, but they're defined as you learn more about your agent acting in in in the environment.

Roland: 我们认为“配方(recipes)”一词可以应用于此,而且我们应该使用相同的名称。因此,一个“agent recipe”实际上是使你能够创建可复现的前沿 AI 系统的东西。它允许你拥有一条随着时间推移不断变得更好的护城河,它不绑定于任何平台或任何提供商。它是由你控制、驻留在你们公司中、且不受限于你使用的模型和提供商的东西。而 loops 应该专注于这一点。Loops 应该是你将这些系统提炼成配方的方式。失败模式应该成为评委(judges)和评估(evals)。重复的行为应该成为技能和提示(prompts)。用户的挫败感、扩展以及记忆应该加入到你的 harness,等等。我们都熟悉这一点,但我们并没有一个合适的术语来描述我们应该如何思考它,以及如何定义它。我们认为 recipes 是将所有东西放入 git 仓库中,并将其视为你构建这些自我改进系统的持续战略的一种方式。

Original English

Roland: We think recipes can be applied to this and we should use the same name. So an agent recipe is really something that enables you to create reproducible frontier AI systems. It's something that allows you to have a mode that keeps getting better over time, which is not tied to any platform or any provider. It's something that you control lives in your company and is agnostic to the models and providers you use. And loops should focus on this. Loops should be the way you distill these systems into recipes. Failure patterns should become judges and evals. Repeated behavior should become skills and prompts. user frustration, extensions and memories to your harness and so on. You we're all familiar with this, but we didn't have the the the right like terminology of how we should think about it and how we should define it. And we think recipes is a way to put everything together into a git repo and treat it as your ongoing um strategy for for uh building these self-improving systems.

Pi Recipes & Portable Agents

Roland: 于是,我们做了 Introspection(内省),但你可以把 introspection 看作是你生成这些配方的方式。所以它们是用于系统内省的配方。我们想要构建一些便携且独立于提供商的东西。因此,我们将对于 recipes 的方法建立在 Pi harness 和用于评估的 Harbor 之上。我们将其融入 git 仓库中,这样一切都可以进行版本控制,并且 agents 可以有一种方法来持续跟踪它是如何改变以及为什么改变的,它旨在归你所有,但由你的 agents 来管理。未来的产品确实应该这样来构建。它将所有者(owner)视为房间里几乎拥有最高品味的人格。但 agents 应该尝试调整自己以适应制造者的品味。因此,我们认为配方实际上应该将制造者的品味编码到你如何构建这些 agents 中。而如果我想使用其他人的配方,我也应该能够带来那种品味。这不仅仅是关于 harness,不仅仅是关于模型,而是你是如何得出这个特定的配方及其背后的原因?这差不多就是关于 agents 的可复现产品和服务背后的理念。

Original English

Roland: So we are introspection but you can think of introspection as the way you generate these recipes. So they're recipes for introspecting on your on your system. We wanted to build something that is portable and provider agnostic. So we built our um approach to recipes on the pi harness and on harbor for evals. We baked it into uh git repos so uh everything could be versioned and agents could have a way to continuously track how this change and why and is meant to be owned by you but managed by your agents. And this is how products should really be built going forward. It's something that treats the owner as the um almost like the the the higher taste um personality in the room. But agents should try to calibrate themselves to to the taste of the of the maker. So we think recipes should be basically encoding the taste of the makers into how you build these agents. And if I want to use someone else's recipe, I should be able to also bring that taste. It's not just the harness, it's not just the model, is how did you arrive at this particular recipe and why? And that's kind of like what uh what is behind uh reproducible um uh products and services around agents.

Roland: 我们有一个配方的早期发布版本,叫做 pi.recipes。它非常类似于 2025 年时的 skills,但更进了一步。这是关于“我需要什么才能拥有一个前沿的 agent”——所有关于我如何将品味编码到评估中?我该如何运行?我们该如何拥有 loops 以便随着时间推移不断改进这些评估?我们如何处理信号并知道使用哪些正确的信号?配合某些模型工作的正确工具是什么?我该如何拥有不同的 harness 配置文件来适应不同的模型?以及中间的一切。所以看看我们在这里一直在构建的东西吧。这还处于早期阶段,但希望能对你们有所帮助。我们觉得这将成长为一些能够真正让你使用不同制造者品味作为你的 agent 配方的东西。

Original English

Roland: Um, we have an early release of recipes is called pi. Recipes. It's very similar to what skills uh used to be in 2025 but is going a step forward. And this is what do I need to have a frontier agent is everything about how do I codify paste into evals? How do I run? How do we have the loops to continuously improve those evals over time? How do we process signals and know what are the right signals to to use? Um, what are the right tools to work with certain models? How do I have different profiles of the harness to work with different models? Um, and everything in between. So have a look at what we've been building here. It's still early uh but hopefully it's useful enough for you guys to to get going. And we feel this is going to grow into something that um really allows you to to use uh different um almost like different the to to be able to use the taste of of different makers as recipes for your agent.

Value per Watt

Roland: 最后一点是,每瓦特的有价值工作(valued work per watt)。为什么这是真正需要优化的指标呢?想一想 Cursor 和 Cognition 是如何从构建最好的产品,转向为该产品构建最好的评估,并最终……

Original English

Roland: And finally, the last point is valued work per watt. And why is this the score to really optimize for? Think of how um cursor and cognition went from building the best product to then building the best evvels for the product and finally

核心理念:以价值为衡量标准的模型构建

Speaker A:基于前两个工件构建最好的模型。我们认为这就像是未来一切发展的秘诀。代码是这种方法取得成功的第一个领域。除了客户支持、法律研究之外,所有事情最终都将归结为这个理念:我用什么代价换取了多少价值?第一步是如何衡量这种价值,第二步是如何知道我在这个价值上是否做了一笔划算的交易。

Original English

Speaker A: building the best models based on the previous two artifacts. We think this is like the recipe for everything going forward. Um code was the first domain where this um was successful. um everything beyond customer support, legal research, um everything is going to come down to this idea. How much value am I getting per what? Um how do I measure the value is the first step and how do I know I'm getting a good deal on that value is the second.

Speaker A:也许这样说会更清楚一些。我们都是从一个基础的测试框架(harness)和一组基础的评估(evals)开始,然后向着前沿(frontier)迈进。而且,你只有通过在生产环境中运行这些系统,才能经历这个过程。在开始之前,你根本无法知道前沿在哪里。但是,这里的最后一步——也是需要大量研究的一步——是,一旦你到达了前沿,我们如何让它在经济上可行?也就是说,我们如何做到在创造这些价值时,不花冤枉钱。并且,我们认为现在已经具备了让这一切变得容易获得且相当高效的基础模块;因为你们已经看到了所有这些微调 API,以及为了让你们执行这个过程而被抽象出来的所有基础设施。现在欠缺的只是相关的诀窍(knowhow)。而这正是我们希望能够推动的:弄清楚如何将“品味”(taste)编码到评估中,以及如何在实验中对其进行验证。

Original English

Speaker A: And maybe this makes it a bit more clear. We've all started from a base harness and a base set of evals and we went to go to the frontier. Um and you only go through that by running these systems in prod. There's no way you you know what Frontier is before you uh you start. Um but the the the last step here which is what is requiring a lot of research um is okay once you've reached frontier how do we make this um uh economically viable which is how do we not spend more than than uh we need for generating this amount of value. Um, and we think we have the building blocks now to make this accessible and pretty efficient in the sense of you've seen all these fine-tuning APIs, all the infrastructure that has been abstracted away for you to do do this process. It's just the knowhow that uh is not there yet. And this is what we we we hope we can like push for the knowhow for knowing how to codify taste into evals and how to validate that in experiments.

将“品味”编码到评估中

Speaker A:你们以前可能听过很多关于评估和实验的说法,但你们并没有真正思考过它们到底是什么——它们不仅仅是测试,它们实际上是创造者的“品味”,而智能体(agents)应该能够复现这种品味,并围绕它进行自我改进。以前没有人想过:我如何让它具有足够的便携性?作为一名艺术家或软件开发者,我如何将我的品味变成任何人都可以下载到大脑中的东西,并且能够成为我的一比一复制品?这有点像当前强化学习(RL)正在做的事情:我们如何将这些“品味塑造者”转化为围绕它们的环境和评估,从而能够将它们融入到模型权重(weights)中?但这还不够,你可以把“工作者(worker)”看作是内部循环,它生成所有的工件;但是你如何审视这些工件并知道该修改什么,这就是“品味”。这创造了你应该改变什么的候选方案,以及你应该如何在此基础上进行调整。而实验则是你自我校准的方式,用来确认“好的,我的品味在生产环境中确实得到了用户的验证”。我们要确保,不仅是制造者通过离线评估感到满意,终端用户也很满意,并且他们赞同我们所认为的“好”。

Original English

Speaker A: Um, and you you've you've heard a lot about evals and experiments before, but you didn't really think of them of like what are they is is just tests is is really what is the taste of the creator that agents should be able to reproduce and self-improve around. And no one has thought of how do I make this as portable enough? how how do I make my taste as an artist or as a software developer um something that anyone can download in their brain and be able to be a one-toone replica to me and this is kind of like what RL is is is about now is how do we uh turn these um taste makers into uh environments and evals around them so then we can move them into the weights but um there's more than that um you can think of the worker as the inner loop And it generates all these artifacts. But how you look at the artifacts and know what to change is the taste. Uh and this is what creates candidates of what you should change and how you should adapt based on that. And experiments is what how you self-calibrate that okay my taste is actually validated in production with users. And we make sure that not only the maker is happy through the um offline evals but the end users are happy as well and they agree with what we consider good.

实践案例:人才寻访智能体

Speaker A:让我们通过一个实际例子来看看它是如何运作的。我们以一个基准的智能体为例,比方说一个“人才寻访智能体”。这是一个非常典型的案例,因为每个人做招聘的方式都不同,而且很大程度上不在于什么是“好的招聘”,而在于主导招聘的人认为什么是“好的招聘”。所以在这种情况下,我们从非常简单的事情开始:一组工具(网络搜索、LinkedIn),一群已经通过诸如 CodeX 和 Claude Code 这样框架普及的子智能体(sub-agents),以及一段关于你的招聘人员的系统指令。

Original English

Speaker A: Let's go through a practical example of how this works. Let's take a baseline um agent which could be a talent sourcing agent. Um and this is a very classical case of everyone is doing recruiting differently and it's very much about not what is good recruiting but who is leading that recruiting that considers recruiting as good. So in this case we're starting with something pretty simple. Um a bunch of tools web search LinkedIn uh a bunch of sub aents that have been pre-popularized by harnesses like codeex and cloud code and a system instruction which is about your recruiter.

Speaker A:第一步是真正理解这些信号。你可以把“模式(patterns)”看作是检查追踪记录(traces)的一种方式,提取一些常见的行为或常见的用户挫折感,并将它们归纳成一个类别。比如说,智能体去接触大量的科技大厂员工。但作为一名招聘人员,你其实并不希望这样。你想要找到那些“蒙尘的明珠(hidden gems)”。你不会试图去挖角约翰·卡马克(John Carmack)。但是智能体会认为,“哦,约翰·卡马克太牛了。我为什么不联系他呢?” 所以,这是一种你可能永远都不会想到要硬编码去限制的行为,但你发现智能体倾向于这样做。模式就是你发现这些信号并指导你下一步该做什么的方法。

Original English

Speaker A: First step is really understand the signals. So you can think of patterns as being a way to look at the traces, extract some common um behaviors or common user frustrations and turn them into like a cluster. So let's say this idea of uh the agent is going uh and reaching out to a lot of big tech employees. As a recruiter, you don't really want that. You want to find hidden gems. You don't want to try to hire John Carmarmac. But an agent would think that's, oh, John Carmarmac is great. why would I not reach out to him? Um, so, so this is a behavior that you you'd never think of codifying, but you discover the agent tends to do that. Um, patterns is how you discover these signals and inform you what you should do next.

Speaker A:校准裁判(calibration judges)和评估(evals),是我们过去用来思考如何将这些行为量化的方式,试图使其能够在各种追踪和执行过程中应用相同的判断标准。所以,假设我们构建了一个智能体,它会审视一条执行轨迹,并准确识别出那种模式:“嘿,这个智能体是不是联系了谷歌的员工,而不是试图在 GitHub 上寻找隐藏的人才?” 其实,校准部分和评估生成部分并没有那么难,应该完全可以由智能体来构建。你只需要让人类参与到这个循环中来说:“嘿,这是我们正在采取的方法。你同意这个判断吗?你真的同意我们应该更多地寻找隐藏的人才,而不是去联系科技大厂的员工吗?” 就这么简单。你不需要人类真正去编写评估,你只需要他们来“校准”评估。而智能体应该成为真正将制造者的“品味”提取出来,并将其转化为代码的角色。

Original English

Speaker A: um calibration judges and evals is how we used to think about how do we qualify these these behaviors into um something that can try to uh apply the same judgment across traces and across uh execution. So let's say we we build an agent that looks at a trajectory and um identifies exactly that pattern. Hey, did did this agent reach out to Google employees instead of trying to uh find hidden gems on GitHub? Um, and the calibration bit and the eval generation bit is not that hard. It it it should be doable by agents to build. You just need a human in the loop to say, "Hey, um, this is the approach we're taking. Do you agree with this judgment? Do you really agree that we should look more towards hidden gems rather than reach out to um um big tech employees? And that's about it. You don't need the human to actually build the evals. You need them to calibrate the evals. And agents should be the ones that really take the the the taste of the maker and and put them in into code.

Speaker A:一旦你有了这些,创建不同的候选方案(recipe candidates)就很容易了。而这些差异(diffs)正是你想要品鉴的东西。你可以在此基础上建立一个非常好的离线评估集,但真正的考验在于当你走向生产环境时。终端用户是否同意你“不去联系大厂员工”的这种品味呢?这其实就是你想要的:你构建了一个能够充分彰显你品味的产品,然后确保你的用户欣赏并重视这种品味。A/B 测试一直都是确保情况符合预期的有效方法。所以,例如在多臂老虎机(multi-armed bandit)场景中,你就能很好地做到这一点。

Original English

Speaker A: Once you have this, it's pretty easy to create recipe candidates. And this should be the the diffs that you really want to taste. Um, and you can have a pretty good offline evil set around this, but the the the test here is when you go to prod. So, do the end user agree with your taste of not hitting up um big tech uh employees, right? And this is kind of like what you want is you build a product that really emphasizes your taste and then you you make sure that your users appreciate and value that taste. and AB tests have been a way to to to make sure that that's the case. Um so with a multi-arm banded um scenario for example you you'd be able to do that pretty well.

Speaker A:因此,一旦你验证了“好的,我有很棒的品味,而且我的用户也认为我有很棒的品味”,那就是你进行推广(promote)的时候了,也是你进入下一个版本的智能体方案的时候。这里的秘诀是,你一遍又一遍地重复这个过程,并且你知道如何不断地将你的品味——以及你所认为的“好”——编码到智能体中,让它能够为其他人复现同样的服务或产品,同时他们也认同你的品味和执行力都很棒。这实际上就是构建良好闭环的秘诀:是否有人能够以某种方式在我的系统上进行迭代。一个很好的例子就像是《穿普拉达的女王》(The Devil Wears Prada)里的米兰达(Miranda):在某些情况下米兰达会怎么做?你会想要将那种思维方式编码到智能体中,让它们能够在更高的层面上做同样的事情。

Original English

Speaker A: So once you validate okay I have great taste and my users believe uh I have great taste as well that's when you promote and that's kind of when you go to to the next version of an agent recipe. The secret is you keep doing this over and over again and you know how to continuously codify your taste and your um what what what good is to you into an agent that can reproduce the same service or product uh for other people and they also agree you have great taste and you have great execution. And this is really kind of like the the secret of building good loops is okay can can someone iterate on my um system in a way as uh you know um a good example here is like Miranda from the Delor product right what would be Miranda do uh in certain cases and you kind of want to codify that that thinking into like agents that can do the same stuff at a higher level.

核心总结

Speaker A:所以,核心总结如下。首先,闭环本身就是产品。你要试图将自己自动化为一个更高层级的裁判,并且确保你的第二层级(second loop)智能体能够对你试图推向生产环境的智能体施加同样的判断。第二点,系统蒸馏(system distillation)是壁垒所在。也就是说,你如何持续地将这种品味注入到这些工作者中,以及它们如何持续地进行自我验证和协同工作,这是你应该关注的最重要的事情;而且你做得越快,你就能越快地建立起一条成为垂直领域 AI 公司的防御性护城河。

Original English

Speaker A: So the takeaways are this. Um the loop is the product. You try to automate yourself as the u as a um higher level judge and you want to make sure your second loop agents are able to apply the same judgment to to the agents you're trying to to to push to prod. Second bit system dissolation is the mode. So, how do you continuously inject that taste into these uh workers and they how how they continuously self-verify and work together is uh the biggest thing that you should focus on and the faster you do it uh the the the faster you you build a defensible um approach to to becoming a vertical AI company.

Speaker A:最后,衡量你是否取得进展的标准应该是“单位成本所产生的有价值工作(valued work per what)”。因此,首先要确保你正在生成的工作是有价值的。其次,要确保经济模型是合理的,而这其中的价格差异,基本上就是人们从 Claude Code 转向使用你们所提供产品的原因。我们一直在深入思考这些理念,并且正在围绕“如何在生产环境中部署这些技术”开发一些非常有趣的产品。我们很乐意听取你们的想法。我们希望能够更多地了解某些垂直 SaaS 公司是如何规划将这些推向生产环境的,或者智能体实验室是如何考虑“围绕他们自己的产品创建这些自动研究实验室”这一理念的。请与我们保持联系。我们会在附近转转,欢迎找我们多聊聊。非常感谢大家!

Original English

Speaker A: And finally, valued work per what is how you should measure um am I making progress or not. So first make sure that uh the the the work you're generating is valuable. Second make sure that the economics makes sense and the um the the difference in price is is basically what um people would would switch away from cloud code to to something you provide. We've been thinking a lot about these ideas and we're building some very interesting products around how to deploy this in production. We'd love to hear from you. would love to get um to to understand more about how how certain um vertical SAS companies are are looking to go to prod with um or how agent labs have been thinking about this idea of um um creating these like auto research labs around their their own products. Um get in touch. We're going to be around the block for for chatting more about this and thank you very much.

机器工厂自学记忆的故事

Rushab:给你们讲一个关于工厂如何教会自己记忆的故事。大家好,我是 Rushab。我在印度经营着一家名为 Machine Craft 的 100 人规模的工厂。我们没有数据科学团队,没有机器学习预算,什么都没有。但不知何故……

Original English

Rushab: tell you a story about a factory that taught itself how to remember. Hi, I'm Rushab. I run machine craft, a 100 people factory in India. No data science team, no ML budget, none of that. And somehow

构建 36 个 AI 代理组成的商业大脑

Speaker A: 最终,我们构建了一个由 36 个 AI 代理组成的系统,它负责运行我们整个进入市场的流程。我觉得这听起来仍然有些不可思议。让我向你们展示这一切是如何发生的,以及为什么你们也能做到同样的事情。所以,关于我们公司,情况是这样的。从表面上看,它似乎全是机器和金属,但公司真正的核心,真正重要的部分并不是机器,而是知识。客户是谁,我们在 2019 年给他们的报价是多少,为什么那台特定的机器需要那种奇怪的定制调整。在整整三代人的时间里,所有这些知识仅仅存在于三个人的大脑中。最初是我祖父的,然后是我父亲的,现在是我的,当你仔细思考这个问题时,这实际上是一种非常令人恐惧的公司运营方式。许多人加入了我们。也有人离开了我们。这扇旋转门从未停止过。而每一次有人离开,我们公司大脑的一部分也就随之离开了。我们并不害怕竞争对手。我们害怕的是遗忘,或者某天醒来时发现,整个公司竟然只存在于两个越来越疲惫的脑袋里。所以,我产生了一个想法。我得说实话,这听起来一开始很疯狂。如果不把这些知识写在没人会读的文件里,而是我们自己“培育”一个能记住这一切的大脑会怎样?不是那种简单的聊天机器人。而是你可以与之交互的,公司的数字孪生体。我没有去雇佣一个销售团队,而是尝试自己构建一个。

Original English

Speaker A: we ended up building a 36 AI agent that runs our entire go to market. I think that's still a little ridiculous. Let me show you how it happened and why you can do the same thing. So here's the thing about our company. From the outside, it looks like machines and metal, but the actual company, the part that matters isn't the machines, is the knowledge. Who the customer is what we quoted them in 2019, why that one machine needed that weird custom tweak. And for three generations, all of that lived in exactly three brains. Initially my grandfather's, then my father's, and now mine, which is a genuinely terrifying way to run a company when you sit with it. A lot of people have joined us. People have left us. The revolving door never stopped. And every single time someone walked out, a chunk of our brain walked out with them. We weren't scared of the competitors. We were scared of forgetting or waking up one day and realizing the whole company only existed inside two increasingly tired heads. So, I had an idea. I'll be honest. Sounded insane first. What if instead of writing the knowledge down in some document nobody ever reads, what if we grew a brain that just held it? Not a chatbot. You poke at a twin of the company. I didn't hire a sales team. I tried to build one.

Speaker A: 我们稍微偏题一下,因为你需要了解这其中的情况有多么复杂。我们制造热成型机。这种机器加热塑料板并将其塑形。核心机器是一样的,但最终它可能制造出水培农场托盘、水疗浴缸、电动汽车车身面板、医疗设备外壳,甚至是包装材料。七个完全不同的领域,七种完全不同的买家。所以,这个大脑不能只是死记硬背一本宣传册。它必须知道某个特定的客户到底生活在哪个“宇宙”里。第一步简单得几乎有些无聊。把所有的东西都喂给它。我的意思是所有的东西。多年的报价单、图纸、付款计划、时间表、电子邮件记录,我们自己私有历史中的数百千兆字节数据。不是公共互联网的数据,而是我们自己的内部网络数据。而这里的剧情反转,也就是让每一个听我讲过这个故事的工程师都感到惊讶的部分在于,我们从来没有训练过任何模型。地下室里没有嗡嗡作响的 GPU,也没有进行微调。我们只是审视了所有的历史数据,将它们切分成易于消化的小块,然后让现成的模型去阅读并提取事实。我们将每个小块的含义存储为向量和关系。谁与什么相关联,构成了一个图谱。这个大脑存在于一个更聪明的模型中。它实际上是一个组织得非常、非常好的记忆库。

Original English

Speaker A: A quick detour because you need to know how messy this is. We make thermopforming machines. They heat up a plastic sheet and shape it. Same core machine, but it ends up making hydroponic farm trays, spa bathtubs, EV car panels, medical casings, and even packaging. Seven totally different worlds, seven totally different buyers. So, this brain couldn't just memorize a brochure. It had to know which universe a given customer lives in. Step one was almost boringly simple. Feed it everything. And I mean everything. years of quotes, drawings, payment schedules, timelines, email threads, hundreds of gigabytes of our own private history. Not the public internet, our internet. And here's the plot twist, the part that surprises every engineer I tell this to. We never trained a model. No GPUs humming in the basement, no fine-tuning. We just looked at all the history, chopped it into bite-sized chunks, and let offshelf models, read it, and pull out the facts. We stored the meaning of each chunk as vectors and relationships. Who's connected to what as a graph? The brain is in a smarter model. It's actually a really, really well organized memory.

Speaker A: 现在,事情开始变得有点奇妙,不过是朝着好的方向。我们不再把 Era 当作一个软件来看待,而是开始把它当成我们正在培养的某种东西。所以我们给它赋予了一个仿生学的身体;通过感官来弄清楚它在和谁说话,通过“肠胃”将文件消化成事实,它还有记忆、梦境循环,以及一个用来抵御不良信息的免疫系统。为什么要模仿生物学?嗯,因为进化已经花了十亿年的时间来解决这个问题。你如何随着时间的推移保持连贯性?我们只是直接抄了作业。那么,最大的问题来了,为什么要用 36 个代理,而不是一个超级天才般的提示词呢?因为——如果你曾经尝试过,你早就知道了——一个被期望能做所有事情的提示词,最终会把所有事情都做得很糟。所以,它并不是一个单一的头脑。它是一个万神殿。一整个由专家组成的团队。每个人都有着极其明确的单一工作。Athena 掌控全场。Prometheus 负责销售。Plutus 处理定价。Hippastus 了解每一个机器规格。Vera 负责核实一切。而 Memon,我最喜欢的那个,负责守护修正过的内容。所以,一旦人类修正了某个错误,它就永远被修正了。一个代理,一份工作。这是一个团队,而不是一个孤胆英雄。这里最酷的部分是,他们会开会。Athena 会召集这些专家。他们实际上会进行争论,最终在另一端得出唯一的答案。这就像拥有一个永远不睡觉、永远不疲倦、而且不知怎么还没有任何自我中心的董事会。

Original English

Speaker A: Now, this is where it gets a little weird in a good way. We stopped thinking of era as a software and started thinking of it as something we were raising. So we gave it a body modeled on biology senses to figure out who it's talking to, a gut to digest the documents into facts, a memory, a dream cycle, an immune system to fight off bad information. Why biology? Well, because evolution already spent a billion years solving. How do you stay coherent over time? We just copied the homework. Okay, so the big question, why 36 agents instead of one genius mega prompt? Because, and you already know this if you've ever tried it, one prompt that's supposed to do everything ends up doing everything badly. So, a isn't one mind. It's a pantheon. A whole cast of specialists. Each one has exactly one job. Athena runs the room. Prometheus owns the sale. Plutus does pricing. Hippastus knows every machine spec. Vera fact checks everything. And Memon, my favorite, guards corrections. So the second a human fixes something, it stays fixed forever. One agent, one job. It's a team, not a hero. And here's the cool part. They hold meetings. Athena pulls in specialists. They actually argue and a single answer comes out the other side. It's like having a boardroom that never sleeps, never gets tired, and somehow has no ego.

Speaker A: 那么,所有这些实际上运行着什么呢?老实说,是整个前端业务,也就是一个陌生人在某个地方存在到他们成为我们的客户之间的所有环节。每天有九项具体的工作。外发包含我们现实世界真实情况的电子邮件。在通话前根据交叉核实的事实构建的客户简报。报价。用于外联的左滑/右滑模式。唤醒死去的线索,我称之为过去的冲击。处理呼入的回复,并在我们浪费一个小时之前弄清楚一家公司是否合适。九项工作,一个永不入眠的接线员。所有这些存在于哪里?存在于一个光标标签页中。这就是全部了。你打字,系统伸出十几只手,搜索知识库,阅读收件箱,起草电子邮件,构建代码,然后在大任何东西实际发出之前展示给你看。在引擎盖下,这确实是一个真正的技术栈,而不是用胶带粘在一起的演示原型。用于向量的数据库,用于 CRM 的关系图谱。三家不同的模型提供商,每一个都是因为最适合某项工作而被挑选出来的,用于谷歌的工具,用于吞咽文件的工具,用于每一个沟通渠道的工具,加上监控以便我们能看到它的思考过程,所有的这一切,构成了 Fabric。

Original English

Speaker A: So, what does all this actually run? Honestly, the whole front business, everything between a stranger exists somewhere and now they're a customer. Nine concrete jobs every single day. Outbound emails that actually reference my real world. Account briefs built from cross-cheed truths before a call. Quotations. A swipe left, swipe right mode for outreach. Reviving dead leads, which I call blast from the blast. Inbound replies and figuring out before we waste an hour whether a company is even a fit. Nine jobs, one operator who never sleeps. Where does all this live? One cursor tab. That's genuinely it. You type and a reaches out with a dozen hands, searches the knowledge base, reads the inbox, drafts the email, builds the code, and then shows you before anything actually goes out. Under the hood is genuinely a real stack, not a demo held together with the tape. databases for vectors for relationship graph for the CRM. Three different model providers each picked for the job it's actually best for tools for Google for swallowing documents for every communication channel plus monitoring so we can see what it's thinking all of it Fabric.

多智能体 AI 村庄中的自动研究

Arena: 好的。大家好。我是 Arena,前微软和 Supercell 的工程师。今天我想谈谈在多智能体 AI 村庄中的自动研究(auto research)。我将使用一个类似电子游戏的 AI 村庄作为贯穿这里的例子,但我认为更广泛的问题是许多 AI 工程师正在开始遇到的。我们如何评估和改进那些在长时间内携带状态的智能体?在深入探讨自动研究层之前,我想先谈谈 Paradox 项目。我们在 Supercell 的 AI 创新实验室开发了 Paradox 项目。我和我的队友 Arnach Manikanden。我们构建了一个模块化的 AI 框架,允许任何开发者在一个电子游戏中插入智能自主的代理,这些代理可以与其他玩家或代理互动、竞争或合作,并使它们成为动态的游戏伴侣。

Original English

Arena: Okay. Hi everyone. I'm Arena, former engineer at Microsoft and Supercell. And today I want to talk about auto research in a multi- aent AI village. I will use a video game like AI Village as a running example here, but the broader question is one I think many AI engineers are starting to run into. How do we evaluate and improve agents that carry state over a long period of time? Before I get into the auto research layer, I want to talk a bit about project paradox. We developed project paradox at supercell's AI innovation lab. Me and my teammate Arnach Manikanden. We built a modular AI framework that allows any developer to plug in intelligent autonomous agents within a video game that can interact, compete or cooperate with other players or agents as well and place them uh and make them into dynamic game companions.

Arena: 现在,为了给出这些代理能做什么的例子,代理可以带着意图移动。它们可以去任何地点或找任何人,而且它们受到自身记忆、情感或好奇心的引导。这些代理可以与世界互动。它们可以拾起物体,把它们扔在任何地方,而且它们也了解自己环境中的上下文,比如物体或其他的角色及代理。我还想指出的是,游戏开发者也可以添加新的动作让这些代理在我们的框架内完成。代理不仅可以丢弃或放置物体,显然也能对周围发生的事情做出反应。而周围发生的这些事件也会实时影响它们自身的信念和情感。当然,如果代理无法开启对话,那就不完整了,对吧?在这种情景下,代理可以接近其他代理甚至玩家,这使得游戏感觉更加栩栩如生。当然,这些对话被储存在它们的记忆中,并符合它们自身的情况……并且会影响它们自己的情感、信念或目标。综合起来,这些代理构成了我们的多代理框架。

Original English

Arena: Now, to give examples of what these agents can do, the agents can move with intent. They can go to any location or person, and they're guided by their own memories, emotion, or curiosity. These agents can interact with the world. They can pick up objects, drop them anywhere, and they're also aware about the context in their own environment, such as objects or other characters or agents as well. I would also like to note that game developers can also add new actions for these agents to accomplish within our framework as well. Instead of just dropping or uh placing objects, agents can also obviously react to what's happening around them. And these events that happen around them affect their own beliefs and emotions on the fly as well. And of course, it wouldn't be complete if agents can't start conversations, right? agents can in this scenario approach other agents or even the player as well and this makes the game feel more alive. And of course these conversations are stored within their memory and is according to their own um and affect their own emotions and beliefs or goals as well. And al together these agents make our multi-agentic framework.

Arena: 嗯,是的,等一下。所以其背后的架构是刻意设计为有状态的。第一个重要部分是每个代理的记忆。每个代理都有自己的记忆命名空间,由 RAG(检索增强生成)提供支持。所以记忆不会在代理之间相互渗透。第二,我们将情感作为一个小向量进行追踪。所以在某个事件或对话之后,系统可以更新诸如快乐、悲伤、恐惧、愤怒或厌恶等数值。第三,代理对其他代理和玩家有信念分数。你可以把这看作一个信任矩阵,基本上在互动发生后,大语言模型(LLM)从根本上决定信任分数是应该上升、下降,还是保持不变。第四,每一条记忆都会收到一个重要性分数。

Original English

Arena: Um yeah yeah one second so the architecture was intentionally stateful behind this. The first important part was per agent memory. Each agent has its own memory namespace backed by rag. So memory did not bleed between agents. Second, we tracked emotion as a small vector. So after an event or conversation, the system could update values like joy, sadness, fear, anger, or disgust. Third, agents had belief scores towards other agents and the player. You can think of this as a trust matrix basically like after the interaction happens the LM basically decides whether the trust score should go up down or whether it shouldn't change at all. And fourth, every memory receives an important score.

Arena: 为了更好地解释这一点,比如说你几天前吃了一顿晚餐,你可能不记得晚餐吃了什么,对吧?但是如果几天前有人被谋杀了,你肯定会记得。所以代理将会评估——或者说大语言模型会评估——一个事件的重要性分数,如果它越过了一个阈值,它就会将那段特定的记忆存储在一个单独的缓存中,以便以后能更好地检索到重要的上下文。这里有一个它正好在运行的例子。我们要去邀请其中一个角色和我们一起去野餐。在这里,我们的角色 Blossom 决定拿起一个糕点去野餐区,因为我们要求她这么做。请记住在后台的对话过程中,她规划了所有的……

Original English

Arena: Um to to explain this better, like let's say you had dinner a few days ago, you probably wouldn't remember what you had for dinner, right? But um if someone was murdered a few days ago, you definitely remember that. So the agent will evaluate or the LM will evaluate uh an important score of an event and if it crosses a threshold, it will store that specific memory uh in a separate cache so that important context can be retrieved better later on. And here's an example of it just working. Um, we going to ask one of the characters to go on a picnic with us. Here, uh, our character Blossom um, decides to pick up a pastry and go to the picnic area because we asked her to do so. Keep in mind during the conversation in the background, she plans all

长时间跨度下的智能体社会一致性问题

Speaker: 要完成这些连续的动作。之后当我们跟她对话时,她也能结合语境进行回复。是的。但这也是一个有趣的问题真正开始的地方。正如你们在上一个例子中看到的,对于短期的游戏体验,我们的架构运行得相当不错。比如一个角色可以制定计划,四处走动,交谈并记住最近的互动,也能回应我们或其他角色。但在更长的时间跨度下,我们注意到这种社会一致性开始变弱。所以在这个例子中,我们让一个智能体把芒果打折的传闻散播给另一个智能体,那个智能体接收到这个信息,然后去告诉了另一个智能体。在中间发生了一系列事件之后,当玩家向其中一个智能体询问关于芒果的事时,它并没有完全像我们预期的那样存储那个语境,或者没有提供我们想要的那种语境。这就是事情自然开始变得混乱的地方。比如系统可能会记住大致的话题,但丢掉了话题的来源。一个传闻可能会变成一种担忧,而不仅仅是传闻,智能体可能会把它当作事实陈述;或者一个智能体可能知道某个事实,但在为自己的行动制定计划时未能执行,未能想起它。

Original English

Speaker: of these sequences of actions to accomplish. And when we talk to her afterwards, she will also reply within context as well. Yeah. But this is where an interesting problem actually started. As you saw in the last example, like for shortterm game play, this our architecture worked pretty well. like a character could make a plan, move around, talk and remember the recent interaction and respond to us or other characters as well. But over longer horizons, this is where we notice the social consistency start to get weaker. So in this example, we have one agent spreading a rumor about a sale on mangoes to another agent and that agent receives that information and goes and tells another agent about it. Later on, after a number of events occurred in between, when the player asks one of the agents about the mangoes, it doesn't exactly store that context that we were expecting or it doesn't give us the context that we kind of wanted to. And this is where things are starting to get messy naturally. Like the system may remember the rough topic but lose the source of the topic. A rumor may be concern instead of just a rumor like the agent might state it as a fact or um an agent might know a fact but fail to execute fail to remember it while creating a plan for its actions.

将 Auto Research 引入 Project Paradox

Speaker: 所以这里的问题变成了,我们如何在长期运行的社交行为中,而不仅仅是在一次响应中,提升多智能体系统。这就是我们想要引入 Auto Research(自动化研究)的地方。正如大家所知,几个月前,Karpathy 发布了关于 Auto Research 的内容,这立刻让我们非常好奇。也许我们可以让系统在自身上运行实验,我们能把这个用于我们的系统吗?所以我们的理解是,与其手动微调提示词或观看一个漂亮的演示,我们不如定义一套场景集,运行智能体,收集轨迹数据,对行为进行评分,然后改变一小部分策略接口,并只保留那些实际提高了分数的改动。这就是我们试图将 Project Paradox 与 Auto Research 结合的地方。所以在这一步,基本上我们的多智能体框架 Project Paradox 更像是一个实验台,而 Auto Research 则成为围绕它的实验循环。重要的是,这不仅仅是为了优化 RAG 检索。更广泛的框架是优化智能体协议,比如智能体如何写入记忆、检索记忆、表达不确定性、更新信任、进行来源归因,以及根据新事实重新规划。

Original English

Speaker: So the question here became how do we improve a multi-agentic system over longrunning social behavior and not just over one response. And this is where we wanted to bring in auto research. As you all know, a few months ago, Karpathi posted out auto research and this this made us immediately very curious. Uh perhaps we can make the system run experiments uh on itself and can we use this for our system as well. So what we understood is instead of manually tuning a prompt or watching one nice demo, we could define a a scenario suit, run the agents, collect traces, score the behavior and change a small policy surface and only keep the changes that actually improve the score. And this is where we're trying to bridge project paradox with auto research. So at this point basically our multi-agentic framework project paradox is more like a lab bench and auto research becomes the experimental loop around it. And importantly this is not only about improving rag retrieval. The broader framing is optimizing the agent protocol like how do agents write memories, retrieve them, communicate uncertainty, update trust attribute sources and replan around new facts.

评估社会层面的行为改善

Speaker: 基本上,是的,在这个语境下……在这个语境下,Auto Research 不是村里的另一个智能体。正如我所说,它是村庄外的一个元系统。村民们只有局部的视角,当然,他们只知道自己看到、听到、记住或推断的事情,因为他们之间没有共享的记忆数据库。只有当其他智能体正确传达信息时,信息才会传递。Auto Research 层在这里有着不同的职责。它读取一次运行的完整轨迹,将发生的情况与场景的事实基准进行对比,对行为进行评分,并提出对智能体协议或认知策略的一个受约束的改动。然后它重新运行该场景,并询问社会层面的行为是否变得更好了。这是我们试图寻找的关键转变。所以我们不再只是评估单个回答。我们是在评估整个运行过程。

Original English

Speaker: basically um yeah in this context uh oh yeah in this context art research is a not another agent in the village like I said it's a meta system outside the village the villagers have local perspectives of course they only know what they saw heard remembered or inferred because there isn't a common memory database in between them. Information only travels once uh other agents communicate them properly. The auto research layer has a different job here. It reads the full traces of a run, compares what happens against the scenario ground truth, uh scores the behavior and proposes a constrained change to the agent protocol or cognitive policy. Then it reruns the scenario and asks society level behavior like did society level behavior get better. This is the key shift we were trying to look for. So we were no longer evaluating one answer. We were evaluating an entire run.

Auto Research 实验循环机制

Speaker: 这是一个循环可能的样子。首先,我们定义一个控制场景,稍后我会详细说明。例如,一个智能体了解到一个公开事实,或者一个智能体听到了一个传闻。那可以是一个控制场景。然后我们运行模拟。在运行期间,我们收集结构化的轨迹:观察、对话、记忆写入、检索、信念更新等,只要在那种情况下对我们相关的,我们都会收集。接着,我们对这种行为进行评分。信息是否如我们预期的那样传播了?来源归因是否保留下来了?比如,智能体是否记得是谁开始的传闻?不确定的事情是否保持了不确定性?智能体是否根据它们确切知道的情况采取了行动?然后这里的 Auto Research 层会提出一个微小的策略更改。这一点很重要。它不应该重写整个应用程序。当然,它应该只修改一个受控的策略层面。然后我们重新运行。如果分数提高了并且护栏起作用,我们就保留这个改进。如果没有,我们就直接回滚。

Original English

Speaker: And this is what one of the loops would look like. Like first we define a control scenario which I'll elaborate a bit more about later. For example, one agent learns a public fact or one agent hears a rumor. Uh that could be a controlled scenario. Then we run the simulation. During the run, we collect structured traces, observations, conversations, memory rights, retrievalss, belief updates, whatever is relevant to us in that case, we collect. Then we score this behavior. Did the information spread as we expected it to? Did the source attribution survive? Such as, does the agent remember who started the rumor? Did uncertainty stay uncertain? Did agents act on what they actually knew? And then the auto research layer here proposes a small policy change. And this is important. It should not rewrite the whole application. Of course, it should only edit a controlled policy surface. And then we rerun. If the score improves and the guard rails hold, we we keep the improvement. And if not, we simply just revert back.

设计控制场景的重要性

Speaker: 谈到控制场景,之所以场景设计很重要,是因为如果不加控制,社会行为通常有点模糊。意思是,如果你只是让我们的环境中的智能体四处游荡,它看起来可能很酷,你可能也会得到不错的互动,但要评估系统是否真正得到了改进,实际上是非常困难的。这就是为什么我们认为你需要控制场景。比如,一个场景可以测试公开事实的扩散。假设智能体 A 了解到了面包店明天要关门。正确的智能体都了解到这件事了吗?他们记得是谁说的吗?他们会根据这个事实改变自己的计划吗?另一个场景可以测试传闻的不确定性。假设智能体 A 听说智能体 C 可能要离开村子。当这个传闻传播时,“可能离开”是突然变成了“正在离开”,还是仍然保持为“可能离开”?它是成了一个事实,还是依然作为一个传闻?另一个场景可以测试重新规划。团队原本有个计划,但有个智能体了解到,比方说,他们想走的路线被堵死了。智能体是否会更新这个信息,并互相传达以避免不恰当的计划或过时的行动?这里的重点并不是说这些具体的场景是通用的。我们想表达的观点是,长跨度的智能体行为需要一整套场景集。

Original English

Speaker: And talking about controlled scenarios, the reason why uh scenario design matters is that social behavior is otherwise a bit fuzzy uh in general in the sense if you just let the agents in our environment wander around, it might look cool and you might get nice interactions, but it's actually very hard to evaluate on whether the system actually improved. So this is why we believe you need controlled scenarios. For example, one scenario could test a public fact diffusion. Let's say agent A learns uh the bakery will close tomorrow. Do the right agents learn it? Do they remember who said what? Do they do they change their plans based on this fact? Another scenario could test rumor uncertainty. Let's say agent A hears that agent C might leave the village. When this rumor spreads, does might leave suddenly become is leaving or does it stay as might leave? Like does it become a fact or does it still stay as a a rumor? Another scenario could test replanning. The group has a plan but one agent learns let's say the route they wanted to take is blocked. Do agents update this and communicate this uh with each other to avoid uh a improper plan or scale actions. The point is not that these exact scenarios are universal here. The point we're trying to make is that long horizon agent behavior needs scenario suits.

建立平衡的评分卡

Speaker: 再谈回我们关于芒果的例子,在运行了我们的一个 Auto Research 循环之后,这一次在经过了很长一段时间后,当玩家最终向其中一个智能体询问关于芒果打折的事时,我们确实发现,与上次相比,这次智能体能够结合语境进行回应了。为了这次演讲,我们认为确切的公式不如评分卡的形态重要。你不想要一个单一且模糊的指标,比如“智能体质量”。这掩盖了所有有意思的失败。相反,你需要一个平衡的评分卡。对于信息扩散,你可能会衡量触达率,比如在 n 步之后有多少个智能体知道了这个事实。对于信息出处,你衡量那些知道该信息的智能体对来源的保留度。有多少人记得它是从哪来的等等。对于传闻,你可以衡量不确定性的保留度以及错误确定率。对于规划,你可以衡量行动一致性以及重新规划的时间。而对于隐私,你可以衡量信息的封锁情况。这很重要,因为只优化一个指标会产生糟糕的行为。比方说,如果你只针对信息扩散进行优化,智能体可能会学会过度分享一切。而如果你只针对记忆召回进行优化,你可能会产生嘈杂的或过时的记忆。所以这种评分卡正是让系统保持诚实,并防止 Auto Research 智能体为了单纯提高某一个分数而去利用系统漏洞的原因。

Original English

Speaker: And talking about our Mango example again, after running one of our auto research loops, this time after uh a a long pro period of time, when the player finally asked one of the agents about the sale on mangoes, we did find that u the the agent was able to respond within context this time like compared to last time. Um yeah and for this talk the form the exact formula we believe is less important than the shape of the scorecard. Uh you do not want a single vague met metric like agent quality. This will hide all the interesting failures. Instead you want a balanced scorecard. For diffusion, you might measure reach like how many agents know the fact after end steps. For provenence, you measure source retention among agents who know it. How many remember it where it came from etc. For rumors, you can measure uncerny preservation and false surn rate. For planning, you can measure action consistency and time to replan. And for privacy, you can measure containment. This matters because optimizing only one metric can create bad behavior because let's say if you only optimize for diffusion, the agents may learn to overshare everything. And let's say if you only optimize for memory recall, you might create noisy or still um like memories. So this scorecard is what keeps the system honest and prevents the auto research agent from gamifying the system to just increase one specific score.

限制修改范围与搜索受控策略空间

Speaker: 我们在整个项目中得出的另一个重要的工程教训是,保持可编辑范围足够小是非常重要的。Auto Research 层不应该有权限随机重写整个代码库。相反,冻结测试框架、场景和指标非常重要。因此,我们只暴露出系统中我们真正想要优化的部分。在 Project Paradox 中,对我们来说这指的是诸如记忆写入策略、检索策略、通信提示词、信念、信任规则、来源归因、重新规划触发器等。这给搜索过程留出了改善行为的空间,但也防止了它像我们之前提到的那样直接对评估进行钻营。这就是大语言模型(LLM)写随机补丁和它真正在受控策略空间内进行搜索之间的区别。这里我给出一些这种循环应该去搜索的更改类型的例子。如果来源归因消失了,策略修改可能是“在记忆中保留来源,并在写入记忆和生成摘要时记录”。如果传闻变成了事实,策略修改可能是“存储置信度,标记一手与二手信息,并在复述不确定的声明时要求使用保留语气的词汇”。如果公开事实只停留在局部,策略修改可能是……

Original English

Speaker: The other important engineering lesson that we learned over this project is that uh it's important to keep the editable surface really small. The auto research layer should not have permission to randomly rewrite the whole codebase. Instead, it's really important to freeze the harness, the scenarios, and the metrics. So, we're only exposing the part of the system that we actually want to optimize. Here in project paradox for us that meant things like memory writing policy, retrieval policy, communication prompt, belief, trust rules, source attribution, replanning triggers, etc. This gives the search pro process room to improve behavior, but it also prevents it from gaming the evaluation directly as we mentioned before. And this is the difference between the LM writing random patches versus the LM actually searching within a controlled policy space. And here here are examples of the kind of changes I want this kind of loop to search over. If if source attribution disappears, the policy change might be preserve source in memory and uh write uh memory rights and summaries. If rumors harden into facts, the policy change might be store confidence, marked firsthand versus secondhand, and require hedging when retelling uncertain claims. If if facts if public facts stay local, the policy change might

微调智能体协议与长周期评估 (Fine-tuning Agent Protocols and Long-horizon Evaluation)

Speaker A: 对有用的公共事实进行不同的分类,并使智能体主动分享重要的来源证据。关键在于,这些是对智能体协议的微小改变,但它们可以在多智能体系统的社会层面行为上产生更大的影响。这也是我想对我们的主张保持谨慎的地方,因为如果没有重复的当前循环结果,我们认为我不会说系统就普遍改善了。我们想说的是,这正是暴露给自动研究层循环的合适类型的表面,因为它足够小,可以控制,但它仍然足够丰富,至少在某种程度上可以改变社会行为。

Original English

Speaker A: be classify useful public facts differently and make agents proactively share important source evidence. The key is that these are small changes to the agent protocol, but they can have larger effects on a society level behavior for multi-agentic systems. This is also where I kind of want to be careful about our claims here because with we believe without repeated current loop results like I wouldn't say the system just generally improved. We're trying to say this is the right kind of surface to expose to an auto research layer uh loop because it is small enough to control but it's still rich enough to change the social behavior to some extent at least.

Speaker A: 对我来说最大的教训也许是,在这里记忆是不够的。你可以给智能体添加 RAG(检索增强生成)记忆,但仍然无法获得你所寻找的当前长期周期的行为。因为智能体有时需要知道这些信息从何而来。你需要保留它是第一手资料、第二手资料、已验证的还是不确定的。有时你还需要将原始的片段性记忆与智能体当前的信念区分开来。你需要通过场景来测试行为,而不仅仅是通过感觉。所以另一个教训是,回滚(roll back)也是必不可少的。当你优化社会行为时,一个改变可能会改善某一方面,却破坏另一方面。因此,一个更快传播公共事实的策略,也可能会泄露私人信息。一个提高召回率的策略,可能会增加陈旧记忆的使用。因此,这个循环基本上应该像一个棘轮(ratchet)。尝试一个改变,对其进行评分,只有在计分卡有所改善且护栏完好无损的情况下才保留它。

Original English

Speaker A: And the biggest lesson for me perhaps was that memory is not enough here. You can add a rag memory to an agent and still not get the current long-term uh horizon behavior that you were looking for. Um because agents need to sometimes know where that information came for uh came from. You need to preserve whether it was firsthand, secondhand, verified or uncertain. Sometimes you need to separate raw episodic memories from what the agent currently believes too. And you need to test behavior through scenarios, not not just through vibes. So the other lesson is that uh roll back also is not optional. When you optimize social behavior, a change can improve one thing and damage another. So, a policy that spreads public facts uh faster might also leak private information. A policy that increases recall might increase stale memory usage. So, the loop should basically be like a ratchet. Try a change, score it, keep it only if the scorecard improves and guard rails whole.

Speaker A: 我们坚信,这不仅仅与游戏智能体相关。因为尽管我用游戏村庄给你举了个例子,我们相信,比方说,支持智能体需要知道哪个策略更新来自哪里,以及它是否取代了旧的答案。例如,个人助理需要记住他们之前做出的承诺,并在用户想要更改这些个人承诺时做出修正。研究智能体需要出处引用、矛盾处理和假设更新。编码智能体需要跨越问题、文件、队友和不断变化的需求的长期上下文。工作流智能体在世界发生变化时,需要访问控制、交接和重新规划。所有这些系统都有相同的潜在问题。它们随时间推移维持状态。而这些状态会影响未来的行动。所以它们需要控制场景,而行为计分卡就是我们所提议的。

Original English

Speaker A: And we we definitely believe this is not only relevant for game agents because although I gave you an example using a game village um we believe like let's say for example support agents support agents need to know which policy update comes from where right and whether it supersedes an older answer. Personal assistants for example need to remember commitments that they previously made and h make corrections if uh if the user uh wants to change those personal commitments. Research agents need pro uh provenence citations, contradiction handling and hypothesis updates. Coding agents need longunning context across issues, files, teammates and changing requirements. Workflow agents need access controls, handoffs, and replplanning when the world changes. All of these systems have the same underlying problem. They maintain state over time. And that state affect affects future action. So they need control scenarios and behavioral scorecards is what we are proposing.

Speaker A: 简言之,再次强调,这是长周期智能体的配方。如果有一个实用的配方我希望你们能带走,那就是:冻结测试工具(harness),定义场景,记录追踪,对行为进行评分,并且只暴露一个很小的策略表面。在这些改变中进行搜索,只保留那些通过了你的测量的改变。我们相信这是一种对长期运行的智能体有意义的工程模式。我们认为真正的问题是:在控制运行中,系统表现得更好了吗?最后,Paradox 项目最初是为了让游戏智能体在 3D 世界中感觉栩栩如生。但对我们来说,更深层次的工程问题不是动画或对话。而是状态问题,比如哪个智能体知道什么,哪个智能体告诉了谁,什么是真实的、不确定的或过时的。以及智能体会根据他们记住的东西采取行动吗?Otter 研究为我们提供了一种稍微更系统地处理这个问题的方法。不是通过信任一个演示,也不是无休止地手动调整提示词(prompts),而是通过运行控制实验,并只保留那些通过了我们测量的改变。长周期智能体需要的是实验,而不仅仅是提示词。我希望这就是你们从这次演讲中获得的收获。是的,请一定与我们联系。如果你有任何问题,我们很乐意交流。非常感谢你们的聆听。

Original English

Speaker A: So again in brief, a recipe for long horizon agents. If there is one practical recipe I want you to take away, freeze the harness, define scenarios, log traces, score behavior, and expose only a small policy surface. Search over these changes, keep only changes that survive your measurement. And this is an engineering pattern that we believe would uh make sense for longunning agents. The real question we believe is across controlled runs, does the system behave better? To close, project paradox started as an attempt to make game agents feel alive in a 3D world. But the deeper engineing problem was not animation or dialogue for us. It was the state such as which agent knows what, which agent told whom, what is true, uncertain or outdated. And do agents act on what they remember? Otter research. Otter research gave us a way to approach this a bit more systematically. Not by trusting one demo and not by endlessly handtuning prompts, but by running control experiments and keeping only the changes that survived our measurement. Long horizon agents need experiments and not just prompts. And I hope that's the takeaway that you get from this talk. And yes, please do connect with us. We'd love to talk if you have any questions. Thank you so much for listening.

用代码智能体生成视觉产物 (Using Coding Agents for Visual Artifacts)

Amole: 大家好,我是 Amole,Nori Agentic 的 CEO。我们部署了一个理解你的公司、你的代码、文档、Slack 以及其他各类数据的 AI 员工。我们花了很多时间思考编码智能体究竟是如何工作的。大多数人认为编码智能体只写代码,但如果你问我,那纯粹是糟糕的营销。暂时忘掉这个名字吧。编码智能体几乎能做任何事情。只有一个诀窍。你必须能够像智能体一样思考,才能让它做你想让它做的事情。今天我们要讨论的是,我们如何利用编码智能体来做一些大多数人认为智能体极其不擅长的事情。制作视觉产物,比如幻灯片、文档,是的,甚至是视频。每天,全世界要投入大约 34,000 个人类年的时间来制作幻灯片。其中大部分时间不是在思考,而是在瞎摆弄。一旦你把所有的格式调整、品牌设计和把东西移来移去的时间都去除了,一个需要 10 小时才能完成的演示文稿,实际上只需要大约 25 分钟。

Original English

Amole: Hi, I'm Amole, CEO of Nori Aentic. We deploy an AI employee that understands your company, your code, docs, Slack, and other kinds of data. We spend a lot of time thinking about how coding agents really work. Most people think coding agents only write code, but if you ask me, that's just bad marketing. Forget the name for a second. Coding agents can do almost anything. There's just one trick. You have to be able to think like an agent to get it to do what you want it to do. Today we're going to talk about how we use coding agents to do something most people think agents are terrible at. Make visual artifacts like slides, docs, and yeah, even video. Every day, the world pours something like 34,000 human years into making slide decks. Most of that time isn't the thinking, it's the fiddling. A deck that takes 10 hours should really take about 25 minutes once you remove all the formatting and the branding and the moving things around.

Amole: 假设你需要制作一张幻灯片。你会怎么做?你会打开一个工具:PowerPoint、Slides、Figma、Canva,然后你开始操作一个画布。所有这些工具都是为人类的双手和人类的眼睛打造的。点击、拖动、放置、调整大小、对齐到网格。所有的动作和模式对于我们关于世界的空间地理视角来说都是合理的。底层有一个数据结构,但它的格式只有应用程序才能读取。当你把这些工具交给一个智能体时会发生什么?嗯,输出结果完全错误。东西以奇怪的方式重叠。你看不到文本。没有对齐。简直就是垃圾。AI 怀疑论者说这不仅仅是工具的问题。智能体从根本上就无法进行空间推理。还有像 Arc AGI 这样的整个基准测试,它们就是完全围绕这个前提建立的。关于这一点,开发者 Simon Willis 有个著名的小测试。他向每一个新模型提出同样的问题:你能画一只骑自行车的鹈鹕吗?但这里有一个陷阱。智能体只被允许使用 SVG。这是一个快速的直觉检查,用于判断一个模型是否具备空间推理能力。这里有一些在这个测试中模型实际给出的结果例子。是的,这些非常糟糕。我是说真正的、极其糟糕。

Original English

Amole: Say you need to make a slide. What do you do? You open a tool, PowerPoint, Slides, Figma, Canva, and then you start manipulating a canvas. Every one of these tools is built for human hands and human eyes. Click, drag, drop, resize, snap to grid. All motions and patterns that make sense for our geospatial view of the world. There is a data structure underneath, but it's in a format that only the application can read. What happens when you hand these tools to an agent? Well, the output comes out all wrong. Things overlap in weird ways. You can't see the text. There's no alignment. It's just garbage. AI skeptics say that it's not just the tools. agents fundamentally can't reason about space. And there are whole benchmarks like Arc AGI that are built exactly around that premise. There's a famous little test for this from developer Simon Willis. He asks every new model the same thing. Can you draw a pelican riding a bicycle? But there's a trick. The agent is only allowed to use SVG. It's a quick gut check for whether a model can reason about space at all. Here are some examples of what the models actually give you on this test. And yeah, these are pretty bad. Like genuinely, deeply really bad.

Amole: 那么,这是否意味着毫无希望呢?智能体注定就不擅长处理图形?不,我不这么认为。如果你问我,问题不在于模型,而在于媒介。如果我要求你,一个大概率是人类的人,去手写一只鹈鹕的 SVG 代码,你也做不到。SVG 仅仅是一面由数字组成的墙。你无法从一堆数字直接想象出一只鹈鹕。你就是无法以那种方式去“看”。人类就不是那么思考的。我们是以图形方式思考的。所以,我们构建了让我们在画布上作画的工具。Figma、MCP、PowerPoint、CLI、截图和替换循环。所有这些智能体工具都有什么共同点?它们都像人类一样去处理问题。但 AI 不是人类。要求 AI 使用画布,就像要求人类徒手编写 SVG 一样。这其实说不通。你需要根据 AI 思考的方式为其提供工具,不是通过像素,而是通过语言。单词、标记 (tokens)、结构。那才是它的原生媒介。

Original English

Amole: So, does that mean it's hopeless? Agents are just doomed to be bad at graphics? No, I don't think so. If you ask me, it's not the model, it's the medium. If I asked you, someone who is presumably human, to handwrite an SVG of a pelican, you wouldn't be able to do that either. SVGs are just a wall of numbers. You can't go from a wall of numbers to a pelican. You just can't see that way. That's just not how people think. We think graphically. So, we build tools that let us draw on a canvas. Figma, MCP's, PowerPoint, CLIs, screenshot and replace loops. What do all of these agent tools have in common? They all approach the problem like a human. But an AI is not a human. Asking an AI to use a canvas is like asking a human to write SVG by hand. It doesn't really make sense. You need to give the AI tools based on how it thinks, not in pixels, in language. Words, tokens, structure. That is its native medium.

Amole: 想象一种极其擅长描述布局的语言,模型已经见过并训练过数十亿个此类示例,它们能直观地理解,它可以渲染成像素并且可以在任何地方运行。噢,对了。HTML 让模型可以用结构来思考。HTML 标签在语言中内置了含义:标题、图表、网格,然后浏览器将所有这些转换成像素。所以,模型实际上从来不需要去放置坐标。你可以免费获得各种视觉效果、图表和布局、字体和动画。还记得之前那只鹈鹕吗?现在让它做完全一样的任务,但使用 HTML。还是那只鸟,但现在它位于一个模型可以进行推理的结构中。你可以读取、设置主题,并编辑它的每一行。

Original English

Amole: Imagine a language that's incredible at describing layout, that models have seen and trained on billions of examples of that they understand intuitively, that renders to pixels and can run everywhere. Oh, right. HTML lets a model think in structure. HTML tags have meanings built into the language, a heading, a chart, a grid, and the browser turns it all into pixels. So, the model never actually places a coordinate. And you can get all sorts of visual effects, charts and layouts, fonts and motion, all of it for free. Remember that pelican from earlier? Now ask it to do the same exact task, but in HTML. Same bird, but now it's in a structure that the model can reason about. And you can read and theme and edit every single line of it.

Amole: 我一辈子都在用 PowerPoint 制作演示文稿。所以我总是以为这两件事——演示幻灯片和 PowerPoint——是同义词。但这真的不是事实,对吧?PowerPoint 只是你用来制作幻灯片的工具。幻灯片本身,那只是演示模式。事实证明,你的观众里没有人在乎你是怎么进入演示模式的。编辑格式是完全任意的。所以你完全可以选择智能体本来就擅长的编辑格式:HTML,并且如果需要,以后再把它渲染成其他格式,比如 PDF。我们利用这个 HTML 技巧构建了我们所有的幻灯片、我们的董事会演示文稿以及我们的

Original English

Amole: I spent my whole life building slide decks with PowerPoint. So, I always thought that those two things, slide decks and PowerPoint, were synonyms. But that's just not really true, is it? PowerPoint is a tool that you use to make slide decks. The deck itself, that's just the presentation mode. And as it turns out, no one in your audience is going to care how you got to the presentation mode. The editing format is totally arbitrary. So you can just pick the editing format that the agents are already good at HTML and if you need to render to a different format like PDF later on. We use this HTML trick to build all of our slide decks, our board decks and our

销售演示与模型思维

Speaker A: 销售演示幻灯片。这些都是我们实际用来展示并不断发送出去的真实材料。我们也将其用于我们的文档。它为我们的文档增添了色彩和活力,同时符合我们的品牌形象。当然,我们也会用它来制作像现在这样的视频。你所看到的只是一些HTML和CSS。从头到尾其实就只是div而已。几乎所有东西加上一点结构和一点颜色都会变得更好。纯文本是一种选择,通常是出于方便的选择,但如果你真的想创造有用的东西,这通常是错误的选择。不过,我想在这里稍微停顿一下,指出一份单独的精美演示文稿通常是没有任何价值的。你仍然需要去获取所有的内容,所有实际填充该演示文稿的素材,对吧?那么,我们可以再次像模型一样思考。如果你直接让模型访问你的数据,比如你的通话记录或电子邮件,你就可以让模型端到端地构建演示文稿。让你的智能体做所有的苦差事,而你专注于愿景和故事。这就是Nory Sessions让你能够做到的。我曾在通勤地铁上用手机制作了完整的董事会演示文稿。为什么?因为我们的Norybot存在于我们公司的架构之中。当然,Nory自带了实现这一切所需的所有功能。所以,不用再重复造轮子了。这就是我的一点长篇大论。谢谢你的聆听。如果你只能带走一个重点,那就是这个:停止像用户一样思考,像模型一样思考。给它正确的语言。对于图形,你需要的只是HTML。

Original English

Speaker A: sales decks. These are real things that we actually present and send out constantly. We use it for our docs, too. It gives our docs color and vibrancy all while following our brand. And of course, we also use it to make videos like this one. What you're watching is just HTML and CSS. It's literally just divs all the way down. Almost everything is better with a little structure and a little bit of color. Plain text is a choice, generally a choice of convenience, but it's usually the wrong one if you're actually trying to create something of use. Now, I do want to take a quick beat here and point out that a beautiful deck on its own is generally not worth anything. You still have to go and get all of that content, all of the things that actually populate that deck, right? Well, again, we can think like the model. If you just give the model access to your data, say your call transcripts or your emails, you can have the model build the deck end to end. Let your agents do all the grunt work while you focus on vision and story. That's what Nory Sessions lets you do. I've built entire board decks for my phone on the subway during my commute. Why? Because our Norybot lives in the fabric of our company. Of course, Nory ships with everything you need to make this all work. So, don't bother reinventing the wheel. That's my little spiel. Thanks for listening. If you have just one takeaway, it's this. Stop thinking like a user. Think like the model. Give it the right language. And for graphics, all you need is HTML.

重塑移动开发工作流

Zion: 大家好,10X。你们感受到了吗?大家好,我叫Zion,过去14年里我是一名移动软件工程师,今天我在这里和大家谈谈10X,重塑移动开发工作流。所以,你知道,回到以前的时代,当时光标还是你用鼠标操作的东西,而AI智能体还是科幻小说或电影里那种反乌托邦的角色(无论哪种符合你的风格),你知道,就在几个月前,我们还认为我们会继续使用我们的IDE,只是可能稍微好一点而已。而现在我们知道,当我们使用Cloud Code、Codex、Cursor或其他什么进行讨论时,我们已经转向了类似于聊天的工程模式。我们只需告诉它们要做什么,而且除非是为了调试,或是智能体无法解决的问题,否则我们不再使用我们的IDE。从理论上讲,这应该能让我们提高10倍的生产力,对吧?大家都是这么说的。那我们的生产力真的提高10倍了吗?你感觉到了吗?我不知道,因为我感觉不到我们的生产力提高了10倍——无论是作为单个工程师,还是作为一个完整的团队,抑或是整个公司。

Original English

Zion: Hi everyone, 10X. You feel it yet? Hi, my name is Zion and I'm a mobile software engineer for the last 14 years and I'm here to talk to you today about 10X, reimagining the mobile dev workflow. So, you know, back in the old times when cursor was that thing you make with your mouse and AI agents were that dystopian character from sci-fi books or movies, whatever fits your style, you know, just a few months back then when we thought that we will still be using our IDE just maybe slightly better. And now we know that we already switched to like chat style engineering when we discuss with cloud code codex cursor whatever and we just tell them what to do and we don't use our IDs unless it's for debugging or something that the agent couldn't figure out and that in theory should have made us 10 times more productive right that's what everybody says right with are we 10 times more productive do you feel it I don't know because I can't feel that we are 10 times more productive not as a single engineer and not as a whole group and not as the whole company.

从蒸汽机到电动机:工作流的变革

Zion: 那么为什么会这样呢?为什么我们没有看到承诺的10倍生产力在现实生活中实现?所以你知道,人们讲过这样一个故事:当工厂从蒸汽机转向电动机时,起初他们并没有看到那么大的收益。是的,电动机更好。它们更有效率,但人们并没有看到承诺中的10倍、20倍、30倍的生产力提升。原因在于,他们只是将蒸汽机换成了电动机。但真正的收益是在几年后出现的,那时人们明白这不仅仅是关于更换发动机,而是关于改变整个工作流。因为你看,过去在工厂里有一台巨大的蒸汽机,所有的机器都是根据它们的耗电量以及与那台蒸汽机的距离来排列的。所以它并没有按照本应如此的工作流(从开始到结束)来组织。不,它是根据与中央发动机的距离来设计的。当他们意识到这一点,并且还意识到他们可以采用电动机,把它做得更小并放进每台机器内部,然后重新布置工厂使其按照应有的工作流运行时,因为这已经成为可能。那时真正的收益才到来。现在,他们的生产力比以前提高了10倍、20倍、30倍。这不仅仅是因为更换了发动机,而是因为改变了整个工作流。这就是今天我想和大家探讨的。让我们思考一下AI是如何让以前不可能的事情现在变得可能的。我们可以改变我们的工作流,从而使生产力提高10倍、20倍。

Original English

Zion: So why is that? Why do we don't see the promise of 10 times more productive came to an actual life? So you know they tell the story about how when factories switched from steam engines to electric engines at first they didn't see that big of a gain. So yeah, the electric engines were better. They were more efficient, but they didn't see that 10x, 20x, 30x more productiveness that they have been promised. And the reason for that was that they only changed the steam engine with the electric engine. But the real gain came some years afterwards when they understand that it's not only about changing the engine, it's about changing the whole workflow. Because you see, they used to have like one giant big steam engine in the factory and all of the machines were rearranged based on their power consumption and their proximity to that steam engine. So it wasn't organized by the workflow that it should have been like from the start to the end of the workflow. No, it was designed by proximity to that central engine. When they realized that and they also realized that they could take the electric engine, make it smaller and put it inside each machine and then they rearranged the factory to make it work as the workflow should because now it was made possible. Then the real gain came. Now they were 10 times, 20 times, 30 times more productive than they were before. Not because of only changing the engine but of changing the whole workflow. And that is what I want to talk to you about today. Let's think how AI make things that weren't possible before possible now. And we can change our workflow and then becoming 10 times 20 times more productive.

迭代与摩擦

Zion: 为了做到这一点,让我们看看当前的工作流。产品经理(PM)有一个想法。他们与设计师进行迭代。他们与用户进行迭代。他们与开发人员进行迭代。然后他们再回到设计师那里。接着,他们与质量保证(QA)进行迭代。然后他们再次与开发人员进行迭代。或许在所有这些迭代之后,你可能会得到一个生产环境中的产品。那么,重复了那么多次的词是什么?对,迭代。这就是问题所在,因为迭代会产生摩擦。每一次迭代都会造成上下文切换、浪费时间,以及需要进行的沟通和同步。而AI并没有消除所有的这些问题。AI加快了编码速度,但并没有消除摩擦,没有消除迭代。为什么会这样呢?所以,让我们重新想象一下我们可以做什么,请稍微忍耐一下。如果,如果在设计时使用一个工具,测试时使用另一个工具,编码时使用另一个,然后发布时又使用另一个,我们能不能改为使用同一个工具,同一个代码库?如果不再是在Figma上设计,然后将设计文档发送给开发人员让他们弄清楚如何实现这些设计呢?如果设计师实际上可以在代码中进行设计,然后向开发人员发送PR(拉取请求)呢?如果QA可以直接与智能体迭代,只需获得一个带有模拟器的链接,他们就可以准确地告诉智能体测试什么,要注意什么,以及如果他们发现问题,确切地告诉它去修复什么呢?如果我们能让开发工作流直接在代码本身上运行呢?如果上帝也是我们其中的一员呢?不,抱歉,我有点扯远了。

Original English

Zion: To do that, let's look at the current workflows. The PMs have an idea. They iterate with the designers. They iterate with the user. They iterate with the dev. They then back with the designer. Then they iterate with the QA. And they iterate back with the dev. And maybe after all those iterations maybe you have something in production. So what was that word that was repeating so many times? Yeah, iteration. And this is the problem because iteration creates friction. Each iteration creates context switch create time waste creates communication that needed to be done syn synchronization that needed to be down and AI didn't eliminate all of that AI sped up code but didn't eliminate the friction didn't eliminate the iteration why is that so let us reimagine what we could do bear with me for a moment what if what if What if instead of using one tool for designing, another one for testing, another one for coding, and then another one for releasing, what if we could use one tool, one codebase? What if instead of designing on Figma, then sending a design doc to the developer in order for them to figure out how to make those designs alive? What if designers could actually design own code and then send the developer a PR? What if QA could iterate with the agent itself, just getting a link with the simulator and they can tell the agent exactly what to test, what to be cautious of, and if they find something, exactly what to fix? What if we could make the dev workflow works on the code itself? What if God was one of us? No, sorry, I got carried away there.

云沙盒的引入

Zion: 你可能会问,我们如何能做到这一切呢?一种方法是告诉每个人只需下载他们的Xcode和Android Studio,然后教设计师、PM和QA如何构建以及如何在模拟器和仿真器上进行测试,然后在他们的笔记本电脑上占用200GB的存储空间,或者无论这对我们的内存造成什么影响。这是一种方法。但我猜大多数人会拒绝这个想法,而且理由很充分。所以我们可以采取另一种方式。也许我们只需将其放入我们的CI(持续集成)中,对吧?因此,我们让智能体与CI进行迭代,这样他们就不必下载Android Studio和Xcode及其他所有东西。但你其实知道,CI构建需要20到40分钟。我们不可能真的让我们的智能体等上40分钟,仅仅为了弄清楚它推送的iOS代码实际上构建失败了。那还能用什么?我们能使用什么?向大家介绍云沙盒(Cloud Sandboxes)。

Original English

Zion: And you're probably asking, how can we do all of it? So one way would be to tell everyone to just download their Xcode and and their Android Studio and teach designers and PMS and QA how to build and how to test on simulators, emulators and blow to their laptops with a 200 GB on storage and whatever they do to the to our memory. That's one way. But let me guess that most of them would reject that idea and for good purposes. So we can make another way. Maybe we just put it in our CI, right? So we let the agent iterate with the CI so they don't have to download Android Studio and Xcode and everything. But you actually know that CI builds take between 20 to 40 minutes. And we can't actually let our agent wait for 40 minutes just to understand that the iOS code that it pushed actually failed to build. So what else? What can we use? Introducing cloud sandboxes.

云沙盒在移动开发中的应用

Zion: 实际上,云沙盒的概念已经存在很多年了,只是还没有用于移动开发。使用云沙盒,你可以告诉智能体,这里有一个CLI。与该CLI对话。创建一个VM(虚拟机),一个只为这次迭代运行的小型VM。该VM在30秒或更短时间内启动。进行构建。在他们处于Cloud Code、Codex、Cursor或其他任何工具中的应用内浏览器里,向他们展示一个模拟器。然后,他们可以在其上进行迭代,告诉它改变那个模式,回去测试一些东西并更改代码,接着他们推送并打开一个PR。然后设计师可以在代码上工作,完成后发送PR给开发人员。开发人员可以进行迭代,创建一、二、三、四个不同的VM来并行运行。他们发送PR进行审查。QA可以接手,准确地告诉智能体测试什么,告诉它修复什么。从那里,它直接进入应用商店接受审核。让我们来看看。让我们看看它应该如何运作。

Original English

Zion: So cloud sandboxes are actually concept that has been around already for many years, just not for mobile development yet. Using cloud sandboxes, you can tell the agent, here's an here's a CLI. Talk to the CLI. Create a VM, a small VM that runs only for this iteration. The VM boots up in 30 seconds or less. Make the build. show them a simulator on their inapp browser in the cloud code, codex, cursor, whatever. And then they can iterate over it, tell you it to change that pattern, to go back and test something and change the code and they push and open a PR and then the designer can work on code, send a PR to the developer after they done. Developers make an iterations make one, two, three, four different VMs to run in parallel. They send the PR for review. QA can take it from there and tell the agent exactly what to test and tell it what to fix and from there it goes straight to the stores for review. So let's see it. Let's see how it should work.

实操演示

Zion: 想象一下你看到了这个屏幕。想象你正身处例如Codex中。你的左侧是聊天界面。右侧是实际的应用程序。设计师正在与智能体进行迭代。准确告诉它想要做什么,想要更改什么,并在屏幕上立即看到更改。构建时间更快了。它是在云端完成的,预览时间也更快了。然后他们在笔记本电脑上直接与智能体进行更多操作,而不是找开发人员,而且无需安装Xcode或Android Studio。一旦他们完成了,他们就可以告诉智能体去

Original English

Zion: So imagine you see this screen. Imagine you're inside Codex for example. You have the chat interface to your left. You have the actual app to your right. The designer is iterating with the agent. tell it exactly what they want them to do, what they want to change and see the changes immediately on their screen. Build time is faster. It's done on the cloud and preview time is faster. Then they some more not with the developer but with the agent on their laptop without the need to install Xcode or Android Studio. And once they done they can tell the agent to

改变工作流

Speaker A: 提取那段代码,发起一个 PR 并将其发送给开发者。这种工作流使我们的生产力提高了十倍。不仅因为使用了 AI,还因为利用 AI 改变了工作流,对其进行了重塑,并消除了我们过去习以为常的各种摩擦。这就是我们变得十倍高效的原因。谢谢大家。

Original English

Speaker A: take that code, open a PR and send it to the developer. This workflow is what makes us 10 times more productive. Not only because of using AI but because of using AI to change the workflow, reimagine it and remove all that friction that we took from for granted in the old times. That is how we become 10 times more productive. Thank you.

OpenGV 与 OG Assist 简介

Gabe Dees Mesa: 大家好,我叫 Gabe Dees Mesa。我是 OpenGV 的一名工程师。今天我们要讨论的是生产环境中的智能体。具体来说,就是 OpenGV 是如何构建和扩展 OG Assist 的。这个演讲将充满各种干货。我们将讨论 AI 智能体,我们的测试台环境(harness),以及评估、可观测性和链路追踪。我们还会讨论工具和技能。这里面会有很多有价值的内容。我们要和大家分享我们在 OpenGV 的做法,以及在我们的生产环境中,如何应对如此大规模的运行。因此,您将能够看到 AI 智能体的一个真实用例和工作负载。事不宜迟,让我们开始吧。

Original English

Gabe Dees Mesa: Hi everyone, my name is Gabe Dees Mesa. I'm an engineer here at OpenGV and today we're going to be talking about agents in production. Specifically, how open gov built and scaled og assist. Uh so um this presentation is going to be jam-packed with just so much good stuff. Uh we're going to talk about uh AI agents. We're going to talk about our harness. We're going to talk about um eval observability traces. We're going to talk about um tools and skills. Um it's there's going to be a lot of good stuff in here. We're going to talk to you guys about uh what we do at OpenGV and how we operate at the scale that uh we operate at um in production. So you'll be able to see a real use case and workload uh with AI agents. Um so without further ado, let's get started.

Gabe Dees Mesa: 好了,议程。我先快速过一下我们今天要讲的高层内容。我会给大家简单介绍一下 OG Assist 以及 OpenGV 是什么。我会讲述这一切是如何起源的。我们将深入探讨 OG Assist 在效果(effect)上的重大押注,以及我们核心的智能体循环。我们将讨论 A2A 协议、评估机制。我们会讲到我们是如何管理长上下文的。我们还会讨论监控和可观测性,我们如何收集反馈,以及如何基于反馈进行迭代。最后,我们还会谈谈工具和技能,以及在 OpenGV,我们如何不仅将 AI 用于外部以服务客户,还用于内部以改进我们的开发工作流。

Original English

Gabe Dees Mesa: Okay, agenda. So just really quickly going to go through uh high level what we're going to talk about today. Uh I'm going to tell you guys a little bit about OG Assist and what uh OpenGV is. I'm going to tell you guys the origin story of how this all kind of came to be. Uh we're going to talk about OG Assist's uh big bet on effect uh a little bit into our core agent loop. Uh we're going to talk about the A2A protocol, eval. We're going to talk about how we manage long context. We're going to talk about um monitoring observability, how we collect feedback uh and how we iterate on that feedback. We're gonna lastly uh also talk about tools and skills and how at open gov uh we use um AI not only externally uh that we uh serve to customers but also internally to improve our development workflows.

Gabe Dees Mesa: 在深入之前,先简单介绍一下我自己。我叫 Gabe。我是 OpenGV 的软件工程师。我在 AI 智能体团队工作,是协助构建 OG Assist 以及大家今天将看到的某些系统的成员之一。简单介绍一下 OpenGV。OpenGV 是一家软件公司,使命是推动政府变得更加高效且负责任。

Original English

Gabe Dees Mesa: Just a little bit about me before we go any further. My name is Gabe. I'm a software engineer here at OpenGV. I work on the AI agents team and uh I'm one of the folks that helped build uh OG Assist and some of the systems that you guys will be seeing today. So, a little bit about OpenGV. OpenGV is a software company uh on a mission to power more effective and accountable government.

会场过渡

Announcer: 请欢迎我们今天下午活动的司仪,Oliver Wright 美洲区技术总监,Deina Delias。

Original English

Announcer: Heat. Heat. Heat. Heat. Heat. Heat. Heat. Heat. Heat. Heat. Heat. Heat. Heat. Heat. Heat up here. Heat. Heat. Heat. Heat. Heat. Heat. Heat. Heat. Heat. Heat up here. Please welcome our MC for this afternoon's programming, director of technology at Oliver Wright Americas, Deina Delias.

Deina Delias: 大家晚上好。天哪,能和大家一起站在这里,我真的很感激。欢迎来到 AIE 2026。感谢各位在现场和线上的参与。非常感谢。抱歉,我是 Deina Delias,来自 Oliver White 美洲区,我们提供综合业务规划和战略咨询服务。很荣幸能与各位共聚一堂。我们涵盖了极其广泛的内容,18 个方向的研讨会、主题演讲、专题讨论、博览会环节、分组会议,最重要的是,你们的社交环节。今晚你们见到所有的朋友了吗?见到了?没见到?太珍贵了。是不是只有我一个人觉得,你知道得越多,就越觉得自己一无所知?请举手。哦,谢谢。什么?同情的举手?我收下了,谢谢。

Original English

Deina Delias: Good evening everyone. Gosh, I am so grateful to be up here with you. House AIE 2026. Thank you for being here live and online. Thank you so much. So um apologies Deina Delias Oliver White Americas we do integrated business planning and strategy consulting. So honored to be here with you all. We covered so many grounds, 18 tracks of workshops, keynotes, panels, expo sessions, breakouts, and most of all, your networking sessions. Have you met all of your friends tonight? Yes. No. Precious. Am I the only one who thinks the more I know, the more I don't know? Show of hands. Oh, thank you. What? Pity hands up. I I'll take it. Thank you.

Deina Delias: 但幸运的是,博览会有众多非常支持我们的赞助商和合作伙伴,他们随时准备帮助您实现业务和个人项目的最佳实践。去和他们聊聊,拜访他们,让他们帮助您实现目标。去看看跳舞的机器人,和它们合影。赢取赠品。今晚去看看初创企业路演(start a battlefield),并探讨最佳实践。接下来的这位演讲者是我非常敬仰的人,我很荣幸能为他做介绍。他的成就如此之多,很难用简短的几句话概括。所以我会用他自己谦逊的话来介绍。他是一位作家、教育家,也是 AI 最佳实践的倡导者。他将复杂的技术概念转化为通俗易懂的学习材料。我非常期待他将为我们带来的分享。请大家致以热烈的掌声,欢迎 Addios Mani。

Original English

Deina Delias: But thankfully for us, the expo has a mass of wonderfully supportive sponsors and expo partners ready to assist you in your business and personal projects for best practices. Talk to them, visit them, let them help you achieve your goals. Check out the dancing robots. Take a picture with them. Win the giveaways. check out start start a battlefield tonight um and talk about best practices. This next speaker is someone I truly look up to and honored to make his introduction. His achievements are so vast it's hard to wrap them all up in a few sentences. So I'll use his humble words instead. He's an author, an educator, advocate for AI best practices. He translates complex technical concepts into accessible learning materials. I am truly excited for what he has to say for us. Give a huge round of applause for Addios Mani.

让“人”留在环路中

Addios Mani: 大家好。下午好,或者不管你在 YouTube 上看这个视频时是什么时间,大家好。我很高兴能来到这里。今天,我想和大家谈谈在软件工程中,如何才能真正将人类保留在决策环路(human in the loop)中。在讨论这里的架构之前,我确实想先从人的方面谈起。我认为,未来的工程师将由那些能够选择哪些事情值得做的人来定义。他们将掌握证据。他们将掌握理解,以及对智能体越来越多地自动完成的工作做出裁决(verdict)。

Original English

Addios Mani: Howdy folks. So, good afternoon or good whatever time it is when you're watching this on YouTube. I'm really excited to be here and um today I want to talk to you about really uh what it takes to keep the human in the loop where engineering is concerned. I really want to start with the human side before we talk about the architecture here. I think that the engineer of the future is going to be really defined by the person who is able to choose what is worth doing. They're going to own the evidence. They're going to own the understanding as well as the verdict around increasingly automated work that's being done by agents.

Addios Mani: 现在,当我使用“裁决”这个词时,我并不是说我们突然都要变成 Judy 法官。我们不会。但我的意思其实略有不同。我的意思是,我们将对生产决策负责。某个东西要发布吗?我们要拦截它吗?我们要重定向它,还是接受风险?质量是我们经常谈论的话题,但质量会产生证据。裁决赋予了责任,而可问责性(answerability)才是真正让我们能够支持某项裁决的基础。当然,这并不是我们行业开始思考角色演变的唯一方式。

Original English

Addios Mani: Now, when I use the term verdict, I don't mean that we're suddenly all going to be Judge Judy. We're not. But what I mean really is something just a little bit different. I mean we're going to be accountable for the production decisions. Does something ship? Do we block it? Do we redirect it or accept the risk? Quality is something that we all talk about a lot, but quality produces evidence. A verdict assigns responsibility and answerability is really what lets us stand behind a verdict. And this, of course, is not the only way that our industry is starting to think about our roles evolving.

Addios Mani: Boris Churnney 最近用了一些有用的语言来描述许多团队开始感受到的变化。传统的工艺界限正变得模糊,角色正在围绕工作本身进行重新洗牌。在这里,重要的问题变得不再是你的头衔是什么,而是你能负责系统中的哪一部分。我很喜欢这种分类法。它很乐观,又不会显得过于空泛。比如原型设计、构建、清理、增长和维护。这些都是真实的工程模式。智能体将帮助完成所有这些工作,但稀缺的不仅仅是执行任务。而是要知道你的产品需要哪种模式,适用什么样的质量标准,以及谁对结果负责。

Original English

Addios Mani: Boris Churnney recently put some useful language around what many teams are starting to feel. The old craft boundaries are getting blurry and roles are rebundling around the work itself. And the important question here becomes a lot less about what is your title and more what part of the system can you own. Now I like this taxonomy quite a lot. Um it's optimistic without being overly vague. So things like prototype, build, sweep, grow, and maintain. And these are real engineering modes. Agents are going to help with all of them, but the scarce thing is not merely doing the task. It's going to be knowing which mode your product needs and what quality bar applies and who owns the result.

Addios Mani: 说到底,过去几天我们一直在讨论测试台环境(harnesses)、循环工程(loop engineering)和软件工厂。我们可以谈谈为什么会发生这种转变。我们已经超越了将大模型作为全部故事的阶段,对吧?在测试台工程中,编码智能体是模型加上其周围的环境,对吧?你的上下文,你的工具,你的文件系统,git。而测试台就是将智能转化为你可以委托任务的东西。下一步是循环工程,我们不再只是提示一次运行了。我们设计的系统会不断地提示、检查、记忆并决定接下来发生什么。正是在那个时候,智能体开始感觉像基础设施了。

Original English

Addios Mani: At the end of the day, now we've been talking about harnesses and loop engineering and software factories over the last couple of days. We can talk why this shift is happening. We move past the model as the whole story, right? With harness engineering, the coding agent is the model plus the harness around it, right? Your context, your tools, your file system, git. And the harness is what turns intelligence into something that you can delegate to. The next move was loop engineering where we weren't just prompting one run anymore. We were designing systems that kept prompting, checking, and remembering, and deciding what happened next. And that's really when agents started to feel like infrastructure.

Addios Mani: 一旦你开始把所有这些东西组合在一起,你就得到了软件工厂。Dex 在他的演讲中很好地涵盖了这一点。但你有智能体在这个内循环中运行,并产出证据。人类仍然在最终负责这个循环中的生产决策。风向并没有让我们远离这一点。我认为,风向正在使人类的判断成为杠杆率最高的检查点。这就是为什么它现在开始变得重要。对我们很多人来说,人工智能生成的和人工智能辅助的代码正在成为普通代码。

Original English

Addios Mani: And once you start putting all of those things together, you get that software factory. Dex covered this well in his talk. But you have agents that are running inside that inner loop and evidence that comes out. Humans still end up making the production decisions in this loop. And the wind really isn't moving us from it. The wind is moving human judgments the highest leveraged checkpoint I think. And this is why it starts to matter now. AI generated and AI assisted code is becoming normal code for a lot of us.

Addios Mani: Sonar 2026 年的一项调查表示,AI 辅助的代码已不再处于边缘地位。它在我们的代码库中越来越扮演着重要的角色。一旦发生这种情况,可问责性(answerability)就不再是一个哲学领域的问题。它成为了一项工程需求。这里还有一个关于质量的观点,对吧?比如我们过去很关心整洁的代码,能让人读懂的代码。但是,更干净的代码实际上不仅能帮助下一个人和团队中的下一个人。它实际上能帮助下一个智能体。Sonar 的另一项研究发现,干净的仓库和混乱的仓库其通过率大致相同,但干净的代码实际上消耗了更少的 token,并且导致了更少的重访(revisits)。因此,可维护性有很多好处,可以为你的软件工厂提高效率。

Original English

Addios Mani: One of Sonar's 2026 surveys said that AI assisted code is no longer marginal. It's increasingly having a large role in our code bases. And once that happens, answerability stops being this philosophical world. It becomes an engineering requirement. And there's a quality point here as well, right? Like we used to care about clean code. code that people could read. But cleaner code is actually not just going to help the next human and the next person on your teams. It actually helps the next agent. Another one of Sonar's research uh studies found that clean and messy repos had roughly the same pass rates, but clean code actually used fewer tokens and caused fewer revisits. So there's a lot of benefit to maintainability that can fuel efficiency for your factories.

Addios Mani: 虽然让代码生成变得更便宜了,但这并不会自动让代码审查也变得更便宜,对吧?我认为我们中的很多人都面临着这个时刻,而且我们知道工程师们并不天真。Sonar 的数据显示,几乎每个人都对 AI 代码持怀疑态度。现在我喜欢在我的软件工厂里工作。我喜欢构建我的工程...

Original English

Addios Mani: Now making generation cheaper does not automatically make review cheaper, right? I think a lot of us are facing this moment and we know that engineers are not naive. The sonar numbers say that almost everybody is skeptical of AI code. Now I love working in my software factory. I love building my engineering

信任、验证与人类判断的作用

Speaker: 但是问题仍然是容量。如果 96% 的人并不完全信任那段代码,但只有大约一半的人在提交之前总是进行验证,我们就会面临这种危险:我们拥有不信任,却缺乏带宽。因此,安全源于让验证变得更便宜、更清晰,并且让人们更难跳过。如果你从单个审查者放大到整个组织,当治理无法跟上,而采用速度已经远远快于任何公司能够去制定其政策的速度时,审查和验证就开始成为一个瓶颈。这意味着我们有一些难题必须处理,比如,一个模型是否真的碰过这个文件。这些难题还包括:是什么约束指导了那项工作?产生了什么证据,接受了什么风险,谁拥有这个结果。现在,代理可以交付的东西比我们任何人能审查的还要多,对吧?那么,我们还有什么用呢?这是许多人都在思考的问题,对吧?你知道,如果说荷马·辛普森(Homer Simpson)在计算机自动化方面的经验能给我们什么启示的话,也许这就是我们的未来。我不认为这就是我们的未来,但它是事情可能发展的一个方向。现在,让我们再试一次。如果变更是人类进入循环的地方,如果生成的速度扩展得比理解的速度还要快,那么稀缺资源就变成了有证据支持的判断。所以问题不再是代理能做多少,而是人类的判断在哪里还能创造杠杆作用。

Original English

Speaker: loops. But the problem is still capacity. If 96% of people don't fully trust that code, but only about half always verify before committing, we have this danger that we've got distrust without bandwidth. And so safety comes from making verification cheaper, clearer, and harder for people to skip. And if you zoom out from the individual reviewer to the organization, review and validation start becoming a bottleneck when governance isn't able to catch up and adoption is already moving way faster than any company can go and set their policies. And this means that we have some hard questions we have to deal with like did a model actually touch this file. And the hard questions are also like what constraints guided that work? what evidence was produced, what risk was accepted, and who owned the result. Now, the agent can ship more than any of us can review, right? So, what are we still good for? It's a question that's on a lot of our minds, right? And you know, if Homer Simpson's experience automating computers can teach us anything, maybe this is our future. I don't think it is, but it's one direction things can take. Now, let's try that again. If change is where humans enter the loop, if generation scales faster than comprehension, the scarce resource becomes judgment that's backed by evidence. So the question is no longer how much can the agent do, but where does human judgment still create leverage.

Alpha、衰变与品味的价值

Speaker: 现在我想和大家谈谈两个术语,我将在这次演讲的职业部分使用它们。Alpha 和衰变(decay)。Alpha 是你今天能做的事情与当前模型能做的事情之间的差距。这种差距是非常真实的,而衰变则是这个差距的倒计时。如果让你与众不同的是某种能力,那么前沿技术最终会赶上并取代它。对吧?围绕这一点有很多讨论。这也是“品味(taste)”这个词不断被提及的原因之一。Paul Graham 在这里提出了一个观点,我认为非常正确。当任何人都能创造任何东西时,选择创造什么就变得非常重要。我赞同这个观点。但我也认为我们必须非常小心,因为“品味”可能会变成一个神奇的词,用来指代那些我们暂时还不想解释的工作部分。Mitchell Hashimoto 给了我们一个更有用的定义版本。品味是在尚不存在客观指标的情况下,做出高质量定性判断的能力。这很重要,因为它将品味置于基准测试之前,置于市场完全投票之前。当你尝试一个模型,看到它构建的那种用户体验(UX)和那种体验时,你通常能分辨出你认为它有品味还是缺乏品味,或者那里是否存在一个人类可以填补的空白。现在,这也只有在我们能够随着时间的推移,将这种围绕品味的概念转化为批判的例子和更好的判断时,才是有用的。所以是的,当生产变得更便宜时,品味很重要。如果任何人都可以生成 10 个选项,那么稀缺的技能其实是知道哪个选项值得存在。但品味并不是什么永恒的护城河。它也是 alpha。现在,有品味的人仍然会很重要。我个人认为在很长一段时间内,他们仍然会很重要。但这项技能的最佳版本不是故弄玄虚。它是做出更好的决策,并留下你的团队在系统中可以学习的例子。

Original English

Speaker: Now I want to talk to you about two terms that I'm going to use for the career part of this talk. Alpha and decay. Alpha is the gap between what you can do today and what current models can do. That gap is a very real thing and decay is the clock on that gap. If the thing that makes you special is a capability, the frontier is eventually going to come for it. Right? And there's a whole conversation around this. This is one of the reasons why taste keeps coming up. Paul Graham had a point here that I think is very right. When anyone can make anything, choosing what to make becomes very important. And I buy that. But I also think that we have to be very careful because taste can become a magic word for whatever part of the work we don't want to explain just yet. Mitchell Hashimoto gave us a more useful version of this definition. Taste is the ability to make high-quality qualitative judgments where no objective metric exists yet. That matters because it puts tastes before the benchmark and before the market has fully voted. When you try out a model and you see the kind of UX and the kind of experiences that it builds, you can often tell when you think it has taste or lacks taste or when there's a gap there that humans can fill. Now, this is also only useful if we can turn some of this concept around taste into critique examples and better judgment over time. So yes, taste matters when production gets cheaper. And if anyone can generate 10 options, the scarce skill is really knowing which option deserves to exist. But taste is not some eternal moat. It's alpha as well. Now the people with taste are still going to matter. I personally think they're still going to matter for a long time. But the best version of that skill is not mystique. It's making better calls and leaving behind examples that your team in the system can learn from.

所有权与工程师的新定义

Speaker: 现在让我们应用衰变测试。嗯,我们过去拥有会衰变的速度。我们过去拥有回忆能力。你知道,工具链(harnesses)是有记忆的。验证正在转移到工具链、评估静态检查和模型批判中。品味。我仍然认为这会衰变得慢得多,但随着模型从例子和偏好中学习,它仍然会重置。甚至判断力在某些方面也是一个斜坡而不是一堵墙。因此,策略不是死守着某一种能力。我们需要不断将我们的优势提升一个层次。所以这就是为什么“代理能做什么”不再是最佳战略问题的原因之一。代理做不到的事情的清单在不断缩小。对我们来说,更好的问题实际上是:只有人类能对什么负责?这不是因为你知道我们中的任何人有什么魔力,而是因为某些决策实际上需要所有权(ownership)。在工作交付后,它们需要上下文、风险承受能力和责任心。这就是为什么“工程师”这个词必须变得稍微严格一点。现在,比以往任何时候都多的人可以让计算机做事情。我认为这真的很棒。构建者的总潜在市场从未如此之大,这太酷了。但这极大地扩大了杠杆作用。工程师不仅仅是一个会写代码、你知道的,并且能把东西做出来的人。工程师能够对系统进行推理。他们思考约束条件。你捍卫权衡。你可以管理风险。而且当事情开始出问题时,你是那个可以被联系到的人。

Original English

Speaker: Now let's apply the decay test. Well, we used to have speed that decayed. We used to have recall. You know, harnesses have memory. Verification is moving into harnesses, eval static checks, and model critique. Taste. I continue to think this is going to decay much more slowly, but it still resets as models learn from examples and preferences. Even judgment in some ways is a slope rather than a wall. So the strategy is not to cling to any one capability. It's for us to keep moving our edges up a level. So this is one of the reasons why what can the agent do is not the best strategic question anymore. The list of things that agents can't do just keeps shrinking. The better question for us is really what can only a human be answerable for? Not because you know any of us are magical in any way, but because some decisions actually require ownership. They require context, risk acceptance, and responsibility after that work ships. This is why the word engineer has to get just a little bit stricter. More people than ever can now make computers do things. And I think that's truly awesome. The total addressable market for builders has never been larger, and that's so cool. But it's a huge expansion of the leverage. An engineer is not merely somebody who can code, you know, and get things to exist. An engineer can reason about systems. They think about constraints. You defend trade-offs. You can manage risk. And you're the person that can be reached out to when things start to break.

工程师应避免的三个陷阱:认知债务

Speaker: 所以,如果我们想在这个时代保持高效和负责任,工程师应该避免什么呢?嗯,第一件要避免的事情实际上是认知债务(cognitive debt)。现在,认知债务是关于如何解决问题的理解和记忆的流失。我认为,随着我们每天越来越多地使用代理,我们很多人开始感受到这一点。我知道我经常有这种感觉,那是因为我们越来越多地把解决问题的事情交给 AI。对于代码来说,这就是你的代码库中存在的代码量与你的团队中任何一个人真正理解的代码量之间的差距。这就是为什么像授权深度(delegation depth)这样的事情最终变得重要。你可以有一个能通过测试的构建,一个你可以合并的 PR,但你的团队仍然可能会失去实际解释他们正在向生产环境发布的系统的能力。现在一个非常真实的压力也在于我们授权了多少。所以代理现在可以在系统中停留足够长的时间,以至于人类会失去线索。所以,一个 30 秒的运行可能感觉像是一次交互,但一个长达一小时或一天的任务,也就是一些长周期的工作流,当任务可能会持续那么长时间时,尤其是当你开始并行运行许多这类任务时,审查就不能仅仅是最后看一眼,它必须变成一个完整的控制系统。

Original English

Speaker: So what are things that engineers should avoid if we want to stay effective and accountable in this moment? Well, the first thing to avoid really is cognitive debt. Now, cognitive debt is the erosion of your understanding and memory around how to solve problems. I think a lot of us start to feel this the more that we're using agents every single day. I know that I feel this a lot and it's because we're deferring more and more to AI to solve our problems. For code, it's the gap between how much code exists in your repo and how much any human on your team genuinely understands. And this is why things like delegation depth end up mattering. You can have a build that passes you know your tests a PR that you can merge but your team can still end up losing its ability to actually explain the system that they are shipping to production. Now a very real pressure is much is also how much we delegate. So agents can now stay inside the system long enough for the human to lose the thread. So a 30 second run right can feel like an interaction but an hour or a day scale task so something long horizon that's a work stream and when tasks can end up you know lasting that long especially when you begin running many of them in parallel review can't just be a glance at the end it has to become a whole control system.

工程师应避免的三个陷阱:认知投降

Speaker: 第二件要避免的事情是认知投降(cognitive surrender)。这是指当你盲目接受 AI 的回答时,比如,授权很重要,因为授权意味着做这项工作,然后向我展示足够的证据,以便我能对其进行判断。在那种情况下,我仍然需要做出判断。投降实际上是在说,嘿,你的答案现在就是我的答案,在我自己形成任何意见之前。现在沃顿商学院(Wharton)做了一项研究,可以说是给了我们一个警告灯。当 AI 犯错时,仍有 73% 的人认为他们,你知道,他们选择了错误的答案,而且他们感觉更加确信。所以,失败模式并不是使用了 AI,而是借来的自信(borrowed confidence)。

Original English

Speaker: The second thing to avoid is cognitive surrender. Now this is when you blindly accept AI's responses like delegation is important because delegation says do the work then show me enough evidence that I can judge it. I still make a judgment in that situation. Surrender is really saying hey your answer is now my answer before I have formed any opinions myself. Now Wharton did a study that kind of offers us a warning light here. when AI was wrong, 73% of people still thought that they, you know, they picked the wrong answer and they felt more sure. So the failure mode is not using AI, but it's borrowed confidence.

工程师应避免的三个陷阱:编排税

Speaker: 第三件要避免的事情是编排税(orchestration tax)。现在,如果你在湾区(Bay Area)待过,你会看到有些人,不管是好是坏,仍然开着笔记本电脑走来走去,或者在和你谈论云端代理。我们越来越尝试并行运行越来越多的东西,或者告诉彼此我们正在用成百上千个代理进行交付。运行更多的 AI 代理并不意味着有更多个你可用。你的认知带宽无法并行化。所以,你创建的每一个循环最终都会导致更多的决策去路由、合并、验证和集成。而解决办法不一定是减少代理的数量,而是要像设计一个系统一样设计你的注意力。比如你从哪里切入,你需要什么,你重用什么。你只是需要对此非常刻意。

Original English

Speaker: The third thing to avoid is orchestration tax. Now, if you've been in the Bay Area, you will see people who, for better or worse, are still walking around with their laptops open or are talking to you about cloud agents. And we're increasingly trying to run more and more and more in parallel or telling each other that we're shipping with hundreds of agents or thousands of agents. More AI agents running does not mean that there is more of you available. Your cognitive bandwidth does not parallelize. So every loop that you create ends up causing more decisions to route, merge, verify, and integrate. And the fix is not necessarily fewer agents, but it's about designing your attention like a system. like where you enter, what you require, what you reuse. You just want to be very intentional about it.

问责制与职业生涯的数学题

Speaker: 现在,对于很多人来说,“问责制”(accountability)可能是一个可怕的词,如果它让你想去躲在灌木丛里,只告诉你的代理去处理,我也不会感到惊讶。但问责制并不是当代理变得优秀后所剩下的东西。它是让整个系统其余部分得以扩展的东西。如果代理可以做更多的工作,如果它们能以并行的方式做得更快,甚至比我们许多人做得更好,那么稀缺的东西就变成了:解释意图、检查证据、接受风险,以及在决策错误时改进系统的能力。现在,这里是职业生涯的数学题。一种优势的半衰期可能就是一个模型的发布周期。速度、回忆、验证,甚至品味,都会随着前沿技术的发展而改变。但是,一个签名的半衰期,你的信誉,你的专业知识,却是长得多的

Original English

Speaker: Now, accountability can be a scary word for a lot of people, and I wouldn't be surprised if it made you want to go hide in the bushes and just tell your agent to deal with it. But accountability is not what remains after agents get good. It's what lets the rest of the whole system scale. If agents can do more work, if they can do it faster in parallel, better than what many of us could do, the scarce thing becomes the ability to explain intent, to inspect evidence, to accept risk, and improve the system when the decision was wrong. Now, here is the career math. The half-life of an edge might be one model release. speed, recall, verification, even taste all move as the frontier moves. But the half-life of a signature, your credibility, your expertise is much

证据与责任

Speaker 1: 并且通过“署名”,我的确切意思是作品上的名字,这个人、这个团队、这个机构,无论谁是实际发布产品的背后支持者。所以技能可以获得杠杆。责任感能把杠杆转化为信任。而这就是我想清楚划定的界限之一。智能体可以做出选择,它们可以进行路由,可以合并,可以升级问题,它们可以在策略范围内操作。在许多系统中,你知道,它们可以而且应该这么做,但执行和责任是非常不同的两码事。智能体可以遵循你的运行手册,但它不能承担后果。当某些事情失败时,问题在于:谁理解这个策略?谁接受了这个风险?以及谁来负责影响范围?这些天,我们很多人都在谈论“高能动性”,把它当作我们在招聘时寻找的特质。高能动性是指积极地对你的结果取得所有权。所以要知道何时委派,何时检查,何时停止,以及何时在结果上署上你的名字。在这个世界里,高能动性不是“我亲自做所有事情”。你知道,那个版本真的无法规模化。它不仅仅是职场表演,它是带有判断力的所有权。这种能动性阶梯试图让它变得更具体一些。在最底层,有个人标记了问题并把它留给了系统。更高层次的人会执行、诊断、提议、推荐并解决。而最顶层罕见的动作是洞察力。你知道,也许你发现了一个问题,并决定它是否值得投资。也许不值得,也许你就继续前进了。但是当智能体使更多路径成为可能时,能动性不是去追逐每一条路径。它其实只是决定哪些路径值得你拥有所有权和关注。所以把这转化为一个运营模式。智能体可以运行更多内部的执行循环。它们可以调查、实施、测试和报告。我认为这其中存在杠杆,但外部循环仍然是工程设计。所以做决定、验证、批准、负责,那个内部循环是能力。外部循环是能动性。而这是我非常关心的一条边界。你的智能体返回证据。它返回差异对比、测试、日志、原理、追踪记录、轨迹、截图,无论工作本身需要什么。但接下来真正的工程才开始。我们决定这项工作是否值得做。我们验证证据是否足够,然后我们批准、重新引导或对进入生产环境的东西负责。不管你只是与少量智能体合作,还是与成千上万个智能体合作,这都不重要。我仍然非常认为这些想法是适用的。所以这条边界不是“人类看着人工智能的输出”。这条边界是“证据与责任”。所以这里有一条操作法则。要么解释清楚,要么不要发布。这不是因为人类必须敲每一行代码或者阅读每一行代码,而是因为必须要有人足够理解这项工作,从而能够捍卫它。如果你曾经在大型代码库或企业代码库工作过,有些代码库有“所有者文件”的概念,或者某些特定的子目录里有专人负责系统中那个部分。你可以用一种非常相似的方式来思考这个问题。在你的代码库中,谁对架构的那个部分负责?你的模型可能会编写代码,但真正的问题仍然是,你是否能解释智能体正在发布的那些变更,你是否有证据证明你理解其中的风险。现在,这是我想让你们在结尾记住的事情之一。自动化提升了我们所有人的底线。工程学继续向上迈进一个层级。而我们的新工作可能是循环设计、证据设计,以及棕地(现有系统)的管理,但减少按键次数并不意味着未来几年工程工作变少。它意味着有更多的表面积需要品味、验证、所有权,并最终需要关怀。我想我从未对这个领域的未来如此兴奋过。每次我们让编写软件变得更容易时,我们都会预测世界将不再需要那么多软件。而事实上,情况恰恰相反。高级语言出现了、框架、云技术、低代码。模式总是朝着另一个方向发展。当你降低成本时,潜在的需求就会显现出来。那些人们原本认为不可行、无法构建并推向市场的想法突然被解锁了。而智能体也将为许多人做同样的事情。它不会消除工程工作。它将把瓶颈从“我们能构建这个吗”转移到“这应该存在吗,我们能为它负责吗”。所以,建设工厂,保持运转,为最终决定负责。我希望这对你们有用。谢谢大家。

Original English

Speaker 1: longer. And by signature, I really mean the name on the work, the person, the team, the institution, whoever stands behind what's actually shipped. So skills can earn leverage. Accountability can turn leverage into trust. And this is one of the lines that I want to draw pretty clearly. Agents can choose, they can route, they can merge, they can escalate, they can operate inside policy. And in many systems, you know, they can, they should, but execution and responsibility are very different things. The agent can follow your runbook, but it can't inherit the consequences. When something fails, the question is, who understood the policy? Who accepted the risk? And who owns the blast radius? High agency is something that a lot of us talk about these days as being like this thing that we're looking for when we're hiring. High agency is actively taking ownership of your outcomes. So knowing when to delegate, when to inspect, when to stop, and when to put your name on the results. High agency in this world is not I personally do everything. You know, that version doesn't really scale. It's not just hustle theater, but it's ownership with judgment attached. This agency ladder tries to make that a little bit more concrete. At the bottom, you've got someone that flags a problem and leaves it for the system. higher up they execute, diagnose, propose, recommend, and resolved. And the rare top movement is discernment. You know, maybe you find a problem and you decide whether or not it's worth investing in. Maybe it's not and maybe you move on. But when agents make more paths possible, agency is not chasing every single path. It's really just deciding which paths deserve your ownership and attention. So translate that into an operating model. agents can run much more of the inner execution loop. They can investigate, implement, test and report. I think that there's leverage in that, but that outer loop is still engineering. So deciding, verifying, approving, owning, that inner loop is capability. The outer loop is agency. And this is a boundary that I really care about. Your agent returns evidence. It returns diffs, tests, logs, rationale, traces, trajectories, screenshots, whatever the work itself requires. But then the engineering really begins. We decide whether the work was worth doing. We verify whether the evidence is enough and we approve or redirect or own what reaches production. It doesn't matter if you're someone that's just working with a small number of agents or whether you're working with thousands of agents. I still very much think that these ideas apply. So the boundary is not human looks at AI output. The boundary is evidence and responsibility. So here's an operational rule. Explain it or don't ship it. And it's not because humans have to type every line or read every line, but because someone has to understand the work well enough to defend it. If you've ever worked in a large codebase or an enterprise codebase, some code bases have this concept of an owner's file or c certain subdirectories where there are people who are on the hook for that part of the system. You can think about this in a very similar way. Who's accountable for that part of your architecture in your codebase? Your model might write the code and the question is really still whether you can explain those changes that the agent is shipping, whether you've got the evidence where you understand the risks. Now, this is one of the things I want you to remember near the end. Automation moves the floor for all of us. Engineering continues to move up a level. And our new work might be loop design, evidence design, and brownfield stewardship, but fewer keystrokes doesn't mean less engineering over the next few years. It means that there is more surface area that needs taste, verification, ownership, and ultimately care. I don't think I've ever been more excited about the future of this field. Every time that we've made it easier to write software, we've predicted that the world would need less of it. And in fact, the opposite happened. Higher level languages happened, frameworks, cloud, low code. The pattern always went the other way. And when you lower the cost, latent demand ends up appearing. Those ideas that people didn't think were feasible to build and get out there are suddenly unlocked. And agents are going to do the same thing for a lot of people. It's not going to remove engineering work. It's going to move the bottleneck from can we build this to should this exist and can we answer for it. So build the factories, keep the lights on, own the verdict. I hope this was useful. Thank you.

Artificial Analysis 介绍

Micah: 现在请 Artificial Analysis 的联合创始人 George Cameron 和 Micah Hill Smith 上台。嘿,大家好。大家下午好。我是 Micah。这位是 George。我们是 Artificial Analysis 的联合创始人。Artificial Analysis 是一家 AI 基准测试公司。今天,我们要和大家探讨智能的成本。几年前,当我们俩都不会做这样的演讲时,我们会花很多时间来证明为什么智能和成本的权衡很重要。今天我打算跳过那部分内容,我们将直接切入正题,因为如果我需要说服在座的任何人为什么在 2026 年中期,智能成本是我们需要讨论的一个重要话题,我会感到非常震惊。所以接下来的安排是这样的。我将向大家简单介绍我们是谁。我们将使用我们的一些数据来简要了解一下 AI 竞赛的现状。然后,我们将把大部分时间花在分析当前 AI 的成本及其驱动因素上。我们将使用我们最新的智能体知识工作评估中的一些数据。这到底是什么意思?我们建立基准测试和评估,来测试 AI 技术栈中对开发者和制定 AI 技术决策的公司来说至关重要的一切。我们测试芯片、云基础设施、模型和智能体。我们试图弄清楚这些模型有多聪明,它们有多快,以及它们要花多少钱。我们在本网站上发布了大量此类数据。希望你们中的一些人已经看过了。我们与整个 AI 技术栈的公司合作,衡量他们的技术,帮助世界上的人们了解他们能做些什么。后面的幻灯片上有一些例子,是我们最近在 OpenAI、谷歌和英伟达的模型上所做的一些工作。让我们来看看竞赛的现状。在展示第一张图表之前,我要谈谈一个对我们思考构建 AI 评估的方式非常重要的观点。我们预见到的大部分希望 AI 去做的事情,现在的模型还太笨,无法完成。当前模型能做的事情是极其深刻的。今天的情况非常疯狂,但因为未来是如此广阔,这几乎可以肯定仍然是事实。所以这意味着在 AI 领域的任何特定时刻,我们都有这样一个我们认为是“智能前沿”的概念——即今天最聪明的模型能做什么。如果我们认为大多数任务都超越了这个范围,或者在能够可靠地完成它们方面肯定超越了这个范围,这就解释了为什么在这个房间里的我们所有人希望用 AI 做的事情,很大程度上都集中在绝对最新的前沿模型在任何特定时间点能做什么上。这也意味着,存在一组处于前沿之内的任务,而且随着新模型的推出,这组任务每个月都在增长。对于这一组任务,在智能与成本之间进行权衡非常重要,因为如果你选择不将最聪明的模型用于每一项事情,你可以少花 10 倍、100 倍、1000 倍的钱让 AI 完成同样的工作。关于竞赛的现状,我们发布了一个称为“Artificial Analysis 智能指数”的指标。我们喜欢说,它是了解 AI 竞赛最佳的单一数字,但如果我们认为你只需要一个数字,我们就不需要发布网站的其余部分了。这个指标实际上是我们运行的九个不同评估的综合结果。我们的指数现在到了 4.1 版本。它包含了一堆关于智能体的内容。它包含了一堆高难度的推理问答类型的内容。我们真的认为这是让你了解当前情况的最佳单一数字。我们有 Claude Fable 5 排名第一。那个……目前没有……

Original English

Micah: Now joining us on stage are the co-founders of artificial analysis, George Cameron and Micah Hill Smith. Hey, hey. Good afternoon everyone. I'm Micah. This is George. And we are the co-founders of Artificial Analysis. Artificial Analysis is an AI benchmarking company. And today, we're going to be talking to you about the cost of intelligence. A couple of years ago, when neither of us would give talks like this, we would spend a bunch of time justifying why intelligence and cost trade-offs matter. Today I'm going to skip that whole part of the bit and we're just going to get straight into it because I would be shocked if I needed to convince anyone in this room why the cost of intelligence is an important topic for us to be talking about in mid 2026. So here's what we're going to do. I'm going to tell you a bit about who we are. We're going to use some of our data to take a brief look at the state of the AI race. Then we're going to spend most of our time breaking down the cost of AI today and what's driving it. We're going to use some data from our latest agentic knowledge work evalu. What the heck does that mean? We build benchmarks and evals to test everything in the AI stack that matters to developers and companies making decisions about AI technologies. We test chips, cloud infrastructure, models, and agents. We try to figure out how smart the models are, how fast they are, and how much they cost. We publish a ton of that data on this website. Hopefully, some of you have seen it. And we work with companies throughout that entire AI stack to measure their technologies, help them in the world understand what they can do. Got a handful of examples on the slide back there from some of our work with OpenAI, Google, and Nvidia on their models. recently. Let's have a look at the state of the race. Before I show the first chart, going to talk about an idea that is very important to the way that we think about building AI evvelts. The vast majority of the things that we foreseeably want AI to do, the models are still far too dumb to do. It's utterly profound what the models can do. Today things are pretty nuts and yet because the future is so enormous this is almost certainly still true. So what this means is that at any given moment in AI we've got this concept that we think of as the intelligence frontier what today's smartest models can do. If we think of most of the tasks being beyond that, certainly beyond that in terms of being able to reliably do them, that explains why so much of what all of us in this room want to do with AI is focused on what the absolute latest frontier models at any given point can do. It also implies that there exists a set of tasks that are inside the frontier and that that set of tasks is growing every month as new models come out. For that set of tasks, playing the intelligence cost trade-off is incredibly important because by choosing to not use the smartest model for every single thing, you can spend 10, 100, a thousand times less to get the same work done by the AI. The state of the race, we publish a metric called artificial analysis intelligence index. We like to say that it is the best one number for understanding the AI race, but that if we thought you only needed one number, we wouldn't need to publish the rest of the website. What this metric actually is is a synthesis across nine different emails that we run. We're at version 4.1 of our index. It includes a bunch of agentic stuff. It includes a bunch of hard reasoning Q&A type stuff. And we really do think that it is the best one number for your sense of what's going on. We've got Claude Fable 5 on top. That little not currently

智力成本与模型演进趋势

Micah: ……现有的东西。我想我们今天之后得去网站上把它撤下来了。在我们的“智力指数(Intelligence Index)”中,我们喜欢做的一件事就是绘制它随时间的变化图。这里这张图展示了过去几年里,各个实验室推出的最聪明的模型。有些情况并没有太大变化。你可以看到 OpenAI 和 Anthropic 在过去几年里互有胜负。你也能大致看到,在 X 轴的右侧,这些点变得越来越密集,因为发布模型的节奏越来越快,尤其是在过去的一年里。你还可以看到,所有公司都在紧追前沿不放,它们已经并且正在发布一些模型,在短短几个月后就能达到与那些前沿模型相同的智力水平。

Original English

Micah: available thing. I guess we get to go remove that from the website after this today. One of the things we like to do with our intelligence index is plot how it's changed over time. This chart here is the smartest model from each one of these labs over the last few years. Some of it hasn't changed that much. You can see OpenAI and anthropic trading blows over the last few years. You can kind of see the dots getting closer together on the right hand side on the X-axis because the pace of releases especially over the last year has gone up and up. You can also see all of the companies hot on the heels of the frontier who have been and are releasing models that achieve the same level of intelligence as those frontier models just months later.

Micah: 如果我去掉其中一些折线,我们只看整体上最聪明的模型,以及在任何给定时间点最聪明的开源权重(open weights)模型,我们可以画出这条线,并且能够观察开源权重前沿与整体前沿之间的差距。在任何一个月里,你可能都会看到新闻头条要么说开源权重模型离前沿越来越远,要么说开源权重模型刚刚赶上了最新的闭源商业模型。我认为,当我们解读这张图时,我们看到的是:不幸的是,这两种极端的说法都不真实,我们看到的是一个稳定的 3 到 9 个月的差距,并且这个差距在过去 3 年里保持了惊人的稳定性。顺便说一句,这依然非常疯狂,因为这意味着在 Mythos 发布后的 9 个月内,我们预测就会有人免费提供一个和 Mythos 一样聪明的开源模型。大家可以记住我们这个预测。如果在未来一年左右的时间里这种趋势消失了,我会感到非常惊讶。

Original English

Micah: If I take some of these lines off and all we look at is the smartest model overall and the smartest open weights model at any given point, we can draw this line and we can look at the gap between the open weights frontier and the overall frontier. In any given month, you can probably find a headline saying that open weights models are further from the frontier than ever or that open weights models have just caught up to the latest proprietary models. I think when we read this chart, what we see is that unfortunately neither of the extreme versions are true and we see a consistent 3 to 9-month gap that's held surprisingly consistent over all of the last 3 years. That's still pretty nuts by the way though because that does mean that within 9 months of Mythos being announced, we are predicting that someone's going to give away a copy of a model as smart as Mythos. You can hold us that prediction. I'd be very surprised if this trend goes away anytime in the next year or so.

Micah: 除了智力之外,我们还可以绘制许多需要与模型智力水平进行权衡的指标。这一个指标非常简单,就是 Token 的价格。在这个我们称之为“智力成本”的演讲中,这个指标可能会让人感到惊讶,因为我们都觉得目前在 AI 上的支出正在飙升。这完全是事实。但是这里展示的趋势也同样是事实。对于每一个固定的智力水平,Token 的价格一直在以每年 5 到 10 倍的速度下降。这里的每一条线代表了智力指数中 10 个百分点的区间。我向你保证,如果你曾经需要在智力指数上高出 10 分的模型和另一个模型之间做选择,你很难在任务的整个分布中找到任何一个任务,是那个笨了 10 分的模型能够超越更好模型的。图表上的每一条线都在以惊人的速度下降。顺便说一下,这张图上的 Y 轴是对数坐标轴。而且,前沿模型的 Token 成本也保持了惊人的稳定性。

Original English

Micah: Beyond intelligence, we can plot a bunch of the metrics that you have to trade off against how smart the model is. This one's pretty simple. This one's the price of the tokens. This one actually might be surprising in a talk that we've called the cost of intelligence because we all have this feeling that the amount we can spend on AI is skyrocketing higher right now. And that's completely true. But this trend here is also true. Token prices have continued to fall by 5 to 10x every year for each fixed level of intelligence. Each of the lines here is a band of 10 points of intelligence index. I promise you that if you ever have to pick between a model that's 10 points higher on our intelligence index than another model, it's incredibly hard to find any task at all in the full distribution of tasks that the model that is 10 points dumber will outperform the better model on each one of these lines goes down incredibly quickly. It's a log axis on the y-axis on this chart, by the way. And the cost of tokens at the frontier has stayed surprisingly consistent.

Micah: 但是,当我们查看我们为智力指数运行的所有电子邮件和任务的每次任务成本时,是的,这个数字正在上升。这是所有任务的平均值,其中包括一些基于 Agent(智能体)的任务,也包括一些非 Agent 的任务。所以,它实际上掩盖了在今天某些情况下每次任务的成本变得多么极端。如果我们稍微拆解一下,这些字有点小,但我们在左侧看到了最高的数字。GPQA Diamond(著名的重要开源评估数据集),这是几年前的一个推理评估,我们不让模型以 Agent 的形式工作。现在它基本上已经被完全解决了。我们看到每个模型每个答案的成本从不到一美分到大约 50 美分不等。

Original English

Micah: But we look at cost per task across all of the emails and tasks that we run for our intelligence index and yeah the number is going up. This is the average across every task which includes some agentic stuff, some non-agentic stuff. So it's actually hiding how extreme cost per task gets in some situations today. If we break it out a little, these are kind of small but we've got the highest numbers on the left there. GPQA diamond famous important open source evaluation data set from a few years ago. It's a reasoning evaluation. We don't let the models work as agents. It's largely solved right solved now. We see from fractions of a cent per answer for each model up to about 50 cents.

Micah: 在我们的编码 Agent 指数以及我们新的 AI Briefcase 代理知识工作(Agentic Knowledge Work)评估中,我们看到在单个任务上的花费甚至超过了 20 美元。AI Briefcase 中最昂贵的任务实际上是这个数字的好几倍。当然,我们这里也有 Claude Fable 5,尽管说个有趣的事,这里的字很小,但你可以看到 Claude Sonnet 5 实际上使用了数量惊人的 Token,因此它在我们的 AI Briefcase 任务中成本接近最高,在这个图表的底部。但这正是我们大家都有的感受,我们正在尝试做这些非常困难的任务,前沿在不断推进,我们可以让模型去做的事情比以前更多了。所以,尽管每个固定智力水平的 Token 成本每年都在以 5 到 10 倍的速度下降,我们在每个任务上的花费依然可能比过去多得多。这些数量级的变化并不是我们的大脑能够凭直觉轻易理解的,这种矛盾也确实有些不可思议。所以,我现在要把时间交给 George,让他来为大家剖析我们是如何理解这些矛盾的。

Original English

Micah: In our coding agent index and in our new AI briefcase agentic knowledge work eval. We see up to beyond $20 being spent on a single task. The most expensive task in a briefcase is actually several times that leading that of course we do have claude fable 5 although fun fact it's kind of small here but you can see claude sonnet 5 actually uses an enormous number of tokens and so it's nearly expensive in our AI briefcase tasks down the bottom there but this is the thing that we're all feeling that we're trying to do these really hard tasks the frontier keeps moving there are more things that we can ask the models to do than there were a while ago So we can spend enormously more per task than we could even though that cost per token for each fixed level of intelligence is falling by 5 to 10x every year. These orders of magnitude are not things that our brains are good at getting intuitively and the contradictions are kind of nuts. So I'll pass off to George now to break down how we understand some of these contradictions.

AI Briefcase 基准测试与 Agent 成本剖析

George: 谢谢 Micah。那么,为什么 AI 感觉比以往任何时候都贵,而同时对于固定的智力水平来说,获取这种智力的 Token 价格却在急剧下降呢?我认为这就是 AI 工程师世界的真理:我们实际上想要投入更多,追求更高的 Token 预算。我现在要做的是,利用我们的 AI Briefcase 基准测试,对这种“智力成本”进行分析。我们的 AI Briefcase 基准测试是我们全新的 Agent 知识工作基准测试。它在真实的专业任务上评估模型。这里有四个私有场景,每一个都代表了相当于人类数周的工作量。我们会要求模型完成这些真实的任务。然后,我们从三个维度根据这些任务的输出对模型进行评分:标准正确性(Rubric correctness)、分析质量(Analytical quality)和展示效果(Presentation)。这很像我们评估人类工作的方式。

Original English

George: Thanks Micah. So why does AI feel more expensive than ever while for fixed levels of intelligence the prices of accessing that intelligence in terms of tokens is falling dramatically and I think this is AI engineer world fair we actually want to spend more higher token budgets when what I'm going to do now is use our AI briefcase benchmark to do analysis of this cost of intelligence Our AI briefcase benchmark is our new agentic knowledge work benchmark. It benchmarks models on realistic professional tasks. There's four private scenarios, each representing weeks of human equivalent work. And do we ask models to complete realistic tasks? Then we grade models on the outputs of those tasks across three dimensions. Rubric correctness, analytical quality, and presentation. Much like we think about assessing human work.

George: 与其他基准测试相比,AI Briefcase 的不同之处在于,我们试图让它尽可能地贴近现实。当你把任务交给你团队中的其他人,或者当你接到一个任务时,不幸的是,完成这个任务所需的确切信息并不会像放在托盘里一样直接端给你。你需要自己去寻找。你需要翻阅电子邮件,查看最新的 Slack 消息。这就是我们对自己和他人的期望。因此,在 AI Briefcase 给模型的任务中,我们试图模拟这一点。模型完成任务的环境包含了数以千计的文件、杂乱的 Excel 文件、非结构化文档、包含数百页的结构化文档和报告、电子邮件、Slack 消息。我们期望并要求 Agent 能够像我们要求自己一样去完成这些任务。

Original English

George: One of the differentiators for a briefcase compared to other benchmarks is we've tried to make it as realistic as possible. When giving a task to someone else on your team or when receiving a task, unfortunately, you're not given it on a platter with the precise information that you need to complete the task. You need to go out and find it. You need to troll through emails, pick up on the latest Slack messages. That's what we expect for ourselves and others. And so, we've tried to mimic this in the task that we're giving models in a briefcase. The environments that models are completing tasks in are thousands of files, messy Excel files, unstructured documents, structured documents and reports with hundreds of pages, emails, Slack messages. And we expect and ask of agents to complete these tasks just like we ask of ourselves.

George: 当我们查看模型完成这些任务的输出时,你可以看到输出质量存在着巨大的差异。这就是我们在这些 Agent 知识工作任务上评估模型质量和智力水平的方式。这也为我们提供了一个视角,让我们看到过去几年里在这个任务(这是一个商业尽职调查任务)上取得的进展。GPT-4o 给出了一个相当基础的幻灯片。o3 是去年初发布的一个突破性模型。想到 o3 仅仅是去年的模型,我就觉得不可思议。你可以看到 o3 生成了几个要点,有帮助,但这并不是我们在完成这种任务时对自己的期望。所以这就向我们展示了取得的进步——当我们看到 Opus 4.8 的输出和 Fable 5 的输出时,它们在分析的严谨性和展示质量方面要深入得多。

Original English

George: When we look at the outputs of models in completing these tasks, you can see vast differences in the quality of the outputs. And this is how we assess the quality and intelligence of these models on these agentic knowledge work tasks. It also gives us a perspective on the progress that's been made over the last couple of years on this task which is a commercial due diligence task. GPT-4o presents a pretty basic slide. o3 a breakthrough model that was released early last year. Thinking about that o3 was only last year is crazy to me. You can see that o3 produces a few bullet points helpful but not what we would expect of ourselves in completing this kind of task. And so this shows us the progress that's been made when we look at Opus 4.8's output and Fable 5's output, which goes a lot more in depth depth in terms of analytical rigor and presentation quality.

George: 那么,让我们来看看模型是如何完成这项任务的,以及花费了多少成本。如果你还记得 Micah 的幻灯片,他展示了有些模型消耗了价值超过 20 美元的 Token 来完成这些任务。因此,让我们来看看驱动因素,了解一下 Agent 任务的成本。我们要看四个驱动因素,这里的关键驱动因素是:Token 价格、Agent 运行轨迹中的轮次数量、模型的 Token 效率和使用情况,最后,但也可能是最重要的一点,就是提示词缓存(Prompt Caching)的影响。首先来看看 Token 价格。在查看缓存命中率价格、未命中缓存或没有缓存时的输入价格以及输出 Token 价格时,我们得出的第一个结论是,模型之间存在着数量级的差异。这是一个关键的驱动因素。像 Claude Fable 5 这样的前沿模型,与像 DeepSeek V4 Flash 和 GPT OSS120B 这样依然非常出色且实用的“工作马(workhorse)”模型相比,在 Token 价格上存在两个数量级的差异。这里的第二个结论是,单个 Token 之间或者不同类型的 Token 价格之间存在差异。你可以看到,在缓存命中价格上存在着巨大的差异……

Original English

George: So let's look at how models completed this task and what it cost. If you remember Micah's slide, he showed that some models are using over $20 worth of tokens to complete these tasks. And so let's look at the drivers to learn a bit about the costs of agentic tasks. Four drivers to look at and the key drivers here are token price, the number of turns in the agent trajectory, the token efficiency and usage of models, and last but potentially most important, the impact of prompt caching. Taking a look to start with the token prices. What we can see as a first takeaway here when looking at the cache hit rate token price the input not considering a cache hit or without a cache hit price and the output token price. Firstly is that there's orders of magnitude differences between the model. This is a critical driver. There's two orders of magnitude difference in terms of the token price between Frontier models like Claude Fable 5 and still good very usable workhorse models like DeepSeek V4 Flash and GPT OSS120B. The second takeaway here is the difference between the individual token or the types of token prices. You can see that there's vast differences in the cache hit price

Agentic Tasks and Token Costs

Speaker A: 还有无缓存命中的输入代币价格和输出代币价格。我们稍后在查看代币使用情况时会讲到它的影响。接下来,我们现在要求模型执行的是长时间运行的代理(agentic)任务,特别是在真实环境中,它们需要浏览数以千计的文件才能得出答案。而模型正在这样做。它们开始真正地探索环境,实际上类似于我们人类在搜索 Slack 以及执行类似任务时的做法。从这里模型工具调用的明细中你可以看到,它们正在进行数百次调用并在探索它们的环境。它们正在查看图像。它们正在读取文件。它们正在写入文件以进行特定的即席分析,这些分析将会反馈到我们刚刚看到的幻灯片输出中。这在每一轮交互中都需要花费输出代币,然后这些输出代币在代理的运行轨迹中又变成了输入代币,我们需要为此付费。

Original English

Speaker A: and the input token without a cash hit price and the output token price. And we'll get to that impact later when we look at token usage. Next, these are longunning agentic tasks that we are now asking of models, especially in realistic environments where they need to navigate all of these thousands of files to get to an answer. And models are doing that. They're starting to really explore the environment actually similar to humans when we search Slack and and and do similar tasks like that. You can see here with the breakdown of tool calls of models is that they're doing hundreds of calls and they're exploring their environment. They're viewing images. They're reading files. They're writing files to do ad hoc analysis that's going to feed into the the slide output that we just saw. And this costs each turn is output tokens and then those output tokens flow into input tokens in the agent trajectory and we pay for that.

Speaker A: 当我们查看完成一项任务所需的输出代币时,我们可以看到这里存在着巨大的差异。你可以看到昨天才发布的 Claude Sonnet 5 每个任务使用了超过 200,000 个输出代币。把这个和你几年前的 chatbt 查询对比一下,那时候你可能只需要几百个代币、几千个代币,最多也就 200,000 个代币就能完成一项任务。你可以看到,不同模型在这里有着数量级上的差异。这是由两个因素驱动的。首先是我们刚才看过的交互轮数。其次是模型输出的冗长程度。无论是它们进行了多少推理、为了完成任务输出了多少推理代币,还是在完成答案时输出的内容。它需要将那张幻灯片以及所有的细节整合在一起。那都需要代币。而我们需要为那些代币付费。

Original English

Speaker A: When we look at the output tokens to complete a task, we can see there's vast differences. You can see that Claude Sonnet 5 released only yesterday used over 200,000 output tokens per task. Compare that to your chatbt query uh a couple of years ago where you might have been doing couple of hundred tokens, couple of thousand tokens, maybe 200,000 tokens to complete a task. And you can see here that models vary orders of magnitude. And this is driven by two things. This is the number of turns that we just looked at. And secondly, it's the output verosity of the model. Both in terms of how much reasoning they're doing, how many reasoning tokens they're outputting to complete a task and also in completing their answer. It needs to put together that slide and all of that detail. That takes tokens. And we pay for those tokens.

Speaker A: 但是,退一步讲,我们不仅要看模型输出的代币,还要看我们实际需要付费的总代币数。我们在这里左边的图表中展示了这一点。这是一个代币分类的公文包图示:回答代币、推理代币、输入代币。有谁能在这里看到任何输出代币吗?它们全是输入代币。完成长时间运行的代理任务所需的绝大部分代币都是输入代币。你几乎看不到任何输出代币。因此,我们首先需要关注的两个代币价格,是无缓存命中的输入代币价格和有缓存命中的输入代币价格。如果我们回忆一下那张幻灯片,那些模型之间存在着巨大的差异。你可以在右边的图表中看到,这是输入代币命中缓存时的缓存折扣。这里通常在 90% 左右,但不同模型和供应商之间也有所不同,有些模型能达到 99%,而有些则在 80% 左右。如果考虑到绝大多数代币都是输入代币,你就能明白,缓存折扣或缓存命中率的不同,可能会让一项代理任务的总成本产生几倍的差异。因此,我认为我们习惯了思考输出代币,但我建议大家,在思考代理任务和代币成本时,让我们从缓存命中价格开始。

Original English

Speaker A: But stepping back not just at output tokens that the model's output but to total tokens that we're paying for. We have that on the left hand chart here. AA briefcase token breakdown answer tokens, reasoning tokens, input tokens. Can anybody see any output tokens here? They're all input tokens. The vast majority of tokens to complete longrunning agentic tasks are input tokens. You can barely see any output tokens there. And so therefore, the two token prices that we want to look at first is the input token price without a cash hit and the input token price with a cash hit. And if we remember that slide, there's vast differences between those models. And you can see that on the right chart here, which is the cash discount for a cash hit of an input token. It's usually around 90% here, but it's also different for models and providers whereby some models here are 99% and others are around 80%. And if we think about all the the vast majority of tokens being input tokens, you can understand that this can change by uh multiples a difference in a cash discount or a cash hit rate the total amount of an agentic task. And so I think we're used to thinking about output tokens, but I'd ask us, let's start with the cash hit price when thinking about the cost of an angentic task and tokens.

The AI Landscape and Cost of Intelligence

Speaker A: 我认为我们要分享并作为总结的最后一个观点,是理解 2026 年 AI 格局最重要的一张图表。在 2025 年,情况比较简单。那是我们的智能指数条形图。现在,我们首先看的是智能与每个任务成本的关系,因为我们现在正在应对智能成本的权衡问题。为了理解这一点,并帮助我们思考如何看待每个任务的成本,以及我们是否应该只使用最智能的模型或最便宜的模型,一个很有帮助的典型做法是将任务分为两种类型。

Original English

Speaker A: I think the last perspective we want to share with you and wrap up with is the most important chart for understanding the AI landscape in 2026. In 2025, it was simpler. It was our intelligence index bar chart. Now we start with the intelligence versus cost per task as we are now wrestling with these trade-offs of the cost of intelligence. And a helpful archetype to understand this and to reason about how to think about cost per task whether we should just use the most intelligent model or the cheapest model is to break down tasks into two archetypes.

Speaker A: 第一种类型是,对于你想多大程度地完成任务的智能水平没有上限。越智能意味着输出越好。这在当今大多数专业任务中的知识工作里都是如此。并非所有人都同意这一点,但这是我们 Artificial Analysis 非常坚定相信的一点。想想你可能在战略、如何节省成本,甚至在写职位描述时所做的分析。它总是可以更好的。作为人类我们总是能做得更好,对于模型也是如此。所以,在这个层面上,我们需要多高的智能水平是没有上限的,但我们确实需要权衡成本。因此问题是,我们愿意为额外的智能支付多少钱?在做这个决定时,你要看一下这里的帕累托曲线。

Original English

Speaker A: The first archetype is a task whereby there's not a ceiling on how much intelligence you could want to complete the task. More intelligent equals better outputs. And this is the case for most knowledge work today in prof in professional tasks. Not everybody agrees with that but that's something that artificial analysis we believe quite strongly. Think about analysis that you might do on strategy or on how we can save costs or on even writing a job description. It can always be better. We can always do a better job as humans and that's the case for models. So there's not a ceiling on that in terms of what level of intelligence we need, but we do need to trade-off costs. And so the question therefore is how much are we willing to pay for the extra intelligence? And you want to look at the paro line here in making that decision.

Speaker A: 第二种任务类型是存在上限的。一个例子是“上个月我在 Stripe 手续费上花了多少钱”。一个更聪明的模型不一定能给你一个不同或更好的答案。这个任务是有上限的,这种情况下你就要考虑完成该任务的最低智能水平是多少。然后,你应该选择最便宜的模型,即这张图表左侧的那些模型。这就是智能的成本。我们是 Artificial Analysis。我们在招聘。非常感谢大家。谢谢。请大家和我一起欢迎 Arena 的联合创始人兼首席技术官,Whene Chiang。

Original English

Speaker A: The second archetype of task is whereby there's a ceiling. An example is how much did I spend on Stripe fees last month. A smarter model doesn't necessarily give you a different or a better answer to that. There's a ceiling on the task and then you want to think about what is the level of intelligence, the minimum level of intelligence that can complete the task. And then you want to choose the cheapest model that which is to the left on this chart. So that is the cost of intelligence. We're artificial analysis. We're hiring. Thanks very much. Thanks. Please join me in welcoming the co-founder and chief technology officer at Arena, Whene Chiang.

Arena Eval Introductions

Whene Chiang: 大家好。非常激动能在这里分享我们在 Arena 构建代理评估体系的经验。我的名字叫 Wayin。我是 Arena 的联合创始人兼首席技术官。简单介绍一下我自己。我在加州大学伯克利分校获得了 AI 研究的博士学位,在那里我的研究重点是为 AI 系统构建稳健且可扩展的评估方法,而这项工作最终成为了我们今天在 Arena 所做事情的基础,也就是在现实世界中测量智能。你们当中的一些人可能听说过我们早期的工作,比如我们在 2023 年所做的“LMS 作为裁判 (LMS as a judge)”。我们做了一些早期的研究,并构建了一个聊天机器人擂台,有幸参与了其中一些评估研究。那么,Arena 是什么?简单来说,Arena 是一家 AI 评估公司。我们的使命是在现实世界中测量智能,超越单纯的静态基准,而是去评估那些真正能为用户和客户交付实际价值的智能。在过去的几年里,我们一直在追踪所有重大的 AI 突破,显然是在 2022 年那个重要的时刻(ChatGPT 时代)之后。在那之后是 GPD4 turbo,它在聊天和多模态能力上取得了突破。然后演进到了带有 openai01 的推理模型、思考模型。而在 2025 年,我们看到了 nana banana 的图像生成突破,它最初在公开发布之前,是以代码名开始在 arena 中进行测试的。我们还看到 Grok 迎头赶上,最近发布了 GPT images 2 成为当前图像模型的前沿,还有视频 AI 的生成,比如 B 和最近的 bid CES。到了 2025 年末,当 Opus 4.5、5、4.6 从一个优秀的编码模型变成了一个真正的代理编码模型,可以执行跨度更长的任务时,它也出现在了我们进行测量的 co- arena(Arena 2)中,我们看到了它比上一代模型的显著提升。以及我们在 Asian arena(Agent Arena)中测量的最新 fabled 突破,我们稍后会多谈一点。还有最近发布的 GLM 5.2,这对于开源模型社区来说真是一个重大的里程碑。

Original English

Whene Chiang: Hello everyone. Uh excited to be uh uh here sharing our experience uh building agent evals in Arena. My name is Wayin. I'm the co-founder and CTO at Arena. Um quick intro on me. Uh I did my PhD in AI research at UC Berkeley. uh where my focus was building robust scalable evaluations for AI systems and that work eventually become the foundation for what we are building today at Arena uh to measure intelligence in the real world. Some of you uh some of you may have heard uh our earlier work uh like LMS as a judge back in uh 2023. We did uh some of the early study as well as building a chapa arena which and some of the um evaluation research I was fortunate to contribute. So what is Arena? Um simply put it Arena is a AI evaluation company. Our mission is to measure intelligence in the real world beyond just static benchmark but uh the intelligence actually delivering real values to the users the customers and over the past couple years uh we have been tracking you know all the major AI breakthrough obviously after you know the chip moment in 2022 after that it was GPD4 turbo able GPD 4 uh having the breakthrough in chat and multimodel capability and then evolving to uh the reasoning model thinking model with uh openi01 and in 2025 we uh saw the image uh generation breakthrough of nana banana uh which was originally uh started testing in arena as a code name uh before it's public release and we are also seeing um Grok catching up GPT images 2 recently released uh to become you know the current frontier of image uh models as well as you know the video AI generations um B and recently bid CES so towards the end of 2025 when Opus 4.5 5 4.6 uh went from being a great coding model to a gen genuinely agentic coding model that can do longer horizon uh task that also showed up uh in arena 2 that where we measure in co- arena uh we see you know significant improvement over the past generational model and the most recent fable breakthrough um where we measure in Asian arena uh we will talk a little bit more later as well as the most recent GLM 5.2 release which is like really a big milestone uh for the open source model community.

Whene Chiang: 所以我们在 Arena 已经大规模地做到了这一点。现在我们看到每个月有 1000 万的访客在使用我们的产品 arena.ai,并且我们已经收集了 7 亿条跨越所有模态的对话数据——文本、视觉、图像、视频,以及如今的编程和代理领域。我们已经达到了一个巨大的里程碑。非常激动能与大家分享,我们就在最近宣布了,在首次发布我们的评估产品仅仅八个月后,我们的年化收入就达到了 1 亿美元。根据 a16z(az U 分析)的数据,按独立的月活跃访客数量计算,我们也被评为全球顶级的生成式 AI 产品之一。今天我想讨论的主题,同时也是我们提供服务的核心,就是实时排行榜(life leaderboard)。它是基于现实世界的评估,由这 1000 万用户和 7 亿条追踪数据提供支持,对过去几年里所有顶级的 AI 模型进行了排名,并且我们涵盖了文本、图像、视频、代码和代理。因此,我们非常希望能构建一个排行榜,帮助大家为任何任务找到最合适的模型……

Original English

Whene Chiang: So we have at Arena we have done this with scale. We now see 10 million monthly visitor going to uh our product uh arena.ai AI and we have collected 700 million conversations across all the modalities text, vision, image, video, coding these days agentic and we have hit a huge milestone. Very excited to share that just we just recently announced we hit 100 million um annualized revenue in just eight months after we first released our evaluation product. We are also uh ranked among the top genai product globally by unique number of monthly visitors according to az U analysis. So the um topic I want to cover today uh and the core of what we are offering um is life leaderboard uh which is based on real world evaluations u powered by the 10 million users 700 million um traces to rank all the top AI models from tier models uh for the past couple years and we cover text image video uh code agent Um so really wanted to build a um leaderboard that can help everyone to find the best model for

评估 Agent 面临的新挑战

Speaker A: 它们的用例,而且它是免费的。任何人都可以在 arena.ai/leaderboard 上查看和使用。你可以看到所有前沿模型的数据分析,比较成本、性能、用例、不同类别以及这些模型在不同模态下的能力。

Original English

Speaker A: their use cases and it's free. It's available for anyone to see to use at arena.ai/leerboard. You can see all the analytics thereof frontier comparing cost performance you know use cases different category different modality of these models capability.

Speaker A: 所以,今天我想讨论的真正问题是分享我们如何评估 Agent(智能体)的经验。我想分享一下我们过去几个月在构建 Agent 评估体系方面的第一手经验,这与过去我们评估聊天机器人的方式截然不同。在深入探讨细节之前,我想先在这里分享一些经验教训。

Original English

Speaker A: So yeah, so the real problem today I want to talk about is to share the experience how we how do we evaluate agents. um wanted to share our firsthand experience uh in the past common month we've been building uh the agentic eval which is very very different from the you know past in the past we evaluate chat bots and I wanted to share some lesson here before we diving into uh the details

Speaker A: 首先,为什么这很重要?我想谈谈当前的趋势。我们看到从聊天机器人到 Agent 范式的转变非常迅速。如果你看一下 OpenAI 关于 Codex 流量的数据,来自 Agent 的输出 token 份额已经飙升。你可以看到在 OpenAI 内部,本质上 100% 的 Codex 输出 token 都来自 Agent。而在其他组织中,现在的平均比例也已经超过了 60%,并且个体用户的比例也在非常快速地攀升。所以毫无疑问,现在的 token 消耗流是由 Agent 驱动的。

Original English

Speaker A: first why does this matter um wanted to talk about the trend so we have been seeing um the very rapid shift from uh the chatbot to Asian um paradigm shift and if you look at the openi's data on codeex traffic the share of the output token coming from agent has just skyrocketed and you can see inside openai essentially 100% of the uh output tokens from agent from codeex and for other organizations you know average is like above 60% now and individual also climbing very fast so there's no question that the token flow is now driven by agents

Speaker A: 我们还看到 Agent 不仅适用于工程师,它不仅仅用于软件工程。如果你看看 OpenAI 各部门对 Codex 的采用情况,工程部门显然是 99%,但财务、招聘、法务等部门的使用率也几乎达到了 90%。正如你所见,来自 Common Sense Advisory 的研究显示,每月的 token 使用量也在飙升,在未来几年内将达到 60 千万亿 token 的规模。

Original English

Speaker A: and we also see that agents are not just for engineers right it's not just for software engineering if you look at codeex adoptions by department at um openai engineering obviously 99% but also finance recruiting legal and so on they are all like almost like 90% and as so as so as you can see you know the studies from common sac the monthly token usage is also skyrocketing towards like you know 60 quadr quad trillion tokens in the next couple years.

Speaker A: 所以经济数据也反映了同样的趋势。如果你看看 REM 的数据,AI 支出正在接近人力支出。你会发现,排名前 1% 的公司,每个员工每月的 AI 支出实际上已经达到了 7400 美元左右,这大约是软件工程师薪水的一半。这真是一个历史性的转变,这意味着选择最好的模型、最合适的模型以及优化你的 Agent AI 工作流变得前所未有的重要。

Original English

Speaker A: So really you know the economics also tell the same story. If you look at the REM data the AI spending is getting closer to people spend right. So if you see like you know the top 1% of the company's monthly AI spend is per employee is actually already like 7 4K um roughly half of the salary software engineer. So this is really like you know historical shift that um meaning also the stack of like choosing the best model the right model and optimizing your agentic AI workflow is you know more has never been more important.

如何衡量 Agent 的产出

Speaker A: 所以这里的关键问题是,我们赋予了 Agent 很大的自主权,我们花费了很多,投入了很多,那么关键问题是:我们实际上该如何衡量 Agent 的产出?这确实是一个瓶颈,对吧?你想要了解这些 Agent 产出和操作的价值。由于几个原因,这被证明是一个相当困难的技术问题。

Original English

Speaker A: So the key question here is like um we give agent lots of autonomy. We spend a lot. We invest a lot. And the key question here is like how do we actually measure agents outcome? So that's really the bottleneck, right? You want to understand the value of these agentic uh output and actions. And this turned out to be a pretty hard technical problem for a few reasons.

Speaker A: 首先,Agent 是多组件系统,对吧?你有模型、Agent 执行循环、工具、测试框架……这些组件中的任何一个都可能导致系统崩溃。其次,Agent 需要在复杂的工作流中运行。现在在真实环境中,你需要构建应用程序、进行调试、做研究、生成文档、制作幻灯片等等,这更像是一项复杂的综合性任务。

Original English

Speaker A: First agents are multi-component systems, right? You got the model, the agent take loop, um the tool, the harness, um you know, any of these pieces can break the system. You also uh have agent operate through complex workflow. Now in a real environment, you build building app, debugging, doing research, producing document, uh slide deck and so on. So it's like more involved task.

Speaker A: 第三,我们能从执行轨迹中收集到的信号也变得稀疏,并分散在更长的时间跨度内。完成一项任务可能需要 100 次工具调用,然后你才知道它是成功还是失败,或者你是否有机会提供反馈来引导它。

Original English

Speaker A: Uh and third the uh signals that we can collect you know in this trajectory are also becoming sparse a spread across longer horizon. Um you know a task may take 100 to calls to to finish right before you know if it's succeeding or failing or you give any feedback of a chance to steer it.

构建 Agent Arena

Speaker A: 为了深入理解这个问题,我们在 Arena 决定亲自动手构建真实的 Agent 产品和应用程序。这样我们就能从真实用户那里收集自然产生的使用轨迹和反馈,以便我们进行研究并深入理解它。

Original English

Speaker A: uh and to deeply understand the problem at Arena we decided to actually firsthand building real world you know agentic product and app to actually source the organic traces and feedback from the actual users for us to you know do research and deeply understand that.

Speaker A: 因此,上个月我们在 Arena 中推出了 Agent 模式,允许任何人进入 Arena 体验和评估 Agent 功能。它现在已经开放给所有人使用,如果你能播放视频的话,我想给你们看一个非常简短的演示。

Original English

Speaker A: Uh so last month we launched uh Asia mode in arena uh to allow anyone to go to you know arena to experience and evaluate agentic capability. So it's right now available for everyone to use and wanted to show you a very quick demo if if I can start the uh is the video moving.

Speaker A: 好的,这就是 Agent Arena,你访问 arena.ai,选择 Agent 模式,这是一个真实的 Agent 产品。你可以来这里评估模型,你可以输入任何你想要的问题。比如在这个例子中,我的请求是:“下载谷歌的第一季度财报,并创建一个 PowerPoint 幻灯片来总结输出内容。”

Original English

Speaker A: Okay so this is agent arena you go to agent you go to arena.aii I you you choose the agent mode and this is a real world you know agentic product you can go and evaluate model you come in and type any question you want in this case um it's like I ask download Google's Q1 earning report uh and create a slide deck summarizing the output in PowerPoint

Speaker A: 你可以看到 Agent 开始工作,搜索网络,抓取正确的网站,开始构建幻灯片的结构,然后使用一些批处理工具编写 Python 代码来生成幻灯片。你能看到,在最后,模型生成了一个可以供用户下载和查看的文件,这真的是一个由模型输出的实体 PowerPoint 文件。在每次交互结束时,我们会询问“这项任务是否成功”,用户可以通过这种方式提供反馈。

Original English

Speaker A: and you can see the agent goes off and and doing work searching the web pulling the right website start structuring the deck and then using some of the batch tool writing Python code to um generate the the slide deck right and you can see that and at the end uh there's like a artifact generated by the model uh that user can download and see and this is like a you know a real powerpoint uh outputed by the model and then user can at the end we ask every turn like we ask was this task successful or not and user can provide feedback that way

Speaker A: 这就是我们用来评估和了解 Agent 是否实际交付成果的信号之一。这只是为了强调我们的面板功能。在底层,我们构建 Agent Arena 的方式是,为模型提供一套工具:文件系统工具(重写、编辑等)、网络搜索、图像获取、生成、以及最近添加的语音功能。基本上就是赋予模型类似于云端协作工具(如 Harness)的能力,以及运行代码以完成工作的终端访问权限。我们很快还会添加越来越多的连接器,比如 GitHub,它可以连接到你的代码库,以完成更严肃的软件工程任务。

Original English

Speaker A: and this one of the signals that we use to evaluate and understand whether agent actually delivers the outcome. So yeah this is just to highlight the panel and under the hood how we build the Asian arena it you know we give model set of tools um file system tools rewrite edit and so on and search web fetching image uh generation speech as well recently added so just really giving the model tools similar to like a cloud co-work like harness and also terminal access to run code to to to to you know do work and we also are adding more and more uh connector soon like GitHub uh which can connect to your repo to you know do more serious software engineering task

Speaker A: 你可以从这个图表中看到这些工具的使用情况。在一个一周的时间范围内,你可以看到有 570 万次工具调用。Bash(终端)是使用次数最多的工具,占比约 46%。这些 Agent 实际上是在使用这些工具为用户做真正的工作。

Original English

Speaker A: um and you can see this plot is the the usage of these tools uh in a in a time in a oneweek time frame you see 5.7 million to calls um you know bash is was the you know the number one used That's around 46% and the these agents are actually using these tools to do real real work for users.

基于真实使用轨迹的评估

Speaker A: 我们也深入分析了数据,发现用户非常努力地尝试执行难度更高、更复杂的任务。在我们观察到的真实会话中,用户正在构建电影观看列表应用、调试自动驾驶车辆的控制系统、构建 RAG(检索增强生成)管道、在微服务中实现功能等等。

Original English

Speaker A: So we also you know dig into the data and seeing users are you know pushing really hard to um trying to do more harder and complex task. Um so real session we've been seeing like you know users are building you know a movie watch list app debugging a control systems for autonomous you know vehicle and and architecting building a rack pipeline you know implementing features in micro and so on.

Speaker A: 这些会话动辄数百轮对话,包含几百次工具调用,都是非常严肃的工作。从这里你可以看出,我们在 Arena 构建的 Agent 实际上是在与用户一起做真实的工作,并为用户提供真正的价值。我们认为,最好的评估应该立足于并衡量像这样真实世界中的用例。

Original English

Speaker A: So these are the sessions like go over hundreds some of them go hundreds of turns and couple hundreds of tool calls very serious stuff. Um and you can from this you can tell that the u the agent that we built uh at arena is actually doing real work with users and giving user real value and we believe the best evaluation should be uh grounded and measured in real world use cases like this.

Speaker A: 所以我们在仅仅一个月前推出了 Agent Arena。在第一个月里,我们收集了超过一百万条 Agent 轨迹,这些任务涵盖了编程、研究、文档编写、头脑风暴和规划。我们看到超过一半的轨迹属于与工作相关的类别,更偏向于专业用途和复杂任务。

Original English

Speaker A: So we launched agent arena uh just a months ago and in the first months over uh we collected over a million agentic traces and these are you in task spending coding research document brainstorming planning and we see more than the half of these uh uh traces fall into work related category more like towards professional use and complex tasks.

Speaker A: 我们还看到,Agent 在 Arena 上已经编写了超过 5000 万行代码,包括 Python、Markdown、HTML、JavaScript 等。这是工具使用的分布情况,你可以看到编程排名第一。在这些任务中,有些更为复杂,使用了更多的工具,而有些使用的工具较少,这就是代码生成的代码行数统计。

Original English

Speaker A: Um and we have seen Asian also written um more than 50 million lines of code uh on arena, Python, Markdown, HTML, JavaScript and so on. This is the tool distributions that you can see the coding is the number one and some of these um task you can see is some of them are more complex using more tool uh some of them use less and this is the the line of code generation.

Speaker A: 回到评估的问题上来,假设我们收集了一百万条 Agent 轨迹,我们到底该如何将这些轨迹转化为一个排行榜,让我们能看懂哪个模型表现得比其他模型好?我们主要从三种类型的信号中挖掘数据。

Original English

Speaker A: So now the going back to the evaluation question, right? So say we collected a million agentic traces. How do we actually turn these traces into a leaderboard that we can understand which model performs better than the others? And we primarily um mine the signals from three type of uh basically signals.

Speaker A: 第一种是我刚才向你们展示的显性信号,用户会直接告诉我们哪项任务成功或失败。另一种是一些隐性信号。我们会观察用户是否实际下载了文件,是否抱怨模型生成的输出,或者对其进行赞扬等等。所以我们通过所有的轨迹来感知这些隐性信号。此外,还有环境反馈,即当代码运行时实际发生了什么,命令是成功还是失败等等。因此,我们基本上是扫描所有这些会话轨迹中的信号:每一条用户消息、助手的操作、工具、反馈结果,然后将它们聚合成各种指标,如成功率、好评率、过度合规率、耐久度、Bash 恢复能力、甚至是幻觉率。这里的每一个信号都可以……

Original English

Speaker A: One is like explicit which I just show you that user will tell us directly like which task succeeded or failed. Some of them the other one is some implicit. Uh we see that if user is actually uh say downloading the file or like um complaining about the output of the generation from the model or praising it and so on. So more like implicit signals we we sense through all the traces and also there's environment feedback where you know what actually happened when the code run whether the command succeeded or failed and so on. So we basically use these you know scans through all these sessions traces every user message assistant action tools resolve feedback and aggregate them into you know some of these signals like success rate praise over compliance durability bash recovery to hallucination and each of these signal can

Agentic App Leaderboard & Analytics

Speaker A: 从而得出排名,对吧。你可以精准地衡量,你知道在这个特定的指标下,哪个模型表现得比其他的更好,然后我们将其整合到最终的、嗯,你在网站上看到的排行榜中。嗯,所以,嗯,这就是它今天的样子。你可以看到,像,嗯,这段视频有五个不同的指标,而且模型在各个方面表现各异。现在,Fable 5 是排名第一的模型,它的净提升你知道大概是比平均水平高 14%,也就是,你知道,所有模型的平均值,紧随其后的是 Call Opus,GPT-5i High。而且关于这个数据,很有趣的一点是,你可以逐个指标去查看,嗯,模型可能在任务成功率上非常非常擅长,但有时候在稳定性方面,也就是在如何控制模型这方面会相对较弱。你可以确切地看到,比如,模型是在哪里失败的等等,而且我们将会添加,你知道,越来越多丰富的指标,来捕获这些失败的模式。所以在方法论上,核心理念基本上就是一个随机对照试验,我们在其中对智能体组件进行干预。我们衡量,你知道,任何给定组件对任务结果的因果效应,比如我们关心的指标。而且我们现在能够看到的主要就是编排模型(orchestrator models)的因果效应,嗯,但这个框架足够通用,所以我们也能衡量不同,呃,组件之间的交互效应。比如说,假设你想要衡量,呃,工具,你想要衡量不同的测试框架,或者是不同的系统提示词,呃,等等。所以所有这些在这个框架内都是可能的,并且我们也将,你知道,呃,对此进行评估。如果你感兴趣的话,更多的技术细节已经发布,呃,在我们的博客文章上了。嗯,所以,嗯,我们一直在追踪,就像我说的,智能体领域的所有主要发布,其中一个发布就发生在几周前,亚洲竞技场(Asia arena)的 Fable 5。嗯,所以如果你想在 X 上关注我们,你将会看到所有,你知道,最新的发布。而这个排行榜最有趣的一点在于,因为这是基于数百万个智能体追踪(agentic traces)的真实数据,对吧,所以你可以将其切分到你所关心的任何任务分布中。所以举个例子,比如说你关心,你知道,GDP 任务,这更像是具有经济价值的专业工作,相较于消费者用例,你可以,呃,你可以做一些数据分析来对数据进行切片。在里面你会看到,像 GPT-5i 实际上相当不错,呃,在 GPT,抱歉,像是 GDP 任务方面,呃,而 Gemini 往往在消费者用例上做得更好。所以基本上,最好的模型通常取决于,呃,你在做什么,你关心什么样的数据分布。嗯,另一方面则是成本,对吧,你知道,成本也很重要。我们可以,我们基本上可以绘制出这些,呃,净提升(也就是性能)与平均成本的关系图,来看看,来帮助你看到这里的帕累托前沿(Pareto frontier)。你可以看到 Fable 是成本效益最好的那个,呃,大约每次会话 10 美元,而 5ifi 依然要稍微便宜一点,而且 GP,GLM 5.2 Gimme 就像是最具效率的一个。所以借助这些数据,你就可以决定哪一个是适合你预算的最佳模型。另一个需要考量的因素是 token 的消耗,呃,表现更高的模型有时候会生成更多的输出 token,比如使用了更多基于思考的模型,嗯,但,呃,并不总是这样。你可以,你可以在这里看到,像 GPT-5 相对来说就比其他模型更高效。而且另一个有趣的地方是,如果你只看目录价格,你可能会发现,呃,一些模型看似是同样的价格,但如果你真正在现实世界中应用它,有些模型在完成相同的任务时会消耗更多的 token,对吧。所以实际上,我们可以在这里展示,比如,GPT-5i 尽管它有相似的价格,这个价格,呃,和 Opus 一样,但在现实世界中,它使用了更少的 token,用更少的 token 实现了相同的任务,呃,这让它比其他的更高效。而且正如你所看到的,嗯,所以总结一下,嗯,如果你正在构建一个智能体应用程序,嗯,显然你绝对应该记录你的智能体运行轨迹,去理解、去记录智能体与用户、与客户之间的所有交互,然后能够,你知道,深入挖掘这些数据以获取洞察,并衡量这些结果与你所关心的任何业务指标之间的联系,然后利用这些数据,这些真实世界的数据来为你挑选最佳的模型。呃,我们接下来的发展方向是,你知道,显然会添加许多不同的连接器,以引入更多的用户上下文,并真正为许多不同类型的智能体,比如真实代码库上的编码智能体,启用轻量级的电子邮件交互(light emails)。嗯,我们还想引入更复杂的专业用户任务,将其划分为不同的类别,以帮助你了解,呃,模型在这些类别中的表现如何。而且还有更多像是为,嗯,开发者提供的更丰富的指标,供他们用来挑选哪个模型是最好的,同时也会有评分标准,以进行更细粒度的,嗯,打分,甚至与用户协作来定义理想状态会是什么样。嗯,所以我这就讲到这里,呃。很乐意听到你们的反馈,或者如果你有任何问题,随时可以,呃,联系我们。你可以在我们的排行榜上,u arena.ai,找到更多的洞察,或者在 X 上关注我们。我们也会,你知道,定期发布技术博客文章。是的,我们还在招聘,所以,你知道,点击这个链接,或者直接在 X 上给我发私信来联系。谢谢大家。

Original English

Speaker A: produce the ranking right you can measure precisely you know which model performs better than other in this particular signal and we combine that into the final um leaderboard that you see on you know on the website. Um so um that's what you looks like um today. You see like um this video has five different signals and model performed differently across board and right now fable five is the number one models that was you know the net improvement of like 14% over the average which is the you know average of all the models followed by call opus GPD fivei high and what's interesting about this data boy is like you can look at the signal by signal um the model may be really really good at test success but sometimes weaker in terms of like you know stability in terms how do you control the model and you can see exactly like where the model is failing and so on and we are going to add you know more and more signal richer signal to capture these failure pattern. So methodologically the core idea is basically a randomized control trial where we intervene on agent component. We measure the causal effect of you know any given component on the task outcome like the signal that we care uh and the mandible basically is is like the causal effect of of the orchestrator models um that you can you know right now but this framework is general enough so we can also measure the interaction effect between different uh components for example let's say you want to measure uh tool you want to measure different harness harness or different system prompt uh and so on. So all these are possible within this framework and we're going to you know uh evaluate that too and if you are interested more technical details are published uh on our blog post. Um so um we have been tracking like I say all the major release in Asian is one of the release happened couple of weeks ago fable five in Asia arena um so if you wanted to follow us on X you will see all the you know latest release and the interesting thing about this leaderboard is because this is real data right based on millions of agentic traces you can slice it into any task distribution you care about so for example like let's say you care about you know GDP tasks this more like economically valuable professional work versus consumer use cases you can uh you can do some of the data analysis to slice the data and one you know inside here what you see is like GPD5i is actually pretty good uh in terms of like GPT sorry like GDP tasks uh and GM Gemini tends to do better in consumer use cases is so basically the the best model generally depends on uh what you're doing what you care the distribution um and on the other side is the cost right you know cost matter too you can we basically can plot these uh net improvement which is performance against the average cost to see to to help you see the parto frontier here you can see fable is the one that's the best uh cost about $10 per session and 5ifi is still very bit cheaper and GP GLM 5.2 Gimme is like the most efficient one. So you can with this data decide which one is the best model for your budget. Another dance is tokens uh higher performing model sometimes generate more output token like using more thinking model um and but uh not always you can you can see here like GPD5 is relatively more efficient than other models. And the other interesting thing here is like if you only look at the list price you may see uh some of the model is like same price but if you actually put it in the real world some of the model would use more tokens to to for the same task right. So actually we can show here like for example GBD5i although it has similar price this price uh as OPUS but in the in the real world it use less token fewer tokens to achieve the same task uh which is more efficient than the others and as you can see um so to summarize um if you are building an agentic app um obviously you should definitely be logging your agentic traces to understand to log all the interactions between agent and the user and the customers and then be able to you know look into the data mind for insights and measure the outcome links to whatever business metrics you care and use that data to real world data to choose the best model for you. Uh and what we are headed next is you know obviously going to add a lot of different connectors to bring in more user context and enable really the light emails for many different kinds of agents coding agents on real repository. Um and we also wanted to bring more complex task professional users slice that into different categories to help you understand uh how model is doing in those category and so as more like richer signal for um developers to use to pick which model is the best as well as rubrics to do more final grand um scoring and even working collaborating with the user to define what could look like. Um so that's it uh for me. would love to hear your feedback or if you have any question feel free to uh reach out. You can find more insights on our leaderboard u arena.ai or follow us on X. We also publish technical blog post you know regularly and yes we are also hiring so you know check out this link or just DM me on X to reach out. Thank you.

Closing Remarks & Sponsor Thanks

Deina Dias: 请欢迎我们的主持人,美洲 Oliver Wright 的技术总监,Deina Dias 回到舞台。大家好,非常感谢你们,也请给自己热烈的掌声,感谢大家留到最后。是的,谢谢大家。我们确实把最精彩的环节留到了最后。那么,关于创业公司大赛,我刚才是骗大家的。不是在今晚,而是在明晚,和闭幕主题演讲一起进行。所以请务必出席。我们期待在那里见到你们。所以,感谢下午那些精彩纷呈的主题演讲,还要向组织者们表达深深的谢意。我们真的拥有非常棒的赞助商。如果没有他们,这场活动就不可能举办。我们非常激动能与这么多出色的组织合作。冠名赞助商,微软。好的。好的。它在哪儿呢?好的。那么,Lav 和白金赞助商,以及我们的金牌赞助商,当然,还有我们的银牌和铜牌赞助商。感谢你们所有人。祝大家度过一个美好的夜晚,我们明早见。

Original English

Deina Dias: Please welcome back our MC, director of technology at Oliver Wright Americas, Deina Dias. Hey everybody, thank you so much and give yourselves a great round of applause for being here till the end. Yeah, thank you guys. We really truly saved the best for last. So, the startup battle, I lie to y'all. It's not tonight, it's tomorrow night along with the closing speaker notes. So please be there. We look forward to be there. So thank you for the incredible sets of talks for our afternoon keynotes and big big thank you for the organizers. We truly have incredible sponsors. The event could not have happened without them. We're incredibly excited to partner with so many wonderful organization. presenting sponsor Microsoft. Okay. Okay. Where where is it? Okay. So, Lav and Platinum sponsor and our gold sponsor and of course our silver and bronze sponsors. Thank you all. Have a marvelous rest of your evening and we'll see you tomorrow morning.

Final Video Voiceover

Speaker C: 当今世界上正在发生的一切真是不可思议。这使得他们能够解锁越来越高的自动化水平。人工智能编写代码的速度已经超过了人类审查代码的速度。包罗万象。是的。

Original English

Speaker C: It's really incredible what is going on in the world today. allows them to unlock more and more levels of automation. AI writes codes faster than humans can review it. Everything. Yeah.

📌 文中提及的人物和组织

关键字: model-evaluation agentic-workflow reliability technical-insights