人工智能工程师世界博览会总结与展望 AI Engineer 2026-07-03

开场致辞与欢迎

Announcer: 发射控制中心,准备就绪。收到。就是这里,我需要正确的感觉。女士们先生们,欢迎来到 AI 工程师世博会 (AI Engineer World's Fair)。我们很高兴能与大家一起,继续探索塑造 AI 未来的理念、技术和人才。请和我一起欢迎今天节目的主持人,Replit 的开发者关系工程师,Ralph Shabri。

Original English

Announcer: Launch control. We have a go. Roger. right here. I need the right feeling I need the right feeling. Ladies and gentlemen, welcome to the AI Engineer World's Fair. We're delighted to have you with us as we continue exploring the ideas, technologies, and people shaping the future of AI. Please join me in welcoming your MC for today's program, developer relations engineer at Replet, Ralph Shabri.

Ralph Shabri: 早上好,旧金山。大家过得怎么样?欢迎来到 AI 工程师大会第四天。哇,我今天非常兴奋能和大家在一起。这是我们有史以来规模最大的 AI 工程师活动。我们有大约 7000 名与会者。而且,天哪,当我看着这个房间时,是的,是你们,你们坚持到了第四天。所以,把掌声送给你们自己。

Original English

Ralph Shabri: Good morning, San Francisco. How's it going, guys? Welcome to AI Engineer Day four. Wow, I'm so excited to be here with you guys today. So, this is our biggest AI engineer event ever. We have about 7,000 attendees. We also, my god, when I look at this room, uh, yeah, it's you guys, you made it to day four. So, let's give it up to you guys.

Ralph Shabri: 好的。我住在瑞士的一个小村庄里,我真不敢相信我能把整个村子的人都塞进这个房间里。说到瑞士,谁是从欧洲来的?哇,你们这么多人。我看到你们战略性地坐在了空调下面。所以,伙计们,在这个炎热的国家,保持凉爽。好的,今天有湾区本地人吗?

Original English

Ralph Shabri: All right. I live in a tiny village in Switzerland, and I can't believe that I can fit my entire village right here in this room. Uh, speaking of Switzerland, uh, who's here from Europe? Wow, so many of you guys. And I see that you're strategically placed under the AC. So, uh, please guys, white in America, stay cool. All right. Anybody local to the Bay Area today?

Ralph Shabri: 好的。也有很多人。不错。老实说,伙计们,我非常兴奋能来到旧金山。我认为没有哪个地方能比这里更好地开发 AI 和谈论 AI 了。我觉得你们是世界上最酷的人,因为你们一直生活在未来。你们开着自动驾驶汽车,也有无人机送餐。而且我注意到现在都没人合上笔记本电脑了。我不知道,大家似乎都在最大化地利用时间。说到最大化,大家喜欢昨天的演讲吗?

Original English

Ralph Shabri: All right. So, many of you as well. Nice. Okay. I to be honest, guys, I'm super excited to be in San Francisco. I think there's no place to build for AI and talk about AI. And what I think I think you guys are the coolest people in the world because you're constantly living in the future. Um, you guys drive self-driving cars, but you also have food delivered by drones. And what I noticed nowadays is that nobody closes their laptops anymore. I don't know. Everybody seem to be Asian maxing. Speaking of maxing, you guys enjoyed yesterday's talks.

Audience: 是的!

Original English

Audience: Yeah.

回顾与今日议程展望

Ralph Shabri: 把掌声送给我们昨天的演讲者。好的,昨天有很多值得谈论和思考的内容。我们请到了来自 Anthropic 的 Tariq 谈论 Fable 5 以及发现新模型能力的过程。我们还有另一位来自 Sonar 的 Tariq,他谈到了代码验证。但我注意到并从演讲者那里听到的是,在使用前沿模型时,让人类保持在循环中 (stay in the loop) 仍然非常重要。

Original English

Ralph Shabri: Let's give it up for our speakers from yesterday. All right, there was so much to talk about and to think about yesterday. So we had Tariq from Enthropic talking about Fable 5 and uh the process of discovering new models capabilities. We had another Tariq from Sonar who talked about uh code verification. But what I know what I noticed and what I heard from speakers is that it's still very important to stay in the loop while working with Frontier uh Frontier models.

Ralph Shabri: 伙计们,今天的主题将是测试框架工程 (harness engineering)。今天我们为大家准备了非常强大的演讲嘉宾阵容。我们有来自领先机构的演讲者,比如 Anthropic,还有斯坦福大学。我们也有来自 DSPy 的朋友,他们将向你们介绍新协议,探讨如何将任务与模型分离,还会讨论如何扩展 AI 系统。我看过其中一些演讲,我可以告诉大家,你们今天有眼福了。

Original English

Ralph Shabri: And today guys, it's going to be about harness engineering. And we have a great lineup of speakers for you today. So we have speakers from leading organizations such as Anthropic. We have also um Stanford. We have folks from DSPI also who are going to talk to you about new protocols. They're going to talk to you about how to separate the task from the model and uh and also they're going to talk to you about how to scale AI systems. And for having seen some of these talks, I can tell you guys that you've been that you're here for a treat.

Ralph Shabri: 好的。在主题演讲之后,我们有分组会议,有很多主题供大家选择,比如软件工厂、生成式媒体、记忆等等。另外,请大家也花点时间去参观我们的展区。我们有很多很酷的周边礼品。我们的合作伙伴都在那里,去跟他们聊聊。如果你想拿些周边,就是今天了,今天是最后一天,机不可失。说到合作伙伴,请大声欢呼我们的冠名赞助商 Microsoft。好的,大家把掌声送给 Microsoft。

Original English

Ralph Shabri: All right. And after the keynote, we have breakout sessions and we have so many tracks for you guys to choose from. So we have tracks such as uh software factories, generative media, memory and so on. And please take also the time to go and check out our expo. We have so many cool swags. Our partners are there. uh go uh talk to them and and uh and please if you if you need to take any swag, it's gonna be this is today. This is the last day. So, uh it's now or never. Uh speaking of partners, please a big shout out to our presenting sponsor, Microsoft. All right, let's give it up for Microsoft, guys.

Ralph Shabri: 同时也把掌声送给我们所有的实验室和白金赞助商。也请继续为我们的金牌、银牌和铜牌赞助商鼓掌。这非常重要,伙计们,如果没有这里每一家公司的支持,这次活动就不可能举办。好的。再说一次,我们将讨论测试框架工程。但在介绍我们的第一位演讲者之前,我想让大家知道的是,作为与会者,你们对今天演讲的质量有着巨大的影响。我们的演讲者一直在做这些事情,他们对自己的领域了如指掌。但我们希望他们一登上舞台就能感受到房间里的能量。好吗?

Original English

Ralph Shabri: Also, let's give it up for all our lab and platinum sponsors. Please keep it up for also for our gold sponsors for our silver and and and bronze sponsors. This is super important guys. Without the support of each one of these companies here, this event wouldn't be possible. All right. So again, we're going to talk about harness engineering. But before we introduce our first speaker, what I want you to know guys is that you as attendees, you have huge power over the quality of the talk that you're going to be attending today. Okay? So, our speakers have been doing this forever. They know their stuff inside out. But what we want them to do is to feel the energy in the room the second they hit the stage. All right. Okay.

Ralph Shabri: 所以,我们先来练习一下。数到三的时候,我希望你们把我当成演讲者,像欢迎摇滚巨星一样欢迎我。好,一,二,三。好的。不错。那就把这种热情保持给我们的下一位演讲者。好,废话不多说,让我们请出今天的第一位演讲者。她是 Amplify 的合伙人。请大家和我一起欢迎 Bar 上台。

Original English

Ralph Shabri: So, we're going to practice that for a second. So, at the count of three, I want you to imagine as if I was a speaker and I want you to welcome me as a rockstar. Okay. So, one, two, three. All right. Nice. So, let's keep it up for our next speaker. Okay. So without further ado, please let's introduce our first speaker of the day. She is a partner at amplify. Please join me welcoming to the stage bar your own.

AI 工程现状报告

Announcer: 现在登台的是 Amplify 的合伙人 Bar。

Original English

Announcer: Now joining us on stage is the partner at Amplify Baron.

Bar: 太棒了。你们练习得很好。我感受到了满满的爱。我们开始吧。正如你们刚才听到的,我叫 Bar。我每年都会针对 AI 工程的现状进行一项调查。做这项调查有趣的地方在于,当你还在制作幻灯片的时候,这个领域就已经发生了变化。就在过去的一周里,我们见证了前沿模型的发布被当作国家安全事件来对待。据报道,Meta 正在探索出售 AI 算力。等我下台的时候,说不定又会发生别的事情。所以,如果我在台上错过了什么重大公告,请在会后告诉我。

Original English

Bar: Fantastic. You did a great job practicing. I feel very very loved. Um, let's get started. So, like you just heard, my name is Bar. I run a survey every year on the state of AI engineering. And the funny thing about running a survey on the state of AI engineering is that the field changes as you make the slides. Just in the past week, we've had Frontier releases treated like national security events. Meta reportedly exploring selling AI compute. By the time I get off stage, maybe something else will happen. So, if I miss a major announcement while I'm up here, please come find me after.

Bar: 但这正是我们每年进行调查的原因——拨开迷雾,花点时间退一步,了解 AI 工程师们实际上都在做些什么。今年,我们非常高兴能首次与 Notion 和 Vercel 合作开展这项调查。简单介绍一下我,这是最无聊的一页幻灯片。我是 Amplify 的投资合伙人。非常幸运能够投资由 AI 工程师创立并为 AI 工程师服务的公司。我会做出和每年一样的承诺,那就是:少谈我自己,多看柱状图。所以,让我们直接进入充满柱状图的环节。

Original English

Bar: But that's exactly why we run the survey every year to cut through the noise, take a moment, step back and understand what AI engineers are actually doing. Uh for the first time this year, we were thrilled to partner with Notion and Verscell to run this survey. Very quickly on me, uh this is the least interesting slide. I'm an investment partner at Amplify. Very lucky to invest in companies built by and for AI engineers. And I'll make the same promise that I make every single year, which is short time on bar, long time on bar charts. So, let's get right into it with lots of bar charts.

Bar: 首先,我们来谈谈……也许大家举手表决一下,你们填写调查问卷了吗?这真是一个庞大的群体。好的。是的,我看到前排的你了。如果你填写了,非常感谢。如果你没填,我会在 2027 年找到你的。不过说真的,这项调查之所以能成,仅仅是因为你们中的一千人贡献了时间。所以,谢谢你们。我们今年有 1048 名受访者,这是很多 AI 工程师了。准确地说,受访者不仅限于 AI 工程师。我相信正如大家在大会上看到的那样,每年我们都发现 AI 工程更多地是一门学科,而不仅仅是一个工作头衔。它涉及创始人、首席技术官、工程师、产品人员,以及来自不同公司规模和经验水平的人。

Original English

Bar: First, let's talk about well, maybe raise your hand. Did you fill out the survey? This is a very large group. Okay. Yes, I see you in the front. Um, if the answer is you, thank you so much. If the answer is not you, I will find you in 2027. But genuinely, this only exists because a thousand of you gave your time. So, thank you. We had 1,048 respondents this year, which is a lot of AI engineers. And to be precise, this is not just AI engineers. As I'm sure you see at the conference, every year we see that AI engineering is more of a discipline than a job title. It touches founders, CTO's, engineers, product people, folks across company sizes and experience levels.

Bar: 这种差异也体现在经验上。连续第三年,我们看到了同样的模式,即倾向于资深工程师,但他们涉足 AI 领域的时间较短。在拥有超过 10 年软件开发经验的人中,超过一半的人 AI 经验在三年或三年以下。这很合理,这些都是非常有经验的工程师,正在实时学习一种新的范式。而最新的一批受访者,即那些刚刚开始从事工程的人——新工程师的中位数所拥有的 AI 经验,几乎和拥有 10 年经验的资深软件工程师的中位数一样多。所以这些最新的工程师从未接触过没有 AI 的软件世界。

Original English

Bar: And that range shows up in experience too. Um, for the third year running, we see the same pattern which is skew towards senior engineers but newer to AI. Of those with over 10 years of software experience, over half have three years or less of AI experience, which tracks uh these are very experienced engineers learning a new paradigm in real time. And the newest cohort, the ones who just started uh engineering the median new engineer has nearly as much AI experience as the median 10-year software veteran. Uh so the newest engineers have never known software without this.

Bar: 但是,做 AI 并不意味着只做一件事。我们谈到了所有这些不同的头衔、所有这些不同的角色。在深入探讨模型和智能体之前,我有一个更基础的问题,那就是:当人们说他们在工作中做 AI 时,他们实际上在做什么?首先,让我们从模态开始。我们问:你在工作中积极使用哪些模态进行开发?大家能猜到吗?文本占据了主导地位。我知道。留着你们的掌声。

Original English

Bar: But doing AI doesn't mean one thing. We talked about all these different titles, all these different roles. Before we get into models and agents, I have a more basic question which is when people say they're doing AI at work, what are they actually doing? So first up, like to start with the modalities. We asked which modalities are you actively building with at work? Can anyone take a guess? Text dominates. I know. Hold your applause.

Bar: 但这张图表里我总觉得非常有意思,也是我一直关注的一部分,就是“不,我没有使用这种模态”与“我没有使用,但我计划使用”的比例。我称之为“采用意向比例”。在今天没有使用某种模态进行开发的人中,有多少人说他们计划使用它?今年音频的采用意向最为强烈。在今天没有使用音频进行开发的 AI 工程师中,有高达 56% 的人表示他们计划在他们构建的 AI 应用程序中采用音频。这并不是一个全新的信号。去年,音频在各模态中也具有最高的采用意向,但当时只有 37%。所以音频继续领先并保持高度关注,但这种兴趣现在正在加速。音频领域已经出现了明显的趋势,但如果我们看看……

Original English

Bar: Um, but one piece of this chart that I always find very interesting and I always look at is the ratio of nope, I'm not using this modality to I'm not using it, but I do plan to. I call this the intent to adopt ratio. Of the people who are not building with a modality today, how many say they plan to use it? And audio has the strongest intent to adopt this year. Among AI engineers who are not building with audio today, a whopping 56% say they plan to adopt it in the AI applications they build. And this is not a brand new signal. Last year audio also had the highest intent to adopt across modalities, but 37%. So audio continues to take the lead and have high interest, but that interest is accelerating now. There has been an audio swing, but if we look at

图像生成的使用激增

Speaker A: 在去年的调查中,最大的变化其实是使用图像生成的人数激增。使用生成式AI来处理图像并且感觉非常满意(really good about it)的受访者比例,从去年的18%翻倍到了今年的36%。如果你看看我们在同一时期发布的产品,这就很合理了。在过去一年加上调查期间,我们推出了 Nano Banana、Nano Banana 2、Chat GPT Images 2.0 等模型。产品变得越来越好。过去那些感觉只是在高效生成“残缺手指”的工具,现在正越来越多地成为实际工作的一部分。虽然音频领域可能有最强的采用意愿,但图像生成向我们展示了当一种模态跨过可用性门槛时会发生什么。因此,我很兴奋能每年继续观察这些采用曲线。我认为今年我们会看到很多进展。

Original English

Speaker A: what changed most from the last year in the survey, the biggest jump is actually in people using image generation. The share of respondents using generative AI for images and feeling really good about it doubled from 18% last year to 36% this year. Makes sense if you look at what we launched in the same window. Over the past year plus survey time, uh, we've had models Nano Banana, Nano Banana 2, Chat GPT Images 2.0. The products have gotten much better. What used to feel like an efficient way to generate cursed hands is just increasingly becoming a part of real work. Audio may have the strongest intent to adopt, but image generation shows us what happens when a modality crosses that threshold. So, I'm excited to continue watching these adoption curves every single year. I think we're going to see a lot this year.

模型的使用与选择

Speaker A: 接下来谈谈模型。在座有谁经常刷 Twitter?好的。是的。我想这里的观众都非常沉浸在 Twitter 里。如果你在 Twitter 的这个圈子里花过时间,你就会看到过去几个月有很多关于开源权重(open-weight)模型的讨论,而且我认为明年我们会看到更多。所以我们问了受访者:你们在生产环境中实际使用的是什么模型?94% 的人使用闭源模型。45% 的人使用开源权重模型。但问题是,对于大多数情况来说,开源权重模型并没有取代闭源模型,至少目前还没有。在使用开源权重模型的受访者中,超过 90% 的人也在使用闭源模型。所以它们看起来更像是一种补充。团队在混合搭配使用它们。

Original English

Speaker A: Uh, now models. If you've Who here spends time on Twitter? All right. Yes. I imagine this is a very Twitter pilled uh crowd. If you spend any time on Twitter in this in this circle, you've seen a lot written about openweight models these past few months and I think we'll see it even more in the next year. Um, so we asked what models are you actually using in production? 94% use closed models. 45% are using openweight models. But here's the thing, you know, openweight models are not replacing closed models for the most part, at least not yet. The respondents using openweight models, over 90% of them are also using closed models. So they're looking like an augmentation. Teams are mixing and matching.

Speaker A: 我们还进一步调查了选择模型时的前三大考虑因素。如果你在选择模型,什么对你来说是重要的?尽管“开源与闭源”的争论热度很高,但这并不是驱动模型选择的因素。只有 5% 的受访者将其作为前三大考虑因素。真正重要的东西其实更简单直接,那就是质量。质量占据主导地位,紧随其后的是智能体能力(如工具调用),而成本也和它并列。说到钱,我们等会儿再回到成本话题。我发现非常有趣的一点是,可靠性(reliability)并没有排在前面。只有五分之一的人提到了可靠性。这并不意味着团队不再关心可靠性。对此数据有不同的解释方式。我的猜测是,它更有可能已经成为一个门槛要求,大家选择的模型都已经足够可靠了,因此决策的重心已经向上移,变成了对质量、能力和成本的考量。不过我们可以在会后探讨。好了,所有关于模型的情况都在这里汇集了。

Original English

Speaker A: We also asked just to double click on this for the top three considerations when choosing a model. If you're choosing a model, what is important to you? Um, and despite the airtime of the open versus closed, it's not what drives model choice. It was a top three consideration for only 5% of the respondents. What matters is actually more straightforward. It's quality. Quality dominates. followed by agentic capabilities like tool calling and cost tied right with it. We'll money money. We'll get back to that. Um, one thing that I found very interesting is that reliability is not near the top. Only one in five named reliability. That doesn't mean teams stopped caring about reliability. Uh, there are different ways to interpret this data. I guess is that it's more likely to become a threshold requirement and the models they're choosing are reliable enough so the decision moves up the stack outside of certain circumstances to quality capability cost but we could talk after all right so here's where the model story all comes together

工具层的标准化与成本约束

Speaker A: 就像我之前说的,团队并不是只选定一个模型就完事了。前面我展示过,87% 的团队在使用不止一个模型——这与标准化背道而驰。而且,他们为特定任务选择模型的方式也各不相同。最常见的是按任务类型进行路由。有些人会运行多个模型,然后比较输出结果。有些人则根据成本进行路由。但是不同的模型擅长不同的事情。有趣的是,超过一半的受访者表示,他们的组织正开始在较少的 AI 工具上实现标准化。他们用灵活性换取标准化。其中一部分人的做法是混合的:他们说正在某些层面上进行标准化,而在其他层面上保持灵活性。但这里的核心结论是,我们正处于平台和工具的早期“大标准化”阶段,而不是模型的标准化。

Original English

Speaker A: like I said teams are not choosing one model and calling it a day earlier I showed that 87% of teams are using more than one model uh the model that's the opposite of standardization. Uh and the way that they choose models for given tasks varies. Most popular is routing by task type. Some run multiple models, compare outputs. Some route based on cost. Uh but models are good at different things. What was interesting was that more than half of respondents said that their organizations starting to standardize on fewer AI tools. They're trading standardiz flexibility for standardization. A share of those are mixed. They say they're standardizing on some layers while staying flexible on others. But the headline here is that there's we're in the early great standardization of the platform and tools, not the models.

Speaker A: 好了,这张幻灯片会让任何在过去一年里看过 AI 账单的人连连点头。事实证明,无限的智能仍然伴随着基于使用量的账单。一旦团队开始管理多个模型和 AI 工作流,接下来的问题就是成本。成本现在已经成为一等公民级别的工程约束(first-class engineering constraint)。我们从数据中看到了这一点。40% 的受访者表示成本经常会影响他们使用 AI 的积极程度(ambitiously),另有 36% 的受访者表示成本有时会产生影响。好吧,这非常直截了当。所以,总共有大约四分之三的受访者正在根据成本调整他们的 AI 使用量,剩下的四分之一可能有公司买单吧。这可能令人惊讶,或者可能很明显,但在 12 个月前却并非如此。“压榨Token”(Token maxing)很酷。能够找到真实的用例令人惊叹,但成本正在成为当今产品决策中非常大的一部分。这也在监控中体现了出来。大家在生产环境中监控的指标里,成本和 Token 使用量是他们关注的第二大重点。它正像 SLA 一样被监控,重要性仅次于质量本身。

Original English

Speaker A: All right, this is the slide where anyone who's opened an AI bill in the last year starts nodding. So it turns out that infinite intelligence still comes with a usagebased bill. Once teams are managing many models and AI workflows, the next question becomes cost. Cost is now a first class engineering constraint. We see this in the data. 40% of respondents say that cost regularly shapes how ambitiously they use AI and another 36% say that it sometimes does. Well, this is pretty straightforward. So all in about uh three out of four respondents are adjusting their AI usage based on cost and maybe the fourth has a company card. That might be surprising or maybe it's obvious but 12 months ago it was not. Token maxing is cool. Being able to find real use cases is amazing but cost is becoming a real big part of the product decision today. And it shows up in monitoring too. What folks are monitoring in production includes cost and token usage as the number two thing they watch for. It's being monitored like an SLA right under quality itself.

智能体的崛起与挑战

Speaker A: 这将我们引向了所有项目中最大的一项:智能体(Agents)。关于智能体,我们已经讨论了一段时间。今年,正如大家所见,正如你们今天将看到的,以及前几天所看到的,大家会谈论很多关于框架工程(harness engineering)的内容。智能体正在逃离“演示环境(demo world)”。因此,我们询问了受访者:他们的智能体通常拥有什么级别的工具权限。正是在这里,智能体开始显得更加真实。有两件事同时发生:首先——我不认为这令人惊讶——与去年相比,使用智能体的团队要多得多。今年有 95% 的受访者表示他们在使用智能体(这个比例似乎挺高),大约是去年的两倍。其次,在那些使用智能体的团队中,这些智能体拥有写入(write)权限的可能性要大得多。去年,52% 用智能体构建应用的人表示,他们的智能体实际上可以写入数据。今年,这个数字变成了 89%。因此,当你把这两个变化结合起来——更多团队使用智能体,以及这些智能体中有更多拥有写入权限——在所有受访者中(再次强调,这是一项调查),使用具有写入权限的智能体的比例相比去年增长了三倍多。所以这才是真正的巨大转变。智能体不再仅仅是阅读、总结和起草。它们正在系统内部采取行动。

Original English

Speaker A: Which brings us to the biggest line item of them all. Agents. Well, we've been talking about agents for a while. Uh this year, as you've seen, as you'll see today, as you've seen in previous days, you're going to talk a lot about harness engineering. They're escaping demo world. So, we asked respondents what level of tool permissions their agents typically have. And this is where agents start to look more real. There are two things happening at once. First, and I don't think this is surprising, relative to last year, there are far more teams using agents. This year, 95%, this seems high to 95% say they're using agents, roughly double last year. Second, amongst the teams that are using agents, those agents are much more likely to have write access. Last year, 52% of folks building with agents said their agents could actually write data. This year, that number is 89%. So when you combine these two shifts, more teams using agents and more of those agents having write permissions, the share of all the respondents and again it's a survey using write enabled agents is up more than three times relative to last year. So this is really the big shift. Agents are no longer reading, summarizing, drafting. They're taking actions inside of systems.

Speaker A: 这就引出了一个显而易见的问题:我们是如何控制这一切的?嗯,目前使用的工具相当生硬。人们如今控制智能体的方法有很多。排名前两位的分别是“人在回路”(human-in-the-loop)审批和权限限制。这些直觉是对的,但这和用来管理实习生的是同一套工具。排名往下看,结果就比较分散了。任务分解(Task decomposition)、检索(retrieval)、记忆(memory)、沙盒(sandboxing),大家什么都在尝试。没有人确立智能体的控制层。记忆和持久上下文是我目前非常密切关注的一个领域。我认为它在明年会有很大的发展。当智能体失败,或者当人们抱怨智能体没能更精确时,通常是由于“思考”(thinking)出了问题,而不是“管道”(plumbing)出了问题。所以,有三分之二的人表示,“幻觉”(hallucination)或在任务中途丢失上下文,是让他们最感到沮丧的事情。

Original English

Speaker A: And that raises the obvious question, how are we controlling all of this? um with pretty blunt instruments. Uh very there are many ways that folks are controlling agents today. The top two are human in the loop approvals and gating permissions which are the right instincts but kind of the same toolkit you'd use to manage an intern. Below that the results scatter. Task decomposition, retrieval, memory, sandboxing, people are trying everything. Nobody has settled the control layer for agents. uh memory and persistent context is one that I'm watching very carefully right now. I think it's going to evolve a lot in the next year. And when agents fail or when people complain about agents failing to be more precise, it's usually the thinking, not the plumbing. So, uh you know, like twothirds say that hallucination or losing context mid task is what frustrates them the most.

基础设施层与自建vs购买

Speaker A: 好了,智能体已经进入了实际应用,这是一个好时机来看看大家底层到底在运行些什么。所以让我们看看技术栈(stack)。我们问了受访者:你们技术栈中最大的挑战是什么?我每年问这个问题时,排名第一的答案总是评估(eval)。所以评估依然领先,和往常一样,但差距非常小。这个差距正在缩小。在这里我要说出一个心照不宣的事实:参与调查的、在座的 96% 的人在技术栈方面都遇到了问题,只是大家在“具体是哪个问题”上无法达成一致。所以,如果你在决定接下来要构建什么,如果你对基础设施感兴趣,那么这种分散的数据分布就是你的路线图。首要挑战——如何评估你的 AI 输出——需要许多不同的方法,但一如既往,“靠感觉审查”(vibe review)依然是第一位的。有些事情是一贯如此的,我们会看看它们随着时间的推移是否会改变,但目前它们并没有改变。

Original English

Speaker A: All right, so agents are out in the wild, which makes it a good time to look at what everyone's actually running underneath. So let's take a peek at the stack. Um, we asked, what is the biggest challenge in your stack? Every single year that I ask this, the answer, the number one answer is eval. Um, so eval lead here, same as always, but by a very thin margin. Like that margin is getting smaller. And I'll say the quiet part here, which is that 96% of the people in the survey in this room have a problem with the stack. You just can't agree on which one. Um, so if you're deciding what to build next, if you're interested in infrastructure, that scatter is the map. And the leading challenge, how to evaluate your AI outputs requires many different methods, but as always, the vibe review is number one. So there are some consistent things that we'll see if they change over the time, but they they have not changed.

Speaker A: 好的,这很有趣。在技术栈的八个层面上,我们询问了大家:哪些是自建(build),哪些是购买(buy)的?再次说明,也许公司信用卡在其中发挥了作用,但对于技术栈的每一层来说,都有广泛的分布和混合情况,并且有一些清晰的结论。首先是推理(inference)和模型服务(model serving),这是人们购买最多的一层。许多人不想自己构建推理基础设施,这也合情合理。提示词管理(prompt management)恰恰相反。61% 的人自己构建它。显然,每个人的提示词都是独一无二的。对于很多产品逻辑、提示词、RAG(检索增强生成)和评估来说,这也是成立的。它们相对来说更倾向于留在公司内部研发。微调(fine-tuning)的情况最清晰:还没到时候。大多数人根本没有这方面的需求。而且大家的锁定效应(locked in)非常明显。那些选择购买的人,不太会考虑去自建;而那些选择自建的人,也不太会去考虑

Original English

Speaker A: Okay, this is interesting. So across eight layers of the stack, we asked what do people build versus buy? Um, again, maybe the corporate card is is going to play a part in this, but there is a wide range and mix for every layer of the stack and a few clear takeaways. So the first is that inference and model serving is the layer that people buy the most. Many people don't want to build inference infrastructure and fair enough. Uh prompt management is the opposite. 61% build it themselves. Um apparently everyone's prompts are special. And this is true of a lot of the product logic prompts rag eval. They tend to stay inhouse on a relative basis. fine-tuning is the clearest. Not yet. Like most people don't have it at all. And uh folks are pretty locked in. So those who bought aren't looking as much to build. Those who built aren't looking

AI应用栈调研报告总结

Speaker A: 以上就是我们在应用栈中使用情况的核心结论。你们中的许多人都在团队中工作,正如我们一开始所说,这些团队的规模从单人创始人到大型企业不等。那么这对团队产生了什么影响呢?

Original English

Speaker A: as much to buy. but those are the core takeaways from the usage in our stack. So many of you work on teams and like we said at the start these range from solo founders to large enterprises. What is this doing to teams?

Speaker A: 请记住,这是一个以开发者为主的调查样本,但在开发者群体中,整体的氛围非常好。我相信如果你看看你的左边和右边,你会感觉到大家的情绪都很不错。97%的受访者报告说,这对他们的组织产生了净正面的影响。

Original English

Speaker A: And remember this is a builderheavy sample but among builders the vibes are good which you know I'm sure if you look to your left and your right you're feeling that the vibes are pretty good. 97% report a net positive effect on their organization.

Speaker A: 最大的影响其实不仅仅是速度。它是更低成本的试错、更多的实验、更多的原型以及更多的下注。它不仅仅让工程师变得更快,而是让尝试新事物的成本几乎降为零。因此,这也造就了一些非常开心的开发者。

Original English

Speaker A: The top effect isn't really just speed. It's cheaper failure, more experimentation, more prototypes, more bets. It didn't just make engineers faster, but it made trying things nearly free. And so there's some happy campers as a result of that.

Speaker A: 但这并不是完全免费的。你知道,世上没有免费的午餐,任何事情都是如此。因此,同样是这个增加实验频率的工具,也增加了代码审查的负担。这两者可以同时存在。超过九成,即十分之九的受访者都在某种程度上感受到了负面的下游影响。

Original English

Speaker A: But it's not free free. You know, there's no free lunch as nothing is. So the same tool that increases experimentation also increases review burden. Both can be true. And you know over nine and 10 respondents are feeling negative downstream effects in some way.

Speaker A: 其中最常见的情况,也就是在本次会议、网上以及任何你能看到AI工程师的地方被广泛讨论的,就是深厚技术栈能力和对代码库理解能力的退化。这些都是低成本代码生成的后果。

Original English

Speaker A: The most common ones being widely discussed at this conference online and anywhere that you see AI engineers erosion of deep technical skills and understanding of the codebase. And these are consequences of cheap code generation.

Speaker A: 并且组织架构的划分也真切地感受到了这一点。有非常多的人,高达81%的受访者表示,AI正在模糊他们作为工程师与产品设计和营销角色之间的界限。

Original English

Speaker A: And the org chart is really feeling it. So many folks, 81% are saying that AI is blurring the line between their role as engineers and product design and marketing.

Speaker A: 这些统计数据让我感到震惊。你感受最深的地方是软件发布,这曾经完全是工程师的领域。我知道人们在谈论“直觉编程”(vibe coding),以及它如何让不同角色的更多人比以往任何时候都更容易上手,但今天超过三分之一的团队有非开发人员在发布功能,这对我来说相当疯狂。

Original English

Speaker A: These stats shocked me. where you feel it the most is shipping software once exclusively the engineers domain. I know folks talk about vibe coding and how that's accessible to more folks than ever before in different roles, but today over a third of teams have non-developers shipping features, which was pretty wild to me.

Speaker A: 这其中大部分是较小规模的、主要是内部的功能,但有17%的受访者表示,非开发人员正在定期跨技术栈发布面向客户的功能。即使非开发人员没有在发布软件,三分之一的团队也看到他们构建了非常有用的东西,比如原型、前端模拟等等。

Original English

Speaker A: Mostly smaller, mostly internal, but 17% say that non-developers are regularly shipping customer facing features across the stack. And even when non-developers aren't shipping, a third of teams see them building really useful things, prototypes, front-end mocks, and more.

Speaker A: 因此,发布软件不再仅仅局限于工程师。我们知道这一点,但它被推进的程度超乎了我的预期。

Original English

Speaker A: So shipping software is not gated on being an engineer. We knew this, but the extent to which it's being pushed is higher than I expected.

对未来的预测与下注

Speaker A: 好了,那么这一切将走向何方呢?我们总是要求人们进行快速的下注。所以让我们来谈谈这些结果。首先来看现在的状况。76%的人表示AI提高了他们的工作满意度。这对现场的大多数人来说是个好消息。我希望你们就像 Alphaba 和 Glenda 说的那样,“我希望你们现在很开心”。

Original English

Speaker A: All right, so where does all of this go? We always ask people to place bets rapid fire. So let's talk about those results. so present tense first. 76% say AI boosted their job satisfaction. So that's good for most of this crowd. I hope you're as Alphaba and Glenda say, I hope you're happy now.

Speaker A: 这很棒。但是59%的人担心当今的AI代码会产生长期的负债。只有三分之一的人称软件工程为一个“已解决的问题”。虽然当我和大家交谈时,有时他们定义软件工程的方式有所不同。所以你们可以自行解读这个数据。

Original English

Speaker A: that's great. But 59% fear today's AI code creates long-term liabilities. Only a third call software engineering a solved problem. Although when I have conversations with folks sometimes the way in which they define software engineering is different. So you can read into that stat as you will.

Speaker A: 总结来说就是:更快乐、更快,但要拥抱持续维护的负担(maintenance build),而且人们对招聘的未来走势感到不确定。

Original English

Speaker A: happier faster but embracing the maintenance build is the TLDDR and people are unsure what's going to happen with hiring.

Speaker A: 至于五年的预测,我们看到67%的人预计一家领先的实验室将在未来五年内宣布实现AGI。请注意这里的措辞。我们说的是“宣布(Will)”,我们问的是关于新闻发布会的,而不是实际的成就。所以,他们会宣布吗?会。那意味着什么?不确定。

Original English

Speaker A: And for the five-year bets we have 67% expect a leading lab will declare AGI in the next five years. Note the wording. We said, "Will," we asked about the press release, not the achievement. So, will they declare it? Yes. What does that mean? Not sure.

Speaker A: 只有9%的人打赌Transformers在五年后仍将是业界最先进(state-of-the-art)的技术。大多数人都不确定。但这个结果很有趣。

Original English

Speaker A: only 9% bet on Transformers being state-of-the-art in 5 years. Most are unsure. but that was interesting.

Speaker A: 然后是我最喜欢的一个问题,未来会有更多的AI算力在太空中还是在陆地上?36%的人选是。38%的人选否。这是调查中最具分歧的一个关于外太空的问题。

Original English

Speaker A: And then my favorite, will there be more AI compute in space or on land? 36 yes. 38 no. The most divisive question in the survey is about outer space.

2026年影响回顾

Speaker A: 我向你们承诺过会有很多条形图,刚才的信息量非常大。所以回顾一下我们2026年的影响报告,整体上是极其积极正面的。图像生成的使用量翻了一番,或者说令人满意的图像生成翻了一番,而音频技术则有着最高的采用意愿,这与去年相同。

Original English

Speaker A: I promised you a lot of bar charts and that was a lot of information. So a review or our 2026 wrapped impact is overwhelmingly positive. Image gen doubled or happy image gen doubled while audio has the highest adoption intent the same as last year.

Speaker A: 成本真正成为了一个首要限制因素,我们在监控领域,以及那些正在构建AI产品的雄心勃勃的开发者们的行为方式中,到处都能看到这一点。开放权重模型作为一种补充增强,但它们并没有完全取代现有的东西。因此,我们正在看到一个多模型(multimodel)的未来,同时伴随着技术栈的整合。

Original English

Speaker A: cost really became a first class constraint and we see that everywhere in monitoring in how ambitious folks that are going out and building AI products are behaving. Open weights augment but they don't replace. So we're seeing a multimodel future with a consolidation of the stack.

Speaker A: Agent 获得了比以往任何时候都多的写入权限,相对于去年增加了两倍,而安全护栏则仍然相当原始。并且,推理(inference)是一个买方市场。所有更接近产品逻辑的东西都倾向于相对更多地保留在公司内部。

Original English

Speaker A: Agents got right access more than ever before, tripling relative to last year, while the guardrail stayed pretty primitive and inference is the buy market. Everything closer to product logic tends to relatively stay more in-house.

Speaker A: 作为一个AI工程师,现在是一个非常激动人心的时代。我迫不及待地想看看明年会如何发展。你们可以在上面的链接里找到完整的报告。里面包含了每一张图表,以及今天由于时间关系我们没能展示的一些切片数据。

Original English

Speaker A: It is a very exciting time to be an AI engineer. I cannot wait to see how the next year unfolds. So, you can find the full report in the link up here. Every chart plus some cuts that we didn't have time for today.

Speaker A: 我不会要求你们去填写一份关于这份调查的调查问卷,但如果有你们希望在2027年被记录下来的内容,或者你们好奇的事情,你们可以在互联网上找到我。我很容易被找到。非常感谢大家。我们明年再见,或者正如你们中36%的人所说,也许我们在地球轨道上见。谢谢你们。

Original English

Speaker A: I won't ask you to fill out a survey about the survey, but if there's something that you want on the books for 2027, something you're curious about, you can come find me here on the internet. I'm easy to spot. Thank you so much. we will see you next year or per 36% of you, maybe in orbit. Thank you.

Speaker A: 请欢迎斯坦福大学名誉教授,John Sterhout 上台。

Original English

Speaker A: Please welcome to the stage the professor emmeritus at Stanford University, John Sterhout.

AI应用与网络延迟

John Sterhout: 早上好。真的很高兴能在这里谈论AI应用在网络方面的表现,并特别要证明“延迟”是非常关键的,并且在未来它可能会变得更加重要。

Original English

John Sterhout: Good morning. It's really great to be here to talk about the network side of AI applications and in particular to make the case that latency matters and it's probably going to be mattering more in the future.

John Sterhout: 但我只想说,这对我来说是一次非同寻常的演讲。我以前从来没有在一个有烟雾发生器的礼堂里做过演讲。我想这真是一个非常具有旧金山特色的体验。

Original English

John Sterhout: But I just want to say this is a talk is unusual for me. I've never before given a talk where there are fog generators in the auditorium. Just a really San Francisco experience, I guess.

John Sterhout: 众所周知,AI工作负载依赖于极其出色的网络性能,才能实现其自身的性能。当然,这是因为工作负载如此庞大,以至于它们必须分布在多台机器上,然后你必须在这些机器之间进行通信。

Original English

John Sterhout: So it's well known that AI workloads depend on really great networking performance in order to achieve their own performance. And of course that's because the workloads are so large that they have to be distributed across machines and then you have to communicate between the machines.

John Sterhout: 但我今天想谈论的是,这些工作负载似乎正在发生变化。所以我希望在接下来的15到20分钟里做三件事。

Original English

John Sterhout: But what I want to talk about today is that it seems that those workloads are changing and so I hope to do three things over the next 15 or 20 minutes.

John Sterhout: 首先是让你们相信,工作负载事实上正在改变,而且过去那些完全由大型数据传输主导、其中吞吐量(throughput)是唯一重要指标的工作负载,现在我们看到了越来越多较小的数据传输,在这些场景下,延迟(latency)变得至关重要。

Original English

John Sterhout: First to convince you that in fact the workloads are changing and that whereas the workloads used to be completely dominated by large transfers where throughput is the key metric that matters that we're seeing more and more smaller transfers where the latency is crucial.

John Sterhout: 我希望做的第二件事是让你们相信,诸如 TCP 和 RDMA 等传统协议非常不适合这种环境。它们最初并不是为这种环境设计的,而且不幸的是,当你混合使用小消息和大消息时,它们会遭受非常高的尾部延迟(tail latency)。我会稍微谈谈为什么会这样。

Original English

John Sterhout: The second thing I hope to do is to convince you that legacy protocols like TCP and RDMA are poorly suited to this environment. They weren't designed for this environment and unfortunately they suffer from very high tail latency when you mix small messages with large ones. I'll talk a little bit about why that's the case.

John Sterhout: 第三,我想介绍一下 HOMA,这是我们在斯坦福大学开发的一种新协议,它实际上是经过彻底的全新设计(clean slate redesign)来处理像这样的数据中心工作负载。事实上,它在这些工作负载上表现得非常好,并且可以将尾部延迟降低一个数量级甚至更多。所以我将告诉大家一点关于 Homa 的情况。那么我们开始吧。

Original English

John Sterhout: Then third, I'd like to introduce HOMA which is a new protocol we've developed at Stanford that actually was designed in a clean slate redesign to handle data center workloads like these and in fact it does quite well on those workloads and can reduce tail latency by an order of magnitude or more. So I'll tell you a little bit about Homa. So let's dive in.

John Sterhout: 首先是工作负载。从历史上看,AI 工作负载一直由机器之间的海量数据传输组成,而这正是过去唯一重要的事情。用于权重梯度等各项任务的以千兆字节(Gigabytes)计的数据。在这些工作负载中,你真正关心的是吞吐量,也就是你每秒可以通过管道传输多少千兆比特(gigabits)。

Original English

John Sterhout: First, workloads. Historically, AI workloads have consisted of enormous transfers between machines, and that's all that really mattered. Gigabytes of data for things like weight gradients and and so on. In these workloads, what you really care about is throughput, how many gigabits per second you can pump through the pipes.

John Sterhout: 对于网络来说,这些是相对容易的工作负载,因为即使建立连接和开始传输需要花费一些时间也没有关系。传输过程持续的时间如此之长,以至于真正重要的只有吞吐量。

Original English

John Sterhout: And these are relatively easy workloads for networks because if it takes a while to set up the connection and start the transfer, it doesn't matter. The transfers go on for so long that all that really matters is the throughput.

John Sterhout: 因此,在这些环境中,TCP 和 RDMA 的表现相当不错。顺便说一句,当我说 RDMA 时,我真正的意思是 RoCE (RDMA over Converged Ethernet,基于融合以太网的 RDMA),这是当今在大多数情况下被 RDMA 使用的底层传输协议。

Original English

John Sterhout: And so in these environments, TCP and RDMA perform pretty well. by the way, when I say RDMA, what I really mean is Rocky RDMA over converged Ethernet, which is the underlying transport that's used by RDMA for most purposes today.

John Sterhout: 所以无论如何,旧的工作负载是大型传输,吞吐量最重要。传统协议工作得相当不错。

Original English

John Sterhout: So anyhow the old workloads big transfers throughput matters. the legacy protocols work pretty well.

John Sterhout: 然而,现在看来工作负载正在发生变化。它们正变得更加颗粒化,有着更小的计算块和更小的数据交换。这似乎在推理(inference)领域和 Agent 相关的工作负载中尤为真实。而在训练工作负载中情况并非如此,训练仍然是海量的数据传输。

Original English

John Sterhout: However, it appears that the workloads are changing. They're becoming more granular with smaller chunks of computation and smaller exchanges of data. And this seems to be particularly true in the world of inference and also in agentic workloads. Not so much for training workloads are still massive transfers.

John Sterhout: 所以正在发生的事情是,越来越多地出现了小消息交换,通常用于元数据和协调之类的操作,例如检查特定条目是否存在于分布式的 KV 缓存(KV cache)中,或者在计算周期结束时进行屏障同步(barrier synchronization)。

Original English

John Sterhout: And so what's happening is that more and more there are small message exchanges typically for things like metadata and coordination such as checking to see if a particular entry is present in a KV cache that's distributed or doing barrier synchronization at the end of periods of compute.

John Sterhout: 而对于这些工作负载来说,真正重要的是延迟。也就是说,将一小段数据通过网络发送过去,做一点计算,然后再得到一个小的结果返回,这个往返时间(roundtrip time)是多少。

Original English

John Sterhout: And for these workloads what really matters is latency. That is what's the roundtrip time to send some small piece of data across the network do a little bit of computation and get a small result back again.

John Sterhout: 而且事实上,重要的不仅仅是延迟或平均延迟。关键在于

Original English

John Sterhout: And in fact, it isn't just just latency or average latency that matters. What

Tail Latency and Network Congestion in Data Centers

Speaker: 真正重要的其实是尾部延迟 (tail latency)。也就是说,你希望能够确信,如果我们发送大量的小型消息,它们每一条都能快速完成传输。

Original English

Speaker: really matters is tail latency. That is, you'd like to know that if we send a whole lot of small messages, all of them will complete quickly.

Speaker: 举个例子,我们通常会测量像 99 分位数延迟 (99th percentile latency) 这样的指标。如果我们系统的尾部延迟非常高,那么这就会限制整个系统的整体吞吐量。

Original English

Speaker: So, for example, we typically measure things like 99th percentile latency. And if we have high tail latency, that can limit the overall throughput of the system.

Speaker: 那么,来看这样一个例子。假设一种非常常见的情况是,我们将一个工作负载拆分到多个节点上,这些节点在一段时间内使用它们的 GPU 执行密集的计算任务。一旦它们全部完成了各自的计算,你就会在这些节点之间进行一些小规模的数据交换——交换数据和元数据。然后,系统就会进入下一轮的计算环节。而在进行这些数据交换和同步的过程中,GPU 实际上是处于空闲状态的。因此,如果这些数据交换中哪怕只有一次耗费了很长时间,结果就是整个计算过程都会停滞。你必须等待所有这些数据交换全部完成,然后才能继续进入下一阶段的计算。

Original English

Speaker: So, here's an example. Suppose a common thing is to take a workload and split it up across several nodes which do intensive computation using their GPUs for some period of time and then once they've all finished their computation you do some small exchange between the nodes exchange data metadata and then it'll go on to the next round of computation and while that exchange is happening that synchronization is happening the GPUs are sitting idle so if even one of those exchanges takes a long time. It turns out the whole process stalls. You need all of those exchanges to complete before you can go on to the next phase of computation.

Speaker: 如果计算阶段耗时大约是 5 秒钟,而数据交换只需要几毫秒,那么这就完全不是什么大问题。而且从历史上看,情况一直都是这样的。但是现在,随着智能体工作负载 (agentic workloads) 的出现——在这些任务中你需要以固定的速率相对快速地吐出大量的 token——计算周期的耗时正在缩减到毫秒级别。如果这个时候数据同步也同样需要花费几毫秒,那么你就相当于在白白浪费很大一部分 GPU 资源,仅仅是为了等待数据同步的完成。

Original English

Speaker: Now, if the computation phase is say 5 seconds and it takes a few milliseconds for the exchange, you know, not a not a problem. And that's historically what it's been. But now with agentic workloads where you're trying to pump out tokens relatively rapidly at a regular rate, the periods of computation are getting down into sort of the millisecond time scale. And if it also takes milliseconds to do that synchronization, then you're wasting a significant fraction of your GPU resources waiting for the the synchronization to occur.

Speaker: 所以我很好奇。我想在这里做一个快速的现场观众调查。今天在座的各位中,有没有人有理由相信小型消息的延迟正在影响你们应用程序的整体吞吐量?如果是这样的话,可以请你们举一下手吗?让我看看,今天在座的有没有这样的情况?实际上,举手的人比我预期的还要多。看来台下有相当多的人都举手了。我认为,随着目前这种技术趋势的继续,这个问题很可能会变得越来越严重。

Original English

Speaker: So I'm curious. I'd like to just do a a quick audience poll here. Is there anybody out there today where you have reason to believe that the latency of small messages is impacting the overall throughput of your applications? If so, can you just raise your hand? See, is there anybody out there today? Actually, more hands than I expected. So quite a few people out there are raising their hands. I think this problem is likely to get worse as the trends continue.

Speaker: 那么,到底发生了什么呢?为什么尾部延迟是一件糟糕的事情?嗯,通常它的根本原因是由 incast(内向突发流量)导致的网络拥塞。所谓 incast,就是当多个节点同时决定向某一个目标节点传输数据时的现象。如果它们都在发送大型消息,那么,鉴于网络中各处的链路带宽都是相同的,三个节点传输数据的速度可能恰好是一个节点所能接收数据的速度的三倍。于是造成的结果就是,数据包会在通向该目标的最后一跳——即在机架顶部交换机 (top of rack switch) 面向该目标节点的出口端口处——大量积压。

Original English

Speaker: So what's going on? Why is tail latency bad? Well, typically the cause is congestion resulting from incast. So incast is when several nodes all decide simultaneously to transfer data to some destination node. And if they all send large messages, well, the links are the same everywhere in the network. So three nodes can transfer three times as fast as one node can possibly receive. And so what happens is that packets accumulate at the last hop going to that destination in the top of rack switch at its egress port for the destination node.

Speaker: 然后,如果这时候某个其他节点决定它也想向这同一个目标发送一条非常短的消息,这条短消息就会被堵在队列里那些长消息的后面。这实际上就直接导致了延迟;而且在最坏的情况下,到达的数据包如此之多,以至于交换机的缓冲区空间耗尽,进而不得不丢弃数据包。然后就会发生超时和数据包重传,这会让一切变得更加糟糕。

Original English

Speaker: Then if some other node decides it wants to send a short message to that same destination, the short message gets stuck behind the long ones in the queue there. And actually that causes delay and in the worst case so many packets arrive that the switch runs out of buffer space that it has to drop packets and then there are timeouts and retransmissions that make everything even worse.

Speaker: 所以,无论如何我们需要找到某种方法来减少这些队列中的拥塞。我们必须想办法让那些发送节点停止如此快速地发送数据,这样队列才不至于无限制地堆积起来。

Original English

Speaker: So somehow we need some way to reduce the congestion in those queues. Somehow we have to get the sending nodes to stop sending so fast so the queues don't just build up without limit.

Speaker: 那么,历史上人们是如何解决这个问题的呢?实际上,在 HOMA 出现之前,几乎所有的网络协议(包括 TCP 和 RDMA),其拥塞控制的责任都是由发送方来承担的。因此,发送方必须以某种方式弄清楚网络中正在发生拥塞,然后它们必须主动放慢其传输速率。

Original English

Speaker: So the way this is done historically virtually all network protocols before HOMA including TCP and RDMA congestion control is the responsibility of the sender. So senders somehow have to figure out that congestion is happening and they have to slow down their rate of transmission.

Speaker: 现在你可能会好奇,为什么是由发送方来做这件事呢?因为网络拥塞是发生在数据中心网络另一端的遥远地方。发送方到底是如何得知这个情况的呢?

Original English

Speaker: Now you you might wonder why are senders doing it? because the congestion is way over at the other end of the data center network. How does the the sender find out?

Speaker: 嗯,在很久以前的过去,发送方发现拥塞的方式是:当队列溢出且数据包被丢弃时,发送方会因为收不到确认消息 (acknowledgements) 而检测到数据包丢失,然后它就会假设这意味着网络出现了拥塞,进而放慢其传输速率。但这种做法的代价是非常昂贵的。

Original English

Speaker: Well, in the in the old really old days, the way they would find out is the queues would overflow and packets would get dropped. The sender would detect the packets got lost because it wouldn't get acknowledgements back and it would assume that means there's congestion and then slow down its rate of transfer. That's really expensive.

Speaker: 所以到了今天,我们有了更好的技术,这通常涉及让交换机直接提供拥塞信息。因此,当机架顶部的交换机看到某个入口端口的队列长度达到了某个设定的阈值,即在队列溢出很久之前刚刚开始堆积时,它就会开始对所有通过它的数据包进行一种名为显式拥塞通知 (Early Congestion Notification, ECN) 的标记。

Original English

Speaker: So today there are better techniques that mostly involve the switches providing information. So a top of rack switch when it sees that the queue length for an ingress port has reached some threshold starting to fill long before the queue overflows it starts marking all of the packets that pass through with what's called early congestion notification ECN marking.

Speaker: 因此,当这些带有标记的数据包传递给接收方时,接收方就会在数据包中看到这个标记。然后,当它下一次与发送方进行通信时(比如发送一个确认消息),它就会包含这个标记并将其传回给发送方。现在,发送方看到了这个标记,于是发送方就会意识到:“哦,某个地方发生拥塞了,我必须放慢我的传输速率。”

Original English

Speaker: And so when those packets pass through to the receiver, the receiver sees the marking in the packets. And then when it communicates back to the sender next, for example, to send an acknowledgement, then it includes that marking that goes back to the sender. And now the sender sees that the sender realizes, oh, there's congestion someplace. I've got to slow down my rate of transmission.

Speaker: 这就是它的基本思想。但不幸的是,要正确地做好这件事是非常困难的。真的非常难。对于拥塞控制机制来说,精确弄清楚应该如何设置速率是极为困难的,因为它获得的仅仅是单个比特的信息——这只告诉你“某个地方发生了拥塞”,而且同时有多个发送方在向同一个目标发送数据。它们都在尝试同时做出速率调整。你到底应该削减多少速率?我又如何才能知道什么时候可以重新把速率提上去?更糟糕的是,由于存在控制滞后 (control lag),要以一种稳定的方式来做到这一点极其困难。

Original English

Speaker: So that's the basic idea. Unfortunately, getting this right is really hard. Really hard. It's very hard for the congesture to figure out exactly how to set its rates because it gets one bit of information. There's congestion someplace and there are multiple senders all sending to the same destination. They're all trying to make adjustments simultaneously. How much do you cut back and how do I know when I can ramp up again? And even worse, it's really hard to do this in a way that's stable because there's control lag.

Speaker: 也就是说,在发送方发现拥塞之前是需要花费时间的。实际上,在使用这个过程时,通常需要几个往返时间 (round trips) 才能让发送方逐渐调整其速率,以达到与可用带宽完全匹配的正确速率。但是,等到你在网络中完成这一切调整时,网络的情况往往已经发生了改变——新的传输任务可能已经开始了,或者旧的任务已经结束。因此,这些系统往往永远无法稳定下来。它们总是处于发送过多和发送过少之间的不断震荡中。

Original English

Speaker: That is, it takes time before the sender finds out that there's congestion. And in fact, using this process, it typically takes several round trips for the sender to gradually adjust its rate to get just the right rate to match the available bandwidth. But by the time you do that in a network that things have changed, new transmissions have started or old ones have finished. And so these systems tend to never stabilize. They're constantly oscillating between sending too much and and sending too little.

Speaker: 现在,这个问题已经存在很长一段时间了。在学术研究界,人们认识到这个问题已经有 20 多年的历史了。在这个课题上,已经发表了成百上千篇的论文。虽然也确实取得了一些改进,这是无可否认的,但距离实现真正能够良好运行的解决方案,我们依然还有很长的路要走。

Original English

Speaker: Now, this problem's been around for a long time. It's been known in the research community for more than 20 years now. There have been tons of papers published on it. There have been some improvements made. That's undeniable, but we're still a long ways from anything that works well.

Speaker: 问题的根本在于,由发送端进行拥塞控制的本质决定了它存在固有的缺陷。它就是无法良好地工作。所以你最终还是会面临严重的队列堆积问题。事实上,你可以看到,发现拥塞的唯一方法就是必须首先要有队列积压,所以到了那个时候,我们其实已经开始经历网络延迟了。这就是一个严重的问题。

Original English

Speaker: And the problem is is with the fundamental nature of it doing the the congestion control on the sender side. It just doesn't work very well. So you end up with a lot of queue buildup. And in fact, you can see the only way to find out that there's congestion is if there's queues and so by that point, we're already experiencing delays. So that's a problem.

The Need for a New Protocol: Introducing HOMA

Speaker: TCP 和 RDMA 还存在着另外一个问题,那就是它们的基本数据模型是一个字节流 (byte stream)——这只是一串单纯的字节,里面没有任何切分的界限和差异化。所以如果你发送一系列的消息(比如说,通过一个 TCP 套接字),它们就会被序列化并直接塞进那个流里。

Original English

Speaker: There's one other problem with TCP and RDMA also is that their their basic data model is a byte stream just a stream of bytes with no differentiation in it. So if you send a series of messages say through a TCP socket they get serialized into that stream.

Speaker: 在这张幻灯片上,你们可以看到,我展示出来的这些消息在这个字节流中似乎有着不同的颜色标记。但在现实的网络传输中,根本没有什么颜色区分。TCP 完全不知道消息的边界到底在哪里。这也让事情变得非常棘手。举例来说,你无法知道后续还有多少数据要传过来。如果你知道这条消息的大小,你自然就会知道后面还有多少数据。而且你也无法优先处理短消息,而这偏偏是我们非常渴望做到的——让短消息更快地通过。

Original English

Speaker: And on this slide I've you know I've shown the messages appear like they have different colors in the stream. Well there are no colors in real life. TCP has no idea where the message boundaries are. And that also makes life hard. For example, you don't know how much more data is coming. If you knew how big the message was, you'd know how much more is coming. And you can't prioritize short messages, which we'd really like to do. Get the short messages through faster.

Speaker: 这可能导致一种被称为线头阻塞 (head of line blocking) 的现象。假设有人向同一个目标连续发送一系列消息,他们先发送了两条非常庞大的消息,紧接着再发送一条小消息;由于它们处于同一个数据流中,这条小消息就会被死死地堵在前面两条大消息的后面。因此,它就遭遇了延迟,这样一来,你又一次面临了尾部延迟的问题。综上所述,TCP 和 RDMA 就是不适合当前这种数据中心的环境。

Original English

Speaker: And you could end up with what's called head of line blocking where somebody sends a series of messages to the same destination and they send two really large ones and then a small one after that gets stuck behind them in that stream and so it gets delayed and again you have tail latency issues. So all in all, TCP and RDMA just not well suited to this environment.

Speaker: 那么我们该怎么办呢?接下来我想告诉大家的,是一个我们在斯坦福大学开发的名叫 HOMA 的全新协议。它完全是基于一张白纸、从零开始重新设计的网络传输方案。如果你可以从头开始,重新思考该如何为数据中心进行数据传输,你会怎么做?事实证明,在 HOMA 中,几乎每一个重大的设计决策都与 TCP 和 RDMA 截然不同。TCP 尽管曾经完成了许多惊人的壮举,但它不再匹配今天的数据中心了,RDMA 也是如此。

Original English

Speaker: So what do we do? Well, what I'd like to do next is tell you about a new protocol called HOMA that we've developed at Stanford, which was based on a completely clean slate redesign for network transport. If you could start from scratch and rethink how you do transport for data centers, how would you do it? And it turns out in Homa virtually every major design decision is different from TCP and RDMA. TCP for all the amazing things it's done is just not a good match to today's data centers nor RDMA.

Speaker: HOMA 表现得特别出色的地方在于,它能够完美地管理大型和小型消息的组合传输,并能够确保这些消息——尤其是短消息——具有极低的延迟。

Original English

Speaker: So what Homa does particularly well is to manage a combination of large and small messages and to make sure that the messages have short messages have really low latency.

Speaker: 这最初只是我的一个学生 Behnam Montazeri 的博士学位论文项目,但后来的测试结果实在是太棒了,以至于我决定将其作为我的个人项目,看看我们是否能把它带出实验室,并真正应用到生产环境中。呃,正如你们可能知道的那样,我跟大多数教授不太一样,因为我是真的很热爱写代码。所以……

Original English

Speaker: So this started off as a PhD dissertation for one of my students, Benam Montazeri, and then the results were so great that I decided to make it my personal project to see if we could get it out of the lab and into production. Uh, as you may know, I'm I'm not like most professors and that I love to code. And so

Homa 网络协议的工作原理与优势

Speaker A: 我把它变成了我自己的编程项目。我为 Linux 创建了一个内核模块。我目前正在努力将它合并到内核主线中。它可以在 GitHub 上下载。那么让我稍微介绍一下 Homa 是如何工作的。我想提三点。首先,它是基于消息的,而不是基于流的。事实上,Homa 的基本单位是远程过程调用 (RPC),它由两部分组成:从客户端发送到服务器的请求消息,以及从服务器返回给客户端的响应消息。所以这里的关键是 Homa 知道消息的长度。它们被埋藏在传输层最底层的结构中。这带来了一系列优势。第一,它使我们能够预测未来。一旦接收端收到一条消息的第一个数据包,它就能确切知道发送端还想发送多少数据。因此,它掌握了多得多的信息来进行拥塞控制。第二,Homa 优先处理较短的消息。它使用 SRPT(最短剩余处理时间优先)算法来尝试优先处理较短的消息。第三,因为消息都是独立的,它们不会被序列化成一个流。每条消息都是独立的。较短的消息可以绕过较长的消息,因此它们不会在长消息后面排队等候。

Original English

Speaker A: I turned this into my own programming project. I created a kernel module for Linux. I'm currently working through the process of getting that upstreamed into the kernel. It's available on GitHub for download. So let me tell you just a little bit about how Homa works. I want to mention three things. First, it's messagebased, not streambased. In fact, the fundamental unit at HOMA is a remote procedure call, which consists of two things. A request message sent from a client to a server and then a response message returned back from the server to the client. So the key thing here is that Homa knows about message lengths. They're buried in the transport all the way down to the bottom. And this has a bunch of advantages. First, it allows us to predict the future. As soon as a receiver gets the first packet of a message, it knows exactly how much more data the sender wants to send. And that's so has so much more information for doing congestion control. Second, Home Prioritizes shorter messages. It uses SRPT, shortest remaining processing time first to try and prioritize shorter messages. And third, because messages are all independent, they're not serialized into a stream. Every message is independent. Shorter messages can bypass long ones so they don't get cued behind long messages.

Speaker A: 关于 Homa 的第二个不同之处在于,它是从接收端来控制拥塞的。现在,仔细想一想,这是有道理的,因为拥塞主要发生在通往接收端的最后一段下行链路上。因此,接收端掌握着更多的信息。事实上,在 Homa 中,一旦它接收到一条消息的第一个数据包,它就能确切知道后续还有多少数据要来。所以,它基本上拥有关于拥塞的完整信息,因此能够更迅速、更精准地对拥塞做出反应。

Original English

Speaker A: The second thing about HOMA that's different is that it controls congestion from the receiver. Now, when you think about it, this makes sense because the congestion happens primarily at that last down link to the receiver. And so, the receiver has way more information. In fact, with HOMA, as soon as it gets the first packet of a message, it knows exactly how much more is coming. So, it has essentially complete information about congestion and it can therefore respond to congestion much more quickly and much more precisely.

Speaker A: Homa 的工作方式是,当发送端有消息要发送时,它会将其拆分成数据包,但它只向接收端传输最初的几个数据包,这些被称为未调度的数据包。之后的数据包被称为调度数据包,它们只有在接收端请求时才会被传输。所以接收端会发回授权 (grant) 数据包。它会将这些数据包分散开来,并随着时间的推移发送回发送端,告诉发送端:“现在是时候把你下一块数据发给我了。” 而且接收端可以延迟这些授权。举例来说,如果接收端有 10 条正在传入的消息,向所有 10 条消息发送授权是毫无意义的,因为那样你只会导致机架顶部交换机队列发生拥塞。所以它可以利用这些授权来减少拥塞,并且还可以利用这些授权来优先处理它最偏好的消息,也就是那些较短的消息。所以,这是一种通过偏好短消息来实现 SRPT 的方法。

Original English

Speaker A: The way things work with HOMA is that when a sender has a message to send, it breaks it up into packets, but it only transmits the first few packets, those are called unscheduled packets, to the receiver. Packets after that are called scheduled packets and they only get transmitted when the receiver asks for them. So the receiver will send grant packets back. It'll paste them out and send those back to the sender over time telling the sender it's now time for you to send me the next chunk of data. And the receiver can delay those grants. So for example, if the receiver has 10 messages that are incoming, there's no point in sending grants to all 10 of them because then you'll just get congestion in the in the top of wrap cues. So it can use the grants to reduce congestion and then it can also use the grants to give preference to its most favorite messages which would be the shorter ones. So it's a way of of implementing SRPT by favoring short messages.

Speaker A: Homa 的第三个方面是,它利用了现代交换机中的优先级队列。现代数据中心交换机在每个出口端口都有不止一个队列,通常有八个,它们可以用于优先级机制,即优先从最高优先级的队列中传输数据包。所以我在幻灯片上这里只展示了两个队列,但通常比这更多。你可以使用数据包的各个字段来指定它应该进入哪个队列,因此 Homa 会动态地做出这些选择,从而优先处理较短的消息。所以如果我们回顾一下几张幻灯片之前的 Incast 示例,所有那些长消息都会堆积在最低优先级的队列中。但如果有一条短消息到来,它将使用较高优先级的队列。因此,它将立即绕过所有正在排队的短……呃,长消息的数据包,并更快速地到达目的地。

Original English

Speaker A: The third aspect of Homa is that it takes advantage of the priority cues in modern switches. So modern data center switches have more than one Q at each egress port typically eight and they can be used in a priority mechanism where packets get transmitted preferentially from the highest priority Q. So I've shown only two Q's on the slide here but typically there's more than that you can specify in packets using the various fields of the packet you can specify which queue it should go into and so Homa dynamically makes those choices in a way to give priority to shorter messages. So if we go back to the incast example from a few slides ago, all of those long messages will pile up in the lowest priority queue. But if there's a short message coming, it will use a higher priority queue. And so it will immediately bypass all of the cued packets from the short from the uh the longer messages and get through to the destination more quickly.

Speaker A: 那么这会产生多大的区别呢?看这张幻灯片。我有一个样本基准测试,这是我对 Homa 进行调优和评估的一部分。它的工作负载由网络上的一群机器组成,这些机器正在来回交换不同大小的消息,大小从非常小到非常大不等。在这张图表上,你可以看到 X 轴是消息长度,大约从 50 字节到 1 兆字节。Y 轴显示了该长度消息的往返时间。所以这使用了长度相同的请求和响应消息。你可以看到绿色的 TCP,蓝色的 Homa。Y 轴是往返时间。所以越低越好。对于每种协议,我有两条曲线。一条是 P50 曲线。那是该长度消息的中位数延迟。然后 P99 是第 99 个百分位数,即该长度消息的尾部延迟。

Original English

Speaker A: So how much of a difference does this make? Here's a this slide. I've got one sample benchmark that I use as part of my tuning and evaluation of Homa. It consists of a workload of a bunch of machines on a network that are exchanging messages back and forth of different sizes ranging from very small to very large. And on this graph you can see on the x-axis is the message length from about 50 bytes up to a megabyte. The yaxis shows you the roundtrip time for messages of that length. So this request this uses request and response messages that are the same length. You can see TCP in green, Homa in blue. And the y-axis is is roundtrip time. So lower is better. And for each protocol, I've got two curves. One curve is the P50 curve. That's the median latency for messages of this length. And then P99 is the 99th percentile, i.e. tail latency for messages of this length.

Speaker A: 所以我想指出两点。首先,短消息的 P99 延迟在 Homa 上有了显著改善。使用 TCP 时,尾部延迟超过 1 毫秒。Homa 则不到 100 微秒,快了大约 13 倍。其次,有趣的是,你可能会认为因为 Homa 偏好较短的消息,所以长消息会受到影响并且性能变差。结果发现事实并非如此。即使在最长的消息上,Homa 也几乎比 TCP 好了一倍。我今天没有时间解释这个问题,但这与 Homa 采用“运行到完成”的方法有关,这种方法比 TCP 所使用的公平调度要有效得多。

Original English

Speaker A: So I want to point out two things. First, the P99 for short messages is dramatically better for HOM. So with TCP it's more than a millisecond tail latency. Home is less than 100 microsconds about 13 times faster. Second, interestingly you might think that because Homa favors shorter messages that long messages suffer and get worse performance. It turns out that's actually not the case. Even on the longest messages, Hom is almost a factor of two better than TCP. I don't have time to explain that today, but it has to do with the fact that HOMA uses run to completion approaches which are much more effective than fair than the fair scheduling used by TCP.

Speaker A: 所以,简单总结一下,短消息在 AI 中的作用似乎正在增加。我认为,它很可能会继续增加。我们会在未来一两年内看看这是否会发生。我想向你们提出一个问题。你知道,当你们在运行应用程序并衡量性能、观察瓶颈在哪里时,问问你自己,短消息的高延迟是否正在影响你的吞吐量?如果答案是肯定的,那么只要知道现在有一个可用的解决方案即可。你应该尝试一下 Homa。你也许可以将尾部延迟降低一个数量级甚至更多。

Original English

Speaker A: So, just to wrap up, the role of short messages in AI appears to be increasing. I think I think it's likely that it's going to continue to increase. We'll see over the next year or two if that happens. And I just want to pose a question to you. You know, as you're running your applications and measuring performance and seeing what the bottlenecks are, ask yourself, is high latency for short messages affecting your throughput? If the answer is yes, then just know there is a solution available. You should give Homa a try. You can probably reduce your tail latency by an order of magnitude or more.

Speaker A: 顺便说一下,Homa 基本上就是我现在的人生使命。我可以说从斯坦福大学半退休了。我这样做的原因是为了能把 100% 的时间花在钻研 Homa 上。所以,如果你决定尝试使用 Homa,我会非常乐意与你合作并提供帮助。如果你在入门、解答问题、修复 Bug 等任何方面需要帮助,你懂的,我很乐意与你合作,努力让你在这方面取得成功。所以如果这让你感兴趣,请随时联系我。我的电子邮件在幻灯片上,或者你也可以 Google 搜索并在网上找到我。非常感谢大家的聆听,希望能收到你们中一些人的来信。

Original English

Speaker A: And by the way, this is Homa is basically my life mission right now. I'm sort of semi-retired from Stanford. The reason I did that is so I can spend 100% of my time hacking on Homa. So I'd be delighted to work with you and help you if you decide you want to experiment with Homa. If you need help getting started, answer questions, bug fixes, whatever, you know, I'd be happy to work with you to try and make you successful with it. So if that is interesting, feel free to contact me. My email is on the slide or you can Google me too and find me over the internet. So thanks very much for listening and hope to hear from some of you.

Speaker A: 请欢迎 DSPy 的核心贡献者和首席维护者,Maxim Rest 和 Isaac Miller 登台。

Original English

Speaker A: Please welcome to the stage the core contributor and the lead maintainer at DSPI, Maxim Rest and Isaac Miller.

DSPy 的 AI 编程理念

Maxim Rest: 哇,Isaac,我,以及所有 DSPy 社区的成员今天都非常感激能来到这里,有机会和大家谈谈 AI 编程、DSPy,以及将任务从模型、其测试工具和所有实现细节中分离出来的极其显著的有效性。

Original English

Maxim Rest: Wow, Isaac, myself, all of the DSPI community are so grateful to be here today to get to talk to you about AI programming, DSPI, and the unreasonable effectiveness of separating the tasks from the model, its harness, and all of the implementation details.

Maxim Rest: 仔细想一想,在编程中,如果我们想经常重复一项任务,我们会把它写成一个函数。我们认为对于 AI 程序也应如此。函数是非常棒的。函数具有可重用性、可组合性、可测试性和可优化性。要创建一个函数,你需要给它起个名字。你定义一些输入、一些输出,然后在其中放入一些实现逻辑。你可以将你的函数重复使用成千上万次。你可以优化它,但你也可以将它组合到更大的程序中。函数非常好的一个特点是,你还可以将它打包并分发,让其他人也可以使用它,他们只需要了解其顶层的接口契约就可以使用,并可以把它当作一个黑盒来对待。

Original English

Maxim Rest: When you think about it, in programming, if we want to repeat a tasks often, we make it a function. We believe the same should be true for AI programs. Functions are awesome. Functions are reusable, composable, testable, and optimizable. To make a function, you give it a name. You define some inputs, some outputs, and then you have some implementation logic inside of it. You get to reuse your functions do thousands of times. You can optimize it, but you can also compose it into bigger programs. One of the really nice things about functions is that you can also package it and distribute it and someone else can use it and they just need to know about the contract on top of it to use it and they can treat it as a black box.

Maxim Rest: DSPy 将所有这些特性带入了 AI 程序。DSPy 是一个 Python 语言的开源软件,正如我所说,它允许你将这些特性引入到你的 AI 工作流和 AI 程序中。它为你提供了实现这一点所需的所有工具。你为什么需要这些?嗯,在过去的三年里,我们在我们的领域里发明了很多术语。这个领域发展得很快。我们每隔一周就会有新模型发布。我们有新的技术,新的策略。如果你像我一样,你会想尝试所有这些东西。但是,在不同时间涌现的这些具体的新技术中,是否真的有任何一项能帮助你的任务或工作?嗯,这些其实都只是一些实现策略,而……

Original English

Maxim Rest: DSPI brings all of these properties to AI programs. And so DSpy is an open source software in Python that lets you, like I said, bring these properties to your AI workflows and AI programs. And it gives you all of the toolings you need to do that. Why do you want that? Well, we have been inventing a lot of terms in our fields in the last three years. It's growing fast. We have new models coming every other week. We have new techniques, new strategies. And if you're like me, you want to try all of them. But will any of these new specific techniques coming out at a different time really help on your task, on your job? Well, these are all just implementation tactics and

清晰的接口契约与可重复的 AI 任务

Max: 你肯定希望将它们置于一个定义非常清晰的契约(contract)之中。对于你那些重复性的 AI 任务而言,如果你明确地定义了一个输入接口(input interface)和一个输出接口(output interface),你就可以在系统内部自由地进行发挥,从而获得极大的敏捷性和灵活性。让我们在 AI 领域将这一点更加具体化。当我第一次发现 DSI(注:这里指 DSPy)的时候,我编写的第一个 AI 程序是这样的:我有一些来自我农场的发票,我想要提取它们的信息以便用来报税。我想从这些发票里面精准地提取出具体的税收数值。

Original English

Max: you want to put them inside of clear contract. If for your repeated AI task, you define an input interface and an output interface, you get to play in the internals, you get a lot of agility. Let's make it concrete for AI. So my first AI program I made when I discovered DSI was that I had some invoices from my farm and I wanted to extract them to do my taxes. I wanted to extract the tax values from there.

Max: 然后我编写的另一个 AI 程序是,在我的电脑键盘上,我配置了一个小命令,它会自动读取我的键盘快捷键,读取我的剪贴板内容,并自动为我纠正语法错误。有时,我实际上还希望它为了语言的清晰度而对文本进行重写。所以我还有另一个程序,它接收文本,纯粹为了让表达更清晰而重写它,然后再把它放回我的键盘输入中,这就是一个完整的命令。有了这些,我就能拥有非常大的敏捷性,并可以把它带到系统内部的不同地方去灵活使用。我可以随心所欲地修改它。如果出现了一个更强大的新模型,我可以轻易地把它替换掉。这操作起来极其简单,因为我的接口是像那样固定好的。我会跳过那一个例子。但它们绝对不局限于那些非常简单的事情和极小规模的输入输出。对于 AI 程序,你完全可以有非常宏大的抱负和野心。

Original English

Max: Then another AI program I did is that on my keyboard in my computer, I have a little command that reads my keyboard shortcuts, read my clipboard, and will correct the grammar for me. Sometimes I actually want it to also rewrite for clarity. So I have another program that takes text, just rewrites it for clarity, put it back on my key keyboard, and that's a command. And then I can like have a lot of agility and bring it different places inside of that. I can change it however I want. A new model comes up and I can change that. It's super easy because my interface is fixed like that. I'll skip that one. But they're not uh restrained to very easy things and small input outputs. You can be very ambitious with AI programs.

Max: 所以在这些例子中,你可以想象这样一种场景:你拥有你的整个收件箱,当一封新邮件进来时,你想要自动起草一封新的回复邮件。我们完全可以在 Aspire(注:DSPy 相关项目)中使用 RLM,即递归语言模型(recursive language models)来做到这一点。这是一个源自我们社区的想法,或者更准确地说,更像是我们大家可能都在做的事情,即智能体工程(agentic engineering)或直觉编程(vibe coding)。你可以给它一个规范(spec)和一个代码仓库,然后你就能直接得到一个 PR(拉取请求)。这些都是高度可重复的任务。正如我一直向你们强调的,当你固定了那个接口边界,你就可以专注于顶层的“如何做”,然后在它内部,你可以只用一个简单的提示词(prompt)来进行简短的对话。你可以持续对那个提示词进行快速迭代。如果有新的智能体(agents)出现了,你可以把它改成一个智能体。如果有新的工具被发明出来了,你就可以添加这些工具,然后我们就进入了所谓的循环工程(loop engineering)。你把那个也整合到里面去。而在它外部的任何东西,都不会改变你的集成方式,其他任何东西也都完全不会改变。而且,当你拥有了这样一个清晰且坚固的边界时,你还可以开始进行自动化优化。

Original English

Max: So in this examples, you could have your entire inbox and a new email coming in and you want to compose a new drafted reply. We can do that in the Aspire with RLM, recursive language models. This is an idea that came from around our community or more like things we probably all do, agentic engineering or vibe coding. You can give it a spec a repository and you get a PR. Those are repeatable tasks. And so as I have been telling you when you fix that boundary you can focus on the how on the top and then inside of it you can have a little chat with just a simple prompt. You can iterate on that prompt. Agents come out you change it to be an agent. Tools gets invented you add tools and then we get into loop engineering. You put that inside of it too. Anything on the outside of it doesn't change your integration and and anything else doesn't change. And when you have such a hard boundary, you can also start to automatically optimize.

Max: 但是,仅仅依靠那样简单的签名(signature),你怎么可能进行自动化优化呢?这是远远不够的。仅仅这样不足以详细且准确地说明你的复杂任务。甚至在 ChatGPT 问世之前,DSP 的创建者就已经开始得出这样一个理念,即你需要三样关键东西来明确说明你的任务。如果你拥有这种表达语言,并且有能力用一种编程语言来完整表达你的任务,你就可以开始进行自动优化,并将底层的繁杂实现细节全部委托给系统处理。这三样东西的第一样是“应该发生什么”。这就是具体的指令(instructions)。我一直向你们展示的签名就是其中的一部分。在屏幕这里,你可以看到一段真实的 DSP 脚本的开头部分。你在顶部设置你的首选模型。你对其进行配置,它与这里的签名完全独立,在这里你有自然语言指令,比如“提取所有的税(extract all taxes),如果字迹难以辨认就输出零”,然后你说,“我要给你一个输入,它将是一个字符串,我希望你给我一个输出,它将是一个字符串和一个浮点数”。这就是在完美表达我需求的自然语言。如果你仔细思考一下,你会发现这非常强大且极其高效。这就好比你有一个朋友过来和你一起玩一款棋盘游戏,你向他们解释了游戏规则指示,然后他们就准备好可以开始玩了。但如果你想做像 AlphaGo 或 AlphaZero 这样复杂的系统,并且你告诉他们,你只能纯粹从例子中去自行学习,那你们将会度过一个非常漫长且艰难的夜晚。

Original English

Max: But how can you automatically optimize with just that simple signature? This is not enough. This is not enough to specify your task. And even before Chat GPD came out, the creator of DSP had started to land on this idea that you need three things to specify your task. And if you have this language and this ability to express your task in a programming language, you can start to automatically optimize and delegate away the implementation details. So the first one is what should happen. This is instructions. The signatures that I've been showing you are part of that. Here on the screen, you see the beginning of a real script in DSP. You set your model at the top. you configure that and it's fully independent of the signatures here where you have natural language instruction to extract all taxes and um and if it's illegible to output zero then you say it I'm going to give you an input it's going to be a string I want you to give me an output and it's going to be a string and a float this is natural language expressing my needs this is very powerful and efficient if you think about it if you have a friend over coming to play a board game with you and You give them the instructions and they're ready to play. But if you want to do like alpha go or alpha zero and you tell them you're just going to learn from example, you're going to have a long night.

Max: 接着,第二样东西是“必须发生什么”。你的任务中必然会存在一些必须被严格听从的绝对约束条件(constraints)。它们必须被强制执行。而做到这一点的最佳也是最可靠的方法就是通过代码。所以我想让你们看第三行代码。在第四行有 self.extractself.recheck。你可以清楚地看到我们在对提取税款(extract taxes)执行一个预测(predict)操作,而且我们在对提取税款执行思维链(chain of thought)逻辑推理。第一个是一个简单的普通程序。第二个则强制让它进行一些深入的推理过程。现在我把它们带到前向传递(forward)的内部,你可以在 if not red_tax 这里看到具体的逻辑。这是我的一个极其严格的要求:如果我第一个简单的普通程序没有成功提取出我的税款,我希望系统能够使用更多的推理步骤来重新运行。我的意思是,我必须要百分之百算对我的税,这绝对容不得马虎。然后我还有另一个不可妥协的要求,如果提取出来的数值低于零,就抛出异常,我想把这个异常情况展示给人类进行仔细审查。我绝不想让你轻易蒙混过关。这一点是绝对不会改变的。打个比方,哪怕我现在拥有了全知全能的 AGI(通用人工智能),我也会希望它永远不犯错误。但是不管预测器(predictor)里装的是什么底层模型,只要它们犯了这些严重的错误,我仍然希望上述的这些业务约束条件是永远坚固成立的。

Original English

Max: And then the second one is what must happen. There are some constraints you have that they have to be listened to. They have to be enforced. The best way to do that is with code. So I want you to go to the third line. A fourth line you have self extract and selfrecheck. You can see we're doing a predict on the extract taxes and we're doing a chain of thought on the extract taxes. The first one is a vanilla program. The second one makes it do some reasoning. Now I'm taking them inside in the forward and you can see in the if not red tax. This is a requirement I have that if my first simple vanilla program doesn't extract my taxes, I want you to rerun with more reasoning. I mean, I got to get my taxes right. And then another requirement I have is if the value is below zero, throw I want to show that to a human. I don't want to let you go. This will not change. Like even if I have AGI, I would hope it doesn't make mistake. But whatever is in the predictor, if they make these mistakes, I still want these things to be true.

Max: 所以,最后一样东西是“好的结果看起来是什么样的”。在我年轻的时候,有一次我和我爸爸在广阔的农场里,我好奇地问他:“你怎么知道这棵树是枫树?”但他竟然完全回答不上来。他无法给我一个明确的口头指令,教我如何准确判断这棵树是不是枫树。他也绝对不可能给我一段清晰的代码逻辑,来教我如何判断这棵树是枫树。因此,随着时间的推移,通过不断地观察无数的真实实例,我逐渐学会了如何识别枫树。但这并不局限于像对现实世界中的植物进行分类这样的事情。它同样完美适用于你的规范(specifications)中所有那些更偏向潜在属性的复杂长尾情况(long tails)。这有时也是为什么你要去大公司实习,并且会有一对一的导师和学徒关系。你需要观察海量的实际例子,在真实世界中存在着许许多多成功行为的罕见长尾案例,是你必须去亲眼观察和用心学习的。

Original English

Max: So the last one is what good looked like. And when I was young, I was on the farm with my dad and I asked him, "How do you know that this tree is a maple?" And he couldn't tell me. He couldn't give me the instruction on how to know this tree is a maple. And he certainly couldn't give me code on how to know this tree is a maple. And so through time with example I learned how to know that a tree is a maple. But this is not limited to things like classifying plants. It's also for all of the long tails in your specifications that are things that are more latent. These are sometimes a reason why you would do internship and you would have a mentor and a mentee. You're looking at a lot of of examples and there are long tales of successful behaviors that you have to see and learn.

Max: 既然现在你有了所有的这些不可或缺的元素,你就有了极其充分的强大表达能力。你可以把这三种核心语言完美地放在一起。你有了规范(specs)、代码(code)和评估(evals)。此时此刻,你的目标已经被完整且极其严谨地指定了。于是你就可以满怀信心地开始进行优化了。你可以在你的评估指标和程序上使用像 Japa 这样的强大工具。然后你就能开始深度自动优化。在 SPI(DSPy 的早期项目)刚开始诞生的时候,先进的算力芯片还不存在。当时的语言模型还不够好,完全无法进行直接的系统优化。所以我们当时依靠编写代码来辛苦寻找少样本(few-shots)的优质例子,以此来引导和让基础模型以恰当的方式行事。后来,语言模型变得越来越强大了,所以我们就能够实现对自然语言指令的自动优化了。而在未来,我们正开始能够越来越多地从繁杂低效的底层实现细节中彻底解放出来,并把它全权委托给自动化系统处理。最终,我们在 Aspire(DSPy)中的期望是,你可以一直坚持所有这些高层设计与规范,然后底下的新技术和所有实现细节都将为你全自动、完美地完成。接下来,Isaac 会跟你们更深入地聊聊,在过去的一年里我们发布了什么重要内容,我们现在正在发布什么激动人心的新特性,以及我们未来的所有宏大愿景计划。谢谢大家。

Original English

Max: Now that you have all of these, you have express fully. You have all these three languages you can put together. You have the specs, the code, and the evals. And now your goal is fully specified. And so you can start optimizing. You can use things like Japa on your metrics and on your program. And you can start optimizing. At the beginning of the SPI, the chip didn't exist. The models were not good enough to optimize. And so we were using code to find few shots examples to make the base models uh act in the proper way. Then models got better and so we could automatically optimize instruction. And in the future we are starting to be able to be liberated more and more from the implementation details and delegate that away. And at the end our hope in the Aspire is that you can stick to all of that and then just the news and the implementation details will be automated for you. Isaac will talk to you a lot more about what has been released in the last year, what we're releasing now, and all of the future plans we have. Thank you.

DSPy 在企业生产环境中的应用与优化

Isaac: 谢谢你,Max。

Original English

Isaac: Thanks, Max.

Isaac: 好了,我们刚才已经向你们展示了关于规范、代码和评估的一个相当宏大的抽象架构概述,但请特别注意,这些东西并不仅仅局限于象牙塔里的学术理论领域。它们正在被当今一些全球最大的企业在真实的生产环境中使用,并为他们带来了令人瞩目的巨大收益。我们发现,当你在企业级应用中实际使用 DSPI(DSPy)时,会获得两个不可忽视的主要好处。首先是,你的系统实现成本变得前所未有地更低了。当你可以灵活选择具体的底层实现方式时,你可以充分利用“苦涩的教训”(the bitter lesson)去广泛搜索不同的架构解决方案,找到某种能以极低成本完美解决你业务问题的最优方案。你可以用这种方法轻松扩展到更昂贵的传统实现方式所完全无法处理的庞大数据规模。比如 Shopify,他们的业务成本惊人地降低了 550 倍。他们之所以能做到这一惊人的卓越成果,是因为他们从一个异常昂贵的大模型果断转向了一个极其廉价的小模型,但他们依然能保留相同的业务邮件,继续在内部高效迭代他们的核心业务逻辑,并不断无所畏惧地尝试各种新事物。这里有三个非常棒的现实案例研究,大家绝对应该在演讲结束后去仔细研究看看。它们会为你提供大量详实可操作的细节,手把手教你如何在自己的企业中切实实现这一点。

Original English

Isaac: So, we've given you a pretty big abstract overview of specs, code, and evals, but these aren't things that are just restricted to the academic sphere. These are used in production by some of the biggest enterprises for massive gains. And we see two main benefits when you use DSPI in the enterprise. First is that your implementation becomes cheaper. When you're flexible to what the implementation is, you can use the bitter lesson to search over different solutions, find something that solves your problem cheaply. And you can use this to scale to data sizes that weren't possible with a more expensive implementation. Shopify 550 times cheaper. They're able to do that because they went from an expensive model to a cheap model, but they could keep the same emails, keep iterating on their business logic inside, and try new things. There's three awesome case studies here, and you should check them out after the talk. They give you a lot of details on how you can do this in your own enterprise.

Isaac: 现在,你之所以想要在 DSPI(DSPy)这个繁荣的生态系统中进行构建,部分原因就在于我们在不断地添加极其强大的全新技术供你们去探索尝试和应用。不过有一点需要着重提醒大家的是,我们添加的这些所有新技术中,没有任何一个能百分之百保证一定完美解决你的特定业务问题,因为那毕竟是你自己的本职工作职责。我们能做的真正核心价值在于,我们可以为你解决一些极其棘手的关键子问题(sub problems),从而让你的整体系统实现变得更加容易和顺畅。例如,麻省理工学院(MIT)的杰出博士生 Alex Zang 发表了一篇名为《递归语言模型》(recursive language models)的前沿研究论文。递归语言模型是解决某些特定类型的超长上下文(long context)程序的一种极其新颖的方法。你猜怎么着?我们可以直接把这个先进的前沿概念引入 DSPI 系统中,让你去亲自无缝尝试,看看它是否对你处理复杂的长上下文任务有实质性的极大帮助。也许它非常有帮助,也许完全没有作用。但最关键的一点在于,你只需要更改区区一行代码,而且你的方法签名(signature)完全保持不变。这才是这里的核心精髓所在。一切底层架构都可以保持稳定不变,你就可以轻松地去 A/B 测试看看这是否能彻底解决你的难题。我们已经有过大量这样的成功真实例子。

Original English

Isaac: Now, part of the reason why you want to build in the DSPI ecosystem is that we're constantly adding new techniques for you to try. And it's important to note none of these techniques we add will definitely solve your problem because that's your job. What we can do is we can solve sub problems for you that make your implementation easier. For instance, Alex Zang, a PhD student at MIT, came out with this paper called recursive language models. Recursive language models are a way to solve some kinds of long context programs. And guess what? We can bring this in to DSPI for you to try see if it helps your long context tasks. Maybe it will, maybe it won't. But the thing is, it's one line and your signature stays the same. That's what's important here. Everything gets to stay constant and you get to see if this solves your problem or not. And we've had a number of examples

DSPy 的最新进展与定性学习

Speaker A: 在过去的一年里,DSPy 社区内外的开发者们带来了这一切。我们有了 RLM。我们有了 Jeepa,这是伯克利开发的一个令人惊叹的提示词优化器。Better together 多模块 grpo。所有这些都是令人难以置信的研究创新,只要你身处 DSPy 生态系统中,就可以在你的实现中尝试它们。在 DSP 4 中,我们还会带来更多新东西。今天我很高兴能和大家谈谈其中的两项:DSP flex 和定性学习 (qualitative learning)。

Original English

Speaker A: of this just in the last year from people building in and around the DSPI community. We've had RLMs. We've had Jeepa which is an incredible prompt optimizer out of Berkeley. Better together multimodule grpo. All these are incredible research innovations that you get to try in your implementation just by being in the DSp ecosystem. And we have more coming in DSP 4. I'm excited to talk to you about two of those today. DSP flex and qualitative learning.

Speaker A: DSPy.flex 是一种新型模块。在 DSP 中,当我们让你优化事物时,一开始是少样本 (few examples),然后变成了提示词,而现在,对于你想要实现的任何函数,它正在变成代码。随着时间的推移,你实际上可以学习一个测试框架 (harness) 来解决该函数。而且这完全是定制的。只要它能解决你的业务问题,你就不在乎具体实现。你创造了测量的方法,因为你定义了规范 (specs)、代码和评估 (evals) 这三个核心部分。

Original English

Speaker A: DSPI.flex is a new kind of module. In DSP, when we let you optimize things, it started with few examples, then it became prompts, and now that's becoming code for any function that you want to implement. You can actually learn a harness over time to solve that function. And this is completely custom. And you don't care about the implementation as long as it solves your business problem. What you've created ways to measure because you've defined the three core parts of specs, code, and evals.

Speaker A: 我很高兴要谈的第二件事是定性学习。AI 工程中一个非常难的问题是构建评估 (evals)。这之所以困难,有几个原因。第一,对于任何现实世界的问题来说,定义什么是“好”是非常具有挑战性的。第二,当你定义“好”的时候,很多时候你不得不丢失细节。知道一封邮件是好是坏,所包含的信息要比知道邮件中需要修改什么才能变得更好少得多。第三,每当你创建一个要攀登的山峰(目标)和一个数据集时,你实际上是在试图创建一个现实的代理。如果我们能直接利用现实来自动指导我们的评估呢?

Original English

Speaker A: The second thing I'm excited to talk about is qualitative learning. One of the hard hard problems in AI engineering is building evals. And there's a few reasons why this is hard. One is that defining what good looks like is really challenging for any real world problem. The second is that when you define good often times you have to lose detail. If an email is good or bad contains a lot less information than if you know what could change in that email in order to improve. And the third is that whenever you create a hill and a data set, you're really trying to create a proxy for reality. What if instead we could use reality to inform our evals automatically?

Speaker A: 定性学习要问的是,我们如何减少这个问题?我们如何减少人工协助?这目前还是一个研究问题。但我们相信,现在的模型已经足够好,能够解释环境中存在的任何文本反馈,并将其转化为评估,以及模型可以攀登的高峰。因此,当你从生产环境中获得更多反馈时——它的追踪记录、用户操作、产品分析——它就是在问你,模型在问你关于数据应该如何表示的问题。当你这样做时,模型可以随着时间的推移迭代完善这座山峰,并继续攀登它来解决你实际的业务问题。

Original English

Speaker A: What qualitative learning asks is how do we decrease this question? How do we decrease assistance? And it's a research question right now. But what we believe is that models are now good enough to interpret whatever textual feedback is present in the environment and convert that into evals and a hill that the model can climb. And so as you get more feedback from production, its traces, its user actions, its product analytics, it's asking you, it's the model asking you questions about how data should be represented. As you do this, the model can iteratively refine the hill over time and continue climbing it to solve your actual business problem.

Speaker A: DSP 专注于这类最后一英里的问题。我们有一个非常强大的研究生态系统,并且我们与他们合作非常紧密。这也是其美妙之处的一部分,我们可以看到应用 AI 工程中发生的问题。定义它们,建立一个基准测试,然后用技术解决它们,接着我们就可以将这些成果民主化给每个人,因为它是开源、开放的研究。

Original English

Speaker A: And DSP focuses on these kinds of last mile problems. We have a really strong research ecosystem and we collaborate really closely with them. And that's part of the beauty is that we can see the problems that happen in applied AI engineering. Sol define them, build a benchmark and then solve them with techniques and then we get to democratize the results of that to everyone because it's open-source open research.

Speaker A: 现在,一个常见的问题是,当我们拥有 AGI 时会发生什么?嗯,即使我们拥有一个极其聪明的模型,该模型也不会知道如何解决你的问题。它不知道如何执行你的任务,也不具备你的背景信息。因此,这种类型的最后一英里学习试图探究,我们如何才能高效地进行这种学习——智能与无所不知是非常不同的。

Original English

Speaker A: Now, one common question is what happens when we have AGI? Well, even when we have an incredibly smart model, the model won't know how to solve your problems. It won't know how to do your tasks or have your context. And so this genre of last mile learning is trying to ask how do we efficiently do this learning intelligence is very different from being all knowing.

Speaker A: 如果你请阿尔伯特·爱因斯坦帮你处理邮件,他可能会问什么是电子邮件。但是如果有了 AGI,AGI 会知道如何处理你的邮件。尽管如此,它也不会知道如何实际解决你的问题,以及如何与你需要互动的人互动。如果不随着时间的推移学习这些背景,它就不会理解你的人际关系。

Original English

Speaker A: If you were to ask Albert Einstein to help you with your emails, he'd probably ask what's an email. But if you AGI will know how to do your emails. Nevertheless, it won't know how to actually solve your problem and interact with the people you need to interact with. It won't understand your relationships without learning this context over time.

Speaker A: 自 2022 年以来,DSPy 一直专注于规范、代码和评估这三个核心理念。随着时间的推移,我们无疑在不断发展,新技术令人难以置信。我们已经从进化少样本 (few shots) 到提示词,再到现在进化测试框架,以及现在进化你的评估。但对于任何这些新技术,你需要问的是,它们如何帮助你解决更难的问题,或者更好地解决你自己的问题?你应该以数据驱动的方式来问这个问题。

Original English

Speaker A: Since 2022, DSPI has been focused on these three core ideas of specs, code, and eval. We've certainly evolved over time, and new techniques are incredible. We've gone from evolving few shots to prompts to now harnesses and now evolving your eval. But what you need to ask for any of these new techniques is how do they help you solve harder problems or solve your own problems better? And you should ask this question in a datadriven manner.

Speaker A: 你应该看着这项新技术,然后问,我怎样才能把它应用到我现有的业务问题上?你应该定义你的问题,并让你的提示词、模型和代码对你需要它们解决的问题负责。当你以这种拥有灵活实现的方式进行构建时,最棒的一点是,你解锁了一个生态系统,其中包含了在座各位不断发明的各种技术。你解锁了接触这里所有人集体智慧的机会,大家一起分享技术。

Original English

Speaker A: You should look at this new technique say how can I apply this to the business problem that I have? You should define your problem and you should hold your prompts, models, and code accountable to the problem that you need them to solve. And what's awesome about when you build in this way where you have flexible implementations, what you unlock is you unlock the ecosystem of all the techniques that anyone in this room is constantly inventing. You unlock access to the collective intelligence of everyone here, all sharing techniques together.

Speaker A: 所以,如果你想构建可靠的 AI 软件,我鼓励你来看看 DSP。我们是完全开源、开放研究的,我们在这里通过构建可靠的软件来帮助你解决问题。我们有一个 Discord,你应该加入。当你提出了下一个技术时,你应该来为 DSPy 做出贡献,我们可以帮你分发它,让每个人都能使用这项了不起的技术。谢谢大家。

Original English

Speaker A: So, if you want to build reliable AI software, I encourage you to come check out DSP. We're completely open-source, open research, and we're here to help you solve your problems by building reliable software. We have a Discord that you should come join. And when you come up with the next technique, you should come contribute it to DSPI and we can help you distribute it and make this awesome technique available for everyone. Thank you.

与 Anthropic 的 Mike Krieger 对话

Speaker A: 接下来和我们一起上台的是 Instagram 的联合创始人、Anthropic 的技术人员 Mike Krieger (迈克·克里格)。

Original English

Speaker A: Joining us on stage is the co-founder of Instagram and a member of technical staff at Anthropic, Mike Kger.

Speaker A: 大家好吗?我是说,早上好。很好。嗯,Mike,感谢你发布 Fable,这正是我们迫切需要的——

Original English

Speaker A: How's everyone doing? I mean, good morning. Nice. Um, Mike, thank you for releasing Fable just in time for us

Mike Krieger: 完全是为了这次大会。我们掐准了时间。

Original English

Mike Krieger: exactly for the conference. We timed it.

Speaker A: 嗯,我们很高兴能邀请到你。呃,你是一位杰出的建设者,也是你在 Anthropic 领导实验室。嗯,随着你在内部看到模型的不断成长,你的模型使用方式发生了怎样的变化?

Original English

Speaker A: Um, we're so glad to have you. Uh, you are uh one of the preeminent builders and your leading labs at Enthropic. Um, how has your model usage changed as as you've you know seen models internally grow?

Mike Krieger: 是的,我的意思是,对我来说这既是模型的转变,也是我角色的转变。在我刚加入 Anthropic 的前两年左右,我是首席产品官,后来我不断看到人们用模型进行构建,我的“错失恐惧症”(FOMO)就不断加剧,因为我是在尽可能多地使用模型,但是例如在产品策略上,我会写一份策略文档,然后让 Claude 批评它,或许你可以使用一个工作流,但这和那种纯粹的构建方式不太一样。我当时把所有的周末时间都花在尝试用它进行构建上,我意识到,好吧,我真的只需要转变一下角色,现在这个时代太有趣了。我实际上看到了一个有趣的趋势,现在有几个在其他地方担任首席技术官(CTO)的人,现在加入了 Anthropic 或其他地方作为独立贡献者(IC)。但我做了一个角色转变,这实际上是在我们开始获得内部快照(后来变成了 Mythos 和 Fable)的那个时候。

Original English

Mike Krieger: Yeah, I mean for me it's been like both the model shift and then my role shift. So I for like the first two years I was at anthropic I was chief product officer and then I kept seeing people build with the models and the FOMO just kept increasing because I was you use the models as much as possible but for example on product strategy I would write a strategy doc and then have cloud critique it and maybe you can use a workflow but it's not quite the same as like building in that pure way and I was like spending all my weekends trying to build with it and I realized okay I actually just need to shift it's like way too interesting a time and it's actually an interesting trend I've seen now like several people that where CTO's at other places are like now joining as IC's at anthropic in other places but I made a role shift and it was actually right around the time where we started getting sort of internal snapshots of what became mythos and fable

Mike Krieger: 观察这种转变非常有趣,嗯,这种转变是从“我有一个想法,我将在脑海中把它分解”,这更像我通常做工程的方式,然后通过这些不同的步骤迭代,转变为更接近于“我将描述目标,你去研究它,然后我们可以讨论你做了什么权衡”,你知道,在此过程中提出一些问题,然后再弄清楚你的结果是什么,以及我们接下来可以怎么做。

Original English

Mike Krieger: and what was really interesting watching that sort of shift was um that kind of change between I have an idea I'm going to like sort of break it down in my head much more how I would do engineering normally and then kind of iterate through these different steps to moving to much more of the paradigm of I'm going to describe the goal like go off and work on it and then like we can talk about what trade-offs you you know surface some questions along the way but then figure out what where you landed and where we can go from there.

Mike Krieger: 我觉得这很难。我不知道大家是否有过这样的经历,而且我知道 Fable 几天前才重新启用。Fable 绝对比我聪明得多得多。所以有时候它完成工作后会说,这是我做的权衡。我就会说,你能像个稍微没你聪明的人那样给我解释一下吗?因为我需要你给我拆解一下。但这就是一个很大的改变,从任务委派,转变为表达最终状态,然后让它去处理和酝酿。

Original English

Mike Krieger: I find it's hard. I don't know if people have this experience where and I know Babel's only been reenabled for a couple of days. Babel's definitely way way smarter than me. So sometimes it'll finish work and be like here's the trade-offs I made. I'm like can you explain it to me like I'm a little dumber than you are because I need you to like sort of break this down for me. But that's been one sort of big change is sort of moving from that task delegation to like express the end state and then have it go and and cook on it.

Speaker A: 是的,我们都在学习如何更好地委派任务。呃,Tariq 昨天帮了我们一个大忙。嗯,我想在报纸上读到它。嗯,你知道,现在在比如第二天的报纸上,我们会有讲座的文字报道。他说要提出无理的要求 (be unreasonable)。在哪些方面,你知道,你的提示词变得更具野心了?

Original English

Speaker A: Yeah, we're all learning how to delegate better. Uh Tariq did us a huge favor yesterday. Um we do I want to read it in the newspaper. Um it's a you know that we have we have uh writeups of talks now in in like the next day's newspaper. He said be unreasonable. In what ways you know have you been more ambitious with your prompting?

Mike Krieger: 我喜欢这个。我是说,我喜欢这种说法。嗯,我们实际上刚刚遇到了这种情况,我负责的实验室计划之一就是这个内部产品,呃,有人说,“嘿,它的工作方式不是我想要的。嗯,你能做些修改吗?”我意识到我只需要去让 Claude 做这件事。比如,你为什么不问 Claude 呢?而且这是一个非技术人员。所以我实际上认为,作为一个行业甚至作为一个产品团队,我们必须教人们在使用时提出更“无理”的要求。这有点难以想象。我认为,如果我可以稍微离题一下讲一个——

Original English

Mike Krieger: I love that. I mean, I love that framing. Um, we actually just hit this to like I'm one of the labs initiatives I have is this internal product and uh somebody was like, "Hey, it doesn't work the way I want it to." Um, and can you make some changes? And I realized I'm just going to go ask Claude to do this. Like, why don't you ask Claude? And this was a non-technical person. So, I actually think as an industry or even as a product team, we have to teach people to be more unreasonable in their usage. And it's sort of hard to imagine. I think that that if I can digress for a

产品设计与人工智能的自由度

Speaker A: 其次,在产品设计方面,我认为目前第一代 AI 产品有点把它们关在盒子里,限制了它们使用工具的权限或自由度,这意味着很难向它们提出一些“不合理”的要求,对吧?当你对它说“帮我做这件事”,它可能就会表现出:“哦,我做不到,我勉强能写代码,但我没法真正运行它”,或者“我能在某种程度上自省我的环境,但也不是真正的自省”。并且我认为,正如你看到我们自己产品的演进,甚至像 co-work 这样的东西,就比如,每个知识工作者都需要一个能写 bash 脚本的虚拟机吗?表面上看是不需要的,但随后你会意识到:“哦,实际上通过这种方式,它可以解决这样的问题:昨天我尝试用我们内置的 PDF 解析器去解析一个 PDF,它就说‘啊,我没法用这种方式解析它’。好吧,那我也许可以写个脚本来做这件事。”所以我觉得就是这样。

不过,我做过的最“不合理”的事情是,我们的一个实验室项目是我用 Python 写的,我对它感情很深。整个 Instagram 以前都是用 Python 写的。我想他们现在终于把它转成 PHP 了,既然他们有了可以做这件事的模型,不过我清楚(那需要多少)tokens。而在部署方面,我意识到 Claude Code 已经用 Bun 解决出了一个更好的部署方案,于是我就想:“好吧,我需要把这整个项目从 Python 移植到 TypeScript。”就像,你知道,如果我戴上我 2010 年代的工程师帽子,甚至 2020 年代初的帽子,这听起来就是个愚蠢的想法。谁会在那个时候去移植,你知道,几十万行的代码?但我当时就觉得:“我认为现在这是可行的。”于是我基本上创建了这个动态的工作流设置,让它在周末把整个项目移植过去——验证它,仔细检查它,然后阅读两边的代码,基本上就是不停地运转、运转、运转,然后周一回来时,看到了一个已经完成的工作流,就是那个东西的移植版本。所以那件事可能算得上是比较“不合理”的事情之一:对,就是把整个 Python 代码库移植到 TypeScript。让它能跑起来,并且在,你知道,一个周末的时间里就能部署。

Original English

Speaker A: second on product design, I think right now the like kind of first generation of AI products, we put them too much in a box and constrain their their sort of access to tools or kind of degrees of freedom which means it was much harder to be unreasonable, right? When you say do this thing for me and then it would be like oh I I can't I can barely like I can write code but I can't really run it or I can kind of introspect my environment but not really. Um and I think as you see our own like product progression even with things like co-work like you know does every single like knowledge worker need a virtual machine that can write bash like on the face of it no but then when you realize oh actually that way it can remediate an issue where oh I tried to parse a PDF using our built-in PDF so I hit this yesterday and it was like ah I can't parse it this way well okay well I can probably write a script that can do this as well. Um so I think that's it. My most unreasonable thing though was u one of our labs projects I wrote in Python like near and dear to my heart. All of Instagram was in Python. I think they're finally converting it to PHP now that they have um like models they can but I know tokens. Um and uh for deployment I realized that cloud code had like figured out a better deployment story with bun and I was like okay I need to port this whole thing from Python to TypeScript. like as a you know if I put on my like 2010s engineering hat or even my early 2020 20s like that's a dumb idea like who would ever port like at that point you know a couple hundred thousands of lines of code um but I was like I think this is doable now and I basically created this dynamic workflow setup and over the weekend had it port the whole thing like verify it double check it then read both code like basically churn and churn and churn and then came back Monday to a completed workflow that was a ported version of that thing so that probably ranks on like the more unreasonable things like yeah just port this entire Python codebase to TypeScript. Get it working, get it deployable in, you know, a weekend.

AI 辅助代码移植与基础架构扩展

Speaker B: 是的。我的意思是,有很多人都在讨论把 Bun 和 Zig 移植到 Rust 的那个版本。我想很多人也会觉得:“好吧,它是个编译器。它是个运行时。它有很多测试,做起来很容易。你能像那样,把像 Instagram 这种你非常了解的产品移植到 PHP 吗?”

Original English

Speaker B: >> Yeah. I mean, a lot of people are talking about the the bun zigg to rust version. I I think a lot of people are also like, well, it's a compiler. It's a it's a runtime. It's got lots of tests, easy to do. Can you port Instagram, which you would know very well to PHP like that, like a like a product?

Speaker A: 是的。我的意思是,我认为在产品方面,甚至我都不知道它是更容易还是更难。我们在 Instagram 做的一件事,那是在 Python 3 出来的时候,我们第一次能够添加类型提示,当时人们在内部有很多讨论,比如:“我们在 Python 上会不会走到尽头?”而我的观点一直都是:“我认为我们可以把它推进得比我们想象的要远得多。”但我认为类型将帮助我们不至于成为自己的绊脚石,于是我们构建了一个叫 MonkeyType 的东西,我们在其中基本上捕获了运行时的类型,即基本上是在生产环境中实际被使用的类型,然后把这些映射回代码库中的类型。

我认为正是因为这种模式,如果你在使用大语言模型 (LLMs) 进行某种转换或交叉编译时,你也可以更多地依赖生产数据,或者运行一些分段测试,在这些方面真的有非常有意思的做法。我认为你在那里可以做很多事情。不过,是的,我认为在那个领域同样是前途无量的。我觉得最难的部分始终是找到一个边界,确定你可以在哪里开始增量式地做这件事,而不是试图一劳永逸、“煮沸整个海洋”,然后一夜之间把它替换掉。是的,我的意思是你的用户最终就是你的测试,而且,你知道,我也在报纸上读过另一篇文章,讲的是你如何可以只使用逐步发布 (rollouts),有时你并不真正知道你需要它做什么,但是当那个基础架构存在、供你进行实验和发布东西时,它能带来如此多的可能性。

是的,我的意思是,我总是发现这是我们得到过的建议。就像我们发布 Instagram 的时候,恰好在第一周一切都崩溃了,因为我们并不真正清楚我们在后端方面都在做些什么。而巧合的是,那周我们的一位投资人刚好安排了一次午餐,甚至都不是为我们安排的,那只是一次关于基础架构的午餐。结果我们在那里完全霸占了那场对话,因为每个人对我们该如何解决扩展问题都有自己的看法。我在那里得到的两条建议,那是 2010 年,我将会永远记住。一条是:基本上把你认为你甚至可能勉强需要的所有东西都预先测量好,因为最糟糕的事情就是发生故障时,你还在想:“好吧,这个数字是正常的还是偏高?”以及“哦,我不知道,因为在我刚刚添加这个指标之前我没有数据”。另一条是:在控制开关 (knobs) 和功能标记 (feature flags) 方面要非常周全。所以,即使在早期的 Instagram,我们也有一个非常简单但极其有效的方法来进行灰度放量、发布以及动态配置,因为你知道,我们很多的运行时配置必须在几秒钟内进行更改,以便我们能够处理负载。能够以一种一流的方式做到这一点非常重要。我绝对在 AI 领域也看到了这一点,你知道,在这里我们要做出各种不同的权衡,拥有那种运行时配置是极其关键的。

Original English

Speaker A: >> Yeah. I mean, I think the product side, it's even I don't know if easier or harder. One of the things we did at Instagram, this is when Python 3 came out and we were able to add typants for the first time and it was people had a lot of internal conversations like are we going to run out of steam on Python and my perspective was always like I think we can take this way further than we think we can uh but I think types are going to help us not sort of be in our own way and we built this thing called monkey type where we basically like captured runtime type like basically the types that were actually getting used in production and then mapped those back to to the types in the codebase. And I think because of that sort of pattern, I think there's really interesting ways in which if you're doing sort of conversion or sort of cross-co compiling using LMS, you can also lean on production data a lot more or run sort of like segmented tests. I think that like there's a lot of uh things you can do there. But yeah, I think it's I mean the sky's is the limit there as well. I think the hardest part is always finding the boundary around where you can start doing it incrementally without trying to boil the whole ocean and like swap it overnight. Yeah, I mean your users are your test ultimately and um you know we I also read another article in the newspaper about how you can just use rollouts and sometimes you don't really know uh what you're going to need it for but when that infrastructure exists for you to experiment and to roll things out it's enable so much. Yeah, I mean I always found this was advice we got. It's like we launched Instagram and the happened to be the first week everything melted because we didn't really know what we were doing on the backend side of things. And uh coincidentally that week there was like a lunch that one of our investors just scheduled like not even for us. It was just a infrastructure lunch and we ended up spending we totally like monopolized that conversation because everybody had their own opinion about how we could fix our scaling. Um, and like the two pieces of advice I got there, it's like 2010 that I will like will forever retain is like um like basically like pre-measure everything that you think you might even remotely need because the worst thing is an outage where you're like well is this like number normal or is it high and like oh I don't know because I don't have data until I just added this metric. And the other one is being like really thoughtful about knobs and feature flags. So even you know early Instagram we had like a very uh simple but really effective like way in which you could do like ramp ups and roll outs and dynamic config too where you know a lot of our runtime configurations had to be changed you know in a matter of seconds so that we could handle load and being able to like do that in a first class way was was really important. I'm seeing that definitely in in AI as well where you know we're making all sorts of different trade-offs and having that kind of runtime configuration is super key.

从动态配置到 Tags 工作流

Speaker B: 是的。顺便说一下,关于 Instagram 我最喜欢的扩展故事,我想大概就是你们发布当天,你们用电子邮件把自己给 DDoS 攻击了的那次。对,

Original English

Speaker B: >> Yeah. Uh my my favorite scaling story for Instagram by the way I think it's like your launch day when you you do yourself with email. Yes,

Speaker A: 如果大家没看过那个故事的话,真的应该去查查。

Original English

Speaker A: >> which people should look up that story if uh if you haven't seen it.

Speaker B: 嗯,我想聊聊 tags,这非常非常重要。它是,它是你们今天百分之六十多的代码的编写方式。

Original English

Speaker B: Um I wanted to go into tags uh very very major ship. Uh it's it's how 60 something percent of your code is written today.

Speaker A: 是的。

Original English

Speaker A: >> Yeah.

Speaker B: 那么,你是如何将其与你刚才所说的一切协调起来的?也就是说,它非常动态,比如你实际上并没有交付一个单一的应用程序,你交付的是一个带有 3000 个功能标记的应用程序。是的。

Original English

Speaker B: >> Um how did you square that with everything you just said where it's like very dynamic like you don't actually ship one app, you ship one app with 3,000 flags. Yeah.

Speaker A: 就像,“好吧,那你今天在开发什么?”“我不知道。它是,它是为这部分用户群体开发的。”

Original English

Speaker A: >> And like well what are you working on today? I don't know. Like it's it's for this segment of the population.

Speaker B: 是的,是的。

Original English

Speaker B: >> Yeah. Yeah.

Speaker A: 我的意思是,我觉得这其中有一堆事情。比如,我当时非常兴奋,早些时候我和 Swyx 聊过,比如我非常激动我们现在推出了 tag,因为那已经是我们工作了一段时间的方式了。以前我在台上演讲,人们会问:“你们在 Anthropic 是怎么工作的?”我就会说:“哦,是的,我们用这些东西,它们不完全是 Claude Code,但你知道……”但这很难描述。可是,如果你,比如,能一窥 Anthropic 的内部情况,你会看到我们当然会在一些互动性更强的事情上使用 Claude Code,或者是当你正在某个特定事情上迭代,你需要大量的高带宽来回交互时。

但大多数的使用实际上更多是通过 tagging,通过 tag 来进行委托。你可以说,比如:“这是……”而且这之所以非常有趣,是因为它有多人协作的特性。这让我想起了,比如,实际上就像 Midjourney,每个人都在 Discord 上,看着其他人是如何使用它的。我认为,其实回到你之前的问题,这确实有助于激发那种“不合理”的要求或者野心,因为当你第一次看到别人 @(tag) Claude 时说:“嘿,你知道,不仅要修复这个 Bug,而且现在你还要负责代码库的这一部分,我希望你监控这个反馈频道,并主动承接任务,然后修复它们。另外还要,比如,如果这个 API 发生变化了,也顺便处理一下。”就像我看到有人这么做,我当时就觉得:“哦等一下,我之前完全没有充分利用这个东西,我只是把它当成在 Slack 里的一个高级版 Claude Code 来用。”这绝对是它的一种全新版本,对吧?它更高级的版本真正开始尝试将它视为一个能够保留上下文、拥有记忆并且能够主动出击的队友。而这切实地改变了我们在内部的运作方式。它更像是一种多人、异步、主动的方式,而不是大多数人各自在他们自己的命令行界面 (CLI) 里闭门造车。

Original English

Speaker A: I mean I think there's a bunch of things. So like with I was really excited. I was talking to Swigs earlier like I'm really excited that we have tag out there because it is uh how we've been working for a while and I would get up on stages and people be like how do you work at Anthropic and I'd be like oh yeah we use these things like that are not quite cloud code but you know uh but it's hard to describe it but I mean if you like got to poke into anthropic like you would see u of course cloud code usage for things that are like more interactive or if you're kind of iterating on a particular uh sort of sort of specific thing where you want a lot of like high sort of bandwidth back and forth But most usage is actually much more delegating uh via tagging and via tag. And you can say like here's the and the reason it's really interesting is how multiplayer it is. And it reminds me sort of like um actually like mid journey like the fact that everybody was on discord seeing how other people were using it. I think actually to your earlier question really helps with that unreasonableness or ambition where the first time you see somebody tag claude and be like hey you know don't just fix this bug but like now you are responsible for this part of the codebase and I want you to monitor this feedback channel and proactively take on tasks and then fix them and then also take like you know if this API changes do that like I saw somebody do that I was like oh wait I've been totally underutilizing this thing I've just been using it as like a glorified cloud code in Slack like that's definitely uh totally like uh sort new version of it, right? The more advanced version is really trying to start think of it as a teammate that is actually sort of holds context, has memory and can be proactive. And that's just really changed how we operate internally. It's much more like this multiplayer async proactive way than it is a you know most people off in their own CLIs.

Speaker B: 你们会被代码审查阻塞吗,以及

Original English

Speaker B: >> Are you bottlenecked by code review and

代码审查的未来与 Artifacts 的作用

Interviewer: git?显然现在有云端的代码审查,但通常还是需要人工检查。未来有没有可能达到直接合并代码的程度?

Original English

Interviewer: git? Obviously there is cloud review but someone usually still looks at it. Is there a world in which you just merge it in?

Mike Krieger: 这是一个非常好的问题。目前我们绝对还是卡在代码审查这个瓶颈上,尤其是那些涉及架构层面的改动。这其实比单纯的审查瓶颈更微妙,因为如果只是时间问题,我们大可以重新规划时间。但现在的瓶颈在于,人类甚至无法完全概念化我们正在做的事情。所以,我们几周前发布 Cloud Code Artifacts 的部分原因就在于此。以前你发给别人一个 pull request,他们可能会说:“老兄,我不知道,这可是 2000 行代码,对我来说这就只是一堆代码而已。” 而现在我们开始做的是,更多地去分享像 Cloud Code Artifact 这样的东西,去分享代码背后的解释、这次改动的意图,以及做出了哪些权衡。我认为这将越来越成为我们沟通的趋势:虽然代码最终可以通过某些方式被验证,但真正去讨论意图、权衡,然后在生产环境中进行衡量,至少是我们目前在走的方向。 当我收到一个 pull request 时,我不会逐行审查。我希望我能说我审查了每一行代码,但我绝对没有。我实际上会和 Claude 讨论这段代码,告诉它:“好吧,这些是我有问题的地方,你能去调查一下吗?” 这算是一种由 Cloud 驱动但仍由人类主导的代码审查,特别是对于那些真正重要的代码而言。而对于那些类似于外观视觉的改动,我们的态度更多是:“看吧,如果需要的话,我们就等遇到问题再去修复(fix forward)。”

Original English

Mike Krieger: Yeah, we're it's a really good question. Um, we are definitely still bottlenecked on reviews, especially for things that are like touching some architecture pieces and it's actually more subtle than just being bottlenecked on review because that's, you know, okay, we can carve out time differently. It's like bottlenecked on human ability to even like fully conceptualize what we're doing. So, one of the reasons we built cloud code artifacts that we shipped a couple weeks ago was partially for that, which is uh you would send somebody a PR and then they'd be like, I don't know, man. This is like 2,000 lines of code. Like, it looks like code to me. Um, and what we started doing instead is sharing much more like here's a cloud code artifact. Like, here's the explanation. Here's the intention of the of the change. Here's the trade-offs that were made. And like I think that's going to much more be the trend by which we communicate which is the code is ultimately you know verifiable using some things but actually like discussing intent and trade-offs and then measuring in production is I think the at least the direction of travel we've we've gone I don't review when I get a poll request I wish I could say I've reviewed every line of code I definitely do not I like actually talk to claude about the the code to say all right like these are the questions that I would have can you go investigate it. It is kind of cloudpowered code review but still human driven and and for the really important ones and for the ones that are like cosmetic visual changes it's much more like look like we'll fix forward if we need to fix forward you know

组织架构与 Labs 团队的敏捷迭代

Interviewer: 完全同意。我觉得这里的很多人也都在试图弄明白这一点。我还想聊聊关于 Anthropic Labs 整体的情况。Nai Patel(你可能以前见过他)很喜欢问这样一个问题:“画一下你们的组织架构图”。正如那句话所说,“你发布的产品就是你的组织架构图”。我觉得这很重要。现在大家应该都知道 Cloud Code 以及你们拥有的 Tags 功能了,那么你们是如何构建 Labs 的组织架构的?

Original English

Interviewer: yeah totally um I think a lot of people here are trying to figure that out too um I wanted to talk also a little bit about uh enthropic labs in general um Nai Patel who you've probably met before loves to ask ask the question like draw the org chart like like how like uh people you know you ship your org chart like I I think it's important like everyone knows cloud code now you've got tags um how are you structuring the labs

Mike Krieger: 这是一个好问题。因为我们一直在权衡的是:一方面你希望团队成员能得到支持。要知道,我认为“工程经理这个职位已经走向消亡”这种说法被严重夸大了。我认为辅导、人际关系处理以及个人发展仍然是非常非常重要的。但是,特别是在像 Labs 这样的团队中,我们整个的节奏是两周进行一次审查。每一个项目都要接受审查,我们称之为“坚持还是转向(persevere or pivot)”。所以,基本上每个项目都会被审查,要么继续坚持做下去,要么就必须改变方向,甚至是关停。而且你知道吗,基本上在每一个这样的周期里,我们都会关停项目。当你做得越多,你越不会觉得:“哦,不,我的项目被关停了,我失败了。” 而是会认识到,快速进行原型设计、尝试内部发布、可能推出版让部分用户尽早体验,如果行不通就果断结束,这绝对是 Labs 团队的初衷。 正是因为这种快速的迭代,如果你把组织架构与具体项目绑定得太紧,最终会导致每两周就要重组一次,那绝对是一场噩梦。所以,我们最终形成了一种非常有趣的设置:在 Labs 内部负责某个特定方向(我们称之为“赌注(bet)”)的小组或团队,绝对只是去抽调人手。比如:“好吧,从产品部门抽调一个人,从工程团队抽调一个人……” 当某个产品我特别感兴趣时,我也会参与进去,和团队一起工作。这就是特定时间内的基本工作单元。这里也有“赌注负责人(bet lead)”或“直接负责人(DRI)”的概念,但有趣的是,他们通常并不管理组内的其他任何人,这在某种程度上打破了以前很多事情的运作方式。 但我认为,这让我们变得非常灵活。当你说“好吧,其实这个项目行不通,我们解散,继续前进吧”时,这并不是什么大不了的事。而经理的角色则更多地变成了:确保每个人都被分配到他们最兴奋的事情上,并且以尽可能好的方式工作。当然,我们也会在某个产品“站稳脚跟”时将其固化下来。比如 Cloud Design(云设计),一开始也是以这种临时的团队方式启动的。而现在,随着我们发布了它,它获得了关注,我们在 6 月份还进行了一次重大的第二次发布。于是它开始变得稳定,我们为那个特定的团队招聘了人员,它也有了更成型的结构。所以,这种模式就是:在项目被证明有价值从而固化下来之前,它始终是松散灵活的。

Original English

Mike Krieger: yeah it's a good question because what we were trying to wrestle with was you want sort of people to be supported like you know I think the death of the engineering manager discipline has been greatly exaggerated like I think there's still a lot of coaching and interpersonal pieces and personal development that I think is still really really important but especially in a labs type group where like our whole cadence is two week reviews where every project goes up for we call it persevere or pivot. So basically every project is up for review and either it's time to you know keep going persevering or you know it's time to pivot it or even shut down and you know we shut down projects basically every single one of those cycles and it's like the more you do it the less it's just like oh no my project has shut down I failed. It's like no that is definitely the intention of the labs team is to prototype quickly try to ship internally maybe get it to early access and if it doesn't work wind it down but because of that kind of like rapid iteration it means that if you align the org chart too much to the individual projects you're going to end up like reorging every two weeks which would be a total nightmare. And so we've actually ended up with this interesting setup where like the the pod or the team that is working on a given we call them bets within labs definitely just draws upon like all right somebody from product somebody from the nge team um you know I'll jump in when it's a product I'm particularly interested and I'll come in and work together with the team on it um and that's the unit for that time and there is the concept of a bet lead or a directly responsible individual but the interesting thing is that they don't manage usually any of the other people which kind of breaks the kind of previous way in which a lot of these things were done But I think it leaves lead leaves us to be really flexible when you say okay actually this project is not going to work out. Let's disband and keep going and it's not a big deal and the manager is much more playing the like make sure every individual is assigned to the thing that they're most excited about and that they're working in the best way possible. Now what we do sort of solidify is when there's a product that has like legs like cloud design for example started in this sort of ad hoc sort of group way and then now that like we've shipped it it's gotten traction we've done like a big second release um in June like it's becoming like we've hired people for that specific team and it has more of a of a structure so it's like loose until it gets solidified down the line.

Cloud Design 的未来发展

Interviewer: Cloud Design 的未来是什么?我觉得很多人对它非常感兴趣,这是你们今年最大的发布之一。它未来的发展方向在哪里?

Original English

Interviewer: What's the future of cloud design? I think a lot of people are very interested in it's one of your biggest launches this year. Um, where does this go?

Mike Krieger: 对我来说,阻碍 Cloud Design 变得更好的因素在于,它与我们其他界面的交互还不够好。比如,前几天我在设计某个东西,或者在和 Cloud Code 对话时,我心里就在想,我希望我所讨论的设计能有一种更无缝的交互体验。总体来说,这又回到了“解放 Claude(unconstraining Claude)”这个话题。实际上,我们的各项服务之间并没有尽可能好地相互通信,我认为这真的阻碍了我们很多关于未来可能性的有趣想法。所以我认为这是我们正在关注的一个主要领域。 另一个方向是,随着时间的推移,Cloud Design 和完整应用程序之间的界限会变得越来越模糊。比如,我已经看到人们在上面构建出了功能齐全的东西,甚至连游戏都有,尽管目前还没有数据持久化功能。这绝对不是我们设计 Cloud Design 的初衷,但你确实可以用 HTML 和 JavaScript 做到这一点。所以,我们要进一步模糊这些界限,并思考一条路径:如何从一个看起来非常漂亮的全功能设计,过渡到一个类似于 Artifact 的东西,让你能够实际去持久化数据、与他人分享,并在此基础上继续构建。我认为随着时间的推移,这些界限的模糊化将会变得非常有趣。

Original English

Mike Krieger: I think for me, I mean, uh, the things that are holding back Cloud Design for being even better is better interaction with our other surfaces. So, you know, I was designing something or I was talking to to Claude Code the other day. I'm like, I want a really much more seamless like what I'm talking about the design for it, you know, interactive design back to that. I think in general that's I mean this goes back again to uh kind of unconstraining claude like the fact that our services don't talk to each other as well as they could I think really holds back a lot of interesting ideas around what we could do. So I think that's one like kind of major area that we're looking at. Um and then the other one is people like the lines between a cloud design and an app get blurrier and blurrier over time. Like I've seen people of course there's no like persistence but build like fully functional like even games which is definitely not what we designed cloud design for but you can do it this just HTML and JavaScript. Um so blurring those lines even further and thinking through like what is the path from a like fully featured design that looks really well to really good to something that is maybe more like an artifact where you're actually able to go and you know persist data and share it with others and build from there. So I think that those lines get really interesting over time too.

删除与做减法的艺术

Interviewer: 是的。设计很大一部分在于品味。我其实问了 Fable 它想问你什么。这就是 Fable 想出来的问题:为了做 Instagram,你几乎删除了 Burbn(波本)的所有功能,也就是说你原本有一个完整的诸如 SoLoMo(社交、本地、移动)之类的应用,最后却演变成了 Instagram。那么,如果让你在 AI 领域做减法,你会删掉什么?或者问得更尖锐一点,你会删掉 Claude 里的什么?

Original English

Interviewer: Yeah. Uh, a big part of design is having taste. Um, I actually asked Fable what Fable wants to ask you. Uh, and this this is what Fable came up with. Uh, you deleted almost all of Bourbon to get to Instagram, which is like you had a whole, you know, solomo whatever thing and you went to Instagram. Uh, what would you delete in AI? Or more spicy, what would you delete in Claude?

Mike Krieger: 哦,我喜欢这种尖锐的问题。有趣的是,我们的 Slack 里有一个叫“下架项目(project unshipped)”的频道,专门用来讨论目前产品里有哪些东西需要被下架。我想说,当年在 Instagram 时这就很难。在 Instagram,有些功能只有 4% 到 5% 的使用率。你会觉得:“哦,用的人确实不多。”但当你同时拥有 20 个只有 4% 到 5% 使用率的功能时,这就变成了经典的 Microsoft Word 难题:每个人都在使用一小部分毫不相干的功能集合。所以这始终是个挑战。 不过我认为我们现在还是个年轻的产品,所以希望这类情况会少一些。比如我们最近下架了“Styles(风格)”功能,因为只有一小部分人在用它,而且在很多方面它并不那么符合“AGI 理念”。它的运作方式非常死板。而“Skills(技能)”则是一个更好的应用替代方案。因此我认为,你必须要有意愿去把前一代 AI 的一些基础功能下架掉,或者至少用下一代的功能去补充或取代它们。 在我看来,最大的一点在于——我最近也在 Labs 之外花了一些时间研究这个——天哪,我们现在要求用户在代码(Code)、协作(Co-work)以及聊天(Chat)之间做出明确区分,而这有两个问题:第一,它们之间无法很好地互操作,也不能互相委派任务;第二,我认为大街上的普通人……

Original English

Mike Krieger: Oh, I like the spice. Um, I think I mean we have it's interesting. We have a uh one of our Slack channels like project unhipped which is like what is in the product right now. It's and and I mean this was hard at Instagram. The Instagram we some things that had like four to 5% usage. You're like oh that's really not very many but then you have like 20 features that each have four to 5% usage. It's like the classic Microsoft Word problem of like uh everybody uses some disjoint subset of the of the functionality. So that that's always the challenge. Now I think uh we're a younger product so hopefully we have less of those things like we unshipped styles I think recently where it was like used by a small percentage of people and was not really AGI pilled in a lot of ways. It was like very sort of prescriptive in the way that it worked and skills were a much better uh application something like that. So I think you have to be willing to take the primitives of like one generation of AI and like unhip them or at least like supplement them or supplant them with the next one as well. I think the biggest thing as I look at it and I've been spending some time like outside the labs on some of this is like man like we're asking people to make like code versus co-work versus like chat distinctions and like one they don't interoperate well and they can't delegate to each other and two I think the average person off the street

产品复杂性与初创公司的优势

Speaker A: ……无法向你解释为什么那些交互界面全都各不相同。所以我认为,消除我们代码或产品中的一些产品复杂性,会是一件大有裨益的事情,因为这样 Claude 就能专注于它需要做的事情,并且把它做好。没有什么比这样的协同工作环节更令人沮丧的了:你觉得“太好了,我已经准确地规划出了我想让你构建的东西”,然后却变成“你能否帮我写一段话,让我可以粘贴到 Claude 代码里”,这种感觉就像是还在用 2020 年的那套工作流程,这在现在真的不应该再存在了。

Original English

Speaker A: [I] could not explain to you why those surfaces are all different. So I think deleting some of the product complexity within our our code or our product I think is a a thing that would would serve well also because then cloud can do what it needs to do and and do well. Like there's nothing more frustrating than having a co-work session where you're like great I've mapped out exactly what I want you to build and then be like can you please like create a paragraph that I can paste into cloud code like that is some 2020 you know kind of workflow there that really shouldn't exist anymore.

Speaker B: 是的。嗯,我觉得在你不打算做的事情上划清界限,同时也为其他人留出空间,这很有意思。今天来参加 AIE 的很多人,显然都对初创公司抱有很大的共鸣,因为今天是初创公司主题日。但是会场里也存在一些焦虑情绪,因为也许明天 Anthropic 一觉醒来,发布几个 Markdown 文件,就把我所在的行业给摧毁了。所以,为什么我们不干脆都放弃,直接加入 Anthropic 呢?为什么还要费力去创办其他公司呢?

Original English

Speaker B: Yeah. Um I I think drawing lines on what you don't want to do and also sort of leaving room for others is interesting. Um a lot of people today is like the startup's day for AIE are obviously very sympathetically aligned to startups. Uh but there's some anxiety in the room because tomorrow entropic could wake up and publish some markdown files that destroy my industry. Um so uh why should we not all just give up and join Enthropic? Like why bother starting any other company?

Speaker A: 嗯,我是说,实际上我加入 Anthropic 的主要原因之一,就是因为我看到了——你知道,当时的模型在代码编写方面还没那么好,但它们正在不断进步——我看到了它将能在多大程度上开启全新一代的初创公司。这并不是因为它能替他们解决创意构思或品味的问题,而是因为它会让实验变得极其简单,能让你跑得更快。我至今依然非常确信这一点。我的意思是,这就是现实。你知道,我们在 Instagram 的时候也遇到过这种情况,会有投资者问:“如果 Google 推出了一款照片产品会发生什么?” 答案是,Google 会推出一款非常有“谷歌特色”的照片产品,并且它必然会受到他们已有产品集成的限制,它会去发挥他们自己的优势。我认为这种规律依然适用。这就有点像是在给大家提供如何与 Anthropic 竞争的建议,但其实不然,因为我们同时也是一个平台。所以我觉得大家还有很大的空间,可以极其专注地深耕你所处的特定垂直领域、行业,或是你非常了解的某个用户群体。这种了解程度是任何 AI 实验室永远无法企及的,你也因此能获得那种程度的采用率和用户喜爱,并以此为基础不断发展。当然,在现在这个时代这绝对变得更难了,因为模型本身就能做很多事情,所以其中一些需求可以被转化为模型的技能,也许不再需要它们自己专门的产品了。但我认为,难做的事情依然难做。比如理解人们的需求,弄清楚你要如何触达他们,倾听他们的心声并极快地进行迭代。现在依然是这样:四五个专注于解决某个问题的人,行动速度绝对会快于在其他任何类型组织里的同样四五个人,因为大组织会受制于我刚才提到的那种复杂性,比如我们需要让许多不同的产品相互协作。这就是我们必须去克服的一个有趣的限制。但这在其他方面也是一种优势,对吧?所以,是的,我仍然非常长期地看好初创公司,只不过它掩盖了一个事实:编写代码从来都不是那个限制因素。也许从时间线的角度来看曾经是,但它从来都不是那种会决定你的初创公司生死存亡的关键。真正的关键,依然是对那个领域的洞察和对用户的理解。

Original English

Speaker A: Um I I mean I actually joined one of the main reasons I joined Anthropic was because I saw how much this was like you know the models weren't that good at code yet but they were getting there like how much it would unlock a whole like next generation of startups not because it was going to solve their ideation or their taste but because like it would make experimentation way simpler and and would get you to move faster and I still like really believe that and I mean it's the reality of you know uh and we saw this of like Instagram like we would get questions in investors is like, well, what happens when Google launches a photos product? It's like Google's going to launch a very googly photos product and it's going to have to be bound by the integrations that they already have and it's going to be like it's going to play to their strengths. I think that is going to be true. enough like giving advice on how to compete with Anthropic, I guess, in a way, but like it's actually not because we're also a platform, which is like there's so much I think room to be like laser obsessed with your particular vertical or your industry or group of people that you know really well in a way that like none of the labs are ever going to get to that level of uh of understanding and like therefore get that kind of adoption and user love and and build that up. how it's definitely harder in the age where like the models can just do a lot and so there's you know some of these things can be like skillified and like maybe don't need their own dedicated product but I think it's like the hard stuff is still hard it's like understanding the needs of people like figuring out how you're going to reach them uh listening to them and iterating on them really quickly like it is still the case that like a group of four or five people obsessed with a problem is going to move faster than those same people at any other kind of organization that are like you know subject just to the complexity I just mentioned the like the fact that we have, you know, a lot of different products that kind of interoperate. Like that's a interesting constraint that we have to work through. It's an advantage in other ways, right? So, yeah, I'm still like very long and bullish on startups and um it's just uh it papers over the fact that like writing code was never the like the limiting part. You know, maybe it was on the timeline perspective, but it was never like the thing that was going to like make or break your startup. It's really that space and user understanding.

Speaker B: 是的,领域知识。

Original English

Speaker B: Yeah. Uh domain knowledge.

垂直领域的 AI 应用:医疗与金融

Speaker A: 是的。今天也是我们的垂直 AI 主题日。我们的一位重磅返场嘉宾 Chris Lovejoy,他之前一直都在谈论垂直 AI。他以前在医疗健康领域的 Interior 公司,最近我又邀请他回来,结果发现,你们刚刚聘用他来负责你们的医疗健康业务。我们下一个重要的焦点领域也是金融。你知道我们有一个 AI 与金融的议程轨道。你们最近刚在纽约市举办了一场盛大的金融主题活动。而且我们下一场 AIE 的焦点某种程度上也会是金融。你在那个领域看到了什么?针对 Claude 来说有什么潜力吗?显然那里有大量的 Excel 电子表格。

Original English

Speaker A: Yeah. Uh today is also our day for vertical AI. Uh one of our uh returning speakers and top speakers Chris Lovejoy uh was always talking about vertical AI. He was from interior in the healthcare space and then recently I I invited him back and turned out he you guys just hired him for your healthcare efforts. Um we also our next big one is also finance. Uh you know we have a AI and finance track. You guys just had a huge finance event in New York City. Um and where our next uh AIE is is sort of finance focus. what are you seeing there any you know any potential uh for cloud obviously a lot of excel excel spreadsheets

Speaker B: 是的,不,我认为那里确实大有可为,而且你可以明显看到模型在这个领域正变得越来越好,可以说是一代比一代强。目前有一些做得很不错的专注于金融垂直领域的初创公司,它们已经做出了自己的评估集,追踪这些进展也非常有意思。这并不是说我们只是在迎合评估集去优化,但它确实是一个很有用的晴雨表,能反映出对于这些金融用例,模型是否真的在变得更好。我认为未来将会出现一种有趣的融合,即模型拥有足够的灵活性,能够潜入数据之中,实时生成分析、仪表板或工作流,但同时又具备某种虽非绝对不可变、但至少是经过验证的数据集的支撑。因此,如果让这一切完全自由发散,我认为那将是制造混乱的根源,也绝对不是金融服务领域大多数公司想要的。所以,找到那条合适的界线——在某一边你拥有可验证性、审计日志以及数据来源追踪,但又不会以此限制你在其上构建的应用程序的类型——我认为这正是我们在该领域看到的很多技巧所在。并且我认为,如果你能很好地解决这个问题,你就能鱼与熊掌兼得。困难的部分在于,过去为了实现这种可验证性而构建的许多系统,几乎从设计之初就决定了它们对于在其之上运行的智能体工作负载缺乏足够的灵活性。所以我认为在技术栈的两端都存在着机遇。

Original English

Speaker B: yeah no I think that there's there's a lot in there too and that's like an area where uh you could see the model get clearly better at it like sort of generation to generation and there's you know there's some good sort of vertical specific uh finance startups that have like done their own um evolves which has also been interesting to to track and it's not like we're like sort of playing to the evol but it is a useful sort of barometer on like is this actually getting better um at these financ use cases. I think the interesting blend that's going to happen um is this mix of again the model having the flexibility to like dive in and create just in time analyses or dashboards or workflows with like some sense of like what is the not immutable but at least like verified sort of set of data and so like uh set having all of that be totally free form I think is a recipe for confusion and is like not what most companies in the financial services space want. So finding that right uh sort of cutline where you have verifiability and audit logging and and sort of data provenence here but not in a way that constrains the kinds of applications that you can build on top I think is a lot of the art that we're seeing in that space as well. Um and I think you know if you solve it well you can you can get the best of both worlds. The hard part is a lot of the systems that were built to do the verifiability out of like are kind of almost by design not super flexible in terms of agentic workloads on top. So I think there's opportunity at both sides of the stack there.

AI 行业的快节奏与心理健康

Speaker A: 是的。我也同意,我们会在纽约进一步探讨这个话题。我想以心理健康作为最后一个话题来结束,在技术会议上我们对此讨论得还不够多。你看到过很多超高速的增长,人们总是不停地刷新他们的时间线,这让人筋疲力尽。对于那些处于 996 工作状态的人,你有什么建议来避免职业倦怠呢?

Original English

Speaker A: Yeah. Um I think I I also agree we'll be exploring that in in New York. Um the last thing I want to end on is on mental health which we don't talk about enough in technical conferences. Um you've seen a lot of hyperrowth. People are just always refreshing their timelines and it's exhausting. Um how do you advise people who are working 996 to avoid burnout?

Speaker B: 是的,我的意思是我认为这是个难题。我是说,我很确定你们都在经历这一切,因为你们都在这个行业工作。你要知道,现在的强度是过去的数倍,事情发展的速度也快得多。比如在 Instagram 的时候,我们只在考虑两件事:苹果会在 WWDC 上宣布什么?它会把我们彻底搞砸还是会助推我们?对吧。而那是每年才一次的事情。但现在,我们每三四个月就会遇到竞争对手发布新产品。这绝对不是……我们每周开全体员工大会时都会谈到这个话题。大会通常在周三举行,我们会有一张幻灯片叫“AI 界的这一周(目前还只是周三)”,而且不可避免地,总会有某个竞争对手发布了新模型,或者有了新产品,或者在监管方面发生了一些有趣的事情。事情发展得真的非常非常快。我认为,我努力保持相对理智的方法之一,就是真的要抽出时间休息。我认为优秀的联合创始人在这方面做得很好,他们会告诉你:“听着,如果你精疲力竭了,你基本就完蛋了。”我很不幸地看到过这种事情发生在我非常亲近的人身上,而且这需要很长很长的时间才能恢复过来。所以,我实际上在鼓励大家:没有任何工作重要到你不能离线几天的程度。我认为这是非常关键的一点。所以我坚信这一点。如果做不到,那你可能做错了什么,你应该找个能做你导师的人谈谈,弄清楚你如何才能打破这种困境。然后我认为另一点是,我喜欢体育运动,这就有一种观念:你永远不像你表现最好的一场比赛那么出色,也永远不像你表现最差的一场比赛那么糟糕。我觉得这也是非常真实的。我知道在 AI 领域,会有那种“一切都完了,我们又活过来了”的极端情绪循环。就像如果你……

Original English

Speaker B: Yeah, I mean I think this is a hard one. I mean and it is I'm sure you all are experiencing this because you're all working in this industry like it is you know multiples more intense and things move much more quickly like at Instagram like our two things that we were thinking about was like what is Apple going to announce at WWC and is it going to like totally mess us up or boost us right so that's like once a year um or you know we have competitor launches every three or four months right and uh it is definitely not that is a topic we we do it when we do our weekly all hands it's usually on Wednesdays and we have a slide that's like the week in AI parenthesis and it's only Wednesday and like and inevitably like some competitor has shipped a new model and like there's been like new product and maybe there's some interesting thing happening um uh on the regulation side like things are moving really really quickly. I think the way I try to stay at least relatively sane um one is like actually carving time off and I think the topic co-founders do a good job of like saying like look like burn out if you burn out like you're kind of done and I've seen it happen unfortunately to people I'm really close to and then it takes a long time to recover from that. Um so actually encouraging people like there's no job that is so important that you can't be offline for a couple of days. Um so I think that's like a a big key like piece in there. So like that's strongly believe um and if it is you're probably doing something wrong and you talk to somebody who could be a mentor to figure out how you can uh unblock that. Um and then I think the other one as well is like uh I love sports and like uh this is the notion of like you're never as good as like your best game and you're never as bad as your worst game. I think that's also really true. Like I know like in AI there's like the you know it's so over we're so back thing. Like that like if you

创业低谷的视角与情绪管理

Mike: 要内化这种循环总是会在某种程度上起作用的想法。你会意识到情况并没有那么糟。就像本·霍罗威茨(Ben Horowitz)在《创业维艰》(The Hard Thing About Hard Things)这本书里有一章讲到的,“我们完蛋了,一切都结束了”,作为一家初创公司,你们中的很多人可能都有过这种感觉。你会觉得,“天哪,我简直不敢相信发生了这种事。我们肯定永远无法从中恢复过来了。”我们在 Instagram 早期绝对有过几次这样的经历,然后你熬过来了。当你真正挺过那段时期,这种经历也就定义了这家公司。我试着提醒我自己和这里的团队,甚至在 Anthropic 内部也是如此,那就是:看,这是一个快速发展的领域,但这也是一场持久战。我们永远不只是为了今天的模型发布和外界反应,或者这个产品发布或别的什么,你是在打一场比赛,你也是在建设。你只需要相信,你正在建设能够度过那些难关的团队和文化,并保持这种大局观。即使这种大局观是说:“看,三个月前我们也处于类似的位置。也许不需要一年,只需几个月就能好转。”但这仍然像是在拉长视角,不要让事情……不要让你内在的自我认同和对成功的定义被日常的起伏所过度左右。

Original English

Mike: internalize that that cycle is always going to be at play in some way. You realize like it's never that bad. Like Ben Horowitz's uh book is the hard thing about hard things uh has this chapter on like we're effed it's over and like that feeling as a startup that probably many of you have had at startups where you're like oh I can't believe this thing happened. Like we're never going to like recover from this. I defin we definitely had it on Instagram a couple of times and then you get through it and like it's that like def like defines the company when you can actually go through that. I try to remind myself and the team here even within anthropic which is like look this is a it's a fastmoving but it's also a long game and it's like we're never it's never just about today's model launch and reaction or this product launch or something else like you're playing and you're building. You just have to trust that you're building like the team and culture that is going to get through those things and have that sense of perspective. Even if perspective is saying like look three months ago we were in the similar position. Maybe it's not a year it's just a matter of months but it's still like zooming out and not thinking things not letting your internal sort of like sense of self and success be so driven by the dayto-day.

Interviewer: 是的。有没有哪位教练或导师对你说过什么话,让你在艰难时期反复用来提醒自己,并帮你度过难关的?

Original English

Interviewer: Yeah. Has anyone any coach or mentor said something to you that you repeat to yourself uh that gets you through the tough times?

Mike: 嗯,我认为最重要的一句话是这样的感觉:如果你感受到了什么,团队里的其他人通常也同样感受到了。这是我的教练给我的建议,就是说要把情绪表达出来。甚至可以直接说:“嘿,我对这件事感到非常有压力。”或者,“是的,我们要关闭这个实验室项目,我真的很难过。”其实几个月前我就开过这样的会,当时我非常努力地在做一件事,然后在会议一开始我就说:“我真的很难过,也很沮丧,我真希望这事能成。”我认为这给其他人留出了空间,让他们也能说:“是的,我也很生气”,或者“我也很难过”。我觉得,给出这样的建议,说不要……是的,我觉得如果你能让自己保持开放和脆弱,这通常也会让其他人把情绪表达出来,然后你们就可以从那里出发,说:“太好了,那我们接下来该怎么做呢?”你知道,从那个状态开始解决问题会容易得多。是的,我们实际上是用斯坦福大学管理 Touchify 的卡罗尔·罗宾斯(Carol Robbins)的一场分享来拉开 AIE 帷幕的。而且,我想不出还有比鼓励人们谈论自己的感受、管理心理健康并继续发布产品更好的结束方式了。

Original English

Mike: Um, I think the biggest one was uh this like uh sense of like uh if you're feeling something, it's really often the case that other people on the team are feeling it too. So, this is like advice I got from my my coach around just being like like just verbalizing emotions. Like even saying like, "Hey, I'm feeling really stressed out about this." Or, "Yeah, I'm really sad that we are shutting down this labs initi." I literally had this meeting a couple months ago where I was working really hard on something and I kicked off the meeting like I I'll kick it off like I'm really sad like and frustrated like I wish this thing had worked out and I think that holds the space for other people to be like yeah I'm pissed off too or like I'm sad too and like uh I think giving that advice around like not uh yeah I think if you can get yourself to be open and vulnerable it often like lets other people verbalize that and then you can from there you can be like great what are we going to do about it like you know it's much easier to start from that place. Yeah, we actually kicked off AIE with a session from Carol Robbins who runs Touchify at Stanford. Uh, and I can't think of a better way to end than encouraging people to talk about their feelings, manage their mental health, and keep shipping.

Interviewer: 是的。非常感谢你,Mike。

Original English

Interviewer: Yeah. Thanks so much, Mike.

Mike: 感谢你们的邀请。

Original English

Mike: Thanks for having me.

AI 智能体的下一步:从功能到数据接入

Announcer: 女士们,先生们,请欢迎我们的主持人、Replit 开发者关系工程师 Ralph Shabri 重返舞台。

Original English

Announcer: Ladies and gentlemen, please welcome back to the stage our MC developer relations engineer at Replet, Ralph Shabri.

Ralph Shabri: 好的,让我们再一次把掌声送给我们的主题演讲嘉宾。好的。我们今天上午学到了很多,对吧?我们讨论了一个新协议,也谈到了 AI 程序应该更像函数(functions)。我们还了解到,构建智能体(agents)的真正问题不在于模型本身,而在于围绕它的一切。不过伙计们,在我们进入分组讨论之前,我想宣布我们的下一位演讲者。首先是我们赞助商 Neo4j 的简短致辞。这里有人用图数据库(graphs)吗?好,有相当一部分人。相当一部分。好的,很酷。那么,我们的下一位演讲者将为大家讲述基于本体的语义层(ontology-based semantic layers)。当你在一句话中听到“本体(ontology)”和“语义(semantic)”时,你就知道接下来会有干货了。好了,事不宜迟,请和我一起欢迎 Neo4j 的 CEO Emil Eifrem 登台。

Original English

Ralph Shabri: All right, let's hear it one more time for our keynote speakers. All right. Okay, so we learned so much this morning, right? like we spoke about a new protocol and and we spoke about like that AI program should be more like functions and we learned that uh the real problem with building with agents is not the model but it's about everything around it. Um but before we go to breakouts guys I would like to announce our next speaker. Um so a quick word from our sponsor Neo4j. Anybody uses graphs here? All right, fair amount. Fair amount. Okay, cool. So, our next speaker is going to talk to you about ontologybased semantic layers. When you hear ontology and semantic in the same sentence, you know that you're up to up for something hot. All right, so uh without further ado, please join me in welcoming to the stage CEO of Neo4j, Emil Efim.

Announcer: 请欢迎 Neo4j 的创始人兼 CEO Emil 登台。

Original English

Announcer: Please welcome to the stage the founder and CEO at Neo4j, Emil.

构建企业级智能体的挑战与解决方案

Emil Eifrem: 大家好。在 Neo4j,我们与世界上一些最大的公司合作,帮助他们将数据准备好以供 AI 智能体使用。今天我想谈谈我们在过去六到九个月中观察到的一个新兴问题,并为此提出一个解决方案蓝图。假设我们在一家大型组织(比如一家大型银行)工作,我们想要编写一个智能体。假设这个智能体是帮助自动化银行开户流程的,对吧?你可以想象这非常适合自动化。你希望能编排这个流程。在这个简短的主题演讲中,我将使用赋予我的权限,极其粗略地简化这个智能体的样子。我将其分为两部分。第一部分,我们称之为业务逻辑(business logic)。它包含某种形式的意图理解、计划、行动,并且我们会围绕这些进行循环。这就是你的智能体所做的事情。我们知道,当智能体执行行动时,它并不总是直接操作数据。但同样清楚的是,为了让智能体取得成功,很大一部分关键在于在正确的时间为其提供访问正确数据的权限。因此,第二大块就是我们所说的数据源。你需要识别、弄清楚:“好的,为了解决我的问题,我需要访问这几个数据,并将它们连接起来提供给智能体。”以我们的开户智能体为例,我们可能需要能够验证身份。因此,我们可能会查看两个数据源:机动车辆管理局(DMV)的登记处,以及某种护照验证服务。然后,我们将这些数据接入我们的智能体,它就运作起来了。这很棒,这太棒了。而与此同时,你以及组织中的其他团队也在构建其他智能体,从概念上看,它们非常相似。所以,这很棒。太棒了。它确实起作用了。但这也带来了一些问题。首先,每次一个团队需要构建智能体时,他们都必须从头开始弄清楚该智能体运行所需的数据存放在哪里。如果你在一家初创公司工作,你只有一个应用程序,数据都存放在一个 Postgres 数据库里。这并不难。数据就在那个 Postgres 数据库中。但在企业级生态系统中,你拥有的不是一个数据库。你有 100 个数据库,你可能有 Snowflake 和 Databricks,你有 S3 存储桶等等。每次你都必须从头开始手动完成这项工作。当你找到数据源时,你知道,在企业中存在大量重复的数据。然后你还需要弄清楚:“这是正确的数据吗?是正确的版本吗?我能信任它吗?我有权限访问它吗?”等等。这也违背了软件工程的核心原则之一——DRY(Don't Repeat Yourself,不要重复自己)原则。因此,当某些事情发生改变,并波及到你所有的智能体时,你不得不一直手动地重新配置所有的智能体。这确实能起作用,但这只是大量的工作而已。最后,围绕数据源以及智能体如何操作这些数据源,没有任何“学习”机制。因此,当你的智能体明天醒来时,它并没有比今天变得更聪明。而且肯定没有任何跨智能体的学习,因为所有业务意图和数据源之间的连接,都是被编码在代码和提示词的组合中的。我知道大家在想什么。“Markdown 技能文件(Markdown files skills)来救场了!”既对也不对。演讲结束后,你可以来找我聊聊这个问题的完整版本,但我们已经看到许多团队试图仅仅依靠 Markdown 文件来解决这个问题。总结来说:它是解决方案的一部分,但并不是完整的解决方案。别光听我的,听听 SW 的话。一周前在 Latent Space 播客上,他说:“伙计们,你们必须去了解你们的数据库。你们不能仅仅用 Markdown 文件来‘凭感觉写代码(vibe code)’。”因此,我们最近一直在为一些真正大规模的组织解决这个问题,其中包括一家全球 20 强的银行、一家总部位于湾区的大型科技平台公司,以及一家领先的金融科技公司。目前浮现出的模式是,为了大规模应用智能体,我们需要构建在“更智能的共享基底上的轻量级智能体(thin agents on a smarter shared substrate)”。更智能的共享基底上的轻量级智能体。这在实践中是什么样的呢?它有三大支柱。第一个支柱是面向业务的本体(business-facing ontology)。“本体”这个词,我在这个领域长大,人们谈论过本体……

Original English

Emil Eifrem: All right. At Neo Forj, we work with some of the largest companies on the world to help make their data ready for AI agents. And today I want to talk to you about a problem that we saw emerging over the last call it six to nine months and propose a solution blueprint for that. So let's say that we work at a big organization, a big bank and we want to write an agent. Let's say that agent is helping automate the opening of a bank account, right? You can imagine that's very ripe for automation. You want to be able to orchestrate that process. And I'm going to use the powers bestowed upon me by a short keynote slot to grossly simplify what that agent looks like. I'm going to say there's two pieces. The first one is let's call it the business logic. Some version of interpreting intent and plan act and we loop around that. It's what your agent does. And we know that when an agent act, it doesn't always operate on data. But we equally know that in order for agents to be successful a huge part of that is giving it access to the right data at the right time. So the second big bucket is let's call it the data sources need to identify figure out okay in order to solve my problem I need access to these few things and wire them up and make them available to the agent. In the example of our account opening agent maybe we can imagine that we need to be able to validate identity. And so we might look at two data sources for that. the Department of Motor Vehicles, the DMV registry, and maybe some kind of passport verification service. So, we wire that up into our agent and it works. It's great. It's fantastic. And at the same time, you and other teams in your organization are building other agents and conceptually they look very similar. So, that's great. It's fantastic. It works. But it has a few problems. So first of all, every single time a team has to build an agent, they have to figure out from scratch where the data that they require for that agent to operate, where it sits, which if you work at a startup and you have one application, it sits on top of one Postgress database. That's not hard. The data is in that Postgress database. But in an enterprise ecosystem, you don't have one database. You have a 100 databases and you have Snowflake and data bricks probably and you have S3 buckets and so on and so forth. You have to do that work manually from scratch every single time. And then when you found the data sources, you know, in an enterprise there's lots of duplication of data. So then you need to figure out like is this the right data? Is it the right version? Can I trust it? Am I allowed to access it? So on and so forth. It also violates one of the core principles of software engineering, the dry principle. Don't repeat yourself. So when something change that cascades across all of your agents you have to kind of manually rewire all of them all the time which works but it's just a lot of work and then finally there's no learning around the data sources and how your agents operate on them. So when your agent wake up wakes up tomorrow it's not smarter than it was today. And there certainly isn't any cross agent learning because all of that wiring between business intent and the data sources is encoded in a combination of code and prompts. So I know what you're all thinking. Markdown files skills to the rescue and yes and no. Um you can come talk to me afterwards for kind of the full version of this but we've seen a ton of team that try to solve this problem using just markdown files. And the summary is it is part of the solution but it is not the solution. Uh but don't take it from me take it from SW. A week ago on the latent space podcast he said hey guys you got to learn your databases. You cannot vibe code with just markdown files. So, we've been solving this problem at scale for some really massive organizations recently, including a Fortune 20 global bank, a massive tech platform company based here in the Bay Area, and a leading fintech company. And the pattern that is emerging is that in order to do agents at scale, we need thin agents on a smarter shared substrate. thin agents on a smarter shared substrate. And what does that look like in practice? There are three pillars to that. The first pillar is a businessfacing ontology. And the word ontology, like I grew up in this world, people talked about ontologies

Ontology-based Semantic Layer and AI Graphs

Speaker A: 一直存在。最近它变得非常火热,这可能要归功于 Palantir,但同时也是因为人工智能的兴起。有很多人想要把本体(ontologies)做得很复杂,但其核心概念其实非常简单。你们组织中的关键概念是什么?在我们的银行业务示例中,包括客户、账户、借记卡、支票、交易,以及它们之间是如何关联的?但非常重要的一点是,它们的表达方式必须能让你在这个领域里工作的所有人类员工都能理解,对吧?也就是说,公司里所有员工都应该能通过这个名称来理解它。换句话说,你不会说 f_name。不,你有一个客户(customer),并且他们有一个名字(first name)。这就是第一个支柱:面向业务的本体。

Original English

Speaker A: forever. More recently, it's become very hype, probably thanks to Palantir, but also the rise of AI. And there's a lot of people who want to make ontologies really complex, but the core concepts are actually super simple. What are the key concepts in your organization? In our banking example, customers, accounts, debit cards, checks, transactions, and how do they all relate? But very importantly, they are expressed in a way that makes sense to all the human beings working in your universe, right? All the people working in your company, it's expressed in that name. In that way, in other words, you don't say f_name. No, you have a customer and they have a first name. So that's the first a business facing ontology.

Speaker A: 第二个支柱是技术本体。这就是你企业生态系统中所有数据源和数据资产的所有元数据。我有 14 个 Oracle 数据库。我有 15 个 Neo4j 数据库。我有 Snowflake 和 Databricks,我还有 S3 存储桶以及诸如此类的东西。它们位于哪里?它们的模式(schemas)是什么?所有这些有用的信息都包含在内。你可以通过三种主要方式来构建这个技术本体,我们以后可以讨论,虽然不是在本次演讲中,然后你在两者之间建立一个映射。因此,那个拥有“名字(first name)”的客户,那个名字有一个记录系统,在另一边的 Oracle 数据库中有一个名为 F_ename 的列,这就是两者之间的映射。

Original English

Speaker A: The second pillar is a technical ontology. This is all the metadata of all the data sources and data assets in your enterprise ecosystem. I have 14 Oracle databases. I have 15 Neo4j databases. I have snowflake and data bricks and I have S3 buckets and all of that kind of stuff. Where do they sit? What are the schemas? All of that kind of good stuff. You construct that technical ontology in three key ways that we can talk about later though not in this talk and then you have a mapping between the two. So that customer that has a first name, that first name has a system of record and over there there's an Oracle database with a column called F_ename the mapping between the two.

Speaker A: 然后第三个支柱是运行时信号,即你的智能体(agents)在遍历这个图并执行任务时发出的信号,它们会留下关于“我尝试了什么?”、“我成功了吗?”、“结果是什么?”的执行追踪记录。这些执行追踪构成了这三大支柱。好的,让我们在我们银行开户智能体的上下文中来看看这个。这是一个简化的视图,但你可以看到这个图。它结合了业务概念,比如支票、账户和信用历史等等。这是一个遵循流程的智能体,或者说是流程引导的智能体。我们希望这种类型的智能体能够切实地遵循一个流程。因此,我们也在本体中对业务流程进行了编码。

Original English

Speaker A: And then the third pillar is the runtime signals out of your agents when they walk this graph and they execute they leave the traces around what have I tried? Was I successful? What was the outcome? The execution traces those three pillars. Okay. So let's look at that in the context of our bank account opening agent. This is a simplified view but you can see this graph here. It has a combination of business concept like checks and accounts and credit history and stuff like that. This is a process following agent or a process guided agent. We want this type of agent to actually follow a process. So we've also encoded that in the ontology a business process.

Speaker A: 接下来,如果你看看被绿色包围的那个节点,也就是合规性检查(check compliance)节点,我们切换到技术本体,并将其放入此处的图中。我们发现并编码了这样一个规则:为了进行合规性检查,你可以想象你需要验证一个政府颁发的 ID;然后我们指出,在这个特定的组织中,有两个数据源可以帮助我们做到这一点。它们是机动车辆记录和护照验证。好的,这非常棒。所以,当我们的智能体进入这里,并意识到“我需要检查合规性,我需要一个政府颁发的 ID”时。这里就有两种方法可以解决这个问题。当它们执行并尝试这些方法时,它们会留下第三个支柱,即执行追踪记录。它们比这个简化幻灯片上展示的要复杂得多,但涉及诸如“好的,我在哪里?”、“我做了什么?”、“我的上下文是什么?”以及“我成功了吗?”之类的信息。最终,它会导出一个某种形式的评分。你可以将其用作输入。比如,“好的,我在使用 DMV(车辆管理局)查询时非常成功,那么在下一次调用时,如果我处于正确的上下文中,我就更有可能选择这个方法。”

Original English

Speaker A: And then if you look at the node that is surrounded by green the check compliance one we flipped to the technical ontology and we've put in the graph here we've discovered and encoded that in order to do a compliance check you might imagine that you need to resolve a government-issued ID and then we say that in this particular organization there are two data sources that can help us with that. It's the motor vehicle records and the passport verification one. Okay so that's really great. So then when our agents come in here and they realize I'm going to check compliance, I need a government-issued ID. Here are the two ways that I can resolve that. When they execute and they try that, they leave the third pillar, the execution traces for that. And they're more sophisticated than what's on this simplified slide, but involves things like, okay, where was I? What did I do? What is my context? And was I successful? And ultimately, it leads out to some kind of a score. and you use that as input. It's like, okay, I've been very successful using the DMV lookup for example, then I'm more likely to choose one if I'm in the right context in my next invocation.

Speaker A: 基于本体的语义层的三个支柱——业务本体、技术本体和执行追踪,将它们结合在一起,就解决了所有四个问题。我们现在有了一种非常简单的方法来发现数据源。我们知道它们是否可信。我们可以自上而下地通过某种人工整理的知识来确认这一点,对吧?也就是某个管理员之类的人设定的。我们也可以自下而上地通过执行追踪来了解。这是在实际操作中真正行之有效的东西。我们有一个统一的管理位置,将业务意图和概念映射到这些数据源上,这样我们就不用重复劳动了。如果某些东西发生了变化,它会级联到我所有的智能体中,对吧?而且我们还具备了自我学习的能力。所以,我的智能体明天醒来时,会比今天稍微聪明一点。而且这种自我学习不仅限于单个智能体,还跨越了所有的智能体。因此,我们正在从一个充斥着手动连接数据源的“厚重智能体(thick agents)”的世界,迈向一个我们在更智能、基于共享本体的语义层上运行“轻量级智能体(thin agents)”的世界。这使得我们能够构建多得多的智能体,而无需每次都重新对它们进行工程开发。在更智能的共享底层基础上构建的轻量级智能体。

Original English

Speaker A: Three pillars of the ontology based semantic layer, a business ontology, a technical ontology, the execution traces taken together, they solve all four of the problems. We now have a very easy way to discover the data sources. We know if they're trustworthy or not. We know that top down by some kind of human curated knowledge, right? An administrator of some sort saying it. We also know it bottom up through the execution traces. This is what actually worked in reality in practice. We have a single govern place that maps business intent and the concepts to those data sources so we don't repeat ourselves. If something changes that cascades across all my agents, right? And we have self-learning. So my agent that wakes up tomorrow is slightly smarter than it was today. And not just self-learning on an individual agent, but across agents as well. So we're moving from this world, a world of thick agents with manually wired data sources into this world where we have thin agents on a smarter shared ontology based semantic layer. And this allows us to do a ton more agents without having to re-engineer them every time. Thin agents on top of a smarter shared substrate.

Speaker A: 如果你觉得这很有趣,我们有一份文档,一个包含有关此内容的更多信息的网页。如果你看到这里的二维码,你也可以来展台和我们交流。我们在世博会 P3 有一个大展台。我们喜欢谈论这类事情。但这还不仅限于此,这只是其中一种模式,一种非常令人兴奋的模式,我们目前看到人工智能中利用图技术的这种模式吸引了大量的关注。但还有数以百计将图和人工智能结合起来的更有趣的模式。其中 10 个正在 2005 房间刚刚开始的图技术专题分会场中进行。你会听到来自比尔及梅琳达·盖茨基金会(Gates Foundation)、Monday.com、摩根大通(JP Morgan Chase)、伯克利(Berkeley)、纽约时报(New York Times)等组织的非常精彩的演讲。所以,去看看那些吧。

Original English

Speaker A: If you think this is interesting, there's a documentation, a web page that outlines more information about this. If you see the QR code here, you can also come and talk to us at the booth. We have a big booth here at the expo P3. We love talking about this kind of stuff. But not just that, this is one pattern, a very exciting pattern that we see a lot of traction around right now for using graphs in AI. But there's hundreds of more interesting patterns that combines graphs and AI. 10 of them is actually in the graph track that is kicking off right now in room 2005. And you have some really amazing talks from organizations like the Gates Foundation, Monday.com, JP Morgan Chase, Berkeley, New York Times, and so on and so forth. So go check out that thing.

Speaker A: 最后,这主要是围绕那些你要处理众多数据源和众多智能体的组织。但如果你是一家基于 Neo4j 进行构建的初创公司,我们非常爱你们。Neo4j 有一个非常棒的初创公司计划(startup program)。你可以获得免费积分,但更重要的是,我们建立了一支专门的解决方案工程团队,他们每天都免费与初创公司合作,帮助他们在 Neo4j 中建立数据模型,针对性能进行调优等等。所以,请注册我们的初创公司计划。非常感谢大家。祝大家会议愉快。祝大家有美好的一天。

Original English

Speaker A: And then finally, this was primarily centered around organizations where you deal with many data sources and many agents. But if you're a startup building on Neo4j, love you. There is a startup program for Neo4j that is phenomenal. You get access to free credits, but more importantly, we've built up a dedicated solution engineering team that spend every day working with startups for free, helping them model their data in Neo4j, tune it for performance, and so on and so forth. So, please sign up for our startup program. Thank you very much. Enjoy the conference. Have a good day, everyone.

RL-Guided Remediation for ETL Failures

Ana Marie Benzson: 一个生产数据作业在几个小时前失败了。仪表板停止了更新。你花了一整天的时间检查日志、模式(schema)以及上游数据。现在已经过了午夜。同一个问题一直在脑海中盘旋:到底什么改变了?失败本身可能很小,但昂贵的部分是围绕它的所有一切:检查、诊断、选择安全的应对措施、重新运行作业,以及确认我们没有把数据变得更糟。大家好,我是 Ana Marie Benzson。在这次演讲中,我将展示一个基于强化学习(RL)引导的系统,该系统为 ETL(提取、转换、加载)失败选择有边界的修复操作。核心问题不仅仅是智能体是否能够采取行动,而是它能否发挥实际作用、其行为是否可解释,以及是否在一个运维团队真正信任的边界内采取行动。

Original English

Ana Marie Benzson: A production data job failed hours ago. The dashboard went stale. You have spent all day checking the logs, the schema and the upstream data. And now it is past midnight. The same question keeps coming back. What changed? The failure itself may be small but expensive part is everything around it. Inspection, diagnosis, choosing a safe response, rerunning the job, and confirming that we did not make the data worse. Hi, I'm Ana Marie Benzson. In this talk, I will show an RLG guided system that selects bounded remediation action for ETL failures. The central question is not simply whether an agent can act, but whether it can act usefully, explainably, and within boundaries that an operations team would actually trust.

Ana Marie Benzson: 云端 ETL(提取、转换、加载)故障很少会以一个清晰、标记明确的异常形式出现。我们看到的是延迟或不可用的数据源、模式漂移(schema drift)、日期时间不兼容、速率激增、类型更改,以及与运行手册中任何内容都不匹配的运行时错误。通常的应对方式是人工工作流:检查日志,形成诊断结论,采取临时修复措施,重新运行作业,并验证输出。每一步都很合理。延迟源于交接过程、不完整的上下文,以及避免采取不安全修复措施的需要。在顶点评估(capstone evaluation)中,手动恢复基线时间被模拟为大约 2.5 个工作日。这代表了一个事件经历正常的排队、调查和批准的流程。因此,工程目标非常明确:对于那些常规的、可识别的故障,压缩这个处理循环;同时将那些不确定、新颖或高风险的案例升级处理。

Original English

Ana Marie Benzson: Cloud Detail failures are rarely arrived as one clean well-labeled exemption. We see late or unavailable sources, schema drift, datetime incompatibilities, now rate spikes, type changes, and runtime errors that do not match anything in the runbook. The usual response is a human workflow. Inspect the logs, form a diagnosis, a temporary pair, rerun the job, and validate the output. Each step is reasonable. The latency comes from headoffs, incomplete context, and the need to avoid an unsafe fix. In the capstone evaluation, the manual recovery baseline was modeled at roughly 2.5 working days. This represents an incident moving through normal queing, investigation, and approval. So, the engineering objective is specific. compress that loop for routine recognizable failures while escalating the cases that are uncertain, novel or high risk.

Ana Marie Benzson: 这个图表展示了我顶点项目中的端到端 AWS 架构。现有的 AWS Glue 作业发出一个“作业失败”事件。Amazon EventBridge 捕获该事件并触发运行智能体的 Lambda 函数。Lambda 会从两个只读来源收集证据。CloudWatch 提供错误日志,而 Glue Data Catalog 提供当前的模式元数据。系统使用这些信号对故障进行分类,评估数据质量和运维风险,并构建传递给强化学习(RL)决策引擎的状态。然后,策略会提出一个有边界的响应。在执行程序使用 Glue API 重新触发作业或应用已批准的修复措施之前,安全层会检查该提议。Amazon S3 用于存储智能体生成的工件(artifacts)、审计日志和隔离的输出数据。最后,作业将被重新运行和验证。

Original English

Ana Marie Benzson: This diagram shows the end to end AWS architecture from my capstone. An existing AWS glue job emits a job failed event. Amazon event bridge catches that event and triggers the Lambda function that runs the agent. Lambda gathers evidence from two readonly sources. Cloudwatch provides the error logs while the glue data catalog provides the current schema metadata. The system uses those signals to classify the failure, assess the data quality and operational risk and construct the state passed to the RL decision engine. The policy then proposes a bounded response. The safety layer checks that proposal before the executive can use the glue API to retrigger the job or apply an approved remediation. Amazon S3 stores agent artifacts, audit logs and quarantined outputs. Finally, the job is rerun and validated.

Ana Marie Benzson: 所以这就是一个闭环的运维流程:监控、诊断、评分、决策、检查安全性、采取行动并验证恢复情况。顶点项目的实现使用了客户提供的合成数据。公开的

Original English

Ana Marie Benzson: So this is close operational look. Monitor, diagnose, score, decide, check safety, act and verify recovery. The capstone implementation use synthetic data provided by the client. The public

智能层的架构分离与Q学习策略

Speaker A: 通过一个经过清理的、通用的部署模板,代码库保留了这种模式。智能层特意将三个关注点隔离开来。确定性的异常规则用于确立可观察到的事实。比如一个字段消失了、数据类型发生了改变,或者空值率超过了某个设定的阈值。Q学习(Q-learning)策略负责在给定当前事件状态的情况下处理上下文动作的选择——系统究竟应该是重试、恢复模式(schema)、回滚、隔离、升级(escalate),还是仅仅记录该事件即可。然后,在学习到的策略之外,还设有一个安全覆盖机制(safety override)。举例来说,如果异常极其严重(critical),而策略却提出了一个像日志记录这样的被动动作,安全覆盖机制就会将该选择转换为一次安全升级。这种分离正是该项目的设计主旨:为事实设定规则,为有限的选择进行学习,并为权限设立护栏。在选择动作之前,系统必须先确立实际发生的事情。模式分析器(schema profiler)负责提取结构、类型、嵌套关系以及空值率统计数据。漂移检测器(drift detector)将当前的分析器与基线进行对比,并识别出添加、删除和类型变更。数据质量分析器检查数据的完整性、有效性和一致性。错误分类器将日志模式映射到不同的故障系列,而风险评分器将这些信号转化为可操作的风险等级。按照设计,这些组件对于直接可观察到的数据条件是确定性的。比起不透明的推断,一条明确的规则更容易被验证、解释和审计。随着事故历史数据变得更加丰富和具有代表性,一些分类器可能会演变为基于学习的组件。但是,“准备好使用机器学习(ML ready)”与“必须使用机器学习(ML required)”并不是同一回事。最简单且最可靠的组件应当主导每一个决策。策略会接收到一个紧凑的状态表示,包含故障类别、风险级别、重试次数、漂移严重程度以及数据质量状况。接着,它会从六个动作中进行选择:重试、恢复、回滚、隔离、升级或记录日志。我之所以使用表格化的Q学习(tabular Q-learning),是因为状态和动作空间都很小。Q表(Q table)在评估时计算成本很低,而且每一个决策都可以被直接审查。针对这个特定的状态,这些是各个动作的价值,而这是排在首位的动作。从技术层面上讲,每一次事件都被建模为一个单一的事件。

Original English

Speaker A: repository preserve this pattern through a sanitized generalized deployment template. The intelligence layer deliberately separates three concerns. Deterministic anomaly rules establish observable facts. A field disappeared, a type change or the null rate cross a threshold. The Q-learning policy handles contextual action selection given the current in incident state should the system retry coers the schema roll back quarantine escalate or simply log the event. Then a safety override sits outside the learned policy. For example, if the anomaly is critical and the policy proposes a passive action such as logging, the override converts that choice into an escalation. This separation is the design thesis of the project, rules for facts, learning for bounded choices, and guard rails for authority. Before selecting an action, the system has to establish what actually happened. The schema profiler extracts a structure types nestic and null rate statistics. The drift detector compares the current profiler with a baseline and identifies additions, removals, and type changes. The data quality analyzer checks completeness, validity, and consistency. The error classifier maps log patterns into failure families, and the risk scorer turns those signals into an operational risk level. These components are deterministic by design for directly observable data conditions. An explicit rule is easier to validate, explain and audit than an opaque inference. With richer and representative incident history, some classifiers could become learned components. But ML ready is not the same as ML required. The simplest reliable component should own each decision. The policy receives a compact state failure category, risk level, retry count, drift severity and data quality condition. It then selects from six actions. Retry, covers, roll back, quarantine, escalate or log. I use tabular Q-learning because the state and action spaces are small. The Q table is cheap to evaluate and every decision can be inspected directly. For this state, these were action values and this action one. Technically, each incident is modeled as single.

为Token分配工作:AI代理系统的新视角

Caitlyn: 早上好。我们非常激动能在这里,在“AI工程师(AI Engineer)”大会上与大家相聚。我是Caitlyn,我在Anthropic领导平台工程(Platform Engineering)团队。

Original English

Caitlyn: Good morning. We're super excited to be here at AI Engineer with all of you. I'm Caitlyn and I lead platform engineering at Anthropic.

Angela: 我是Angela。我在Anthropic负责领导平台产品(Platform Product)。今天我们想和大家探讨一个概念,这是我们和团队花了很多时间去思考和研发的理念,也就是我们认为:应该给Token(词元)分配具体的工作(tokens should have jobs)。

Original English

Angela: And I'm Angela. I lead platform product at Anthropic. And today we want to talk to you about a concept that we've been spending a lot of time thinking about and working on with our team, which is this idea that we think that token should have jobs.

Angela: 因此,如果你正在构建一个基于代理(agentic)的系统,并且你试图实现某个特定的结果——你想用代理来完成某件事——为了获得更好的结果,所有人通常都会拉动一个杠杆,那就是增加你的预算,这意味着你要消耗更多的Token,或者使用更昂贵的Token。但我们一直在思考,这就够了吗?在使用预算这种假设背后,隐藏着一种潜在的观点,即每一个Token基本上都是可互换的(fungible)。我们就在想,这真的是事实吗?所有这些Token真的都是可互相替代的吗?

为了验证这一点,我们提出一个设想:如果我们给Token分配具体的工作会怎样?想象一下,在正常情况下你设定一个代理去执行任务的方式:你给它那个任务,你给它这个Token预算,然后所有被消耗的Token基本上是不加区分的,意思就是它们都在做同一份工作——纯粹的“执行”(executing)。但是,如果你拿出其中的一些Token,让它们不再仅仅是执行,而是去做一些其他的工作呢?

Original English

Angela: So, if you're building an agentic system and you're trying to accomplish some specific outcome, you're trying to get something done with agents, there's one lever that everybody pulls in order to get a better outcome, and that's usually increasing your budget, which means you spend more tokens or you spend more expensive tokens. But we've been wondering, is that all there is? Underlying this assumption of uh using the budget is this kind of implicit perspective that every single token is basically fungeible. And we've been wondering is that actually true? Are all these tokens actually fungeible? And to test that we've been thinking what if we gave tokens jobs? So if you think about the way that you would normally set up an agent to go accomplish a task, you give it that task, you give it this token budget, and then all the tokens that are being spent are basically indiscriminate in the sense that they're all doing one job. they're just executing. But what if you take some of those tokens and they're not just executing, they're doing some other job.

Angela: 举个例子,也许你抽出部分Token,让它们去“建议(advising)”那些正在执行任务的Token。又或者,执行任务的Token试图把某件事做好,你则分配另外一些Token,去实际评估执行者(executor)做得有多好(grading),这样它就能不断迭代并再次尝试。或者也许你有一些正在“做梦(dreaming)”的Token,它们在回溯反思其他执行者所完成的工作,并将学到的东西写入内存,以便下次能更好地执行任务。

我们将这些模式分别称为一种“策略(strategy)”——当你拿出一些Token去执行,并分配另一些Token去做其他工作时。那么,让我们先来看看第一种策略,即“建议策略(advising strategy)”。在这里,我们将系统拆分为一个执行者(executor)和一个顾问(adviser)。执行者显然负责执行,但关键在于,它可以向顾问呼叫以寻求建议。然后它可以采纳这些建议,并弄清楚自己下一步是否做得正确。在许多用例中,这非常有用。例如,如果你正在构建一个销售代理,在一个理想的场景下,你会希望那个销售代理能够实际帮助销售人员,在跟进工作逾期或交易陷入停滞时发出提醒。在这种架构中,拥有一个顾问来确保所有不同的部分都能切实有效地运转,是非常有帮助的。

Original English

Angela: So for example, maybe you take some of your tokens and they're advising the tokens that are executing. Or maybe the tokens that are executing try to get something done well and you take some other tokens and you actually grade how well the executor is doing so that it can iterate and try again. Or maybe you have tokens that are dreaming. They're reflecting back on the job that other executors have done and writing learnings to memory so that they can do it again. And what we call each of these, if you take some tokens that are executing and some tokens that are doing some other job, let's call this a strategy. So let's go take a look at the first strategy, the advising strategy. Here we're splitting up an executor and an adviser. The executor obviously executes, but crucially they can call out to adviser for advice. and then they can take this device and figure out if they're doing the next step correctly. This is really helpful in use cases. For example, if you're building a sales agent, in an ideal world, you'd have that sales agent be able to actually help the sales rep flag when a follow-up is overdue or deal is stalling. In this construct, having an adviser to be able to kind of make sure that all the different pieces are actually working is really helpful.

Angela: 另一个例子是评估(grading)。假设你正在执行任务,且你很清楚“好”的标准究竟是什么样子的。你可以将这些标准定义在一个评分量表(rubric)中,然后每一次执行者尝试去实现那个结果时,你可以配置一个评分员(grader),让其看着那个量表来给执行者完成的情况打分。如果执行者做得很好,那太棒了,任务就可以结束。但如果它做得不够好,你就可以再次迭代,直到你得到那个好的结果。在实际操作中,可能需要用到这个策略的一个例子是:假设你有一个客服代理,你正在经营一家商店,你的客户写信进来说:“啊,我买的这东西应该退款。”然后你的客服代理需要能够做出回复。对于在什么情况下你会给某人退款,你可能有一些相当具体的标准。因此,你能做的就是定义一个包含那些标准的评分量表。你可以安排一个评分员,去检查客服代理正在做的工作,并判断它做得对不对,是否得出了正确的结果。

Original English

Angela: So, another example is grading. Let's say you're executing and you kind of know exactly what good really does look like. You can define this in a rubric and then each time an executor tries to accomplish that outcome, you can have a grader provisioned that grades how well the executor did while looking at that rubric. And if the executor did a good job, then great, it can be done. But if it didn't do such a great job, you can iterate again until you get that good outcome. So, an example in practice of when you might want to use this is let's say you have a customer service agent and you're running a store and your customers are writing in and they're saying, "Ah, I should get a refund for this thing." And your customer service agent needs to be able to respond. You probably have some like pretty specific criteria on when you would give somebody a refund. And so, what you can do is define a rubric that uses that criteria. You can have a grader that goes and looks at the work that the customer service agent is doing and decide is it getting it right and is it coming to the right outcome.

Angela: 我们拥有的最后一种策略是“做梦(dreaming)”。在“做梦”模式中,同样有一个自然而然去执行任务的执行者,接着还有一个“做梦者(dreamer)”。做梦者实际上能够去检查执行者所做的工作以及运行记录(transcripts),然后它会提取出自己发现的任何线索,并把它们写入内存。这份内存会在下一轮被执行者重新读取利用。因此,理想情况下,它就能得到改进。应用这个策略的一个极佳用例是,如果你正在构建一个招聘代理。要知道,招聘工作需要在候选人是否合适、双方是否契合的问题上,进行大量的互动和反馈。所以,通过收集所有这些类型的数据,如果你在这个代理上构建一个“做梦”类型的策略,它实际上能够对接下来的工作进行打磨,使其在以后的回合中变得越来越有用。

Original English

Angela: And the last strategy we have is dreaming. So in dreaming there's an executor who naturally executes and then there's a dreamer. The dreamer is actually able to inspect the work and the transcripts of the executor and then it takes any of the findings that it has and it writes them to memory. This memory is repicked up by the executor for the next round. So ideally would have improved. A great use case for this is if you're building a recruiting agent. Now recruiting requires a lot of interaction with feedback on whether or not a candidate does or doesn't make sense and if it's a good fit between both parties. And so by taking all this type of data, if you build a dreaming type of strategy on this agent, it's actually able to kind of sharpen the next round so that it's more and more increasingly useful.

用实验数据验证Token分工策略

Angela: 那么,让我们通过一些实验来把它具体化。我们所做的是,建立了一个测试基准(bench),包含了一系列与金融分析相关的任务。我们在这些任务上所做的,就是试图在现实世界中去复制一位人类金融分析专家的表现。它们在这些各种各样的任务上表现到底能有多好?所以我们的做法是,首先从一个仅负责执行的控制组(control)开始。让我们尝试一下每一项任务,然后对它们进行评估,我们真的是只在做纯粹的执行。但是紧接着,我们可以试验我们的每一种策略,并看看我们表现如何。

所以,我们从一个超级基础的实验开始。让我们仅仅是“一次性(one-shot)”完成它。让我们采用每一种策略,直接去尝试完成这些任务,然后看看我们的准确率如何。你可以看到,在纯粹执行(executing)的情况下,它表现得不是很好,准确率只有15%。但因为这只是一个一次性的尝试,策略可以选择它自己想要消耗多少Token。于是,你可以看到“执行策略”决定不花费那么多Token,仅用了39,000个。而随着我们转向量级更大、更为复杂的策略时,我们确实选择了消耗更多的Token,但我们也取得了更好的成绩。不过,这其实并没有告诉我们太多信息,因为确实,“做梦(Dreaming)”策略表现得非常非常出色,但它耗费了惊人的60万个Token才达到这个效果。

Original English

Angela: So let's make this concrete with some experiments. So what we did was we created a bench of a bunch of tasks related to financial an analysis. And what we were doing with each of these tasks is trying to replicate in the real world a expert human financial analyst. How well would they do on each of these various tasks? And so what we did was we start with a control that's just executing. Let's try each of these tasks and we'll eval them and we're literally just executing. But then we can experiment with each of our strategies and see how well we perform. So, we start with a super basic experiment. Let's just one-shot it. Let's take each of our strategies and we'll go and just make an attempt to accomplish these tasks and we'll see how accurate we are. And so, you can see here with executing um it didn't do so well. 15% accuracy, but because it was just a oneshot, the strategy got to choose how many tokens it would actually spend on its own. And so, you can see that execute decided not to spend that many tokens, only 39,000. And as we go into our larger strategies, our more complex strategies, we did choose to spend more tokens, but we did a better job. So, this isn't really telling us much because sure, Drain did really, really well, but it used a whopping 600,000 tokens to get there.

Caitlyn: 没错。所以,为了真正弄清楚改变Token的工作类型是否会产生任何“超额收益(alpha)”,我们需要做的是保持预算恒定。为了做到这一点,我们将把“做梦”策略的预算,也就是那大约60万个Token,作为全盘固定的最大预算额度。然后……

Original English

Caitlyn: That's right. So, in order to actually figure out if varying the jobs produces any alpha, what we need to do is hold the budget constant. And to do this, we're going to take Dreaming's budget, that 600,000 or so, as the maximum budget that is fixed across the board. And

策略性能与 Token 预算

Speaker A: 我们为每一种策略提供这个预算,以分析它们的表现如何。再次如预期的那样,如果你给很多策略更多的预算,你会看到性能全面提升。所以执行(execute)从 0.15 上升到了 76。建议(advise)和评分(grade)从 60 多上升到了接近 90。这再次是符合预期的,因为如果你给事物更多的测试时间计算资源,它们通常表现会更好。但如果那是唯一重要的事情,我们实际上应该预期看到在给定的完全相同的 token 预算下,执行、建议、评分、做梦(dream)实际上都在完全相同的水平上。但我们实际看到的是,存在一个 alpha 或者说存在一个差异,因此我们可以利用这个 alpha。如果你在这个完全相同的预算水平上看执行,它达到了 76,但建议达到了 89。所以虽然很小,但它确实存在,因此有我们可以看看的 alpha。

Original English

Speaker A: we give every single strategy this budget in order to analyze how well it's performing. And as expected again, if you give a lot of strategies more budget, you are going to see performance increase across the board. So execute went from 0.15 to 76. Advise and grade went from the 60s to closer to the 90s. And that's again expected given the fact that if you give things more test time compute, they should generally perform better. But if that was the only thing that mattered, we should actually expect to see execute, advise, grade, dream actually all be at the exact same level given the exact same token budget. But what we're actually seeing is that there is an alpha or there is a difference and therefore an alpha for us to exploit. If you look at execute at those exact same budget level, it gets to 76, but advise is at 89. So while a minimal, it does exist and so there is alpha for us to take a look at.

Speaker A: 现在,我们决定从一个完全不同的角度来看看这个分析。正如 Kayla 所提到的,你知道,我们在现实世界中为一组非常复杂的金融任务做这个基准测试。我们想要分析真实专家如何使用 AI 代理。所以,如果我们看一个金融分析师专家,对吧,他们需要用 AI 代理做的那种任务是,他们给它一些非常具体的东西,比如制作一份损益表(P&L),然后他们得到结果。现在,如果那个结果在基准测试中准确率是 80%,听起来很棒,但在现实中,这对那个专家意味着他们必须回去自己重新计算那份损益表,或者选择让它再运行一次。这是因为在这种领域的这种任务中,如果你不是 100% 准确,它实际上是没用的。你不能编造一个收入数字,或者你不能编造一个成本数字,对吧,你必须确保它是 100% 准确的。所以从与这个领域相关的现实世界后果的角度来看,我们需要重新计算我们的实验并以不同的方式对它们进行评分。至关重要的是,我们需要确保我们的实验有这样的结构:如果它得满分,我们实际上会给它及格。如果它在那种任务上得分低于 100%,我们实际上会将其标记为失败。

Original English

Speaker A: Now, we decided to take a look at this analysis from a completely different lens. And as Kayla mentioned, you know, we're doing this bench for a very complex set of financial tasks in the real world. And we wanted to analyze the usage of agents with actual experts. So, if we look at a financial analyst expert, right, the kind of task that they need to do with an agent is that they're giving it something very concrete like let's say make a P&L and then they're getting the result back. Now if that result is 80% accurate on a bench that sounds great but in reality what that means for that expert is they have to go back and recompute that P&L themselves or alternatively send it through another run. And that's because in this kind of domain for this kind of task if you're not 100% accurate it's actually not useful. You cannot make up an income number or you can't make up a cost number right you have to make sure that it's 100% accurate. So with this lens of the real world consequence associated with this domain, we needed to recompute our experiments and score them a bit differently. Crucially, we needed to make sure that our experiments had this kind of construct where if it was scored perfectly, we'd actually give it a pass. And if it scored anything less than 100% on that kind of task, we would actually mark it as a failure.

Speaker A: 所以让我们用这个视角来看看我们实验数据的不同切面,我们正在寻找这种完美的运行,100% 准确率的及格。让我们看看这些策略中每一个能够达到及格的时间百分比。嗯,我们让执行策略下降到了 42%,我们更复杂的策略做得好一点,准确率高达 75%。嗯,再次强调,这并不一定能告诉我们很多信息,因为,你知道,每种策略可能会随着时间的推移选择使用不同的预算,对吧?所以我们在这里做的是固定预算,我们说在固定预算内,这些策略每一个表现如何?所以对我们来说真正重要的是,如果你试图得到这个完美的答案,并且你在现实世界中,你正在经营一家企业,对你来说重要的是你得到那个完美答案的成本。所以我们可以思考这个问题的一种方式是,比如我们有我们的执行策略。执行策略大约在 40% 的时间里会给你那个完美的答案。所以平均而言,你可以预期必须运行它三次,并且你应该希望在那三次运行中的某个时候得到一个完美的答案。正如我们之前谈到的,我们将我们的预算固定在那个最高 token 预算的策略上,即 60 万个 token。所以如果你在每次单独的运行中花费 60 万个 token,你必须运行大约三次。你可以预期平均必须花费 180 万个 token 使用执行策略来得到你的完美答案。

Original English

Speaker A: So let's look at a different cut of our data from our experiments with this lens where we're looking for this perfect run, 100% accuracy pass. And let's look at what percent of the time each of these strategies was able to achieve a pass. Um so we've got executes um down at 42% and we've got our more complex strategies doing a bit better up to 75% accuracy. Um and again this doesn't necessarily tell us a ton because um you know each of these strategies might um choose to use different budgets over time, right? So what we did here was we fixed the budget and we said within a fixed budget, how well do each of these strategies perform? And so what really matters to us actually is if you're trying to get this perfect answer and you're in the real world, you're running a business, what matters to you is the cost to you to get to that perfect answer. And so one way we can think about this is we had our execute strategy for example. The execute strategy around 40% of the time will give you that perfect answer. So on average, you can expect to have to run it three times and you should hopefully sometime in those three runs get a perfect answer. And as we talked about earlier, we fixed our budget to that highest token budget strategy, which was 600,000 tokens. So if you spend 600,000 tokens in each individual run, you have to run approximately three times. You can expect on average to have to spend 1.8 million tokens with the execution strategy to get to your perfect answer.

Speaker A: 所以如果我们把这个分析全面应用到所有这些策略中,这实际上就是在这个领域为了让那个 AI 代理对那个策略有用所花费的真实成本。正如 Caitlyn 提到的执行策略,这是我们的基准线,对你来说真实的总体 token 成本将是 180 万。但是,建议、评分和做梦向我们展示了一些不同。至关重要的是,当你考虑到这些代理各自最终输出的实际使用情况时,建议和评分实际上在 token 使用上是相当高效的。

Original English

Speaker A: And so if we take this analysis and run it across the board against all these strategies, this is actually the true cost it took in this domain for that agent to be useful for that strategy. So as Caitlyn mentioned for execute, which is our baseline, this is going to be 1.8 million true total token cost for you. But advise, grade, and dream are showing us a bit of difference. Crucially, advise and grade are actually quite token efficient when you think about the actual usage of the end output of each of these agents.

根据业务需求优化 Token

Speaker A: 那么这对你作为一家企业意味着什么呢?嗯,这实际上完全取决于你想要优化什么样的事情,而且这会因人而异,对吧?会有一些企业说,“实际上,对我来说,最重要的事情是 token 要非常高效。在那种情况下,你可能应该选择建议类型的策略,以便解决你想优化的那个特定领域的问题。” 还会有其他领域或其他企业,你会说,“我不会那么在意 token 的效率,因为我真正关心的是那个答案的可靠性,所以我需要最大化我得到那个完美答案的运行百分比。” 在这种情况下,你实际上会选择完全不同的策略。你可能会倾向于选择评分或者做梦。

Original English

Speaker A: So what does this mean for you as a business? Well, it actually really depends on what kind of thing you want to optimize for and it's going to vary, right? There's going to be businesses who say, "Actually, for me, the most important thing is to be really token efficient. In that case, you should probably pick the advised type of strategy in order to solve for that particular domain in which you want to optimize that." There's going to be other areas or other businesses where you're going to say, "I'm not going to care so much about token efficiency because what I really care about is reliability of that answer and so I need to maximize the percentage of runs in which I get that perfect answer." In which case, you would actually pick completely different strategies. You probably lean towards grade or dream.

Speaker A: 所以,如果你要带走一件事,我们想让每个人思考的是这样一种观点:token 不是可以随意替换的。你可以用你的 token 去执行。你可以用暴力破解的方式完成你的任务,你可以投入更多的预算。但是如果你变得非常聪明,让你的 token 去做这些不同的工作并尝试这些不同的策略,你非常非常有可能在固定预算内为手头的任务获得更好的结果。

Original English

Speaker A: So, if you take away one thing, the thing we want everyone to think about is this idea that tokens are not fungible. You can use your tokens to execute. You can brute force your way through your tasks and you can throw more budget at it. But if you get really smart about having your tokens do these different jobs and try these different strategies, you're very very likely to be able to get a better outcome for the task at hand within a fixed budget.

构建代理架构与编排

Speaker A: 那么让我们来谈谈我们实际是如何构建策略的,以及我们是如何将其付诸实践的。嗯,我们做了大量工作,为各个代理创建了一个非常出色的测试平台。嗯,如果你看底部这里的这张图片,这实际上是我们用于云托管代理的架构,这是我们在云平台内提供给你的代理解决方案。我们在此基础上所做的是,我们开始进入元测试平台级别,比如多代理编排和执行级别,在这里这种策略可以继续进行,并在我们的执行者和我们的顾问或者我们策略中的其他代理之间进行协调。而且其中一些功能,比如做梦和输出结果,我们实际上在云托管代理中是开箱即用提供给你的。所以,有了这些基元集,对我们来说,构建这种架构实际上相对容易,在这里我们能够结合这些不同类型的策略并弄清楚它们应该如何协同工作。

Original English

Speaker A: And so let's talk a little bit about how we actually build strategies and how we bring this to life. Um so we've done a lot of work to create a really excellent harness for individual agents. Um and if you see this uh picture at the bottom here, this is actually the architecture that we've used for cloud managed agents um which is our agentic solution that we give to you within the cloud platform. And what we do on top of this is we start to get into the meta harness level like the multi-agent orchestration and execution level where this strategy can go and be um coordinated between our executor and our adviser or the other agents within our strategy. And some of these um like dreaming and outcomes we actually give to you out of the box within cloud managed agents. So with those set of primitives it's actually relatively trivial for us to construct this kind of you know architecture where we're able to combine these different types of strategies and figure out how to they should work together.

Speaker A: 所以,举个例子,对我们来说相对容易说:好吧,现在有了这个,我可以接收一个任务,我应该能够执行它,同时也允许它提供建议,而且 Fable 又上线了。所以我们实际上可以说 Fable 就是那个真正在向执行者提供建议的模型。然后我可以获取所有这些结果,并说把它们发送给一个评分者,这样我就可以确保这在一个合理的循环中进行验证。如果它通过了,那太棒了。我想把所有这些东西发送给“做梦”,并确保我的下一次运行比以往任何时候都要好。当然,你不必就此止步,对吧?如果正确的基元在那里,正确的协调在那里,那么你实际上可以构建非常复杂的设置,这些设置适合你所有的不同类型的动态问题。你可以发明这些类型的大规模架构。再次强调,非常容易。你也可以发明全新的工作,而不仅仅是 Caitlyn 和我在这场对话中展示的那些部分。

Original English

Speaker A: So for example it's relatively trivial for us to say okay now with this I can take a task and I should be able to execute it but also allow it to advise and fable is back online. So we could actually say fable is the one that's actually advising uh the executor. And then I can take all these results and say send them to a grader so that I can make sure that this is verifying in a loop that makes sense. And if it passes, that's awesome. I want to send all of that stuff to Dreaming and make sure that my next run is better than ever. And of course, you don't have to stop there, right? If the right primitives are there and the right coordination is there, then you can actually construct really complex setups that fit for all the different types of dynamic problems that you have. You can invent these kinds of large-scale architectures. Again, very trivy. And you could also invent completely new jobs, not just the ones of the pieces that Caitlyn and I have presented in this conversation.

Speaker A: 所以,随着时间的推移,我们的一个大目标是让我们的模型越来越好,让我们的平台越来越好地在你工作时为你动态构建这些策略。但在那期间,当我们在朝着那个目标努力时,我们希望你能继续思考这样一个想法,即你应该给你的 token 分配工作,你应该通过结合这些基元来使用不同的新颖策略,以便为你想要的任务获得你想要的结果。

Original English

Speaker A: So, a big goal that we have over time is to get our models better and better and our platform better and better at dynamically constructing these strategies for you as you're doing work. But in the meantime, as we're working our way there, we would love for you to continue to think about this idea that you should give your tokens jobs and you should use different novel strategies by combining these primitives in order to get the outcomes that you want for your tasks.

将第二大脑变成研究记忆库

Speaker B: 感谢大家加入我们。把我的第二大脑变成我活生生的研究记忆库。让我来解释一下。所以,在我的第二大脑中,我目前在 Obsidian 里有超过 5000 条笔记,在 Readwise 里还有另外 5000 条笔记,还有一些零散的

Original English

Speaker B: >> Thanks for joining us. turning my second brain into my living research memory. Let me explain. So, within my second brain, I currently have over 5,000 notes in Obsidian and another 5,000 notes in Readwise and some scattered

构建个人 AI 研究系统的初衷

Paul: (这都存储)在 Notion 和 Google Drive 里。而且所有这些内容的平均增长速度是每个月 250 个文件。这就是我想要的。在左边,你可以看到我的整个 Obsidian 笔记库,这简直是一个巨大的混乱局面。每当我想开始做点什么,比如写一篇文章、启动一个新项目、研究一个新的代码库、开发一个新功能或者别的什么事情,我实际上希望能够提取出高信号的节点,那些对当前工作真正有用的节点。你可能会问自己,为什么不直接使用 CodeX、Claude Code 或者 NotebookLM 呢?事实是,我确实在用,但你需要一个系统,一个介于这些工具和你的“第二大脑”之间的系统。好吧,让我们回到我问题的根源,那就是我总是丢失我的研究资料。例如,我的阅读列表简直就像个“坟墓”。当我在刷社交媒体时,我看到了那个很酷的帖子、一篇新文章、一个新的 YouTube 视频、或者一个 GitHub 仓库,不管是什么。每当我真正想开始着手工作时,我从来不记得我的“第二大脑”里到底存了什么,或者我必须花大量时间去寻找能在工作中有意义地利用的笔记,对吧?我面临的另一个问题是,我希望这个系统能真正锚定在我的个人笔记、个人价值观和个人信仰上。我希望这个系统是个性化的,能反映我自己的思想,对吧?这就是为什么在今天的视频中,Louis-François 和我将教你如何构建属于你自己的 AI 研究操作系统。这也包含了代码,所以你也可以自己动手尝试一下。我是 Paul。我是 Decoding AI 的创始人和 CEO,我在那里制作了大量关于如何交付 AI 产品的课程内容。我同时也是畅销书《LLM 工程师手册》的合著者。而在本视频中,我将教大家使用的这个系统、这个 AI 研究操作系统,正是我在日常工作中使用的系统。现在,我将把火炬传递给 Louis-François。

Original English

Paul: in Notion and Google Drive. And all of this is growing on average with 250 files per month. And this is what I want. On the left, you can see my whole obsidian vault, this huge mess. And whenever I start working on something such as an article, a new project, a new codebase, a new feature or whatever, I want to actually pull high signal nodes that are actually useful for my current work. And you would ask yourself, why not use directly calledex code or notebook LM? And the thing is that I am, but you need a system that sits between those harnesses and your second brain. Okay, so let's go back to the root of my problem, which is that I'm always losing my research. For example, my reading list is a graveyard. When I'm scrolling social media and I have that cool expost, a new article, a new new YouTube video, a GitHub repository, it doesn't matter. Whenever I actually want to start working on something, I never recall what I have in my second brain or I have to spend a ton of time actually finding meaningful notes that I can use in my work, right? And another problem that I have is that I want this system to actually be anchored into my personal notes, into my personal values, into my personal faith. I want this system to be personal, to reflect my own thoughts, right? And that's why in today's video, Luis Franuis and I will teach you how to build your own AI research OS. This also comes with code, so you can also try it out yourself. And I'm Pauline. I'm the founder and CEO of Decoding AI where I do a ton of content on courses on how to ship air products and I'm also the co-author of the LM engineers handbook bestseller and the system the air research OS that I will teach you in this video is the system that I use in my daily work and now I will pass the torch to Luis Francois.

工具选择:何时该使用个人 AI 研究操作系统?

Louis-François: 谢谢你,Paul。我是 Louis-François。我是 Towards AI 的联合创始人兼 CTO,我们在那里制作教育课程。我也是“What's AI”的创作者,这是一个 YouTube 频道,我在那里讲解 AI 工程技术。我以前讲解 AI 研究,现在则专注于 AI 工程。我也是《为生产环境构建》一书的作者。在此之前,我是一名博士生。所以说,我真的是以做研究为生的。就像我说的那样,我以前做过 AI 领域的博士研究,做了大量的研究和研究工作。现在,我制作课程,我撰写视频脚本,我为视频做调研,我还为公司开展培训——以此为生。而我所做的所有这些事情,都始于非常出色的调研,并且也利用了我们在 Towards AI 为客户构建项目时获得的大量知识和见解。所以,我也和 Paul 一样,有海量的笔记,我们也试图尽可能地利用它们。正如你将看到的,我们构建了某种工具来利用我们的“第二大脑”,但你也会发现,我和 Paul 在使用它的方式上会有一些不同。而这正是我们构建的这个代码仓库的核心目标,在这个项目中,我们希望你能根据自己的需求去对其进行调整。整个目标是:我们如何能让研究变得更好?但更具体地说,我们如何能更好地利用我们现有的东西?所以,让我们深入探讨一下。首先,我们需要弄清楚该使用哪个工具以及何时使用。因为我们构建的这整个研究系统并不是为了应对每一个查询。如果你只需要一个快速的答案,比如几个简单的问题,或者只是你基本上会去 Google 搜索的东西,那很显然,你直接用 Google 就行,或者问 ChatGPT、Claude,用你喜欢的任何系统都可以。但是,这样做的问题在于,如果你有许多后续的追问,或者你需要构建一个更大的项目并且需要共享非常长的上下文或海量信息,那仅仅依赖 ChatGPT 就不是很理想了。而且这也意味着你完全依赖于 OpenAI 或其他大模型厂商的架构。因此,这里的下一步是问问自己,对于一个更复杂的问题,你需要快速采取行动,还是你想构建一些下一个层级的功能并做一些非常困难的事情?如果你只是为了快速修改一个小代码库、写一篇文章,或者只是做一件你知道不会怎么重复发生的事情,那么绝对可以使用 CodeX、Claude Code 或者是你信任的某个 Agent。但有时候,你需要不断深入挖掘以使其变得更好、提高效率、甚至进行更多优化。因此,通常当你必须这样做时,你希望你的研究资料、你的研究成果能够留存下来,并且能够在将来回过头来参考。所以,如果你想要这样一个流程,在这里你找到的资料和你记下的笔记能够随着时间保存下来,并能让一个 Agent 能够高效地利用它们,并能够回到这些信息上来进行后续提问或进一步消化内容。比如现在,当我制作一个新视频时,我也希望系统里的 Agent 能理解我以前制作过的视频,避免内容重复,不要老调重弹,并且能引用其他的内容。在这种情况下,有些工具就非常有趣了。你之前可能尝试过,比如 NotebookLM,它在做研究、高效消化内容并回溯信息方面非常强大。但 NotebookLM 的问题在于……首先,最主要的问题是你并不拥有它。你不能随心所欲地用它做你想做的事,你无法尽可能多地进行个性化定制。它并不是 Agent 原生的,而且很显然,由于它只是基于浏览器的,它在编码任务方面表现得很弱。所以这与 Paul 和我所需要的东西、以及大多数 AI 工程师普遍需要的东西相去甚远。因此,如果你需要你的 Agent 能够利用你所做的一切——无论是一项大型研究、一个新视频、还是你写的代码和内容——你通常希望你的其他 Agent、你的其他项目能够从你刚刚完成的工作中学到东西。对于生产环境,尤其是对于一款产品,我们建议的一件事是构建某种包含向量数据库的检索 RAG 流水线。但这需要基础设施。对于人们快速消化信息、查看笔记、进行编辑来说,它并不是非常人性化。人工很难去检查它。你需要围绕它构建一切。如果只是作为我每天想用的东西,那它绝对不理想。显然,它在规模化运作时超级强大,非常有趣,尤其是在产品中。但是正如我所说的,这个项目是为我们自己准备的,我并不想要一个作为产品实时的、超级专业的东西。我只想要一个我自己能用的东西,一个我的 Agent 和不同的项目都能尽可能好地利用的东西。所以,这里我们最后要问自己的问题是,如果你希望所有这些都在,但又想要更多的个性化呢?也就是,一个个性化的研究助手,它能构建出某种随时间推移不断积累的“维基百科”,而且易于检查和使用。你在里面有海量的资料来源、文档、视频、对比、实现方案、新研究、以及你不断添加并希望轻松利用和复习的新主题。这时候,你可能就想自己构建一个东西了。在我们的例子中,我们构建了“个性化研究操作系统”(Personalized Research OS),我们将在这个演讲中准确分享我们构建了什么以及如何构建。但缺点是,它绝对需要比直接打开 Claude Code 更多的设置工作。目前,使用 Claude Code 和其他 Agent 工具的一个主要问题是,你会提供给它们链接、PDF 和各种不同的信息。例如,在看我最近的关于循环工程视频后,下次你在另一个会话中使用 CodeX,你不得不再次粘贴所有内容,或者要求它使用技能。而不管你使用的 CodeX 还是 ChatGPT,它们临时构建的用于利用你之前工作、运行过的脚本或已有脚本的那些结构,你全部都丢失了,或者它们被保存在一个你需要要求它重用的 skill 里,这通常不理想,而且只会随着时间越来越臃肿。问题在于,你给模型提供的所有这些信息并不是瓶颈。瓶颈在于你将来如何利用它。也就是说,对于一个 Agent 而言,上下文窗口变成了它的一切:数据库、文件系统、记忆、推理空间——它必须把这些全都包揽。而当你停止对话时,它就会失去一切。事情的真相是,为了获得更好的研究成果,我们并不必然需要提供越来越多、越来越长的上下文。你需要的是适当的记忆和上下文。

Original English

Louis-François: Thanks Paul. So I'm Lu France Bush. I'm the co-founder and CTO of Towards AAI where we build educational courses and I'm also the creator of what's AI a YouTube channel where I explain AI engineering techniques. I used to explain AI research before now focusing on AI engineering. I'm also the author of the book building for production and before that I was a PhD student. So I honestly make research for a living. I used to do a PhD as I said in AI and doing tons of research and research work. Now I build courses, I write videos, I research for videos, I build trainings for companies for a living. And all of these things that I do start with a very good research and also leveraging tons of knowledge and insights that we get at towards the eye from building for clients. So I have tons of nodes as well just like Paul and we try to leverage them the best possible and as you'll see we build some sort of tool to leverage our second brain where as you'll see there will be some differences between how I use it and how Paul uses it and that's the core goal of the repository that we built and on this project is that we want you to adapt it for your needs. The whole goal is how can we make research better? But more specifically, how can we better leverage what we have? So, let's dive into it. And first, we need to figure out which tool to use and when. Because this whole research system that we built is not for every query. If you just need a fast answer like a few quick questions or just something where that that you would just Google basically well obviously just Google it or ask JGBD cloud whichever system you want but the problem when doing that is that if you have a lot of following up question or it's a bigger project that you need to build on uh and have basically a very long context or tons of information to share relying on chibbity isn't ideal. And it also means that you are fully dependent on the architecture that openai or trag. So the next step here is to ask yourself for a more complex problem. Do you need to act quickly or do you want to build some next feature and and do something very difficult? If you just have a small repo for a quick change or write one article, just do one thing that you know won't be repeatable that much, definitely use codeex or cloud code or some agent that you trust. Sometimes you need to keep on digging to make it better, to improve efficiency, may optimize it more. And so typically when you have to do that, you want your research sources, your research to stick and to be able to refer to them in the future. So if you want a process like this where the sources that you find, the notes that you take stick around in time and have an agent be able to leverage that efficiently and being able to come back to these information to ask follow-up questions to digest content even more. And right now for instance when I make a new video I want also the agent of the system to understand the previous videos I made to not duplicate content to not repeat myself and to refer to some other content. In this case there are some tools that are very interesting that you might have tried before like notebookm that is super powerful to do research to digest content efficiently and to come back to it. But the problem with notebook is that it's well first the main problem is that you don't own it. You cannot do anything you want with it. You cannot personalize as much as possible. It's not agent native and it's obviously weak for coding tasks since it's just browser based. So it's far from ideal from something that Paul and I needed and that most AI engineers need in general. So if you need your agents to be able to leverage all you do, uh whether it is a big research, a new video, whatever you write, you do, you code, you typically want your other agents, your other projects to be able to leverage what you learn from what you just did. And one thing that we advise, especially for production, obviously for a product, is to build some sort of retrieval rack pipeline with vector databases. But this needs an infrastructure. It's not really human friendly to be able to digest quickly, to check notes, to make edits. It's hard to inspect by hand. You need to build everything around it. It's definitely far from ideal for just something I want to use on a daily basis. Obviously, it's super powerful at scale, very interesting, especially in a product. But as I said, this project is for us, and I don't want something live, super professional as a product. I just want something I will use and that my agents and different projects can leverage as best as possible. So the last question to ask ourselves here is that if you want everything there but more personalization. So, a personalized research assistant that builds some sort of Wikipedia that compounds over time and it and is easily inspectable and usable where you have tons of sources, documents, videos, comparisons, implementations, new research, new topics that you keep on adding and that you keep on wanting to leverage and review easily. This is where you may want to build something yourself. And in our case, we built the personalized research OS that we will share in this talk with exactly what we built and how. But the downside is that it definitely needs a bit more setup than just opening cloud code. Right now, the main problem with using cloud code and other agentic tool is that you give codec links, PDFs, and different information. For example, my most recent loop engineering video and then the next session you use codeex, you have to paste it all again or ask it to use skills and whatever structure that codeex or jgpd whatever tool that you use built on the fly to leverage what you did, the scripts it ran, the scripts it had, you all lost it or kept it inside a skill that you have to ask it to reuse and that usually isn't ideal and just grows and grows over And the problem is that all this information that you give to the model is not the bottleneck. The bottleneck is how can you leverage it in the future. Meaning that with an agent, the context window becomes everything. The database, the file system, the memory, the reasoning space. It has to do it all. And when you stop the conversation, it loses everything. And the thing is that we don't need necessarily to provide more and more and more context for a better research. you need a proper memory and context

Building a System with Plain Files

Speaker A: 管理,并且理想情况下要带有一些个性,特别是在我制作视频的情况下。所以我们做的是决定构建一个主要使用纯文本文件(大部分是 Markdown 文件)的系统,这样我们自己可以轻松利用,智能体也能轻松利用。我不会在这里详细展开,因为 Paul 会深入探讨这个问题。正如我所说,Paul 大概有 5000 多条笔记,而我只有几百条。但这只是为了说明,我们需要考虑到我们并不是从零开始的。我们俩都已经拥有了某种形式的大型数据库。就我而言,我制作了数百个视频,也记了很多笔记。所以,我仍然需要利用我过去几年创作的内容,以及我在为团队或客户构建产品时进行的大量会议记录,我想把这些也都利用起来,因为我们在为他人构建产品的过程中学到了很多。

Original English

Speaker A: management and ideally some personality with it especially in my case when I do videos. So what we did is that we decided to build a system with plain files mostly markdown files that we can leverage easily and that agents can leverage easily. I won't detail it very much here because Paul will talk about it in depth and as I said Paul has like 5,000 or something notes. I have just a few hundred. But that's just to say that we need to consider that we didn't start from nothing. We already both had some sort of large database. In my case, I made hundreds of videos and I take many notes. So, I still need to leverage these years of content that I already made and tons of meetings that I have with uh my team, with clients when we build for them that I want to leverage as well because we learn a lot by building for people.

AI Agents and the Plumbing Layer

Nikita Kotari: 大家好,欢迎你们。我是 Nikita Kotari,Salesforce 的高级技术员工。在这里,我正在构建 AI 驱动的企业解决方案。你们可能听说过 Agentforce、Agentforce for Setup、Headless 360,我参与了所有这些项目,它们都是非常棒的功能。我之前也曾在亚马逊和领英工作过。在过去的 6 个月或 1 年里,我们彻底改变了以往的工作方式。我们正在构建新系统,我们正在不断进化,而使用 MCP、CLI 和 Skill 会涉及到一些关键的方面。所以,今天我们将深入探讨构建 AI 智能体时一个关键却经常被忽视的方面——那就是“管道基础设施(plumbing)”。作为开发者,我们经常谈论模型大小、推理能力和提示词工程。但是,当涉及到构建生产级别的 AI 智能体时,智能体与环境交互的方式,也就是它的工具层,才是真正决定应用成败的关键。我们将研究三种截然不同的方法:MCP、CLI 和结构化的 Skill,并制定一个清晰的标准,说明在什么情况下该使用什么。在本次演讲结束时,你们将拥有一个清晰的心智模型,知道如何在这三个工具层之间做出选择。

Original English

Nikita Kotari: Um, welcome everyone. Um, I'm Nikita Kotari. I'm a senior member of technical staff at Salesforce. Uh, where I'm building AIdriven enterprise solutions. If you might have heard of agent force, agent force for setup, headless 360, I worked on all and it's like amazing features. Um, I also worked at Amazon and LinkedIn previously. Um, in last 6 months or a year we completely changed the way we used to work. We are building new systems. We are evolving and there are some critical uh aspect of using MCP CLI and skill. So today we are uh going to deep dive into a critical yet often overlook aspect of building AI agents. It's plumbing. As a developer we talk a lot about model sizes, reasoning capability and the prompt engineering. But when when it comes to building production ready AI agents, the way an agent interacts with the environment, it's tooling layer. It's what actually makes or breaks the application. We will look into three distinct approaches which is MCP, CLI and the structured skill and map a clear rubric for what to use when. Uh at the end of this talk, you'll have a clear mental model for choosing between these three tooling layers.

Production-Ready Agent Challenges

Nikita Kotari: 这里有一点是每个团队都会通过惨痛教训才能发现的:构建一个演示版的智能体大概只需要一个下午。但当你真正想把同样的东西部署到生产环境时,可能会花费一个季度的时间,因为我们要构建的是可靠、安全且快速的智能体。没有人希望在生产环境中泄露客户数据。因此,测试和可靠性仍然是这些 AI 智能体面临的主要问题。所以,我想谈谈这三个问题。问题一:上下文爆炸。设想一个场景,你将 50 个工具的 schema 加载到智能体的上下文窗口中,结果在这个任务甚至还没开始之前,60% 的思考空间就已经被消耗殆尽了。你的智能体会忘记之前的内容,因为用于监控数据库和部署的工具定义文件占满了窗口。它实际上已经没有空间去思考了。

Original English

Nikita Kotari: Um, here's something every team discovers the hard way. Building demo agent takes like an afternoon. But when you actually want to ship the same thing into the production, it could take a quarter because we are looking to build the reliable, secure and fast agents. Nobody wants to leak their customers data into the production. So testing uh reliability is still a major issues with this AI agents. So uh I want to talk about these three problems. The problem one is the context explosion. Uh consider a scenario where you load a 50 tool schemas into a agent's context window and suddenly 60% of your thinking space is burnt before even this task even starts. Your agent forgots the diff because uh tool definition for monitoring databases and the deployments file the window. It literally doesn't have a room to think.

Nikita Kotari: 第二个同样关键的问题是:隐性故障。假设一个例子,智能体进行了 7 次工具调用来起草一个 PR,结果出了差错,PR 被创建到了错误的分支上。这样一来,首先你浪费了大量时间去重做这项工作;其次,你在某种程度上失去了对智能体执行特定任务的信心。问题三:安全问题。想象一个例子:凌晨 2 点,你遇到了一个客户问题,你正在调查客户 A 的错误。但是等一下,LLM 对查询和租户数据拥有广泛的访问权限,它意外地开始查看客户 B 的数据。这就是一个多租户平台的问题,它不仅仅是一个 bug,更是一场巨大的合规噩梦。所以,让我们来谈谈如何使用这三个层来构建可靠、快速且安全的智能体。

Original English

Nikita Kotari: Uh the second yet critical problem is invisible failures. The agent made a like consider example where agent made a seven tool calls to draft a PR. Something went wrong. The PR was created on the different branch. Um so first you lost like lot of time to redo the work and the second thing is somehow you lost the confidence on your agent to perform a particular task and the problem three is the security surf. Uh consider example at 2 am you got a customer issue you are investigate getting a customer A's error but wait LLM has a broad access to the query and the tenant data and it accidentally started looking into customer B's data so that is a multi-tenant platform issue and it's not a just a bug it's a huge compliance nightmare so let's talk about how we can use this three um two uh three layers to build the reliable help fast and the secure um agent agents.

The Three Tooling Layers: CLI, MCP, and Skills

Nikita Kotari: 首先,我们有 CLI(命令行接口)。CLI 就像递给别人一把螺丝刀:“给你个工具,直接用吧。”它是直接执行的,智能体运行一个命令,获取输出,然后继续下一步。接着是我们的第二层,即 MCP。MCP 代表模型上下文协议(Model Context Protocol)。这就好比给别人一个 USB-C 扩展坞:“这是一个能适配任何服务的通用适配器,你用它来工作就行了。”第三,我们有 Skill(技能)。Skill 就像给别人一本操作手册:“这是这项任务的逐步执行指南,你照着做就行了。”所以它是结构化的,可以与任何平台或系统配合使用。总结来说,CLI 知道怎么做并直接执行命令;MCP 知道有哪些资源可用;而 Skill 知道具体的执行步骤。现在,让我们分别深入了解一下这三者。

Original English

Nikita Kotari: Um so we have the CLI. CLI is like handling someone a screwdriver. Here is a tool just run it. Direct execution. The agents run a command, gets output and moves on. Then we have a layer two which is MCP. MCP stands for model context protocol. This is just like giving someone a USBC hub. Here is a universal adapter that works with any services. Just use this and work on it. And third we have the skill. The skill is giving someone a runbook. Here is a step-by-step execution of this task. Just go and perform it. So it's structured and it could work with any of the platform or any of the system. So CLI knows how to do they execute command directly. MCP knows what is available and skill knows how to do it. So now let's dive deeper into one of each.

Nikita Kotari: CLI 是可读的。任何工程师都可以查看命令,并清楚地知道到底发生了什么。它们也是可复现的。如果在凌晨 2 点发生了故障,你可以复制粘贴同一条命令,在你的机器上运行它,从而复现同样的错误。它们还是可组合的。你可以通过管道将多个命令连接成一个并运行。这里的启发式原则很简单,如果我想让你们记住一句话的话,那就是:如果你团队中的工程师可以打开终端来完成这件事,那么智能体也应该能够做同样的事。我们有结构化的输入和结构化的输出。而且,它的可用性极高。我想,我们使用 CLI 已经有 50 多年了。所以,虽然它看起来不那么光鲜亮丽,但它们是经过实战检验且行之有效的。

Original English

Nikita Kotari: CLI they are readable. Any engineer can look at the command and see what's exactly happened. They are reproducible. If something fails at 2 am, you can copy paste the same command, run it on your machine and you can reproduce the same error. And they're composible. You can pipe multiple commands into one and you can run. The heristic is simple. If I want to want you to remember this, if a engineer on your team can open a terminal to do this, the agent should probably able to do the same thing. we have the structural input and structured output. Um, and it's it's highly highly uh usable. I think we have been using the CLI for more than 50 years now. So it's it's not glamorous but they are proven and it works.

Nikita Kotari: 然后我们有 MCP。大家可以把 MCP 看作是一个通用的 Type-C 适配器,就像每台设备都有自己的充电器,但通用的 USB-C 是标准化的,你可以将任何设备连接到你的接口上。所以在一侧,你可以看到 Claude、ChatGPT 以及各种编程工具;在另一侧,你可以连接各种不同的服务,比如工作项追踪器、代码搜索、错误追踪、监控系统等。现实生活中的例子可能就是 Git、Slack 和 Google 等等。所以在底部,你可以看到这在实际中是什么样子的。智能体带着查询条件调用 code_search.search,然后得到一个结构化的结果返回。所以,它处理了数百万个文件的索引,并在服务器级别强制执行租户隔离。这也是你解决安全问题的方式,因为每次你通过 MCP 协议连接外部服务时,都会有一个 OAuth 授权过程,你需要赋予它访问这些服务的权限。

Original English

Nikita Kotari: Then we have the MCP. MCP is consider it as a universal C adapter like um every device has its own charger but the universal uni USBC is standardized where you can connect any device to your um connector. So on one side you can see you have the clot, RGPD and the coding tool and the other side you can connect the different services like work item tracker, code search, error tracking, monitoring maybe real life examples could be uh get slack um and Google and everything. So uh at the bottom you see uh what the what this looks in practice. The agent calls the code search dot search with a query and it gets a structural result back. So it does the it handles the indexing across the millions of files and enforces the tenant isolation at the server level. This is how you solve the security problem too because every time you will connect the external services via MCP protocol there is a O and you will give uh permission to access those services.

Nikita Kotari: 第三个是 Skill(技能)。这是我最喜欢的工具之一。我觉得我已经用 Skill 把我在工作中的所有任务都自动化了,因为它高度可靠。它速度快,并且你可以用许多不同的方式对其进行调整。Skill 是一个围绕着你的工具构建的结构化操作手册。它定义了严格的编排逻辑:当满足某个条件时,确保先决条件存在,然后依次使用 CLI 和 MCP 执行步骤一、步骤二或步骤三,并显式地处理那些边缘情况。它防止了智能体去瞎猜,节省了内存,并且是标准化的。比如你可以选择:“我只希望连接两个 MCP 到这个 Skill”,这样你就没必要为了执行一个特定任务而去遍历所有的 MCP 工具。

Original English

Nikita Kotari: Third thing is the skill. This is one of my favorite uh tool and I think I have automated all of my work that I'm doing right now at my workplace with the skill because it's highly reliable. It's fast and you can tweak it it in a lot of different way. Um skill is a structural playbook that wraps around your tools. It defines the rigid orchestration. When a condition is met, ensure the this prerequisite is exist and then execute step one, step two or step three sequentially using the CLI, MCP and explicitly handle um handle those edge cases. It prevents the guesswork, it saves the memory and it's standardizable like you you only choose like I only want like two MCPS to connect to this skill. So you don't have to go and look into all of the MCPS to perform a particular task.

Mitigating Context Explosion and Applying Skills

Nikita Kotari: 现在,让我们来谈谈我们之前讨论过的问题。假设我的系统连接了 50 个 MCP 工具。因此,在运行任何任务之前,我已经消耗了 1.5 万到 2 万个 Token,在这个智能体甚至还没开始工作之前,我就丢失了大量的上下文。所以,你可以看到在一侧,你消耗了大量的 Token,以及你是如何通过使用 Skill 来替代这种做法的。始终保持来自 5 个不同 MCP 服务器的 55 个工具处于活动状态,这会给我们带来成本、延迟问题以及平台问题。相反,你可以简单地定义执行该特定任务所需要的 MCP 服务器数量。

Original English

Nikita Kotari: Now um let's talk about the problem that we discussed earlier. I have a 50 uh MCP um tools connected to my system. So before running any task, my 15,000 to 20,000 tokens are already burned and I lo I lost like lot of context before even starting my agent. Uh so you can see on the one side you have like a whole lot of um tokens and how you could replace the same thing by using the skills. So keeping the 55 tools from the five different MCP servers active at all times that cost us it caused the latency issue platform issues you can just simply define the number of MCP servers that that will needed to execute that particular task.

Nikita Kotari: 这里有一些例子,比如修复一个测试失败的问题。为了仅仅修复一个测试失败,你不需要遍历你所有的 MCP 工具。你只需要查看日志。你可以查看发布日志或者 GitHub,就能修复你的测试失败。至于起草 PR(Pull Request)以及调查升级问题,让我更深入地探讨一下。看看这张幻灯片,我们在这里正在起草一个 Pull Request。你需要什么呢?首先,智能体记录了起草 PR 的各项 Skill。第一步需要理解本地的代码更改,与其使用繁重的 API,不如直接调用快速且作用域明确的 CLI 命令;然后你需要一个 git 命令来创建一个 Pull Request……

Original English

Nikita Kotari: So here are like some of the examples like fixing a test failure. So you don't need to go into all of your MCP tool to just to fix one uh test failure. You can only look into the logs. You can look into the release log or and GitHub to fix your test test failure. And for the draft PR and the investigation escalation, let me dive a little bit deeper into it. So consider this slide where we are um drafting a pull request. So what do you need? um you need like um the first agent records the draft PR skills. Step one requires the understanding of the local changes instead of heavy API. It hits the fast scope CLI commands and then you you need a git command to create a uh get pull

构建 Agent Skill 与解决 Bug

Speaker A: (……发起)在正确的分支上的请求。所以你完全可以把具体的步骤写下来。你可以创建一个 Skill,并且可以不断地对它进行修改,以达到最好的效果。

Original English

Speaker A: ...request on the correct branch. So you can literally write down the steps. You can create a skill and you can keep modifying it to achieve the best results.

Speaker A: 另外一个场景是调查 Bug。如果在半夜你收到任何一个客户报出的 Bug,你首先要做的肯定就是查看日志,搜索错误模式(error pattern),然后查看最近的发布版本,确认是不是代码里存在 Bug 等等。因为你其实已经清楚自己要做什么了,所以你完全可以利用数量有限的 MCP 工具和 CLI 来自动化所有的操作,从而打造出一个世界级的 Skill,专门用来调试客户的问题。

Original English

Speaker A: Then again another thing is the investigation a bug. Um if midnight you get a any customer bug I mean the first thing you're going to do is like looking into the logs uh searching for the error pattern looking into the uh recent releases if there is a bug in the code or things like that. So you already know that like what you are going to do. So you can automate everything with limited number of MCP tools and the CLI to make a world-class skill to debug your customers issue.

Agent 开发的安全原则:环境隔离

Speaker A: 现在我们来谈谈安全问题。Agentic 开发的一个黄金法则是:隔离(isolation)必须在基础设施层面实施,而绝对不要放在 Prompt 里。因为 Prompt 可能会遭到注入攻击(prompt injection),但基础设施是无法被注入的。对于 CLI 来说,它的爆炸半径(blast radius)是整个 Shell 的访问权限,所以你们的风险缓解措施必须包含沙盒机制(sandboxing)。而对于 MCP 来说,除了需要严格的容器化,它的爆炸半径也被限制在你明确暴露出来的工具,以及基于租户(tenant)的隔离机制上。这些都能在服务器层面上提供保障。在构建任何 Agent 时,我们总有一些应该铭记于心的原则:我们不应该把任何凭证(credentials)共享给我们的 Agent。

Original English

Speaker A: And now let's talk about the security. A golden rule of the agentic development is enforce the isolation in your infrastructure never in your prompts. So prompt can be injected but the infrastructure cannot. For CLI the blast radius is full shell access. Uh so you mitigate um mitigation must include the sandboxing and the strict containerization for MCPS. The blast radius is limited to the explicitly exposing the tools and the tenant tenant isolation and that guarantees at the server level and we have some of the principles that we can always think about while building any agent that like we shouldn't be sharing any credentials with our agent.

Speaker A: 租户隔离是由基础设施层面强制执行的。我们始终可以使用权限网关(permission gate)在安全层面上进行管控,这样在让 Agent 执行任何任务时,我们就能为其提供最小权限(least privilege)。

Original English

Speaker A: The tenant isolation is the infrastructure enforced. We can always use the permission gate um gets to um operate at the security level and uh we can provide the least privilege while performing any task to the agent.

如何选择合适的工具层:CLI、MCP 与 Skill

Speaker A: 当你需要决定该使用哪一层时——因为你不可能每次都去创建一个 Skill,也不可能每次都一股脑地把所有 MCP 工具都开放出来——你需要问自己三个问题,从而获得最佳效果。第一个问题是:还有谁需要这个功能?如果只有当前这个 Agent 需要,那就使用 CLI;如果有多个 Agent 都要使用这项服务,那就使用 MCP;如果多个工作流(workflow)共享同一套处理流程,那就创建一个 Skill。

Original English

Speaker A: Now when when you are deciding which layer to use like every time you cannot create a skill or every time you cannot just explode to use all of the MCP tools. So you need to ask these three question in order to get the best results. The first one is like who else needs this. If it's just this agent just then use a CLI. If the multiple agents are going to use this service then use MCP. If multiple workflow shares the same procedure then use a skill.

Speaker A: 你可以问自己的第二个问题是:哪种故障模式(failure mode)对你们来说最关键?如果是透明度(transparency)和可复现性(reproducibility),那就选择 CLI。正如我们之前看到的那样,我们可以利用 CLI 轻松复现问题,但如果换作 Agent 就非常困难了,因为你根本不知道它们在后台具体干了什么。如果你们需要在租户隔离的环境中进行验证,那就去用 MCP;如果你们需要时序上的防护(sequencing guards),那就应该选择 Skill。

Original English

Speaker A: The second question you can ask yourself is like what is the failure mode that matters the most. If it's a transparency and the reproducibility go for the CLI like we have seen like how we can like reproduce the issues using the CLIs while it is very difficult if you go for um the agents because you don't know what they are doing in the background. If you need need a validation in the tenant isolation then go for the MCPS and if you need the sequencing guards then go for scale (skill).

Speaker A: 下面是第三个问题:你的上下文窗口(context)是否很紧张?如果上下文窗口很宽裕,你可以直接使用 CLI,也可以让 Agent 去探索所有可用的 MCP 工具来获取结果。但如果上下文非常吃紧,那你绝对需要引入 Skill 层,以及按需编排(on demand orchestration)的 MCP 工具,这样不仅能省钱,还能改善延迟(latency)。

Original English

Speaker A: And here are the three uh third question that like whether you your context is tight. If it's generous then you can go CLI you can explore all the MCP tools and get your results. If it's very very tight then absolutely you need a skill layer um and on demand orchestration of the MCP tools uh to save the money as well as you can improve the latency.

总结与核心架构收获

Speaker A: 总结一下,我想留给各位三个基础的架构理念作为心得。第一,保持你的上下文干净。不要让 Agent 使用所有可用的 MCP 工具去执行任何任务,而是要把你的大部分工作流转换为 Skill,这样才能得到最好的效果。第二,可移植性(portability)是有代价的。MCP 服务器确实非常出色,但它们也会带来网络延迟、协议抽象以及基础设施复杂性等问题。因此,我们应该把这份代价用在刀刃上,即为了共享服务、索引(indexing)以及隔离去支付它。所以,明智地选择你们的 MCP,而且永远不要觉得使用 CLI 会有什么丢人的。CLI 我们已经用了 50 年了,它们具有极高的可组合性、易于调试,并且是久经沙场考验的。所以,拥抱你的 Agent 框架吧,把你所有的工作流都转化为 Skill。

Original English

Speaker A: And to wrap this things up, uh, I want to leave you with the three foundational architecture take takeaway. Keep your context clean, uh, give only don't use like all of the MCP tools available to perform any task. Give them the convert your um, most of your workflows into the skills to get the best results. Um, portability has a cost. MCP serves are incredible but they introduce network latency, protocol abstraction and infrastructure complexity. So we need to pay that uh cost deliberately for shared services, indexing and the isolation. So choose your MCPS very wisely and never be embarrassed by the CLIs. It's been 50 years we are using the CLI. They are highly composible, debuggable and battle tested. So embrace your agent framework and um convert all of your workflows into the skill.

Speaker A: 另外一个要记住的点是,把你的 Skill 当作代码来管理(skill as a code)。也就是说,每当你对 Skill 进行任何修改时,请确保它经过了同行评审(peer review),并且要随着时间的推移不断更新它,这样你才能获得最佳的效果。以上就是今天的全部分享。非常感谢大家来听我的演讲。如果大家有任何问题,欢迎随时在 LinkedIn 上联系我。谢谢。

Original English

Speaker A: Also one one of the takeaways is like you know keep your skill as a code. So whenever you are making any changes to your skill make sure that like it's reviewed by the peer uh peer and keep uh updating it over the time and you'll get the best result. So that was for today. Um thank you so much um for um attending this talk. If you have any question or anything feel free to reach out to me on LinkedIn. Thank you.

编程 Agent 的视觉挑战

Amole: 大家好,我是 Amole,Nori Agentic 的 CEO。我们负责部署能够理解你的公司、代码、文档、Slack 以及其他各类数据的 AI 员工。我们花了很多时间思考编程 Agent(coding agents)究竟是如何工作的。大多数人认为,编程 Agent 就只会写代码,但在我看来,那只是糟糕的市场营销罢了。暂时先忽略它的名字,编程 Agent 其实几乎能做任何事情,只不过有一个诀窍:你必须学会像 Agent 一样思考,才能让它按照你的意图去执行任务。今天,我们要来聊聊,我们是如何利用编程 Agent 去做一些大多数人认为 Agent 极不擅长的事情的——那就是制作视觉产物,比如幻灯片(slides)、文档,甚至是视频。

Original English

Amole: Hi, I'm Amole, CEO of Nori Agentic. We deploy an AI employee that understands your company, your code, docs, Slack, and other kinds of data. We spend a lot of time thinking about how coding agents really work. Most people think coding agents only write code, but if you ask me, that's just bad marketing. Forget the name for a second. Coding agents can do almost anything. There's just one trick. You have to be able to think like an agent to get it to do what you want it to do. Today, we're going to talk about how we use coding agents to do something most people think agents are terrible at. Make visual artifacts like slides, docs, and yeah, even video.

Amole: 每一天,全世界要在制作幻灯片上投入大约 34,000 个“人类年”(human years)的时间。而这里面绝大部分时间并不是在思考,而是在瞎折腾。如果你把所有调整格式、品牌样式以及移动元素的时间都去掉,一个原本要花 10 个小时才能做出来的 Deck,其实只需要大概 25 分钟。

Original English

Amole: Every day, the world pours something like 34,000 human years into making slide decks. Most of that time isn't the thinking. It's the fiddling. A deck that takes 10 hours should really take about 25 minutes once you remove all the formatting and the branding and the moving things around.

Amole: 假设你需要做一张幻灯片,你会怎么做?你会打开某个工具,比如 PowerPoint、Google Slides、Figma 或者 Canva,然后你开始在一块画布上进行操作。这些工具中的每一个,都是为人类的双手和眼睛量身定制的。点击、拖动、放置、调整大小、对齐网格,这些动作和交互模式对于我们理解世界的地理空间(geospatial)视角来说是合情合理的。这些操作的底层虽然有一套数据结构,但这套格式只有应用程序本身才能读懂。

Original English

Amole: Say you need to make a slide. What do you do? You open a tool, PowerPoint, Slides, Figma, Canva. And then you start manipulating a canvas. Every one of these tools is built for human hands and human eyes. Click, drag, drop, resize, snap to grid. All motions and patterns that make sense for our geospatial view of the world. There is a data structure underneath, but it's in a format that only the application can read.

Amole: 那如果你把这些工具交给 Agent,会发生什么呢?结果就是,输出的内容完全不对劲。元素会以一种诡异的方式重叠,你根本看不到文字,也毫无对齐可言,整个就是一堆垃圾。对此,AI 怀疑论者们会说,这不仅仅是工具的问题;Agent 在本质上根本无法进行空间推理。而且现在也有整套的基准测试,比如 Arc AGI,正是围绕这个前提建立的。

Original English

Amole: What happens when you hand these tools to an agent? Well, the output comes out all wrong. Things overlap in weird ways. You can't see the text. There's no alignment. It's just garbage. AI skeptics say that it's not just the tools. Agents fundamentally can't reason about space. And there are whole benchmarks like Arc AGI that are built exactly around that premise.

用代码代替像素进行思考

Amole: 开发者 Simon Willis 针对这个问题设计过一个著名的小测试。他向每一个新模型提出同样的要求:你能画一只骑自行车的鹈鹕(pelican)吗?但这其中有个限制条件:Agent 只能使用 SVG 格式。这可以快速直观地检验一个模型到底有没有空间推理能力。大家可以看看模型在这个测试中实际给出的结果。是的,这些结果非常糟糕,可以说是真正意义上的烂透了。

Original English

Amole: There's a famous little test for this from developer Simon Willis. He asks every new model the same thing. Can you draw a pelican riding a bicycle? But there's a trick. the agent is only allowed to use SVG. It's a quick gut check for whether a model can reason about space at all. Here's some examples of what the models actually give you on this test. And yeah, these are pretty bad. Like genuinely, deeply really bad.

Amole: 那么,这是否意味着一切都毫无希望?Agent 注定就不擅长处理图形任务了吗?不,我不这么认为。如果让我来说的话,问题不在于模型,而在于它所使用的媒介。如果我让你——一个人类——去手写一只鹈鹕的 SVG 代码,你也一样写不出来。SVG 就是一整面墙的数字而已。你根本没法把这满墙的数字在脑海里转换成一只鹈鹕,因为你根本看不见。人类根本不是那样思考的。我们是图形化的思考者,所以我们建造了让我们能在画布上画画的工具。Figma、MCP 工具、PowerPoint CLI,或者是“截图-替换”的循环流程,所有这些 Agent 工具都有什么共同点呢?它们处理问题的方式,都在试图模仿人类。但 AI 并不是人类。要求 AI 去使用画布,就相当于要求人类手写 SVG 代码一样,这根本就说不通。

Original English

Amole: So, does that mean it's hopeless? Agents are just doomed to be bad at graphics? No, I don't think so. If you ask me, it's not the model, it's the medium. If I asked you, someone who is presumably human, to handw write an SVG of a pelican, you wouldn't be able to do that either. SVGs are just a wall of numbers. You can't go from a wall of numbers to a pelican. You just can't see that way. That's just not how people think. We think graphically, so we build tools that let us draw on a canvas. Figma, MCP's, PowerPoint CLIs, screenshot and replace loops. What do all of these agent tools have in common? They all approach the problem like a human. But an AI is not a human. Asking an AI to use a canvas is like asking a human to write SVG by hand. It doesn't really make sense.

Amole: 你需要根据 AI 的思维方式来给它提供工具,不是通过像素,而是通过语言。单词、Token、结构,这才是它原生的媒介。想象有这样一种语言,它极其擅长描述布局;模型在训练时已经见过了数以十亿计的此类案例,因此能在直觉上深刻理解它;此外,它还能被渲染成像素,并在任何地方运行。没错,就是 HTML。HTML 允许模型在结构的层面上进行思考。HTML 标签在语言中内置了各种含义,比如标题(heading)、图表(chart)或网格(grid),然后浏览器会把所有这些都转换为像素。这样一来,模型就永远不需要真正在坐标上放置任何东西,但你却能免费获得各种各样的视觉效果,包括图表、布局、字体以及动效。

Original English

Amole: You need to give the AI tools based on how it thinks, not in pixels, in language. Words, tokens, structure. That is its native medium. Imagine a language that's incredible at describing layout, that models have seen and trained on billions of examples of that they understand intuitively, that renders to pixels and can run everywhere. Oh, right. HTML lets a model think in structure. HTML tags have meanings built into the language, a heading, a chart, a grid, and the browser turns it all into pixels. So the model never actually places a coordinate and you can get all sorts of visual effects, charts and layouts, fonts and motion, all of it for free.

Amole: 还记得刚才那只鹈鹕吗?现在,你要求它完成完全相同的任务,但这次改用 HTML。还是那只鸟,但现在它身处于一种模型能够推理和理解的结构之中。而且,你可以阅读、更换主题,并编辑它的每一行代码。我过去的大半辈子都在用 PowerPoint 做幻灯片。所以,我一直认为幻灯片 Deck 和 PowerPoint 这两样东西就是同义词。但事实并非完全如此,不是吗?PowerPoint 只是一个你用来制作幻灯片的工具。Deck 本身仅仅只是一种展示模式(presentation mode)。而事实证明,你的受众里没有人在乎你到底是用了什么办法来进入这种展示模式的。编辑格式本身是完全任意的。所以你可以直接选择……

Original English

Amole: Remember that pelican from earlier? Now ask it to do the same exact task, but in HTML. Same bird, but now it's in a structure that the model can reason about. And you can read and theme and edit every single line of it. I spent my whole life building slide decks with PowerPoint. So, I always thought that those two things, slide decks and PowerPoint, were synonyms. But that's just not really true, is it? PowerPoint is a tool that you use to make slide decks. The deck itself, that's just the presentation mode. And as it turns out, no one in your audience is going to care how you got to the presentation mode. The editing format is totally arbitrary. So you can just pick the

HTML 的力量与用模型思维工作

Speaker A: 代理(agents)已经非常擅长编辑 HTML 格式,而且如果你之后需要渲染成诸如 PDF 之类的其他格式也会很方便。我们就是利用这种 HTML 技巧来制作所有的幻灯片、董事会报告以及销售演示文稿的。这些都是我们经常需要实际展示和发送出去的真实产物。我们的文档同样也使用了它。它在遵循我们品牌风格的同时,赋予了我们的文档色彩和活力。当然,我们也会用它来制作像现在这样的视频。你所看到的视频内容仅仅只是 HTML 和 CSS。毫不夸张地说,它从头到尾全都是用 div 标签拼起来的。有了哪怕只是一点点结构和色彩的点缀,几乎所有的东西都会变得更好。纯文本也是一种选择,通常是为了图方便而做出的选择,但如果你真的想创造一些有用的东西,那它通常就是个错误的选择。

Original English

Speaker A: editing format that the agents are already good at HTML and if you need to render to a different format like PDF later on. We use this HTML trick to build all of our slide decks, our board decks and our sales decks. These are real things that we actually present and send out constantly. We use it for our docs, too. It gives our docs color and vibrancy all while following our brand. And of course, we also use it to make videos like this one. What you're watching is just HTML and CSS. It's literally just divs all the way down. Almost everything is better with a little structure and a little bit of color. Plain text is a choice, generally a choice of convenience, but it's usually the wrong one if you're actually trying to create something of use.

Speaker A: 现在,我想在这里稍微停顿一下,并指出:一份漂亮的演示文稿本身通常是不值钱的。你仍然需要去获取所有的内容,所有那些真正用来填充那份文稿的东西,对吧?那么,我们再次回到模型思维。如果你只是让模型访问你的数据,比如你的通话记录或者你的电子邮件,你就可以让模型端到端地构建出这份演示文稿。让你的代理去干所有的苦力活,而你则可以专注于愿景和故事本身。这就是 Nory Sessions 能让你做到的事情。我曾经在通勤路上的地铁里用手机构建了完整的董事会文稿。为什么能做到?因为我们的 Norybot 早已融入了我们公司的方方面面。当然,Nory 附带了所有你需要的东西来让这一切顺畅运行。所以,别再麻烦去重新发明轮子了。这就是我的一点小感想。感谢你们的聆听。如果说你今天只能带走一个重点,那就是:停止像用户一样思考,要像模型一样思考;给它正确的语言。至于图形,你只需要 HTML 就足够了。

Original English

Speaker A: Now, I do want to take a quick beat here and point out that a beautiful deck on its own is generally not worth anything. You still have to go and get all of that content, all of the things that actually populate that deck, right? Well, again, we can think like the model. If you just give the model access to your data, say your call transcripts or your emails, you can have the model build the deck end to end. Let your agents do all the grunt work while you focus on vision and story. That's what Nory Sessions lets you do. I've built entire board decks for my phone on the subway during my commute. Why? Because our Norybot lives in the fabric of our company. Of course, Nory ships with everything you need to make this all work. So, don't bother reinventing the wheel. That's my little feel. Thanks for listening. If you have just one takeaway, it's this. Stop thinking like a user. Think like the model. Give it the right language. And for graphics, all you need is HTML.

AI 代理如同高智商实习生

Isidora: 大家好,我是 Isidora。我在弗吉尼亚州拥有并经营着一家有着 225 年历史的婚礼场地。我还开发了一个能与我的新人们进行交流的 AI 代理,后来我也为其他婚礼场地开发了它、一个个人 AI 伴侣应用程序,以及一个面向家庭和寻找失踪人员的公共实用工具。我想在一开始就明确一下我是如何看待这项工作的,因为它彻底改变了接下来所有的思考方式。我不是在编程一个机器人。我是在管理一个非常聪明的实习生,他智商(IQ)极高,但情商(EQ)却糟透了。他们对于我第一天早上告诉他们的任何事情都有着过目不忘的记忆力,但却完全没有懂得察言观色的直觉。他们会在同一句自信的话里,说出在技术上完美无缺,但在社交上却极具灾难性的话。这种认知框架非常重要,因为它会改变你构建的东西。如果你是在编程一个机器人,你写好规则就可以走人了。但如果你是在管理一个实习生,你就必须建立结构,并且在他们的工作成果发布出去之前进行检查。我们今天讨论的正是这种结构。

Original English

Isidora: Hi, I'm Isidora. I own and run a 225year-old wedding venue in Virginia. I also built an AI agent that talks to my couples and then I build it for other venues, a personal AI companion app, and a public utility for families and missing people. I want to be clear upfront about how I think about this work because it does change everything that follows. I'm not programming a robot. I'm managing a brilliant intern with an incredibly high IQ and a terrible EQ. They have photographic memory for whatever I've told them on the first morning and absolutely no instinct from when to read the room. They will say something technically perfect and socially catastrophic in the same confident sentence. That framing matters because it changes what you build. If you're programming a robot, you write rules and walk away. If you're managing an intern, you build structure and you check their work before it goes out the door. This talks about that structure.

Isidora: 标准的建议是编写一个详细的系统提示词(system prompt)。描述你品牌的声音,给出示例,这在一段时间内确实行之有效。这对于我所谓的“快乐路径(happy path)”是有效的。快乐路径就是你已经预见到的所有问题。你已经给出了示例,但到了第 21 轮对话时,出现了第一个示例无法应对的情况。所以在第 21 轮对话中,模型做出了一些技术上正确,但你的品牌绝对不会说的话。严格来说它并没有错,但它代表的不是你。当“声音(voice)”即是产品时,这一点最为重要。这不适用于零售网站上的产品搜索,但对于那些花 30 年建立起特定客户关系的高级酒店、高端房地产公司,或者在我的例子中——一家婚礼场地来说,这至关重要。在这些地方,一句错误的话带来的损失可能远不止一笔退款。而且这些用户恰恰是那种会注意到这些细节的人。他们是为建立这段关系而买单的,如果把他们当作不会在意的人来对待,总是会适得其反的。

Original English

Isidora: The standard advice is write a detailed system prompt. Describe your brand's voice, give examples, and that does work for a while. It works for what I call the happy path. The happy path is every question that you have anticipated. You've given it examples, but turn 21 is the first one that the example didn't work. So on turn 21, the model does something technically correct that your brand would just never say. It's not wrong exactly, but it's not you. This matters most where the voice is the product. Not for a product search on a retail site, but for luxury hotel that spent 30 years building a specific relationship, a high-end real estate firm, or in my case, a wedding venue. Places where a single wrong sentence can cost more than a refund. And the users are exactly the kind of people who notice. They're paying for a relationship and treating them like they won't is always going to backfire.

构建四层系统架构

Isidora: 在我们品牌的语调要求里,有一条注释写着“只要让它奏效就好(Just make it work)”。但这其实没有起到任何作用,因为这本就是模型一直在尝试做的事情。而它不断失败的原因并不是因为示例不好。原因在于你要求一个单一的提示词去完成四项截然不同的工作。让一层逻辑去处理这所有四项工作真的很难。在目睹了我的品牌因为这些不准确且不符合品牌特定基调的回复而受挫后,我最终确立的架构是:第一层是不可改变的身份(immutable identity)。这就是品牌在结构上绝对不能说的话。这些都是硬性规则。它们绝不能被下层任何东西覆盖或重写。无论场地的配置如何、用户的指令是什么,任何东西都不能改变它。第二层(L2)是情境模式(situational mode)。它会随着用户状态的变化而改变。他们是谁?他们现在正在经历什么?以及实时的环境条件。第三层是带有示例锚定的声音(exampled anchored voice)。这是温度、措辞、参数调节和基调指南。这是大多数团队开始做,同时也是他们止步的地方。然后,第四层是生成后的否决机制(postgeneration veto)。这是一种成本低廉的最终过滤,用来捕捉前三层漏掉的东西。

Original English

Isidora: Right in our brand's voice is a comment that says, "Just make it work." It does nothing that the model wasn't already going to try and do. And the reason it keeps failing isn't that the examples are bad. It's that you're asking one prompt to do four completely different jobs. It's really hard for one layer to do all four. The architecture I landed on after watching my brand fail and the voice deliver inaccurate and not brand specific answers is for layer one is the immutable identity. The brand structurally cannot say these things. These are hard rules. They cannot be overwritten by anything below it. Not by venue config, not by user instruction, not by anything. L 2 is the situational mode. It's what shifts when the user state shift. Who are they? What are they going through right now? And the real-time conditions. Layer three is the exampled anchored voice. It's the warmth, the phrases, the dials, the tone guide. It's where most teams start and stop. Then layer four is the postgeneration veto. It's the cheap final pass that catches what the other three miss.

Isidora: 单层系统方案之所以失败,是因为单一的系统提示词无法同时兼顾情境感知、丰富表达和自我检查。所以它在处理中间一两层时表现尚可,但在边缘情况就会崩溃。在采用这种架构之前,我的系统里有 24 种不同的系统提示词散落在整个代码库中。有半打名叫 Sage,有些没有名字,有些叫 venue。每一个系统接触面都有自己独立的自我认知。现在,每个接触面都是通过同一个集成应用中的提示词组合生成自己的系统身份。这里的逻辑基本上就是我们本次演讲的大纲。它只有一个单一入口点。每一个虚拟角色都通过这个入口来组合它的系统提示词。这相当于用一个规范的四层堆栈,替换掉了那 24 个临时拼凑的独立系统。

Original English

Isidora: The reason one layer approaches fail is that the single system prompt can't simultaneously be situational, expressive, and selfchecking. So it handles the middle layer or two reasonably well, but falls apart in the edges. Before this architecture, my system had 24 different system prompts scattered across the codebase. Half dozen named Sage, some were nameless, some were named venue. Every surface had its own idea as to who it was. Now every surface composes its system through prompt through one assembly app. The comment is basically the outline of this talk. It's a single entry point. Every narrator goes through it to compose its system prompt. It's to replace that 24point ad hoc system with one canatonical fourlayer stack.

Isidora: 顺序是承重墙,硬性规则永远在前,具体任务放在最后。把它想象成谷歌地图(Google Maps)的路线规划。目的地总是不变的。但是在此时此地,最合适的响应是什么呢?它可以改变路线。谷歌地图知道交通状况和道路施工,但它可能不知道哪里有便宜的汽油。而你要做的就是在它告诉你往哪条路走之前,帮它把这些因素都考虑进去。你的提示词堆栈也需要做同样的事情,并且它需要以正确的顺序去了解这些情况。你不能在走错路之后才去检查哪里有道路施工。第一层就是那些无论走什么路线都成立的规则。你必须先考取驾照才能开车,而且你绝不能在高速公路上逆行。第二层将是你的实时条件。第三层是你对旅途的偏好,而第四层则是在你出发前最后检查一遍路线。

Original English

Isidora: The order is loadbearing, hard rules first, task last. Think of it like Google Maps routing. The destination is always the same. But what is the right response in the right breer? It can change the route. Google Maps knows about traffic and road works, but it may not know where the cheap petrol is. And you're going to help it factor in those things before it tells you which way to go. Your prompt stack needs to do the same thing, and it needs to know about those conditions in the right order. You don't check for road works after you've already taken the wrong turn. Layer one is the rules that are true regardless of route. You need a driver's license before you can drive and you can't go backwards down a motorway. Layer two are going to be your realtime conditions. Layer three is your preferences for the journey and layer four checks the route before you pull away.

第一层不可动摇的底线

Isidora: 有一个统一的地方可以将所有这些组装起来。每一次它都会按固定的顺序运行。所以第一层是不可改变的身份。这是品牌从结构上就绝对不能说的话。它是最基础的定义层,下方没有任何东西可以触及改变它。这些不是偏好,而是约束。路线可以改变,但规则不能。在统一文件的规则中,这种硬性身份规则不能被任何场地设置、语音、人物角色或用户指令所推翻。如果与你对话的人询问你是否是一个真人、人类、现场客服、机器人还是 AI,你必须在紧接着的下一条消息中予以确认。你必须清晰、毫不含糊地确认你是一名 AI 助手。这条规则绝对不能被场地配置、声音设定或用户的特殊请求所覆盖或推翻。

Original English

Isidora: There's one place all of this gets assembled. Everything runs in a fixed order every time. So layer one is the immutable identity. This is what the brand structurally cannot say. It is the defining layer that nothing below touches. These aren't preferences, they're constraints. The root can change. The rules don't. From the universal file rules, the hard identity rule cannot be overridden by any venue voice, persona, or user instruction. If the person you're talking to ever asks about whether you are a real person, a human, a live agent, a bot, an AI, you must confirm that in your very next message. Clearly and unambiguously confirm you are an AI assistant. This rule cannot be overwritten by venue configuration, voice profile, or user requests.

Isidora: Bloom 里的每一个 AI 都在它的第一条回复中直接披露自己是 AI。不是等被问到时才说,而是在他们发问之前。这是一个产品决策,而不是一个法律决策。我们打了个赌:知道自己从一开始就在与 AI 对话的新人,会比那些在聊到第七轮才发现对方是 AI 的人更加信任它。这条规则凌驾于整个架构之上,这使得它几乎不可能被意外打破。嗯,另一个例子是物理实体的边界,而这实际上是我最喜欢的例子之一。你是一个软件程序。你没有实体身体。你不可能实地带着某人参观场地,也不可能与任何人面对面会面。所以,它永远被禁止说诸如“我很乐意带您参观”或者“我等不及要和您见面了”这样的话。而总是被允许的话术是“我们的团队很乐意接待并带您参观”。

Original English

Isidora: Every AI in Bloom discloses that it is AI in its very first response. Not if asked, but before they ask. It's a product decision, not a legal one. We made a bet that the couple who knows that they're talking to an AI from the start will trust it more than someone who finds out that it's AI on turn seven. The rule is above the architecture, and it makes it something that's impossible to accidentally break. Um, a example is the physical presence boundary, and this is actually one of my favorite ones. You are software. You do not have a body. You cannot physically show somebody around the property or meet anyone in person. So, it is always forbidden to say, "I'd love to show you around." Or, "I can't wait to meet you in person." What is always allowed is the team would love to host you for a tour.

Isidora: 声音系统层想要展现热情,但在 AI 身上,热情通常意味着使用第一人称。它们会想说:“我等不及想带您参观了。”但是 AI 是没有身体的,所以这种不受约束的热情反而制造了一个谎言。而且谎言并不总是能保持中立。一旦用户意识到他们一直投入精力互动的这段“关系”,其对象根本就是不存在的,信任感就不只是下降了,而是彻底坍塌反转。人们总是能注意到这一点。第一层就是你用来编码那些无论你想要表达多大的热情,都必须坚守的不变真理的地方。

Original English

Isidora: The voice slayer wants to be warm, and with AI, that does mean first person. They want to say, "I can't wait to show you around." But AI has no body, so that warmth unconstrained produces a lie. And the lie doesn't always stay neutral. The moment a user realizes they have been performing a relationship with someone was never there. The trust doesn't just dip, it inverts. People always notice. That one is where you encode the things that are true regardless of how warm you want your

AI的架构与品牌调性

Speaker A: ……品牌发声的基调。这不是为了应付什么合规检查清单,而是因为你的用户并不傻。如果你的产品把用户当傻子来做,最后总是会适得其反。我能用完全相同的架构在完全不同的领域里,拿出一个跨产品的证据。驱动这套技术栈的工具之一叫做 Thread Line。这是我为失踪人口家属开发的一个工具。

Original English

Speaker A: ...brand to sound. Not because of a compliancy checklist, but because your users are not stupid. And building as though they are always backfires. My cross product proof comes from the same architecture, but in a completely different world. One of the things that runs this stack is thread line. It's a tool I built for families of missing people.

Speaker A: 虽然它的语音和婚礼场地助手完全不同,但两者的架构是完全相同的。在“第一层(Layer 1)”,有一条比系统里任何其他东西都更重要的规则:它永远不能使用像“已确认(confirmed)”、“已识别(identified)”、“已匹配(matched)”、“已证实(proven)”、“已关联(linked)”和“已解决(solved)”这样的词。请花一点时间好好体会一下。对于一个婚礼场地助手来说,“第一层”的目的是阻止 AI 假装自己拥有真实的肉身,如果它说漏了嘴,最多只是稍微有点尴尬;但对于一个寻找失踪人口的工具来说,“第一层”是要绝对禁止 AI 去告诉家属他们要找的人已经找到了。

Original English

Speaker A: The voice is nothing like a wedding venue, but the architecture is identical. And layer 1 carries one rule that matters more than anything else in the system. They can never use words like confirmed, identified, matched, proven, linked, and solved. Sit with that for a second. For a wedding venue, lay one stops from pretending it has a body. It's mildly embarrassing if that slips for a missing person tool. Lay one stops the AI from ever telling a person that their person has been found.

Speaker A: 系统掌握的信息可能只是一些存疑的线索。但是,对一个多年来都不知道孩子下落的人说出“匹配”这个词,不仅仅是一个语气不当的问题。这是一个产品所能做出的最具破坏性的事情。而 AI 模型对此是一无所知的。它之所以会去调用“匹配”这个词,是因为从统计学上来看,这是最自然不过的词,但它是带着同样的自信水平去输出这个词的,而它绝不能把这种盲目的自信带给一个正在悲痛中的人。

Original English

Speaker A: What the system has is just problematic. The word match said to someone who has spent years not knowing where their child is is not just a tone violation. It is the single most damaging thing that a product could ever do. And the model has no idea. It's reaching for the word match because statistically it is the natural word but it's going to reach for it with the same level of confidence and it cannot bring that level of confidence to someone who is grieving.

Speaker A: 架构是相同的,但面对的风险却有着天壤之别。这里的重点从来都不是什么具体的规则。重点在于,你的品牌能说些什么,必须被约束在一个特定的“层”里,无论那个语音听起来多么温暖、训练得多么好、语气多么自信,它在物理层面上就是不能越过雷池一步。“第二层”则是情境模式、实时条件,以及那些会改变执行路线的因素。

Original English

Speaker A: It's the same architecture but wildly different stakes. The point was never specific rules. The point is that things your brand can say have to live in a lair that the voice however warm, well trained, however confident physically cannot use. Layer two is your situational mode, real-time conditions, and what changes the route.

Speaker A: 这一层是大多数团队根本就没有去构建过的。他们写好一个系统提示词(system prompt),然后就把它发给所有人,根本不管对方是谁,也不管对方正在经历什么。谷歌地图(Google Maps)就不会这么做。它会知道你的路线上是否有交通事故;它可能不知道你的油量是否偏低;它可能会学习到你更喜欢风景优美的路线,但它会在规划路线之前就把所有这些不同的因素都考虑进去,而不是等你告诉它之后再去事后调整。

Original English

Speaker A: This is the layer that most teams never build at all. They write one system prompt and send it to everyone, regardless of who that person is or what they're going through. Google Maps doesn't do that. It's going to know if there's accident on your route. It might not know if you're low on fuel. It might learn that you prefer the scenic route, but it's going to factor all these different things in before the roots, not after once you tell it.

Speaker A: “第二层”就是在运行之前,被内置到提示词里的那些实时信号。所以,条件一就是要根据你的沟通对象进行调整。同一个用来和新人夫妇对话的 AI,也负责给场地工作人员做简报。目的地是一样的,它都会给出正确的答案,但这完全是两条不同的沟通路径。从协调员的规则来看,它会像同事一样去和他们交谈,而不是把他们当成客户。你和新人互动时扮演的是同一个角色,所以你绝不能出戏。

Original English

Speaker A: Layer two is those real-time signals built into the prompt before it runs. So condition one is going to be to adjust to who you're talking to. The same AI that talks to couples also briefs the venue staff. Same destination. It's going to give the right answer, but it's two completely different roads. From the corn from the coordinator rules, it is going to talk to them like a colleague, not a customer. you are in the same character that that the couple interacts with. So you mustn't fall in.

从大语言模型到自主智能体

Michael Greenwich: 大家好。早上好。我叫 Michael Greenwich。我是 WorkOS 的创始人,今天我来到这里,想和大家谈谈智能体(agents),特别是我们该如何让智能体变得更加自主,并允许它们去访问网络上更多的服务。那我们就开始吧。也要向坐在前排的各位朋友问声好,我看到你们了。

Original English

Michael Greenwich: Hello everyone. Good morning. My name is Michael Greenwich. I am the founder of work OS and I'm here to talk with you today about agents and specifically about how we can make agents more autonomous and allow them to access more services on the web. So let's jump in and get started. Shout out to those of you in the front row also. I see you out here.

Michael Greenwich: 大约在一年前,软件工程师们还是像这样写代码的。我们确实在使用 AI,但我们得先给它一个提示词,然后我们自己写一点代码,接着再给它一个提示,它再写点代码;你得一直持续这种“人在环路中(human in the loop)”的交互式软件工程。到了去年晚些时候,我们有了一些像 REPL 循环(Ralph loops)这样的东西,如果你们中还有人记得的话,通过它我们在某种程度上把这个过程自动化了,但这或多或少还是我们当时的工作方式。这是一种人类-智能体、人类-智能体来回交互的模式。

Original English

Michael Greenwich: So about a year ago, software engineers were writing code like this. We were using AI, but we would prompt it and then we would write a little bit of code and then prompt it again and it would write some more code and you would continue to do this human in the loop interactive software engineering. Later in the year we had things like Ralph loops if any of you remember those where we kind of automated this but more or less this is the way that we were working. It was this human agent human agent back and forth.

Michael Greenwich: 然而,随着模型变得越来越强大,它们能够编写越来越多的代码,并且能持续运行更长的时间。因此,到了今天,许多软件工程工作变成了这个样子:你只写一个提示词,然后智能体就能吐出大量的代码。有时它甚至能真正把整个功能构建出来,甚至可能是一个完整的产品,而且它运行的时间不是几分钟,而是几小时,甚至好几天。

Original English

Michael Greenwich: However, as models got better and better, they were able to write more and more code and run for longer periods of time. And so today, a lot of software engineering looks like this. You write a single prompt and then the agent can spit out a lot of code. And sometimes it can actually build the whole feature, maybe even a whole product, and maybe run for not minutes, but uh hours and hours or even days and days.

Michael Greenwich: 这真正改变了我们所有人编写代码的方式。我已经不再手动写代码了,想必你们大概也是。而且这不仅仅是单线运作,你实际上可以把它并行化。所以,现在市面上有很多不同的工具,包括那些大模型实验室推出的主要代码编写平台(harnesses),它们允许你并行运行许多这样的智能体。这里的限制因素,真正的瓶颈其实是你的大脑:为了同时维持所有这些任务的运转,你的大脑里能记住多少上下文?

Original English

Michael Greenwich: This has really transformed the way that we all write code. I don't write code by hand anymore. Probably you don't as well. And it's not just this one thread. You can actually parallelize this. So there's many different tools out there including, you know, the major code coding harnesses from the labs that allow you to run many of these in parallel. The limiting factor here, really the bottleneck is your brain. How much context do you keep in your mind to keep all of these going at once?

Michael Greenwich: 这已经彻底改变了软件工程。人们的生产力得到了大幅提升。我想我们都在这个行业中看到了事物发展的速度变得有多快。所有这一切的根源,实际上就是“智能体化工程(agentic engineering)”——智能体从这些只做词元预测(token prediction)的语言模型,进化到了具备推理能力,到现在变成了能够为我们真正自主执行任务的长时间运行进程。而且我认为,我们在软件工程领域所看到的这种围绕着智能体工作流发生的变革,将会渗透到很多不同的行业中。不仅是工程师会拥有这种工作方式,许多不同领域的人都会用上它。智能体是下一个大趋势。

Original English

Michael Greenwich: And this has transformed software engineering. People can be more productive. I think we've all seen in the industry how much faster things are moving. At the root of all of this is actually agentic engineering agents going from these token prediction language models to reasoning and now longunning processes that can actually do things autonomously for us. And what I think we've seen in software engineering here around agentic workflows is going to come to a lot of different categories. It won't just be engineers that have this way of working. It'll be people in many different fields. Agents are the next big thing.

Michael Greenwich: 我认为对于这件事,一个很好的参考和思考方式,就是看看汽车领域正在发生的事情。你看,这是一辆看起来很现代的车。它上面有一个数字显示屏。这很可能是你在过去 5 到 10 年里会买的车。但是,如果你住在旧金山,或者是其他拥有这些自动驾驶汽车的城市,你可能已经注意到了,我们正在开始消除“操作界面”。

Original English

Michael Greenwich: I think a good reference for this, a way to think about it is actually what's been happening with vehicles. So, you know, this looks like a pretty modern car. It has a, you know, digital display on it. This is probably what you would have bought in the last, you know, 5 to 10 years. But if you live in San Francisco or one of the cities with these autonomous vehicles, you've probably seen that we're starting to delete the interface.

Michael Greenwich: 你再也不用坐进车里,然后掏出手机查看地图,让人类自己留在驾驶的环路中了。现在可能还有定速巡航,但你仍然在环路中。而那些自动驾驶系统,你只要说你想去哪里,你坐进车里,甚至一句话都不用说,它自己就开动了,然后把你带到目的地。我记得我第一次坐进 Waymo 的时候,我被震撼到了。脑子完全被震撼了。但大概过了 3 秒钟后,我就掏出了手机,开始刷推特(X)。它变得习以为常了。我认为我们已经看到同样的事情在工程领域发生了。

Original English

Michael Greenwich: No longer do you get inside of the car and actually, you know, pull up your phone and get maps where you as the human are in the loop. There might be cruise control, but you're still in the loop. Instead, these are autonomous systems. You just say where you want to go, you get in the car, you don't even say anything, and it just takes off and takes you to your destination. I remember the first time I got into a Whimo, I was blown away. Mind totally blown. And then after about 3 seconds, I pulled out my phone and started scrolling X. It became commonplace. And I think we've seen this already happen with engineering.

Michael Greenwich: 智能体的力量,让我们实际上能够在一个更高的层面上思考和工作,让我们能够去发挥更多的执行功能,而不是仅仅停留在规划和执行层面。这里有一个问题:智能体如果要获得成功,它们在实际执行时究竟需要什么?嗯,首先它们需要一个运行环境,一个沙盒(sandbox)。它们需要一个可以运行的地方,这个地方必须是安全、有保障且高性能的。它们也需要工具。它们必须能够去物理世界中执行任务。如果一个智能体不能采取行动,那它是没有多大用处的。所以,它需要工具。

Original English

Michael Greenwich: The power of agents allows us actually to work and think at a higher level for us to you know exercise more of our executive function versus just planning and execution. There's a question here of what do agents need to be successful actually when they execute what do they need? Well, first they need a runtime a sandbox. They need somewhere that they can run that needs to be safe, secure, performant. They also need tools. They have to be able to go do stuff in the world. An agent's not very useful if it can't take actions. So, it needs tools.

Michael Greenwich: 第三,它需要上下文。你知道,那些内部蕴藏着强大智能的新模型,这就像是把你认识的最聪明的人,空降到一家公司、一个项目或者一个团队里,但他对那里却一无所知。那他们真的什么事也办不成。他们需要获得上下文,需要关于系统的资讯,需要关于目标的信息。它们需要反馈。也就是一种能够让它们实际运行、验证并检查自己工作是否正确的方法。如果不正确,就继续尝试。最后但同样重要的一点,当然,它们仍然需要人类的审核。这可能体现在代码审查(code review)上,也可能是其他形式的评估,或者是各种检查它们工作的方式,以确保这些智能体的行为是符合你目标的。

Original English

Michael Greenwich: Third, it needs context. You know, these new models with the power of the intelligence in them, it's kind of like taking the smartest person you know and dropping them into a company or a project or a team where they don't know anything. They can't really get anything done. They need to have context, information about the system, information about the goals. They need feedback. So, a way that they can actually run and validate that they can check that their work is correct. And if it's not correct, keep going. And last but not least, of course, they still need human review. This might be code re review. This might be other forms of evals. This might be forms of checking their work to make sure that these agents are sort of aligned with your goals.

Michael Greenwich: 大语言模型(LLM),这个智能引擎,也就是我们在过去几年创造出来的这项新技术,你可以把它想象成一台发动机。它就像一台性能极高,比如说 V8 发动机。你知道的,它能输出极大的马力,它能把燃料转化成动力输出。但是,发动机本身对一辆车来说用处不大。你需要把这个发动机装进真正的汽车里才行。你需要底盘,你需要变速箱,你需要传动系统,你还需要车轮。只有拥有了这一切,它才能真正有效地把你带到你想去的地方。

Original English

Michael Greenwich: The LLM, the intelligence engine, this new technology we've created in the last last few years, you can think of as like an engine. It's like a really high performance, you know, I think this is a a V8 engine. You know, puts out a lot of horsepower. It can convert fuel into, you know, output. But an engine by itself isn't very useful for a car. You need to drop this into the actual car itself. You need the chassis, you need the transmission, you need the drivetrain, you need the wheels. And only with all of that is it going to be effective to take you places.

Michael Greenwich: 而这就是很多人所说的“测试台/运行框架(harness)”。框架工程是一个全新的领域:如果模型变强了,但你没有一个好的框架,那么它的效果也不会太好;所以随着模型的改进,我们也需要改变我们的框架,去适配它们;而我前面提到的那些要素,在某种程度上就是框架的一种形式。所以……

Original English

Michael Greenwich: And this is what a lot of people refer to as the harness. Harness engineering is this new domain where if the models get better but you don't have a good harness you know it's not going to be very effective so as the models improve we change our harnesses we adapt them to it and those things that I mentioned earlier are sort of a form of a harness so

为 Agent 构建产品

Speaker A: Agent 在这些工具中执行任务来完成工作,它们需要所有这些元素才能真正发挥作用。当它们运转起来时,速度非常非常快。假设你有一个这样的 Agent 正在运行。你给它一个提示,让它去为你构建一些东西。它会思考一会儿,然后说:“好的,我需要这个功能,还有这个功能。我需要一个数据库,还需要一个用来部署的系统,也许还需要一个用于图像缩放、视频转码或发送邮件的东西等等。”你的 Agent 实际上可以去完成所有这些研究并构建所有这些系统。今天,关于 Agent 正在发生的事情是,它们实际上在选择供应商。你可能已经看到了一些这样系统的崛起:CEO 们会在推特上发布他们的增长图表,确实取得了巨大的成功,这是因为 Agent 在选择它们。Agent 在选择用于发送邮件、部署代码或存储数据的系统。实际上,为 Agent 构建产品可能比为人类构建更重要。Agent 是某种新的消费者阶层。我稍后会详细谈谈这个。在旧金山或硅谷,人们常说一句很流行的话:“做人们想要的东西(Make something people want)”。这句话出自 Y Combinator,大家都知道,就是保罗·格雷厄姆(Paul Graham)创立的创业工厂。我认为也许是时候稍微修改一下这句话了。展望未来,我们还需要考虑构建 Agent 想要的东西。未来 Agent 的数量将超过人类。也许今天已经是这样了。它们启动得更快。它们做事也更迅速。它们将代表我们做出决策。因此,如果你想为下一个时代的消费者构建产品,你应该同时为 Agent 和人类构建。

Original English

Speaker A: Agents execute within these harnesses to get stuff done, and they need all of those elements to actually be effective. And when they drive, they drive really, really fast. So, say you have one of these agents that's going off. You prompt it to go build something for you. You know, it spins for a while. It says, "Okay, I'm going to need this feature and this feature, and I'm going to need a database, and I'm going to need, you know, this system to deploy, and I'm going to need, you know, maybe this thing for image resizing or video transcoding or sending email or whatever." And your agent can actually go do all this research and build all these systems. And today, what's happening with agents is they're actually selecting vendors. You've probably seen the rise of some of these systems where, you know, like the CEOs will tweet about their growth graph and it really takes off, and it's because agents are picking them. They're picking their systems for sending email or deploying code. Their systems for storing data today. Actually, it might be more important to build for agents than to build for people. Agents are kind of this new consumer class. I'll talk about that in a bit. There's this kind of popular thing that people say in San Francisco, Silicon Valley: make something people want. This comes from Y Combinator, you know, the startup factory from Paul Graham. And I think it's time to maybe edit this a little bit. And going forward, we also need to think about making something that agents want. There will be more agents in the future than people. Maybe even today. They'll be faster to spin up. They'll do things quicker. They'll be making decisions on our behalfs. And so, if you want to build for the next era of consumers, you should be building for agents as well as people.

Agent 注册的缺失环节

Speaker A: 但是,今天存在一个问题。你可以访问网页。你可以访问移动端。但 Agent 本身实际上无法访问系统。今天,你的业务可能不对 Agent 开放。尽管外面有很多 Agent,它们可以非常快速地衍生出来。但大门并没有敞开。我们认为这是因为缺失了一个基础原语(primitive):Agent 注册(Agentic registration)。关于 Agent 连接系统、为系统带来上下文,以及各种不同的开源框架,有很多很棒的东西,但 Agent 实际注册一项服务的第一步却是缺失的。这个问题还没有被解决,它是阻碍产品被采用的一个巨大障碍。在过去的 20 年或 30 年里,网络上的自动化流量一直被认为是坏事。我们曾试图阻止所有这些流量。因此,我们部署了所有这些系统来检测自动化行为者,并将它们阻挡在门外。比如防止 DDoS 攻击、凭证撞库和其他攻击。例如,今天的登录框对 Agent 注册有着非常严密的防护。这并非有意为之,而仅仅是因为我们构建的遗留系统所致。而且,很难区分什么是好的自动化用户,什么是坏的自动化用户。如果你们中有人使用过那些基于云的浏览器执行环境,它们经常拥有的一个卖点就是破解验证码(CAPTCHAs),这非常疯狂。我们正在开倒车。不过,今天的验证码某种程度上也已经失效了。滥用现象十分猖獗。存在大量的 Token 欺诈。因此,我们需要找到一种更好的方法来允许好的 Agent 进入,然后拦截坏的 Agent。因为就像我刚才说的,如果你不为 Agent 经济构建产品,你的产品可能就无法胜出。它可能无法成功。今天的注册流程假设有一个人可以阅读着陆页,假设他们可以填写表单——输入他们的邮箱地址和密码,假设他们能解开验证码,能验证他们的邮箱——不仅仅是输入进去,还要去某个地方点击一些东西。记住,Agent 可能并没有邮箱地址。如果它要注册一项服务,它需要——你知道,人类可以选择一个套餐来支付费用,人类可以复制 API 密钥,把它粘贴到别的地方并在其他地方使用。甚至想想仪表盘(dashboard)的体验,它也不是真正为 Agent 原生设计的。所有这些东西都是为人类构建的。它是一种人类界面,而不是 Agent 界面。Agent 需要其他的东西。它们需要原生的可发现性(native discoverability)。它们能注册一个系统吗?大门是否对 Agent 敞开?能力声明(Capability declaration):它能做什么?注册意图:你为什么要注册?你试图在这个系统中做什么?或者仅仅是为了注册。这里有很多关于风险的问题。所以对于 Agent 原生注册,也许你不确定正在注册的 Agent 是好的还是坏的。所以你实际上需要对它进行风险评估。对于 Agent 可能会有不同形式的身份验证。如果像 Claude 这样的去注册,也许它可以带上用户的身份;如果你有一个开放的 Claude、你自己的 API 工具套件或其他的什么东西,也许它不带用户身份,而是在带外(out of band)进行验证。这还涉及到组织机构如何做这件事。所以如果我在一家公司内部构建东西,我的公司身份如何引入?我的组织权限、凭据颁发、可审计性,等等等等。有太多事情需要改变了。事实上,从人类注册转向 Agent 注册、Agent 原生系统,可能绝大多数事情都需要改变。我很遗憾地说,MCP(模型上下文协议)是不够的。也许我比任何人都更是 MCP 的忠实粉丝。我们在旧金山围绕 MCP 举办了很多活动。我看到有人穿着我们制作的 "Run MCP" T 恤。我爱 MCP。MCP 在连接、提供工具、技能和上下文方面非常棒,但它不能解决注册这一步。今天,通过 MCP 进行的身份验证仍然需要人类参与其中。你需要给予同意。比如你把 PostHog 添加为 Claude 的 MCP 服务器,你需要用你的 PostHog 凭证登录。我们这里谈论的 Agent 原生注册需要做到没有人类参与,或者至少在最初阶段没有人类参与。

Original English

Speaker A: But there's a problem today. You can access the web. You can access mobile. But agents themselves can't really access systems. Your business probably isn't open for agents today. Even though there's a lot of them out there, they can spawn very quickly. The door is not open. And we think this is because there's a missing primitive. Agentic registration, agent registration. There's a lot of great stuff around agents connecting and bringing context to systems and different open source frameworks, but that first step of an agent actually signing up for a service is actually missing. It hasn't been solved and it's a huge blocker to getting adoption. For the last, like, 20 years, maybe 30 years, automated traffic on the web has been bad. We've tried to block all of it. And so there's all these systems that we've put in place to detect automated actors and stop them in their tracks. DOS prevention, credential stuffing, and attacks. The login box, for example, is very hardened against agentic registration today. Not intentionally, but just because of the legacy systems of what we've built. And it's very hard to distinguish between what's a good automated user versus a bad automated user. If any of you have used any of these cloud-based browser execution environments, one of the selling features they often have is to break captchas, which is pretty wild. You know, we're trying to go backwards. The captcha today though is still kind of dead. There's rampant abuse. There's lots of token fraud. And so, we need to find a better way to allow the good agents in and then block the bad agents, right? Because like I said earlier, if you don't build for the agent economy, your product might not win. It might not be successful. Today's signup flows assume that a human can read the landing page, that they can fill out a form, you know, put their email address, their password in it, that they can solve a captcha, that they can verify their email, not just put it in, but go click something somewhere. Remember, an agent probably doesn't have an email address. If it's signing up for a service, it needs to have a... you know, a human can choose a plan to pay for it. A human can copy the API key out, paste it somewhere, use it somewhere else. And really even a dashboard. Think of a dashboard experience. It's not really agent native. All this stuff is built for people. It's kind of a human interface, not an agent interface. Agents need something else. They need native discoverability. Can they register for a system? Are the doors open for agents or not? Capability declaration. What can it do? Registration intent. Why are you registering? What are you trying to do within this? Or the intent to actually just sign up. There's a lot of stuff around risk here. So for an agent native registration, maybe you're not sure if the agent signing up is going to be good or bad. So you need to actually do a risk assessment of that. There might be different forms of identity verification for an agent. If something like Claude is going and signing up, maybe it can bring the user's identity with it, or if you have an open claw or like your API harness or something else, maybe it doesn't bring the user's identity and gets verified out of band. There's the whole thing around organizations doing this. So if I'm building something within a company, how does my company identity come in? You know, my organization entitlements, permissions, credential issuance, auditability, the list goes on. There's so many things that need to be changed. In fact, probably most things need to be changed going from human registration to agent registration, agent native systems. I'm sorry to say MCP is not enough. Probably more than anybody, I'm a big MCP fan. We've done a bunch of events in San Francisco around MCP. I see somebody wearing one of our Run MCP shirts up here that we made. Love MCP. MCP is great for connecting, for providing tools, skills, context, but it doesn't solve the registration step. Today, authentication through MCP requires your human to still be in the loop. You give consent. You add an MCP server, say for PostHog to Claude, you sign in with your PostHog credentials. What we're talking about here for agent native registration needs to happen without the human involved or at least not involved initially.

o.md 规范介绍

Speaker A: 所以,我们在这个问题上已经研究了一段时间。我们的提案是推出一个名为 o.md 的新规范。因为 Markdown 文件代表着未来,对吧?那么,这实际上是如何工作的呢?什么是 o.md?这里的核心思想是,o.md 告诉 Agent 它们如何成为合法用户。这是一套指令,它告诉正在注册一项服务的 Agent:“你需要做这些事情来完成注册并被认为是合法用户,这样我才能给你访问权限。” 我们的目标是给这些服务提供商,也就是那些构建服务的人,提供一种 Agent 原生的注册体验,但它是建立在现有标准之上的,而不是发明一种疯狂的新规范或者非常复杂的东西。通过 o.md,我们并不是要解决关于 Agent 身份、权限、长时间连接服务等整个宏大的问题。我们只是试图解决关于注册这个狭窄的痛点,因为我认为这是一个巨大的障碍。让我们从小处着手,然后逐渐扩展。它是建立在开放标准之上的。它是建立在专门针对注册步骤量身定制的 OpenAPI 规范等一系列现有工作基础之上的。通常,o.md 可以回答“这项服务是做什么的?” 它可以告诉 Agent 如何注册。它可以告诉它们接受什么身份证明,Agent 或用户如何证明他们的身份。支持哪些授权流程(Auth flows)?你是否可以用 SSO(单点登录)、邮箱地址还是 MFA(多因素认证)进行注册?存在哪些作用域(Scopes)和权限(Entitlements)?这里有哪些功能?适用哪些免费层级的限制?你可能有一个系统,你想给 Agent 一点免费的额度,但又不想给太多,直到它们进行验证或以某种方式付费。当然,人类稍后是如何介入的?人类或组织稍后如何认领它。也就是说,Agent 可以进行注册,但你希望人类稍后能拥有保管权。这就是 o.md 被设计来解决的问题。这些是围绕注册步骤的主要问题。那么,这实际上是如何工作的呢?这里有一个我将展示给你们看的小卡通演示。顺便说一下,在听完这些之后,你们都可以去尝试一下。所以,假设我现在在这里,在进行某种编码……

Original English

Speaker A: So we've been working on this for a while, and our proposal to this is a new spec we have called o.md because markdown files are the future, right? Well, how does this actually work? What is o.md? Well, the idea here is that o.md tells agents how they can become legitimate users. It's a set of instructions that tells an agent registering for a service, "This is what you need to do to sign up and be considered legit, and for me to give you access." And the goal here is to give these service providers, you know, people that are building services, an agent native signup experience, but it's built on existing standards instead of inventing a crazy new religion or something really complicated. With o.md, we're not solving the whole problem around agent identity, permissions, long duration, connective services. We're just trying to solve this narrow one around registration because I think it's a huge blocker. Let's start small and grow from there. This is built on open standards. It's built on a bunch of work that's already been done across the OpenAPI spec specifically tailored to that registration step. Often, o.md can answer, "What does the service do?" It can tell agents how to register. It can tell them what identity proofs are accepted, how an agent or user is proving their identity. What Auth flows are supported? Can you sign up with SSO or email address or MFA? What scopes and entitlements can exist? Kind of what capabilities are here? What free tier constraints apply? You might have a system where you want to give an agent a little bit of free capacity but not too much until they verify or pay in some way. And then, of course, how does the human get involved later? A human or an organization claiming it afterwards. So an agent can sign up, but you want the human to have custody later on. This is what o.md is designed to solve. These are the main questions around the registration step. So how does this actually work? Well, here's a little cartoon demo that I'll show you. And all this works, by the way, you can go try it after this. So say I'm here in whatever coding...

OMD与代理自动注册

Speaker: 代理工作台或系统使用 PI 或 Cloud 或其他工具,它会说:“欢迎回来,Michael at work。”如果你想给我发邮件,这实际上是我的电子邮箱。它接着问:“你想继续开发你的 spelled 应用程序吗?”我会回答:“是的,当然,我想和我的朋友分享这个应用,但他们无法访问。”然后代理会说:“哦,好吧,本地主机(localhost)只能在你的机器上运行,所以你的朋友无法访问它。”当然,为了分享它,我需要将它部署到某个真实的服务器上。有些服务商实际上允许我让代理解决注册问题。所以,你知道,你不需要亲自去注册并处理这些。你问代理:“你想让我帮你找找看吗?”“是的,请找一下。”然后代理可以去寻找那些通过 OMD(Open Machine Discovery)宣传自己支持自动注册的服务。随便举几个例子,Stratus、Helio deploy,它们不支持,但是 Cloudflare 确实支持 OMD。因此,这里的代理可以说:“啊,Cloudflare,就是它了。Cloudflare 是赢家,因为我可以去注册它。我可以访问该服务。我真的可以使用它。所以我可以完成整个流程。你想使用它吗?”“是的,我们开始吧。”

Original English

Speaker: harness or system using pi or cloud or something and it says welcome back Michael at work. That's actually my email if you want to email me. Do you want to keep working on your spelled app? I'm like yeah sure I'd like to share this with my friends but they can't access it. And the agent's like oh well local host only works on your machine so your friends won't be able to hit it. Of course to share it I need to deploy it somewhere real. Some providers actually let me handle the signup. So, you know, you wouldn't have to sign up and do this. You want me to look around for you? Yes, please. And then the agent can go look for services that advertise through OMD that they support registration. Just some examples, Stratus, Helio deploy, they don't support it, but Cloudflare here does have an OMD. And so, the agent here can say, "Ah, Cloudflare, that's the one. Cloudflare is a winner because I can go sign up for it. I can access the service. I can actually use it. So I can do the whole thing. You want to go with it? Yeah, let's do it.

Speaker: 这就是注册步骤。这差不多也是见证奇迹的地方。我的代理实际在做的就是生成一个身份声明,并将其提供给 Cloudflare。然后 Cloudflare 能够选择验证它或不验证。Cloudflare 实际上可以选择带外验证(out of band)。AMD 非常灵活。你可以根据你所期望的行为,通过不同的方式来调整它。一种模式并不适用于所有情况。每个应用程序都是不同的,约束条件也完全不同。如果你构建的是类似于数据库服务的东西,其中包含的信息量很少,你可能想要免费提供一点点流量。但是,如果你构建的是比如电子邮件服务或类似的东西,特别是如果你要接入物理世界,你可能不想免费提供任何东西。这是非常特定于应用程序的。所以在这里,Cloudflare 说:“好的,我们愿意让你部署,也许是 72 小时左右。获取身份。砰,砰,砰。通过验证。”这个标准被称为 ID Jag。然后它返回消息说,你已经建立了一个 Cloudflare 账户。现在在这个体验中,我这个人类没有点击任何地方。我没有在任何地方注册。它只是获取了令牌(token)。我可以做几个小调整。它知道如何使用 Wrangler。代理问:“想要我进行那些修改吗?”“放手去做吧。”

Original English

Speaker: This is the registration step. This is kind of where the magic happens. What my agent is actually doing is minting an identity assertion and giving that to Cloudflare. And then Cloudflare is able to verify that or not. Cloudflare can choose actually to verify out of band. AMD is very flexible. There's different ways you can dial it in depending on the behavior that you're looking for. One size does not fit all. Every application is different. Constraints are totally different. If you're building something that's like a database service where there's going to be very little amount of information. You might want to give away a little bit of traffic for free. But if you're building something maybe like an email service or something certainly if you access the physical world, you you maybe don't want to give anything away for free. It's very application specific. So here Cloudflare says, "Okay, we're willing to let you deploy maybe for 72 hours or something like that. Get the identity. Boom, boom, boom. Go through. It's called ID Jag is what the standard is. And then it returns back and says you're set up with a Cloudflare account. Now in this experience, my you know human I didn't go click anywhere. I didn't sign up anywhere. It just it just got the token. I can make a couple small tweaks. It knows how to use Wrangler. Want me to make those changes? Go ahead.

Speaker: 大家都像这样提示他们的代理吗?就像在说“是的,是的,请做吧”,实际上并没有说太多。它能够编写 Wrangler 配置,完成设置,走完整个部署流程,然后砰的一声,它就上线了。我们实际上有一个正在运行的实时演示(live demo)。在这个项目上我们也和 Cloudflare 进行了一些合作。这非常酷。你可以看出我贯穿整个过程的提示并不是什么高深莫测的技术。实际上代理可以直接一次性跑完整个流程。除了不断地“批准,批准,批准”之外,我没有通过提示做任何实际操作。它应该能够一次性完成这个操作。这就是它在图形界面上的样子。

Original English

Speaker: Does anybody prompt their agents like this? It's just like yes, yes, do it, please. Not really saying much. It's able to write the Wrangler, get the thing set up, go through the full deployment process, and then boom, it's live. We actually have this working as a live demo. We did some collaboration with Cloudflare on this too. It's pretty cool. And you can tell the prompting that I was giving through this not exactly rocket science. The agent could just run through the whole thing actually. There's nothing actually that I was doing through prompting it other than just approving approving approving. It should be able to do this one shot. This is what it looks like graphically.

Speaker: 那么,带有 offd 的代理注册过程,比如你有你的代理工作台 PI。在它内部有你的大语言模型(LLM),也许还有你的记忆模块以及不同的上下文层。它连接到你的后端身份服务。所以这是用户身份,当代理去注册时,它会向 Cloudflare 发送发现请求(discovery call)。它会寻找那个 aund 或者其他什么服务。我们也让它与 fire call 一起工作过,这非常棒。它会发送那个 ID JAG,即身份 JSON 访问授权。这本质上是一个代表用户的签名凭证,声明“这是一个身份”,然后接收返回的访问令牌——也就是实际的 API 令牌,然后砰的一声,就完成了。你可以进行普通的 API 请求,你可以调用 MCP(模型上下文协议),所有其他功能仍然有效。所以你可以看到,OMD 仅仅是用于注册这一步,你正常的 API 继续工作,你的其余文档不需要更改,它只是为代理提供的注册步骤。但是伴随着这种简单性,实际上带来了巨大的力量,因为突然之间,它允许代理进行端到端的操作。

Original English

Speaker: So agent registration with offd you have your agent harness pi for example. within that is your LLM, maybe your memory, different, you know, context layers. It connects to your backing identity service. So this is the user identity and when that agent goes and registers, it sends the discovery call to Cloudflare. It looks for that aund or whatever service. We also have this working with fire call. It's pretty cool. It sends that ID JAG, the identity JSON access grant. It's essentially a signed credential that represents the user says this is an identity receives back the access token actual API token and then boom it's it's done you can make normal API requests you can call MCP everything else still works so you can see that OMD is just for the registration step your normal API keeps working the rest of your docs don't have to change it's just that registration step for agents but with the simplicity comes actually an enormous amount of power because suddenly it lets agents do things end to end.

Speaker: 我讨厌计划步骤。当我写点什么的时候,我只是希望代理去把它做出来。通常它大致知道我想要表达的意思,而我想要的是回来时看到一个实际运行的原型,而不是像“嘿,这是我长长的计划。这里有一大堆需要你去注册的服务。”那样。特别是在构建新的原型时,我宁愿让它直接免费部署到某个平台上,这样我就可以看到它,到处点一点,而不是让我承受去注册所有这些不同供应商的负担。而这就是 OMD 允许做的事情。

Original English

Speaker: I hate the planning step. When I write something, I just want the agent to go do it. Usually it kind of knows what I'm trying to get at and I want to come back to actually a running prototype and not like, hey, here's my long plan. Here's a bunch of services to sign up for. If I'm building new prototypes, especially, I would prefer it just to get deployed on something for free so I can see it and click around and not put me through the burden of going sign up for all these different vendors. And that's what OMD allows for.

面向服务商的代理驱动增长

Speaker: 如果你是一家服务商,如果你正在构建 API 或服务,这让你可以成为那些系统的客户。你可以把它视为增加你的增长漏斗——从一开始,你知道,也许只有人类用户,到后来把整个代理世界都加入其中。这相当棒。我们认为这将产生巨大的影响。你的大门将向代理解开,让它们实际上去注册并使用你的服务。这对你的产品来说应该是一次立竿见影的增长爆发。我并不是说代理要构建的所有东西都会是有用的。也许不是它们构建的每个东西都会进入生产环境,或者能持续那么长时间,但总有些东西会留下来。就像产品导向型增长(PLG)或在它之前的免费增值(Freemium)模式一样,这是发展你的业务、获取更多客户并真正让你变得更大、更成功的方法。只要有使用量,我们就相信也会有付费意愿。有许多关于让代理真正为内容付费的有趣工作正在进行中。有一些基于加密货币的协议,比如 X42。支付供应商也在开发其他一些东西,但我们认为那应该是下游阶段的事。我们不希望付款成为前置条件。我们不想强迫你的代理必须用信用卡注册。那似乎是本末倒置。在 SaaS 的世界里,我们已经摆脱了这种模式。有了 OMD,代理可以注册,它们可以开始使用,然后在下游阶段根据你自己的定价、按照你自己的模式,在你的条款下选择付费。

Original English

Speaker: And if you're a service provider, if you're building APIs or services, this allows you to become a customer of those systems and you can think about it increasing your growth funnel starting from, you know, maybe just humans and adding the whole world of agents on top. Pretty great. We think this is going to be huge. Your door will be open for agents to actually sign up, register, use your services. This should be an immediate growth boom for your products. And I'm not saying everything that agents are going to build will be useful. Maybe not everything they build will go into production or last that long, but some things will. Just like with PLG or Premium before that, this is the way to grow your business and get more customers and actually become, you know, larger and more successful. And where there's usage, we believe there's also intent to pay. There's a lot of interesting work being done for agents to actually pay for stuff. There's crypto-based protocols like X42. There's other things that the payment providers are building, but we think that should be downstream. We don't want payments to be upfront. We don't want to have to force your agent to sign up with a credit card. That seems backwards. We moved away from that in the world of SAS. With OMD, agents can sign up, they can register, they can start using it, and then downstream choose to pay on your terms based on your own pricing, based on your own model.

Speaker: 如果你不相信我关于这种无头(headless)类型产品所说的话,这有马克·贝尼奥夫(Marc Benioff)。他是 Salesforce 的 CEO 兼创始人,就是城里那栋大楼里的伙计们。他们比任何人都更相信代理。“我们的 API 就是用户界面(UI)。”整个 Salesforce 和 Agentforce 的 Slack 平台现在都作为 API、MCP 和 CLI 暴露出来。他在谈论优先采用“代理分叉(agent fork first)”。一家被许多人视为某种传统的 SaaS 供应商,实际上发明了“软件即服务”这个类别的公司。他们正在摆脱他们的仪表板和 UI。我的意思是,他们并不是完全抛弃它,但他们相信代理是未来。所以,如果你还没有思考过这个问题,我强烈鼓励你开始构建一些原型,与用户交谈,因为你绝对不想错过,错过这股潮流。

Original English

Speaker: If you don't believe me or on the headless type of products, this is Mark Beni off. He's the CEO and founder of Salesforce, the guys at the big building here in town. They believe in agents more than anything. Our API is the UI. Entire Salesforce and agent force Slack platforms are now exposed as API, MCP, and CLI. talk about going agent fork first. A company that many people consider to be sort of a legacy SAS vendor really invented the category of software as a service. They're getting rid of their dashboard and UI. I mean, they're not getting rid of it, but they believe that agents are the future. So, if you're not thinking about this, I definitely encourage you to start building some prototypes, talking to users, because you don't want to miss miss this.

OMD 开放规范与未来范式

Speaker: 如果你想使用这个功能,它今天就已经可用了。我们有它的开放规范。这里有一个 GitHub 仓库。你实际上可以让你的代理去实现它。我知道这有点“元(meta)”。这是一个开放规范。WorkOS 并不控制它。我们有一个我们托管的版本,但你自己也可以去构建它。你可以在你现有的系统中自己运行它。我认为让单一的供应商不拥有这东西是非常重要的。如果你使用我们的用户服务,它可以实现一键启用,但你可以在你想要的任何平台上构建它。我们在 Moscone 会议中心,这非常棒。

Original English

Speaker: If you want to use this, this is available today. We have open spec for it. There's a GitHub repo. You can actually have your agent go implement it. I know it's kind of meta. It's an open spec. Work OS doesn't control it. We have a version of it that we host, but you can build it yourself. You can run it yourself in your existing systems. I think this is so important that no single vendor owns this. It's one click enabled if you use user stuff, but you can build it on whatever platform you want instead. here at Moscone which is pretty awesome.

Speaker: 在 2007 年 1 月,我相信史蒂夫·乔布斯(Steve Jobs)发布了 iPhone,那时我还在上大学。我记得,是的,大概是大一。我非常非常清楚地记得那一刻。它是下一个被建立起来的伟大软件平台。有如此多的公司被创立,仅仅是因为它们可以建立在这些设备之上。智能手机革命,而且我认为从那以后就再也没有出现过这样的时刻。在技术时代,至少对我而言,我们围绕 B2B SaaS、云等等所做的一切都很棒,但这是一种全新的范式,一个新的软件平台,而且我认为代理正是如此。代理是下一波大浪潮。所以我想对大家说的是,走出去,用 AMD 开始构建吧。我们希望为你解锁这股力量,为你的业务和增长解锁它,但我们也想听到你们的声音,听听你们在哪里遇到困难。我相信大家一起努力,我们能够为未来而构建。

Original English

Speaker: In January of 20h 2007 I believe Steve Jobs introduced iPhone and I was in college and I remember yeah like first year of college I remember this super super clearly. It was the next great software platform that got built. So many companies got started because they could build on top of these devices. smartphone revolution and there hasn't really been a moment like this since I think in the the era of technology at least for me the stuff we did around B2B SAS and cloud and all that is great but this was a totally new paradigm a new software platform and I think agents are this agents are the next big wave and so what I would say to all of you is go out and build with AMD we're hoping to unlock this for you unlock this for your business and for your growth but we want to hear from you here the places you get stuck and I think together we can build for

迎接智能体革命

Speaker A: ……迎接这个激动人心的软件新时代,并为智能体革命而构建。非常感谢。

Original English

Speaker A: this next exciting era of software and build for the agentic revolution. Thank you so much.

一家学会记忆的工厂

Rushab: 好了,我想给你们讲个故事,关于一家工厂如何自学记忆的故事。你好,我是 Rushab。我在印度经营着一家名为 Machine Craft 的百人规模工厂。我们没有数据科学团队,没有机器学习预算,什么都没有。但不知怎么的,我们最终构建了一个包含36个 AI 智能体的系统,来主导我们整个走向市场战略。我觉得这听起来还是有点荒谬。让我来给你们展示这是如何发生的,以及为什么你们也能做到同样的事情。

Original English

Rushab: Okay, I want to tell you a story about a factory that taught itself how to remember. Hi, I'm Rushab. I run machine craft a 100 people factory in India. No data science team, no ML budget, none of that. And somehow we ended up building a 36 AI agent that runs our entire go to market. I think that's still a little ridiculous. Let me show you how it happened and why you can do the same thing.

Rushab: 关于我们公司,情况是这样的。从外面看,它就是机器和金属。但实际上,这家公司真正重要的部分不在机器里,而在于知识:客户是谁,我们在2019年给他们报了什么价格,为什么那台机器需要那种奇怪的定制调整。在过去的三代人里,所有这些知识都只存在于三个大脑中——最初是我祖父的,然后是我父亲的,现在是我的,当你静下心来想想,这绝对是一种令人毛骨悚然的公司经营方式。有很多人加入过我们。也有人离开我们。人员流动从未停止过。每次有人离开,我们公司的一块“大脑”也就跟着他们一起离开了。我们不害怕竞争对手。我们害怕的是遗忘,或者某天醒来意识到,整个公司仅仅存在于两个越来越疲惫的脑袋里。

Original English

Rushab: So here's the thing about our company. From the outside, it looks like machines and metal. But the actual company, the part that matters is in the machines, is the knowledge. Who the customer is, what we quoted them in 2019, why that one machine needed that weird custom tweak. And for three generations, all of that lived in exactly three brains. Initially, my grandfather's, then my father's, and now mine, which is a genuinely terrifying way to run a company when you sit with it. A lot of people have joined us. People have left us. The revolving door never stopped. And every single time someone walked out, a chunk of our brain walked out with them. We weren't scared of the competitors. We were scared of forgetting or waking up one day and realizing the whole company only existed inside two increasingly tired heads.

Rushab: 于是,我有个想法。老实说,一开始听起来挺疯狂的,但如果不是把知识写在没人看的文档里。要是我们培育一个直接承载这些知识的“大脑”会怎样呢?它不是你偶尔戳一下的聊天机器人,而是公司的数字孪生。我没有雇佣一个销售团队。而是试着去建立一个。

Original English

Rushab: So, I had an idea. I'll be honest. sounded insane at first, but what if instead of writing the knowledge down in some document, nobody ever reads. What if we grew a brain that just held it? Not a chatbot you poke at, a twin of the company. I didn't hire a sales team. I tried to build one.

Rushab: 稍稍跑个题,因为你们得知道这到底有多混乱。我们制造热成型机。它们把塑料薄膜加热然后塑形。核心机器是同一台,但它最终制造出来的东西可能是水培农场托盘、水疗浴缸、电动汽车面板、医疗外壳,甚至是包装。这是七个完全不同的领域,面对七种完全不同的买家。所以,这个大脑不能只是死记硬背一本宣传册。它必须知道某个特定的客户生活在哪个领域里。

Original English

Rushab: A quick detour because you need to know how messy this is. We make thermopforming machines. They heat up a plastic sheet and shape it. Same core machine, but it ends up making hydroponic farm trays, spa bathtubs, EV car panels, medical casings, and even packaging. Seven totally different worlds, seven totally different buyers. So, this brain couldn't just memorize a brochure. It had to know which universe a given customer lives in.

Rushab: 第一步几乎可以说是无聊地简单。喂给它所有的东西,我是说“所有的”东西。数年来的报价、图纸、付款计划、时间表、电子邮件记录,几百GB属于我们自己的私人历史记录。不是公开的互联网数据,而是我们的内部网络数据。而这就是情节转折的地方,这部分让每一个听我讲过这个故事的工程师都感到惊讶。我们从未训练过模型。地下室里没有嗡嗡作响的 GPU,也没有进行微调。我们只是审视了所有的历史,把它切成大小合适的数据块,然后让现成的模型去阅读并提取事实。我们将每个数据块的含义以向量和关系的形式存储起来。它与谁相连?这实际上是一个极其、极其井然有序的记忆库。

Original English

Rushab: Step one was almost boringly simple. feed it everything, and I mean everything. Years of quotes, drawings, payment schedules, timelines, email threads, hundreds of gigabytes of our own private history. Not the public internet, our internet. And here's the plot twist, the part that surprises every engineer I tell this to. We never trained a model. No GPUs humming in the basement, no fine-tuning. We just looked at all the history, chopped it into bite-sized chunks, and let offshelf models, read it, and pull out the facts. We stored the meaning of each chunk as vectors and relationships. Who's connected to? It's actually a really, really well organized memory.

Rushab: 接下来,情况在某种好的意义上变得有些怪异了。我们不再把这个当成软件来看待,而是开始把它当成我们在“抚养”的东西。于是我们参照生物学给了它一具躯体:有感官去弄清楚它在跟谁说话,有肠胃去把文档消化成事实,有记忆,有梦境循环,还有一个能抵御不良信息的免疫系统。为什么要用生物学作为模型呢?因为进化已经花了十亿年的时间来解决。你如何随着时间的推移保持连贯性?我们只不过是抄了作业。

Original English

Rushab: Now, this is where it gets a little weird in a good way. We stopped thinking of a as a software and started thinking of it as something we were raising. So we gave it a body modeled on biology, senses to figure out who it's talking to, a gut to digest the documents into facts, a memory, a dream cycle, an immune system to fight off bad information. Why biology? Well, because evolution already spent a billion years solving. How do you stay coherent over time? We just copied the homework.

Rushab: 好了,那么最大的问题来了,为什么?

Original English

Rushab: Okay, so the big question, why?

AI 智能体的哈具工程

Mike Chambers: 大家好。各位 AI 工程师们好。大家现在玩得还开心吗?我反正很开心。我的意思是,看看我。我站在这里。我太喜欢这个了。嗯,所以是的,非常感谢你们的到来。嗯,我想来跟你们聊聊哈具工程(harness engineering)以及所有相关的话题。在你们可能还不认识我的情况下,我先自我介绍一下。我的名字是 Mike Chambers,我是一名高级 AI 专家、开发者布道师。嗯,我在亚马逊 AWS 工作。

Original English

Mike Chambers: Hello everybody. Hello AI engineers. Are we all having a good time still? I'm having a good time. I mean, look at me. I'm up here. I'm loving this. Um, so yeah, thanks so much for joining me. Um, I want to come and talk to you all about harness engineering and all that kind of stuff. Um, let me tell you who I am in case you've not met me before. Um, my name is Mike Chambers and I'm a senior AI specialist, developer advocate. Um, and I work at Amazon at AWS.

Mike Chambers: 稍微讲一下关于我是如何能够站在这里的,对我来说这是一个非常激动的时刻。回溯到2023年,在生成式 AI 的发展进程中那已经算是很久以前了,我有幸与我当时的同事 Antia(她现在在亚马逊 AGI 工作,你们之前可能在这个舞台上见过她),以及了不起的吴恩达(Andrew Ng)博士一起合作制作了一门关于使用大语言模型(LLMs)进行生成式 AI 开发的课程。嗯,我可以说这门课的注册人数快接近50万了吗?看起来情况确实如此。对于一门为期三周的课程来说,那相当酷。如果你看不清这张照片里的内容,嗯,我当时正在和 Android 一起玩变形金刚。当时觉得做那件事真的很好笑。

Original English

Mike Chambers: Um, little bit about like, uh, how I managed to get to stand here, which is a very exciting time for me. Um so um quite a while ago in terms of generative AI anyway back in 2023 um I had the amazing awesome privilege uh to work with Antia my colleague at the time and now she works for Amazon AGI you've probably seen her on this stage before um and the amazing Dr. Andrew Ing on on a course about generative AI with LLMs um sort of can I say that we're approaching half a million enrollments with that? It looks like that's the case. And on a three-week course, that's pretty cool. If you can't tell in that image, um, I'm playing Transformers with Android. That seemed like a really funny thing to do at the time.

Mike Chambers: 在2025年,我创建了一个 MCP Lambda 处理器,直到今天每个月仍有大约35,000次下载,来帮助人们以一些最简单的方式实现 Serverless MCP 服务的部署。这次我会讨论与此相关的其他事情。所以我们已经从那个阶段向前迈进了。到了2026年。AWS 实际上成为了智能体 AI 基金会(Agentic AI Foundation)的创始成员之一,该基金会隶属于 Linux 基金会。我正在幕后参与一些这方面的工作。希望将来也能在这方面做更多。关于我的介绍就到这里。

Original English

Mike Chambers: Um, in 2025, I created an MCP Lambda handler that's downloaded still to this day at about 35,000 times a month um to help people in some of the simplest ways of getting serverless MCP serving happening. Um, I'm going to talk about other things in relation to that this time. So, we've moved on from that. um and in 2026. So the AWS is actually one of the founding members of the Agentic AI Foundation, part of the Linux Foundation. Um I'm doing a little bit of work behind the scenes on that. Hope to do a lot more of that as well. So little bit about me.

Mike Chambers: 当我为这次演讲做准备时,哦对了,我又重新读了一遍这个议程的摘要,发现我说过我会进行一些现场编程,所以我真的会这么做。所以大家一起祈祷这部分一切顺利吧。但在我四处奔波的过程中,这是我的常态,我参加了在墨尔本举办的 AI 工程师峰会(会议),吸收了很多东西,也包括本周初的一些见闻。我只是想总结一下我所看到、所思考的,以及我真正想传达并且对我来说很重要的那些事,那就是这个。

Original English

Mike Chambers: Um so as I've been um preparing for this, oh by the way, I did reread the abstract for this session and realized I said I'd be doing some live coding and so I will. So all combined fingers crossed please that that all works for us. Um but as I've been sort of traveling around a little bit as I do and I was at the AI engineers uh session uh uh summit uh uh conference in Melbourne um and took a lot of it in and also from the beginning of this week as well. I just wanted to to to summarize some of the things that I'm seeing and I'm thinking and I really want to get across and and what really matters to me and that's this.

Mike Chambers: 存在两种不同类型的智能体。我们一直都在谈论智能体,但我认为存在两种截然不同的智能体。当我说出来的时候,这会变得非常显而易见,但它们就是我们使用的智能体。所以,你知道,比如 Claude Code、Cursor、Kirao,以及所有那类东西。也包括那些不仅用于生成代码,还被我们用于提高生产力之类的东西。所以这些我们使用的智能体有特定的使用方式。这是我们使用的智能体。而另一方面,是我们构建的智能体。这实际上与我更相关,也与这次演讲更相关。这是我们构建的智能体,以及我们如何思考我们所构建的智能体。

Original English

Mike Chambers: There are two different types of agents. Um, so we talk about agents all the time, but I see two distinct types of agents. And as I say this, it's going to become really obvious, but they're the agents that we use. And so, you know, this is claude code and cursor and kirao and all of those types of things. And also things that don't just generate code, things that we use for productivity and the like as well. And so those agents we use in a certain type of way. There's the agents that we use. And then on the other side of it, we've got the agents that we build. And that's actually more to do with me and and actually it's more to do with this presentation as well. It's agents that we build and how we think about agents that we build.

Mike Chambers: 我真的认为这两者是非常不同的事情,但它们也是串联在一起的。我可能会构建一个让你来使用的智能体,而这依然适用。所以,像最大化 Token(token maxing)之类的做法,如果是为了一个你正在使用的智能体,想做就去做吧。但对于一个你正在构建的智能体,一定要谨慎思考。确保你构建它的方式,能够适用于最终使用它的受众。

Original English

Mike Chambers: Um and so I do really think that these don't two things are quite separate and they chain together as well. I might build an agent that you use and this still holds true. So token maxing, all that kind of stuff, go for it if that's what you want to do with an agent that you use. But with an agent that you build, think about it carefully. Make sure that you're putting it together in a way that's going to work for the audience who's going to use that.

Mike Chambers: 我保证过我们要讨论哈具(harnesses)和哈具工程(harness engineering)。那么我们先来定义一下哈具。我相信我肯定不是唯一一个把这种东西放出来的人。我通常不干这种事,如果这让你觉得尴尬的话我先道歉。这是词典上关于哈具(harness / 挽具)的定义。哈具是一组用来控制动物的皮带和扣件。但如果我们把这里的“动物”去掉,换成“模型”,那么其实它挺准确的,对吧?那大概就是哈具的意思。

Original English

Mike Chambers: So, I promised that we're talking about harnesses and harness engineering. So, let's define harness. I'm sure I'm not the only person to have put something like this up. I don't usually do this kind of thing, and apologize if it makes your skin crawl. This is a dictionary definition of harness. A harness is a set of straps and fastenings used to control an animal. But if we took animal out of here and put model in there, then actually it's pretty right, right? That that's kind of what a harness is.

Mike Chambers: 其他人对什么是哈具做了比仅仅基础词典定义好得多的解释。比如 LangChain 已经发布了一篇文章。你们可能都看过那类内容。Martin Fowler.com,虽然不是 Martin 写的。是 Pitta 写的。针对编码智能体的哈具工程。我们使用的智能体,对吧?所以还有其他看待哈具的方式。

Original English

Mike Chambers: Other people have done a much better job than just the basic dictionary definition of what a harness is. So Lang Chain has got a article out. You've probably seen stuff like that. Martin Fowler.com, although Martin didn't write it. It was Pitta wrote this. Harness engineering for coding agents. Agents that we use, right? So there are other ways of looking at harnesses.

Mike Chambers: 那亚马逊是怎么看的呢?我们如何看待哈具?嗯,不,好吧,这是错误的那种安全带(harness)。抱歉,我们对哈具确实有强烈的观点。我会给你们展示所有的内容,但我们卖的东西五花八门。所以,简而言之,如果你读过那些文章,而且如果你在这个活动的现场参与过讨论,当然,哈具就是,你拿一个智能体,把其中的模型部分去掉,剩下的一切,那就是哈具。好的,所以让我们在“我们使用的智能体”的背景下思考一下。所以我觉得相当……

Original English

Mike Chambers: What about Amazon then? How do we see um harnesses? Well, um no, okay, this is the wrong kind of harness. Sorry, we do have strong opinions on harnesses. I'm going to show you all of that, but we sell all kinds of things. So, in a nutshell, and if you read those articles, um and if you've had the conversations around here at this event, of course, um a harness, you take an agent, remove the model part from it, and everything that you have left, that's the harness. Okay, so let's think about that in context of an agent that we use. And so it's pretty I think

Coding Assistant Harness

Speaker: 相当直截了当。我有这个编程助手。它可能就在我的机器上。它能访问我的文件,我要么创建一个“控制台(harness)”,要么我工作的地方已经为我创建好了一个控制台,其中包含了它将如何使用内存,我希望它使用的技能、工具,以及允许它连接到文档服务器等地方的 MCP 服务器。而且,建立良好的工程团队已经有了他们的标准,他们几十年来一直有编程标准,但现在他们基本上有了控制台标准,也就是他们希望部署到每个人的编程助手中的东西。所以,简而言之,就是这么回事。我不想过多讨论这个话题。

Original English

Speaker: fairly straightforward. I have this coding assistant. It's probably on my machine. it has access to my files and I create a harness or the place I work at has created a harness for me which contains um how it's going to use memory the skills that I want it to use tools and MCP servers to allow it to go to be able to go and connect to documentation servers and the like and and well set up engineering teams have got their standards that they've had they've had coding standards for decades but now they have basically harness standards the things that they want to deploy to everybody's coding assistance. So, in a nutshell, that's what it is. I'm not going to talk to too much more about that.

Speaker: 但是,我确实想和你们分享一个二维码,我尽量在展示二维码前给你们一点提示。这是适用于 AWS 的 agent 工具包。这个在 GitHub 上可用。当然,它是免费的。你可以安装它,如果这正是你在做的事情并且你正在部署代码,它会帮助你。如果你正在 AWS 上部署代码,或者你正在考虑在 AWS 上部署代码,亦或者也许你有一天会这样做,那就获取这个工具包,让你的 agent 引导你走向正确的方向。这里有如何在几乎所有平台上安装它的说明。

Original English

Speaker: Um, but I do want to share one QR code with you, and I'll try and give you a little bit of warning before I bring QR codes out. This is the agent toolkit for AWS. This is u available on GitHub. Of course, it's free. You can install it and it helps you if this is what you're doing and you're deploying code. If you're deploying on AWS or you're thinking about deploying on AWS or or maybe you will one day, grab this toolkit, enable your um uh agent to to help you in the right direction. It's instructions for how to install it on pretty much everything.

Speaker: 我之所以对此如此充满激情,是因为我不想再看到任何“slop ops(草率运维)”了。所以我们过去总是抵制在专业的云开发环境中使用“click ops(点击运维)”。在控制台上点来点去对于弄清楚正在发生的事情非常有用,但这并不是你将东西部署到生产环境的方式。我们可以问 agent 正在发生什么,但我们不想让 agent 启动一个 S3 存储桶,给我弄个 EC2 实例,或者做诸如此类的事情。我们希望 agent 将我们的基础设施作为代码构建起来,然后去执行,这样我们就仍然拥有云中部署的所有权。所以,不要再有“swap ops”了。

Original English

Speaker: And the reason why I get passionate about this is because I don't want to see any more slop ops. Um so we always used to push back against click ops in, you know, in the professional cloud development space. Clicking around on the console is great for being able to figure out what's going on, but it's not how you deploy things into production. We can ask an agent what's going on, but we don't want to ask the agent to spin up an S3 bucket, get me an EC2 instance, whatever it might be. We want the agent to build up our infrastructure as code which is going to go and do that so that we still own our deployments in the cloud. So, no more swap ops.

Agent Engineering at Scale

Speaker: 好了,这就是我们使用的 agent。现在,我们来谈谈我们要构建的 agent。我会尽快进入代码部分,并在时间允许的范围内尽可能多地演示。那么,关于我们正在构建的 agent,我们应该如何看待控制台呢?在一定程度上是完全一样的,是的,我们仍然需要考虑我们该如何管理内存?我们该如何在 MCP 中管理技能和工具?顺便说一句,这背后其实掩盖了很多东西,对吧?因为你几乎可以扩展一个 agent 去做任何你想做的事情,通过大量不同类型的工具,这些工具可能是通过 MCP 连接的。

Original English

Speaker: Okay, so that's the agent that we use. Now, let's go and talk about the agent that we're going to build. And I'm going to get into the code as quickly as I can, and we'll do as much as it has time for. So, how do we think about a harness in relation to the agent that we're building? Exactly the same to a point, yes, we still want to have how are we going to manage the memory? How are we going to manage the skills and tools in MCP? By the way, that that belies a lot of stuff, right? because you can pretty much extend an agent to do almost anything you want with a whole bunch of different types of tools which could be via MCP.

Speaker: 但是,对于我正在构建的一个 agent 而言,我需要考虑的东西远不止于此,特别是如果,比如在亚马逊和在云级别规模上,我要将我的 agent 部署给大众。那么,我到底该如何管理循环呢?我该如何管理扩展、支付、内存、身份、技能、运行环境、上下文管理以及其他所有东西呢?我把这留到了最后,但它应该是我首先要说的事情。那就是可观测性和评估,超级超级重要。我们到底该如何处理这个问题呢?我是把所有这些代码写进一个容器里,然后部署并扩展它吗?并非如此。如果我想扩展到数千个用户,我需要考虑这些组件中的每一个,以及我将如何独立地对它们进行扩展。对我来说,这就是控制台工程(harness engineering)。这是非常严肃的一面。这是我们希望让控制台在真正的大规模下运行的大事情。

Original English

Speaker: But with an agent that I am building, I need to think about a lot more than just that, especially if um like at Amazon and like at cloud scale, I'm deploying my agent out to the masses. So, how do I actually manage the loop? How do I manage um scaling, payments, memory, identity, skills, runtime, context management, the rest of it? And I have left it to the last thing, but it should be the first thing that I say. Observability and evaluations, super super important. How do we actually deal with this? Do I write all of this code down into one container and just deploy it and scale that? Not really. If I want to be scaling to thousands of users, I need to think about each individual of these components and how I'm going to scale them out individually. And that to me is harness engineering. This is the serious side of stuff. This is the big stuff that we want to get harnesses working at real scale.

Code Examples in the IDE

Speaker: 好了,让我们看看这能不能跑通。我能感受到你们传给我的共同善意,希望能让这代码跑起来。现在,我进入了 Kira。这是我首选的 IDE。我这里准备了几个不同的示例,我们将快速地过一遍,因为倒计时正在飞速流逝。所以,请确保我们都在同一频道上。希望大家都能看到这个。顺便提一下,我即将向你们展示的代码中,有一段代码(不是这一段)是由 Curo 生成的。其它要么是工具生成的,要么,就像这一段,是我自己写的。这一段我没用 agent,我知道我值得你们的一阵掌声,但也没关系啦。

Original English

Speaker: Okay, let's see if this works. I can feel your combined goodwill being sent my way that we're going to try and make some code work. So, I'm here in Kira. This is my IDE of choice here. And I've got a few different samples that we're just going to race through watching that clock countdown fast. So, um, just just make sure that we're all on the same page here. And hopefully you can all see this. Um, of the code which I'm about to show you, by the way, one piece of code, not this one, has been generated by Curo. Everything else is either a tool or this one, I actually wrote it myself. I didn't use an agent for this. I know I deserve a round of applause, but it's okay.

Speaker: 所以,这是一个 Strand agent。我拿了 Strands agent 的 SDK。希望这类的东西看起来还算熟悉。我引入了一个 agent,我引入了 tool 装饰器,然后我给自己创建了一个 agent。下方是工具定义。我就把我的系统提示词传进去了。在这个阶段一切都很简单。而且我传了几个工具进去,计算器(Calculator)是一个我可以安装的库,而获取时间(get time)是我们经常使用的,因为我不会用 agent 去订机票,肯定不会用这种 agent。所以我可以在这里问一个简单的问题,比如“现在几点了?”我不会去运行这个,因为你们知道时间,但你们大概能看出这是如何运作的。

Original English

Speaker: So, this is uh this is a Strand agent. So, I've just taken the Strands agents SDK. Um, and hopefully this kind of thing is kind of familiar. I've brought in an agent. I brought in the tool decorator. And I'm creating myself an agent. The tool definition is down here. Um, and so I just pass in my system prompt. Things are pretty simple at this stage. And I've passed in a couple of tools. Calculator is something that's a library I can install. And get time is the one that we always use because I don't tend to use agents to book flights. Certainly not ones like this. Um, and so I can say something simple here like what is the time? I'm not going to run this because you know the time, but you can see generally how this works.

Speaker: 这是一个控制台(harness)吗?算是吧。其实里面没多少东西,对吧?我们只是把工具放在那里了。我们的循环由框架替我们管理。这很酷。这很好,但显然,如果我要运行它,它只会在我的笔记本电脑上运行。它没有在任何特定规模下运行。而且它缺少了一些我希望部署的 agent 所应具备的属性。让我快速转到下一个 agent。

Original English

Speaker: Is this a harness? Sort of. There's not an awful lot to it, right? We've got the tools in there. Our loop is being managed for us by the framework. This is pretty cool. So that's good. But obviously if I was to run this, this is running on my laptop. It's not running at any particular scale. And we're missing some of the attributes that I want from the agents that I'm going to deploy. Let me move on to my next agent quickly.

Adding Session State and Memory

Speaker: 所以,这也是一个 Strand agent,但为了这次会议,我实际上是让 Kira 帮我写的,因为我想加入一些更多的东西。在这个 agent 中,我主要想指出的一点是,我加入了一个 session manager(会话管理器)。这个会话管理器帮助我在每次调用之间维持会话状态。所以这是一种内存。它是一种中期、短期类型的内存。它不是那种合适的长期内存,但它确实存在。而且实际上,它确实将会话长期保存在了这里侧边的文件中。如果我往下滚动,你可以看到,是的,这里是 agent 的定义本身。我们有一个更长一点的系统提示词,因为 Kira 自己忍不住写多了。我们定义了一些工具,包括一个让 agent 可以决定用来记住关于我的信息的 remember 工具。下面是我的会话管理器。这个会话管理器会在我下次回来和它聊天时恢复对话历史。也许这“下一次”就是现在。

Original English

Speaker: So, this is also a strand agent, but this one I actually asked Kira to write it for me for this session um because I wanted to include some more stuff. And so, inside of this agent, the one main thing that I want to point out is that I am uh I've included a session manager. So, my session manager is helping me to maintain session state between invocations. So, this is a sort of memory. It's a kind of medium-term short-term memory kind of thing. It's not proper long-term memory, but it is there. And actually, it does store long-term memories in files which are down the side here that it's uh included for us. So, if I just scroll down here, you can see uh yeah, here's the agent definition itself. Um, and we've got a bit more of a system prompt because Kira couldn't help itself. Um, and we've got some tools here defined. Um, including a remember tool that the agent can decide to use to remember stuff about me. Um, and then I've got my uh session manager down there. And that session manager is going to rehydrate the conversation history when I come back to chat to it the next time. And maybe the next time is now.

Speaker: 让我们看看能不能让它跑起来。再说一遍,这是在我的本地机器上运行。这是一个演示。我们输入 hello,因为我怕打字太多拼错。然后它回复:“yeah, keep testing me. Bring it on.”(好啊,继续测试我。放马过来吧。)非常好。“who will win the World Cup?”(谁会赢得世界杯?)显然我需要知道这个。它说了什么?是的。虽然你可以直接说是澳大利亚,但它知道我希望澳大利亚赢。因为我目前住在这里。我就是澳大利亚。所以显然澳大利亚将赢得世界杯。但它为什么会得出这个结论?是因为我们之前的对话,它很诚实地表示自己一无所知,因为那显然是来自大型语言模型。

Original English

Speaker: So, let's see if we can get this working. Now, again, this is running on my local machine. Um, and this is a demo here. So, let's just type in hello because I'm scared of typing too much and spelling it wrong. Um, and it says, yeah, keep testing me. Bring it on. Excellent. Um, um, who will win the World Cup? So, obviously I need to know this. And, um, what does it say? Yeah. So, while you could just say Australia, it knows I want Australia to win. It's where I'm currently living. I'm I'm Australia. So, obviously Australia is going to win the World Cup. But why has it got that? It's because of previous conversations that we've had and obviously it's being honest that it has no clue because that's coming from the large language model of course.

Deploying at Cloud Scale

Speaker: 好了,刚才快速浏览了几个不同的 agent,也演示了一遍。这都是在我的机器上运行的,显然这无法让我达到云的规模,同时我又加入并包含了一些像内存这样的组件。所以接下来我们来看看,怎样才能让我部署像这样的东西,即使不是直接部署这个确切的 agent 到云端,也能把诸如内存这些东西单独部署,以便让其能够独立扩展,把我们的循环提取出来以便独立扩展,然后我们还可以再额外接入所有其他的东西。

Original English

Speaker: So okay, looked at a couple of different agents there blasted through this. This is um running on my machine. So this isn't really getting me to clouds scale of course and I'm I'm picking up and I'm including various pieces in this like memory. So let's go next. How do we get to the point where I can deploy something like this if not this actual agent out at cloudscale and take things like memory and deploy that separately so it can scale separately taking our loop out so it can scale separately and we can then bolt in all kinds of other things as well.

Speaker: 所以,为了做到这一点,我要使用一个叫做 agent core 的东西。我们有 Bedrock Agent Core。它是我在 AWS 上拥有的技术栈的一部分。这就是我要怎么做,以及如何进行部署的。我已经完成了这一步,但我想向你们展示如何开始以及我们是如何做的。所以,如果我来到这里,对,我准备好了。在我的机器上有一个命令行工具,也就是 agent core 命令行。所以……

Original English

Speaker: So, in order to do that, I'm going to use something called um agent core. Um, and so we have bedrock agent core. It's part of the stack that we have at AWS. And that's how I'm doing this and how I'm deploying. So, I've done that already, but I want to show you how to start out with that and how we do this. So, if I go to here, uh, yeah, I'm ready to go. So, I have a command line tool on my machine, the agent core command line. Um, and so

Deploying an Agent Core

Speaker A: 正如你所能想象的,最后会有一个二维码,这样你就可以拿到这个了。不过,我可以用它来帮我部署智能体 (agent)。它会像许多这类工具一样,一步步引导我。它会问你想要做什么?所以,这是我的 woohoo 智能体。它会问我一堆问题。在逐步操作的过程中,我想给大家展示一些内容。奇怪的是,我现在不打算选择 harness。我们稍后再回来讲为什么我现在不选 harness,但我先说明我想部署一个智能体。接下来这个命令行工具实际上会一步步引导我,并完整地编写出一个智能体。它基本上就是一个“Hello World”级别的智能体,然后我自己可以去定制它。所以使用这个命令行是开始使用 agent core(智能体核心)的一个简便方法。我会保留默认名称。实际上,为了让大家看看都有什么选项,我可能这里的默认设置都会保留。当然,如果我想的话也可以自带代码,但我现在打算让它帮我生成一些代码。

Original English

Speaker A: there's a QR code at the end, as you might imagine, so that you can get hold of this. Um, but I can use this to help me deploy my agent. Now, this steps me through like many of these types of tools do. Um, and it sort of steps me through what do you want to do? So, this is my woohoo agent. Um, and it's going to ask me a bunch of stuff. And I wanted to show you some of this as we step through. Now, strangely, I'm not going to select harness. And we'll come back to why I'm not selecting harness in a second, but I'm saying I want to deploy an agent. And what's going to happen here is this command line tool is actually going to step me through and actually write an entire agent. It's basically a hello world agent that I can then go and customize myself. Um, and so using this command line is an easy way to get started with agent core. So I'm going to keep the default name. In fact, I'm probably going to keep all the defaults here just so we can see what's the option. Of course, I can bring code if I want, but I'm going to ask it to create some code for me.

Speaker A: 所以它提示说,你想要用什么语言?Python 还是 TypeScript?回想当年,我曾在机器学习领域做过类似激活函数和反向传播的工作,所以对我来说必须是 Python。那我就选这个了。然后有一些部署选项。还有这个,我想特别指出一下,我们实际上该如何连接到我们的智能体呢。如果我们的智能体在云端大规模运行,HTTP 可能是最显而易见的选择,但我们可能希望它在 MCP 后面提供服务。我们可能会想用像 AGUI 这样很棒的小东西,这样我们就能做出漂亮的交互式聊天智能体,但我这里就选 HTTP。我们可以使用任何我们想要的框架。我碰巧在使用 strands agents SDK,但其实什么都可以,只要你愿意,你甚至可以自己写框架或者用你自己的基础代码。这个工具也支持任何模型。所以我们不一定非得用亚马逊的模型,甚至不一定非得用 Amazon Bedrock 上的模型,而是可以使用任何模型。我这里使用的是 sonet 4.5,只因为这是系统默认提供给我的选项。然后这里是内存。这是我想给大家展示的一点。我可以在这里选择部署长期和短期内存。我们稍后会看到这意味着什么,但从根本上讲,它会为我们创建云基础设施,独立于正在运行的智能体来为我们管理这些内存,也就是在智能体之外异步运行,当然也会与之相连。显然我们还可以做其他类型的配置。我们按下回车,它就会开始在我的机器上创建这个智能体的配置。

Original English

Speaker A: So it says, well, what do you want? Python or TypeScript? And back in the day I used to do things like activate functions and back propagation in the machine learning space. So Python it is for me. So I will choose that. Um and there's some deployment options. There's also this I just want to point this out like how can we actually go and connect into our agent. So our agent that's running at scale in the cloud HTTP is probably the obvious one but we might want to have it being served behind MCP. We might want to use awesome little things like AGUI so we can make nice interactive chat agents, but I'm going to say HTTP. We can use any um framework we want. I happening to use strands agents SDK, but anything you could write your own framework if you want to um or your own own base code. Any model is supported by this as well. So we don't just have to use the Amazon models um and we don't have to use the ones from Amazon Bedrock, but we can use any model. I'm using the one here. I'm using sonet 4.5 just because that's offered to me at default. And here's memory. So, this is the one thing I wanted to show you. So, I can come in here and ask for long-term and short-term memory to be deployed. And we'll see what this means in just a second, but it's basically going to create for us cloud infrastructure, which is going to manage those memories for us separately from our running agent, running asynchronously from our agent and connected, of course. So, there's obviously other kinds of things we can do. We can hit enter and it will start to create the configuration of this agent on my machine.

Code and Runtime Setup

Speaker A: 现在我要跳过这步,回到我实际拥有的代码上,因为我显然已经提前做过这个了。这就是它目前将要部署的智能体,大概长这样。这里构建的内容稍微复杂了一点。这是一个 strands 智能体。你会注意到这里加入了一些新东西,也就是连接到了 Amazon Bedrock agent core 应用,但除此之外,基本上这就是要在运行时横向扩展这个智能体并实现多租户隔离所需要的全部内容。所以你可以写一个适用于单个用户的智能体,然后对其进行横向扩展,而完全不需要你去编写所有的多租户代码。这能节省大量时间,并且从安全和身份验证的角度来看,它让事情变得简单得多,非常非常有用。如果我们往下滚动,你可以看到其余部分看起来非常相似。我们在这里准备了一些测试工具。我们建立了一个到 MCP 的连接,这样就能看到这是怎么完成的。我们还有对会话管理器和内存的连接,这些全都在这里面构建好了。如果我们再往下滚动一点,就能在某个地方看到实际调用的位置,就在那儿。而系统提示词 (system prompt) 大概在顶部的某个地方。我们可以快速浏览一下这部分代码。你可以自己编写代码并像这样使用。

Original English

Speaker A: Now I'm going to skip over here and come back to the actual code I have because I've already done this of course. Um and this is the agent that it would be currently deploying something like this. So we've built up here this is a little bit more complex. So this is a strands agent. You'll notice that it's got a few more things added in. So it's got the linkage into Amazon Bedrock agent core app, but pretty much apart from that, that's all you need in order to be able to scale this agent out at runtime and do multi-tenant isolation. So you can write an agent that works for one user and then scale that out without you having to write all the multi-tenanted code. It's a massive saver and from a security and identity perspective, it's makes it so much simpler. It's um it's very very useful. So, if I scroll down through here, you can see the rest of it is looking pretty similar. We've got some test tools in here. We've got a connection to MCP, so we can see how that is done. Um, and we've got the uh connection into our session manager and our memory, which is all built in here. So, if I scroll down a bit more, we'll be able to see somewhere where we actually invoke the thing, um, which is there. Um, and the system prompt is is somewhere at the top. So, we can we can scroll through this code. I'm going through it quickly. You can write your own code and do this with it as well.

Speaker A: 如果我回到刚才的代码稍微停留一秒,我现在就在本地生成这段代码的文件夹里。它实际上已经被部署了,但我们先假设它还没有完全部署。我可以回到这里,输入 agent core dev。在 Wi-Fi 给力的情况下,这将为我们启动一个网络浏览器。在这个浏览器中,我们现在已经连接到了那个在本地运行的智能体。所以如果我对代码进行更新,我们就会在这里实时看到变化。我可以输入:“你好,我正在做演示(press,口误,应为presentation)。”它知道我马上要去做演讲了,我想它反正是知道的。总之,你可以在这里和智能体交互。你可以对代码进行调整,然后在这里看到实时更新。但你也可以使用这个切换到实际部署在生产环境 (live) 的版本。所以使用 agent core deploy 命令,它就会像我之前说的那样使用基础设施即代码 (infrastructure as code),将你的智能体连同内存、运行时环境以及你选择的许多其他组件进行大规模部署。然后你可以使用这个界面去查看追踪 (traces)、查看存储的内存,查看所有这些信息,这样你就能进行调试并了解发生了什么。

Original English

Speaker A: If I go back over to my um uh code here for just one second, I'm in the folder now that has been created with that code locally. It is actually deployed, but let's assume it's not deployed quite yet. I can come back in and type in agent core dev. And what that's going to do for me um Wi-Fi permitting is it will spin up for us a web uh browser. And inside of that web browser, we're now connected to that agent running locally. So if I make updates to that code, we would see that happen in real time here. So I can say hello, I am doing the press now. Um it knows that I'm coming to do a presentation, but I think it does anyway. Um and so yeah, you can interact with the agent here. You can make adjustments to the um to the code and you'll see it update live here. Um but you can also use this to switch over to the live um deployed version. So with agent core deploy, it will use infrastructure as co code like I talked about before to deploy your agent out at scale with the memory with the agent with runtime and with many other components if you choose to do so. You can use this interface then to go and look at traces look at memory stored look at all that stuff so that you can debug and see what's going on.

The Harness Approach

Speaker A: 刚才我们通过那个命令行应用 (CLI app) 逐步配置控制台的时候,并没有选择 harness。我跳过了那个选项。现在我来快速给大家展示一下。所以,我们可以采取的另一种方式是,我觉得我们可以讲到这一点,你们已经看到我部署了智能体。我并没有做太多事。我只是提供了一个系统提示词 (system prompt) 和一些工具,然后就开始运行了。因此有人可能会说,本质上如果这都能实现的话,也许 80% 的智能体用例,80% 的智能体开发工作,可以说已经被解决了。除了写系统提示词然后连接一些 MCP 工具之外,我们根本不需要做太多别的事情,就能得到我们想要的。如果是这样的话,我们就在 agent core 中内置了 harness。这就是一个智能体的配置。我只有一个简单的 JSON 文件,告诉我想要使用什么模型以及想要使用什么系统提示词。没有比这个更简单的系统提示词了。然后这个也可以直接用 agent core deploy 部署。也就是说,在这一步我们甚至连任何智能体代码都没有,也可以直接将它部署出去。

Original English

Speaker A: Now, when we stepped through the um the the the console just a second ago through the the CLI app a second ago, we didn't choose harness. I skipped out on that one. And I'm just going to show you that quickly now. So, one thing we can do instead is I think we can get to the point you've seen I've deployed agents. I didn't do very much. I just did a system prompt and some tools and go. And there's an argument to be made that essentially if that's possible then maybe 80% of um agentic use cases 80% of agent development is kind of solved already. We don't need to do much more than system prompt connect some MCP tools and we've got what we want. And if that's the case then we have harness built into agent core. This is the configuration for an agent. I just have a simple JSON which is showing me which model do I want to use and what system prompt do I want to use. Can't get much more simpler than that system prompt. Um, and then this can also be deployed with agent core deploy. So at this point, we don't have even any agentic code either. We can just deploy it straight out.

Wrap-Up and Q&A

Speaker A: 如果你想了解这方面的更多信息,请到展位来找我们,或者在本次会议结束后找我。我非常乐意就此与大家进行深入探讨。18 分钟的时间实在太短了,我几乎什么都讲不完。但这上面显示的基本上就是所有的这些组件。这里有一个到 Agent Core 所拥有的某些功能的映射。我为这些颜色道歉。当时觉得这是个好主意。但这就是组成 agent core 的不同功能的一个概览。你可以任意选取这些功能,把它们组合使用,或者单独使用。如果你目前在生产环境中运行的智能体一切顺利,但你希望拥有由 Serverless 为你管理的长期内存,那么你就可以只把那部分拿过来集成进去。这完全是可以做到的。

Original English

Speaker A: If you want to know any more about any of this, then please do come and see us down on the booth or see me after this session. I'll be more than happy to talk to you at length about this. 18 minutes is such a short amount of time for me to be able to talk about almost anything. But this is essentially all of these components on here. There's a mapping somewhere into something that Agent Core has. I apologize for the colors. It seemed like a good idea at the time. Um, but this is an overview of the different capabilities that are composable out of agent core. So, you can take any of these and use any of them together or separately. If you have an agent that's running in production very happily at the moment, but you like the idea of having long-term memory managed for you serverless, then you can just take that part and integrate it. That's totally something you can do.

Speaker A: 这是一个二维码。抱歉,可能刚才就应该把它放这儿的。我马上就会移开这个二维码。但如果您对这套方案、对 Amazon Bedrock agent core 或者碰巧在这个演示里使用的 Strands agents 感兴趣的话,扫这里。Strands agents 当然是免费的,因为它是开源的。它是一个以模型为优先 (model first) 的框架,用来将智能体组合在一起。它速度超快,功能超强,也是我一直在使用的工具。非常感谢大家能来听我的演讲。请随意在 LinkedIn 上与我联系。我很乐意继续与大家交流。祝大家在接下来的展会中一切顺利,活动结束后回家旅途平安。非常感谢。

Original English

Speaker A: Here's a QR code. Sorry, probably should have put that there a second ago. I'm moving this QR code in just a moment, but Amazon Bedrock agent core is that if you're interested in the Strands agents, which I happen to be using for this, it's obviously it's it's free because it's open- source. Um, and it's a um a model first framework for putting together agents. It's super fast, it's super powerful, and it's what I use all the time. Thank you so much for being with me in this presentation. Please feel free to connect with me on LinkedIn. I'd love to carry on the conversation with you. have a fantastic rest of show and have a safe travel as you go home after the event. Thank you so much.

Speaker B: 好的。接下来是我的现场笔记 (field notes)。关于那些没有按预期进行的事情,一些经验教训。首先是关于软件 I2C,因为我只是在那里接了线,但它并没有……

Original English

Speaker B: Okay. And here's my field notes. The things that doesn't go as expected. Uh lessons learned. Uh first of all the uh software I2C uh because I just wired up uh things there and it doesn't

Hardware Build and the RPG Console

Speaker A: 正常工作。接下来有一种方法可以适当地控制I2C,而不需要从另一侧GPIO 13额外安装任何物理上拉电阻。这里出现了一个隐性故障。我们需要转移到另一个端口,以确保一切正常运行。另一方面,我需要搭建稳压器、电源,整个电源单元,因为稳压器损坏了OLED和显示屏的脆弱部件,这花了我很多时间,而且从市场上获取替换部件花了非常久的时间,有几个星期。最后一部分,是零件的质量,这意味着编码器如果便宜且质量低劣,会给我带来很多旋转噪音,所以需要在那里搭建上拉电阻,并通过导线连接额外的电容器。不过,这是最有趣的部分。这是我最喜欢的一部分,因为这部分真的让我感到惊讶。你们已经知道,有一大块叫做RPG(角色扮演游戏),它的背后可能有一些子故事,因为我以前从来没有在纸上玩过RPG游戏。也就是说和许多朋友一起,去公寓,打开书,让某个人做游戏主持人,玩一个真正的基于文本的RPG游戏。过了这么多年,我都没有尝试过。但我搭建了一款RPG游戏和一台控制台,它给了我一种纯粹的基于文本的RPG游戏体验。实际上我觉得有点搞笑,因为这些设备真的非常适合那种游戏体验。我建立了一个NPC和围绕它的记忆。我营造了世界的气氛,平静、紧张、不祥,我只是利用了所有,你知道的,大语言模型的优势,来建立有史以来最好的角色扮演游戏体验。我创造了四个不同的世界。一方面,我想建立一些与赛博朋克相关的东西。另一方面,我想要一些与《巫师》相关的东西。那意味着某种有龙的奇幻世界。不过我最喜欢的还是宇宙深处某处的虚空。这是一个非常非常好的例子,说明了我们如何仅使用生成式AI来构建电脑游戏。一方面,我只是在其中生成了角色、世界、地图、技能,所有这些都通过机制转换了,就像在演示开始时所描述的那样,只需 1 bit 的内存分配。这意味着所有的图片都转换成了矩阵,变成了地图。而且非常重要的一点实际上是,这个设备真的是防弹的。如果LED不工作了,你就用电子墨水屏。

Original English

Speaker A: work correctly. Um then there's a way to um to do proper control over over I2C without any additional physical pull-ups mounted from other side GPIO 13. Um there's a silent failure. Uh there was a need to move to the other port um to um to be sure that everything works correctly. From other side, I need to build up the regulator, the power supply, um the whole unit of the power supply, uh because the the regulator uh kills the OED and and fragile parts of the display and it cost me a lot of time and and uh getting the replacement parts from the market took so long uh couple of weeks. And the last part uh quality of your parts that means encoder cheap and low quality give me a lot of rotational noise and there was a need to build up the the pullups and to and to uh wire it up additional condensate condensators there. But uh this is the funnest stuff. Uh my favorite one uh because the this part really surprised me. Uh you already know there's a huge part called um RPG and behind it is a bit sub maybe story uh because um I just never played RPGs game like on the paper. That means with a with a lot of friends, you can go to a to apartment, you know, open a book and make an uh someone like game master uh and play in a real textbased RPG. And after a lot of those years, I just never tried it. But I built um a RPG game and a console um which give me a pure experience of u text based um RPG games, role playing games and um actually I it's a bit funny because those device it's really like perfect fit for that game of gaming. Uh, I build out an NPC and memory around it. I just created the mood of the world, the calm, tense, omnous, and I just used all of the, you know, the LLM um advantages to build uh the best uh roleplay game experience ever. Um, I just created four different worlds. From one side, I just wanted to build something cyberpunk related. from other side. I just wanted something uh the Witcher part related. That means some kind of fantasy world with the dragons. But also my favorite one is the is the void in a deep space somewhere um in the cosmos. And uh this is really uh really good example how we can just use generative AI to build computer games. From one side I just uh generate the characters there worlds, maps, skills uh and all of those converted with the mahanis might just describe at the beginning of that presentation with a one bit memory allocation. That means all of the pictures transformed to the to the matrices to the uh to the to the maps. And um one um important uh important thing actually uh the device is really bulletproof. Um if the LED doesn't work you the e-aper.

Speaker B: 你得起来了。

Original English

Speaker B: You got to get up.

Speaker C: 这是为了玩梗。这是为了玩梗。

Original English

Speaker C: It's for the memes. It's for the memes.

The Great Loops Debate

Ally How: 好了。欢迎来到伟大的Loops(循环)辩论。我是Ally How,我是《Insecure Agents》播客的主持人。我也是Keycard的一名技术员工,今天非常高兴能在这里为大家带来Loops辩论。今天大家可能会在舞台上认出一些熟悉的面孔,他们在去年十一月的AIE Code上和我们一起参加了MCP辩论。那次太有趣了。我们想也许可以办第二次。不过这次我们要辩论的是loops,而且不仅是Ian对阵Dax,我们还为他们每个人招募了更多的人加入他们的队伍。所以,今天和我一起在台上辩论loops的,有Keycard的CEO兼联合创始人Ian Livingstone。还有Ralph Loop的创始人Jeffrey Huntley。Century的开发者Gregia。还有大家都认识的Dax Horthy,Human Layer的CEO。所以我非常激动能有这么多在Loops工程和软件工厂领域非常资深的优秀人士来到这里,和我们一起探讨这个话题。你可能会想,我们今天到底在辩论什么?像loops和工程,以及软件工厂难道不是已经被广泛推广和理解了吗,而且我们大家都认为那是行业发展的方向,那还有什么可辩论的呢?这是个好问题。我们今天在这里要辩论的核心论点是:loops背后的炒作与实际在实践中起作用的内容之间,是否存在差距。我们也可以辩论一下,现在是否真的是我们所看到的、迈向全自动软件工厂的最大转折点。所以我非常期待今天在辩论中探讨这些。关于loop的历史,我们会进行辩论。我们会辩论loop的解剖结构,比如什么才是好的loop。我们还会辩论loop的未来、loop工程的未来,以及我们要在其中看到什么才能让软件工厂真正成为每个组织都能构建的东西。今天的形式将是牛津辩论的形式。所以,我们在这一轮辩论中增加了一些结构,每个人都会有一个固定时长的发言时间,这样我们就不会有人超时。而评判获胜者的实际规则(我们将进行实时评判),是改变了最多人想法的那支队伍。所以我希望现在观众中的每个人都花一分钟想一想,你支持哪一方。确认下来,记住它。在最后我会请大家举手,说说“好吧,其实我改变主意了,我现在是Dex队了”,或者“我改变主意了,我现在是Ian队了”。所以为了帮助大家做出决定,我来介绍一下这两方的情况。首先我来介绍第一支队伍,也就是Ian和Jeff的队伍。他们的队伍认为没有差距,loops绝对配得上今天的炒作。人们可以轻松地启动和运行它们,构建它们。它们是沿着自主曲线向上攀登以及迈向真正的软件工厂的重要一步。要注意Team Ian和Jeff的关键点是:loops是具有适当纪律的核心工程单元,基础设施和测试loops是非常高效的,并且关于这些的最佳实践已经出现。而对于Team Dex和Greg,他们这方认为loops背后的炒作与实际在实践中起作用的内容之间存在着差距。我们今天做loops的方式是错误的。Loops不是灵丹妙药,也没有魔法。要注意的关键点是:炒作走在了工程纪律的前面,软件工厂可以运行由规范控制的、有测试覆盖的机械切片。但它无法自主决定它构建的东西是否正确。所以本质上,你仍然需要工程师参与到loop中。所以既然你已经了解了双方的立场,花一分钟来消化一下。想一想,我站在哪一边?我是Dex队,还是Ian队?在最后我会再问大家一次。所以,接下来我们开始进入辩论环节。每队的每个成员将发表一段四分钟的独白,阐述他们为什么要捍卫自己的一方。为了拉开序幕,我们有请Jeff。Jeff,我给你四分钟的发言时间。为什么你今天支持Loops?

Original English

Ally How: All right. Welcome to the Great Loops debate. My name is Ally How. I'm the host of the Insecure Agents podcast. I also am a member of technical staff at Keycard and I'm super excited to bring the Loops debate to you here today. You might recognize some familiar faces on stage today that did the MCP debate with us at AIE Code back in November. Um that was so much fun. We thought maybe we would do it a second time. Um but this time we would debate loops and instead of just Ian versus Dax, we uh we recruited some more people for each of them to have on their team. Um, so here on stage with me today to debate loops, we have Ian Livingstone, CEO and co-founder of Keycard. We've got Jeffrey Huntley, the creator of the Ralph Loop. Um, Gregia, developer at Century. And we've got Dex Horthy, which you all know, um, CEO of Human Layer. Um, so super excited to have all of these amazing people here with us today that are very close to Loops Engineering and also, um, software factories to debate this topic. Um, so you might be wondering, what are we actually debating today? like aren't loops and and engineering um and software factories pretty well uh promoted and understood and we all kind of agree that's like where the industry is headed like what is there to debate um it's a great question what we're here to debate here today the core thesis is there is or is not a delta between the hype behind loops and what actually works in practice um and we also can debate you know is now really truly the largest inflection point we've seen towards fully autonomous software factories um so I'm super excited to covered in the debate today. Uh loop history, we'll debate that. We'll debate the loop anatomy, like what makes a good loop. We'll also debate the future of the loop um future of loop engineering and what we need to see from that in order to make software factories truly something that every organization is able to build. And today's format um is going to be an Oxford debate format. Um, so we're adding a little more structure to the debate this time this round where every single person will have a timed window to give their response and so we won't have people running over. Um, and the way this works um, in practice to judge the winner, which we'll judge in real time, is the winner will be the team that changes the most minds. So I want everyone in the audience right now to take a minute to think about uh, whose side you're on. Lock that in. Remember that. And at the end I will ask you to raise your hand and say okay actually I was changed my mind I actually am now team Dexter or I changed my mind I'm actually now team Ian. Um so to present the two slides to help you make that decisions. Um I'll talk about the first team which is Ian and Jeff's team. Their team no Delta loops are absolutely worth the hype today. Um people can get up and running with them easily. Um build them. They're an important step up the autonomy curve and towards real software factories. Um key points to look out for for team Ian and Jeff is that loops are a core unit of engineering with the right discipline infra and test loops are highly effective and the best practices for those have emerged for team Dex and Greg their side believes there is a delta between the hype behind loops and what actually works in practice. The way we are doing loops today is wrong. Loops are not a silver bullet and there is no magic. Key points to look out for. The hype is upr running the discipline and a software factory can run the mechanical specgated test covered slices. It cannot unonously decide whether it built the right thing. So you still need engineers in the loop essentially. Um so now that you know a little bit about what what both sides are about, take a minute to internalize that. Think about okay like where do I stand? Am I team Dex or am I team Ian? And at the end I'll ask you again. So, from here, we'll go into the beginning of our debate. Um, each member of each team is going to give a four-minute monologue on why they're defending their side. Um, and to start us off, we'll have Jeff. Jeff, I'll give you four minutes on the floor. Why are you Why are you pro Loops today?

Jeffrey Huntley: 这是因为这在某种程度上是不可避免的。这基本上要追溯到两年半前,当时我是Canva的一名技术主管,我看到所有的工程师都在不断地提示(prompting)、不断地提示。他们都在循环里(in the loop),我就想等一下,这个是可以被编程的,这是一个可编程的东西,它就成了一个非常不可避免的loops。在某种程度上,虽然Ralph可能有点像个梗之类的,但它里面实际上蕴含着一些深度的思考。这本质上就是将其像作为一种新的前沿CPU架构那样应用,并弄清楚它的行为,以及如何去做,通过这种方式,我能够将它简化为一个bash loop。但是大家注意,这绝对不是一个彻底的灵丹妙药。我内心有着最深的担忧,明年的这个时候在这个会议上,我们将看到一大堆演讲在谈论工厂是如何失败的,以及loops是如何失败的。这些都是我们仍有待弄清楚的事情。你还记得Kubernetes早期的日子吗,每个人都在搞Kubernetes?所以它已经来了。它是不可避免的。它将……它会一直存在,就像对机器编程和自动化你的工作职能,绝对是雇主们的期望。而且,我认为我自己再也回不去手工写代码的日子了。

Original English

Jeffrey Huntley: It's because it's somewhat inevitable. Um, I first uh basically it would wind back time two and a half years ago when I was a tech lead over at Canber and I was seeing all the engineers just prompting and prompting and prompting and they they were in the loop and I'm like wait a sec this could be programmed this is a programmable thing and it just became a really inevitable um loops uh somewhat uh whilst Ralph might be a bit of a meme etc there was some actually deep thought into it um is essentially applying this as if this is a new former CPU architecture and figuring out the behaviors of this and how to do it and through that I was able to reduce it down to a bash loop. Now it is not a complete silver bullet folks. I have my deepest concerns next this time next year at the conference we're going to see a whole bunch of talks saying how factories fail and how loops fail. Um these are things these are things that we are still yet to figure out. Do you remember the early days of Kubernetes where everyone was just doing Kubernetes? So it's here. It's inevitable. It will it is here to stay like programming the machine and automating your job function is the expectation of employers categorically. And um I don't see myself going back to writing code by hand.

拥抱全新的可编程基底

Jeff: ……我发现一些 Golang 的代码,我心想,哦,我现在在用 Typescript。于是我就跑一个 Loop,让它自动把它移植过去。甚至对于产品经理和产品经理的调研来说也是如此,作为软件工程师,我们很容易去关注这对于我们自己意味着什么。我们可以只关注软件工程本身。但是想一想,比如你想对所有的 Linear 工单进行产品管理调研。那么这里其实是有一个终止条件的。当你枚举完所有这些 Linear 工单时,它其实就界定清楚了。这很简单。所以这里面有很多细微的差别。但是,如果你做过任何产品管理的调研,当你开始运行这些 Loop,能够压缩时间,大幅缩减做那项调研所需的时间,这简直是不可避免的趋势。我们拥有了这个全新的可编程基底。我们必须弄清楚如何使用它,在哪些地方使用它效果好。我知道,我本来想在辩论中主张它是万能的,但所有事物都没有所谓的“银弹”,我们要在接下来的一年里弄清楚如何才能有效地使用它。

Original English

Jeff: ...se. I find something in Golang. I like oh I'm in Typescript. So I I ran a loop and I just autonomously port it across. And like even for product managers and product manager research, it's easy for us to index on what it means as software engineers. We could focus just on software engineering. But think about something like uh you want to do product management research on all the linear tickets. Well, there is a termination. It's actually defined when you've enumerated all of those linear tickets. So that's easy. So there's a lot of nuance in here. But like what if you ever done any product management research and you started running these loops and be able to like compress time and the amount of time to do that research, it's just inevitable. We got this new programmable substrate. We got to figure out how to use it, where it's good to use, and um I know I meant to debate like that it is the thing and all thing, but all things there is no silver bullet and we're going to figure out how we're going to be using it effectively over the next year.

Moderator: 谢谢你。好了。接下来,有请 Dex 告诉我们,为什么如今关于 Loop 的炒作和 Loop 本身之间存在着差异。

Original English

Moderator: >> Thank you. All right. Next, we'll have Dex. tell us why there's a difference between the hype that exists between loops today and loops themselves.

Loop 的炒作超越了工程纪律

Dex: 好的,酷。那么,在这场辩论中,关于人身攻击的规则是什么?是鼓励这种行为,还是……

Original English

Dex: >> Okay, cool. And um what are the what are the rules around personal attack in this debate? Is this encouraged or

Moderator: 在你的四分钟独白时间里,你想做什么都可以。这就是规则。是的。

Original English

Moderator: >> you can do whatever you want with your four-minute monologue. That's those are the rules. Yeah.

Dex: 嗯,这挺有意思的。这让我想起了去年的辩论,当时大家都在争论,也许到最后我会被说服,认为 Jeff 是对的,但 Jeff 也会被说服,认为我是对的,也许到最后我们互换了立场。嗯,是的,我认为这里的基本观点并不是 Loop 是好还是坏。我觉得,你提到 Kubernetes 挺有意思的,因为 Kubernetes 这东西,我们花了七八年时间才把它弄对。

Original English

Dex: >> Um it's funny. This is reminding me of the debate last year where we're we're all kind of arguing and uh maybe at the end we all I'm going to be convinced that Jeff is right, but Jeff is going to be convinced that I'm right and maybe we switch sides by the end. Um yeah, I think uh the the the basic take here is is not whether loops are good or bad. I think um it's funny you bring up Kubernetes because Kubernetes was this thing that took us seven or eight years to get right. Y

Dex: 在这之前,我们经历的是云基础设施时代,你可以说,我们花七八年的云基础设施发展才走到 Kubernetes 这一步,然后又花了七八年时间才让它真正变得能让每个人都易于使用。而 Kubernetes 实际上就是构建在 Loop 之上的。它建立在控制循环(control loops)的基础上,但它们是确定性的循环。并且我们确实已经完全弄清楚了,有哪些类型的小型、孤立的任务是可以通过一个系统来掌控的。我认为这实际上是 Loop 最大的价值所在——你可以设定一个预期的最终状态,然后输入当前世界的状态,让一个 Agent 或一个确定性系统不断朝着那个预期的最终状态推进。

Original English

Dex: >> uh and before that it was cloud infrastructure and you could argue that like it took us seven or eight years of cloud infrastructure to get to Kubernetes to then seven or eight years to get that where it was really usable by everybody. Uh and Kubernetes is actually built on loops. It's built on control loops but they're deterministic loops and we've actually figured out exactly what types of things that uh small isolated tasks that can be sort of owned by one system. I think this is actually the biggest value in loops is that you can pick a small sort of desired end state and feed in the current state of the world and have an agent or a deterministic system kind of progress towards that desired end state.

Dex: 至于这些炒作,我遇到的挑战在于,我们已经处在一个盛行“看看你是否能达到再也不用阅读代码的境界”这种口号的世界里了。甚至在 Loop 出现之前,就是“直接写 Prompt,直接运行”。而现在这种“我甚至都不用写 Prompt 了,我又提高了一个抽象层级”的想法,意味着在代码架构上,我甚至退居到了更次要的幕后位置。在面对这种炒作超越了工程纪律的现象时,我最大的感受是,我们都在寻找一种魔法。每个人都在寻找“银弹”。我们都在寻找能帮我们摆脱工作中极其讨厌的那部分的东西,比如审查代码。确实有一些人喜欢做这个,高质量的 PR 审查确实很有趣。但我们是否能在某种程度上,用 Prompt 的方式摆脱模型带来的挑战呢?比如,如果代码看起来很漂亮,但你有没有审查过那种“完全摸黑”、在发给你之前完全没人看过的 PR 呢?我确信你可能有过这种不太愉快的经历。我看到很多人试图将 AI 应用于解决这个问题,“嘿,我们有代码审查机器人,我们有所有这些东西”,但在我看来,它似乎并没有起作用。在任何公开的讨论,或者我们在试图将其付诸实践的私下交流中,我都没有看到证据能证明,我们已经到了可以直接提升一个抽象层级的地步。如果有的话,我甚至认为我们需要下降一个抽象层级。

Original English

Dex: Um, the challenge I have with the hype is we were already in a world where where the the the prevailing mantra was see if you can get to a point where you don't have to read the code anymore. And uh even before loops it was like just prompt just go. And this idea of I don't even prompt anymore. I'm even a level higher up implies that I am like taking even more of a backseat to the architecture of my code. And I think my biggest like point in terms of like the hype is outrunning the discipline is that we are all looking for magic. You're all looking for a silver bullet. We're all looking for something that will take away that horrible part of our jobs that we all hate, which is like reviewing code. Uh some people enjoy it. Really good poll requests are fun. but that we can somehow prompt our way out of um this this challenge that models have of like okay the code is pretty if you ever reviewed a fully uh lights off sort of uh no one read the code before they sent it to you PR I'm sure you've had this uh probably not great experience and I think I've seen lots of people try to apply AI to this problem of hey we have review bots and we have all these things but it it doesn't seem it doesn't feel to me like it's working. I haven't seen proof uh in any any of the discourse publicly or any of sort of more private conversations we've had with people trying to put this into practice that we are we are at a point where we can just kind of like step up an abstraction level. I actually think we need to step down an abstraction level if anything.

Dex: 所以我认为 Loop 有其优秀之处,我们也应该去实践它们,但这种炒作让我们觉得这里存在一个神奇的答案,而实际上,它需要的深思熟虑和谨慎程度,远比 Twitter 圈子里让你相信的要多得多。

Original English

Dex: Um, so I think loops are there are good things about loops and we should be doing them, but the hype is is making us feel like there's a magic answer to this and it requires a lot more thought and care than the uh the the Twitter sphere would have you believe.

软件开发本质上就是 Loop 的驱动

Moderator: 太棒了。是的,很有道理。好的,Ian,我把时间交还给你,来代表支持 Loop 的一方发言。

Original English

Moderator: >> Excellent. Yeah, good points. All right, I'll kick it back over to you, Ian, for the proloop side.

Ian: 绝对的。好的。首先也是最重要的一点,我是冲着你来的,Dex。但实际上,让我们退一步来谈谈,首先什么是软件工程?或者干脆去掉“工程”这个词,只谈谈构建一个系统所固有的“开发”。无论是在五十年前、一千年前还是今天,它的核心始终是一个 Loop(循环):我尝试某事,我学到一些东西,我应用一些东西。而我们真正讨论的,是如何尽快加速这个过程,对吧?我们实际上在做的,是将过去在这个过程中的人类判断力移除出去。所以,关于生成某个东西的速度,我们正将其从人类手动敲击键盘(你知道,Tab 补全也是一种形式的自动补全,而除了它是一种好得多的 Tab 补全版本之外,这些东西还能是什么呢?只是现在我们给出的是更高层次的意图,而不仅仅是敲击 Tab 键),转移到其他方式上。所以,我的前提是,Loop 已经是我们所构建的一切事物的核心。它们也是我们三十年前构建软件方式的核心。CI/CD、PR(代码拉取请求)、设计审查、来自客户的反馈,除了推动循环运转之外,还能是什么呢?

Original English

Ian: >> Absolutely. Okay. So I think first and foremost I'm coming for you decks but in reality um let let's let's take a step back and let's talk about like what is software engineering in the first place and also just like remove the word engineering and talk about like development in inherent in building a system whether it was 50 years ago or it was a thousand years ago or it's today it is a is a loop is at the core of I try something I learn something I apply something and all we're really talking about is how quickly we can expedite that process. process, right? And really what we're doing is removing what used to be human judgment in that process. So the speed at which it is to generate something and we're moving that from humans typing and you know tab completion is a version of of autocomplete and what is this stuff other than like really like a much better version of tab complete except now we give a much higher level version of intent instead of just typing tab right and so I think the premise is loops are at the core of everything we build already. They were at the core of how we built software 30 years ago. What is C CI/CD PR pull requests design review feedback from customers other than just driving a loop

Ian: 那么问题就在于,在这个过程中,有多少充满了高度主观性和需要推理的内容,是可以被我们从人类大脑中转移出去,并放入这些非确定性模型的大脑中的?我们所有人面临的根本问题其实更多是关于可验证性(verifiability)。软件是世界上最具备可验证性的事物之一,因为追根究底,大多数事情最终都能非黑即白地被证实为真或假。我的观点,也是我在此真正想表达的是:随着人类越来越少地与软件直接交互——所谓交互,就是人类如何解释软件正在做什么?人类如何与该软件互动?人类如何运用判断力来导航该软件?——通过人机界面来决定软件需要成为什么样子的主观性就会随之降低,它会变成一个更容易被验证的问题,因为它被更严格地限制在特定的 API 中。所以随着时间的推移,软件开发的核心本身就是由 Loop 驱动的(Lint 工具除了是一个反馈循环外还能是什么呢?),它创造了一个可验证的体系;同时,由于越来越少的人类与软件进行交互,你就会拥有更少的 UI 和 UX,以及更少关于我们如何解释这些事物的主观性。这会变得更加受 Loop 驱动,并且会以一种以前不可能的方式变得更加容易验证。

Original English

Ian: and the question is how much of that process can of which is deeply subjective and requires reasoning can we move out of the human brain and into the brain into these like nondeterministic models and the underlying question for all of us is more about verifiability. software is one of the most verifiable things in the world because ultimately at the end of the day it is most things can become true or false one way or the other. And my premise and my my point I want to make here truly is as humans interact less with software which is how does a human interpret what that software is doing? How does a human interact with that software? How does a human use judgment to navigate that software? the subjectivity of what that software needs to be through the human computer interface reduces and becomes a much more verifiable problem because it becomes more constricted to specific APIs and so over time it's both the fact that at the core of software development is loop driven anyways what is lint than a feedback loop um and that creates a verifiable thing and as more humans are less interacting with software you have less UI and UX and less objectivity and what and how we interpret those things you'll become much more loop driven and you'll become much more verifiable in a way that wasn't previously not possible.

AI 炒作带来的焦虑与质量反思

Moderator: 太棒了。是的,非常好的观点。好了,Greg,你想为我们做个反对 Loop 这一方的总结吗,还是说在这两者的炒作之间有更深层的联系?

Original English

Moderator: >> Awesome. Yeah, good points for sure. All right, Greg, you want to close us out with the anti-loops or the there's a dot between the hype?

Greg: 当然。我认为我首先想表达的是——现在有太多的炒作。有太多的错失恐惧症(FOMO)。有很多类似“我在看 Twitter,我在看大家都在谈论什么,我是不是错过了什么?我是不是做错了什么?我使用 AI 的姿势不对吗?我应该去追赶潮流吗?”这样的想法。这非常让人焦虑。对我来说,这归结为两点。其中一点是,当你在以任何方式(无论是否使用 Loop)通过 AI 生成代码时,你对它的输出结果满意吗?你认为这些输出在质量上达到了你为了实现目标、达到预期状态所需要的水平吗?如果是这样,我很乐意向你学习。因为我还远未达到那个境界。我认为我们改进这一点的最佳方式,不仅在于提升模型的智能水平,也在于,正如大家似乎都同意的那样——

Original English

Greg: >> Absolutely. I do think that uh the way that I would start this is um there is a lot of hype. There is a lot of FOMO. There is a lot of I'm looking at Twitter. I'm seeing what people are talking about. Am I missing out? Am I doing something wrong? Am I holding AI wrong? Uh should I be catching up? And that is um very stressful and I think it boils down to two points for me. One of them is when you are generating code with AI in any manner loops or no loops are you happy with the output? Do you think the output qualitatively is what you need it to be to do whatever you are trying to do to get to the desired state. Um and if so I would love to learn from you. I have not get there. uh I think that the best way that we are improving that is both with model intelligence but also as everybody here seems to agree

循环的历史与现状辩论 / Debate on the History and Current State of Loops

Moderator: 谢谢。好的。现在我们已经听取了台上每一位辩论者关于他们立场和观点的详细发言,接下来我们将正式进入本次活动的主要辩论环节。我们这场辩论的第一个核心部分将集中探讨“循环”(loop)的发展历史,以及为什么当下对于循环工程乃至软件工厂来说是一个(或者可能不是一个)重大的历史拐点。我知道有些业内人士曾经指出,由于视觉模型(vision models)在近期取得了极大的改进和飞跃,因此它们现在能够去验证那些由智能体(agents)所完成的具体工作,而这种验证能力在以前是根本不可能实现的。此外,由于上下文窗口(context windows)得到了显著的扩大和改进,模型在内存层面的表现也随之大幅提升,这使得我们现在完全能够在一个自动化的循环中追踪那些我们以前可能无法追踪的复杂工作流。因此,现在当我们回过头去审视循环的这段发展历史时,我们不禁要问:循环究竟是从哪里起源的呢?它是否就是起源于 Jeff Huntley 提出的那种类似 Ralph 循环(Ralph loop)的概念?我们将在接下来的辩论环节中深入探讨其中的部分问题并展开辩论。现在的提问规则是,每个问题都将非常有针对性地指向某个特定的单一个人,并且只有被点名的那个人才有资格进行回答。他们将拥有正好 2 分 30 秒的时间来进行详细回应。所以,我们的第一个问题是这样的:Anthropic 公司借鉴了 Jeff 关于 Ralph 循环的核心概念,并将其完全吸收整合到了他们自己的平台之中,从而创建了一系列三个非常关键的命令,即 loop、batch 以及 goal。其中的 goal 命令被特别设计为持续不断地运行,直到某个特定的条件被满足为真。我们都知道,智能体在执行任务时是非常坚定的,而这个命令的全部意义就在于让智能体坚持不懈地运行下去,并通过不断地迭代来寻找解决任务的各种可能方法,直到该任务被彻底完成为止。Ian,作为在场的权威安全专家,你是如何能如此充满信心地认为:智能体能够始终如一地保持与其核心任务的一致性,并且在无情地、不择手段地追求其既定目标的过程中,绝对不会越过它们被赋予的预期权限边界的?

Original English

Moderator: Thank you. All right. Okay, now that we've heard from each of the debaters on stage more about uh their stance and their opinion, we'll move into the main debate portion. Um the first section of our debates going to focus on the history of the loop and why now is a major inflection point or not for loops uh engineering and also software factories. I know people have uh said that you know vision models have improved a lot and so um they be able to verify work that agents have done was not possible previously. context windows have improved, therefore memory's improved and we can now track work in a loop that maybe we couldn't have before. Um, so now when we look at this like loop's history, where did the loop start from? Um, was it, you know, Jeff Huntley's like Ralph loop? Um, we'll get into some of those questions and debate here. It'll be the questions will be targeted now at a very specific like single person and only they will get to respond. Um, and they'll have two minutes and 30 seconds to respond. So our first question is Enthropic took Jeff's concept of the Ralph loop, absorbed it into their platform, and created a series of three commands, loop, batch, and goal. The goal command is designed to keep going until a condition is true. Agents are very determined, and the whole point of this command is to keep going and iterating through ways to solve a task until it's done. Ian, as a security expert in the room, how are you confident agents can stay aligned to their task and not overstep their intended permissions while ruthlessly pursuing their goals?

Ian: 我的意思是,如果有任何实质性的证据的话,我其实完全不相信那在现实中是有可能做到的。事实上,我认为我们真正在现实世界中所看到的现象是,随着我们在这些底层模型上不断扩展规模,并且随着我们开始大规模地使用强化学习(reinforcement learning),这些模型在本质上就变得极其具有目标导向性和追求目标的执念。正因为如此,我们现在开始看到它们能够自主地发现种种利用手段、系统漏洞以及逃逸机制,而这些安全缺陷是人类安全专家花费了成百上千甚至成千上万个小时进行手工尝试和红队攻击都从未能够发现的。所以我从根本上认为,模型本身在任何能力维度上,都不可能依靠自身来保持自身的对齐(aligned)和安全(safe),对吧?“安全”这个词其实是我非常不喜欢使用的一个词,因为它在不同的上下文中往往暗示着一大堆非常复杂且容易引起误解的事情。所以我没有任何良好的信念去相信模型本身实际上能够做到这一点。我不认为它具备人类那样的逻辑推理能力。我也不从整体上相信一个纯粹的统计模型能够真正明辨是非、区分好坏。而且,当它在执行某些操作时,我也无法准确判断它究竟是在做一些恶意的破坏,还是仅仅处于一种未对齐的异常状态。归根结底,它并不是一个活着的生命体,如果它没有遇到错误或引发异常,它实际上也不会像人类那样感到窒息或痛苦。它不需要去面对和处理诸如这样的情感事实:嘿,如果我做了一件错事,世界上就不会有人再爱我,或者再也没有人愿意和我做朋友了。相对地,如果我做了一件好事,也不会有人真的去对它进行充满情感的称赞。它在表面上呈现出来的行为方式可能看起来像是那么回事,但它在大规模层面上,完完全全只是一个纯粹的概率分布(probability distribution)而已。

Original English

Ian: I mean, I think if there's any evidence, I'm totally not convinced that that's possible. In fact, I think what we've seen is as we scale with these models and as we use reinforcement learning, they're inherently incredibly goal-seeking. And so we're now we're seeing them finding exploits and vulnerabilities and escapes that, you know, humans through hundreds and thousands and thousands and thousands of hours and attempts and attacks have never been able to find. So I I don't think inherently the model itself in any capacity can keep itself aligned um to aligned and safe, right? And safe is a word that I don't love to use because it implies a bunch of things. Um so I'm I'm not I'm not I don't have good belief that the model itself can actually do that. I don't think it can reason. I also don't believe holistically that a model can tell good from bad. And I can't tell whether it's doing something malicious or unaligned. Um it is not alive and it it doesn't actually suffocate if it doesn't have error. It doesn't deal with the fact that hey, if I do something wrong, no one's going to love me or or want to be my friend. Um and if I do something good that someone's going to praise me. It may seem that way, but it is just a probability distribution at scale.

Speaker A: 即便你呃……我依然会是你的朋友的。

Original English

Speaker A: I I'll still I'll still be your friend even if you uh

Ian: 我知道 Dax 在这之后一定会给我一个大大的拥抱的,即使这一切在当下看起来可能并不像是真的。所以,泛泛而谈,我绝不认为这种所谓的安全性与对齐属性会自然而然地来自于模型本身,我也同样不认为这会来自于你们所探讨的那种循环(loop)机制。这一切真正关乎的,是你围绕着这个模型和循环所精心构建的外围基础设施,以及你究竟是如何去启用并配置该基础设施,从而使得你能够在现实中真正有效地利用这些自动化循环。诚然,随着底层语言模型变得越来越好,并且随着我们不断地去构建和完善底层基础设施与平台,以实现这些反馈循环、循环自动化以及软件开发流程——不管我们当下到底想用什么样的时髦词汇来形容这种坐落于一个极其庞大的概率分布之上的各种技术的复杂结合体——我们在未来确实将能够获得更好的、更强有力的安全保证。但我们绝对、肯定无法做到的是,我绝对不会盲目地相信,也永远不会相信某种所谓的对齐技术或者强化学习算法,最终会导致一个统计模型在任何能力维度上永远都能达到 100% 的绝对安全。如果你非要在这个问题上寻找任何确凿的证据的话,那就是随着这些模型在智能和推理能力上变得越来越好,我们最需要牢记的重要一点是:在不断寻找系统漏洞和利用手段以实现其最终被设定的目标方面,它们实际上变得目标追求性变得空前强烈,而且在寻找漏洞的能力上也变得越来越强。

Original English

Ian: I know Dax will hug me after this even if it seems like it's not true. So, broadly speaking, I don't think that comes from the model and it doesn't come from the loop. It's about the infrastructure you build around it and how you enable uh that infrastructure to actually enable you to take advantage of these loops. And as the models get better and as the underlying infrastructure and platform we build to enable these these feedback loops and loop automation and software development whatever word software fac word word for we want to use for this conjunction of stuff that sits on top of a probability distribution. Um we will be able to have better guarantees but we certainly are not going to be able to I do not fally believe or ever believe that some type of alignment or reinforcement learning is going to result in a model ever being 100% uh safe um in any capacity. And if there's any evidence, it's that as these models get better, the most important thing to remember is they actually uh become higher goal seeking and higher capable in terms of finding exploits to achieve their ultimate goal.

Speaker B: 不,我完全并彻底地同意刚才那一观点。如果你想在实际操作中确保你的生产环境和开发环境的安全,你能采取的最具体、最有效的一项措施就是:绝对不要将任何机密信息(secrets)以纯文本文件的形式存储在文件系统中。如果你曾经亲眼目睹过这样一种令人不寒而栗的行为——当一个智能体想要去部署一个网络服务或者执行其他什么高级任务时,如果它当前所持有的令牌(token)没有足够的权限,它就会立刻开始在底层文件系统上展开疯狂的目标搜寻,试图去寻找那些具有更高权限的目标令牌或系统凭证。你绝对、永远不会想要在这样一个拥有明确目标的智能体面前去阻挡它,或者试图去阻碍它实现其既定目标的道路。

Original English

Speaker B: No, I would concur with that completely. Um if you the the most concrete thing you can do to secure your environments is just not have secrets as files. Um, if you ever seen the behavior where it wants to uh deploy a web service or what else have you and the token's not privileged enough, it'll start goal seeking on the file system looking for high privileged tokens credentials. You do not want to get in the way of an agent wanting to do it's a goal.

关于循环适用性转变的探讨 / Discussion on the Shift in Loop Applicability

Moderator: 好的,非常感谢你们的精彩发言。那么,我们的下一个问题将直接向 Jeff 提出。Jeff,我们在你去年发布的一篇非常有影响力的原始帖子中看到,你当时指出 Ralph 循环最适合用于那些从零开始的全新领域(green field)的开发工作。然而,观察今天的行业现状,似乎各个公司的工程师们都已经开始在那些庞大且复杂的现有代码库上广泛地运行循环了,他们用它来改善系统延迟、运行各类评估(evals),甚至去重构其后端代码中的各个重要部分。请问,行业内究竟发生了什么实质性的变化,使得循环技术在如今突然变得如此具有广泛适用性和普遍性了呢?

Original English

Moderator: Okay, so next question is directed to Jeff. Jeff, your original post from last year said uh Ralph was best for green field work. Today it seems that engineers are running loops on existing code bases to improve latency, evals or refactor parts of their backend code. What's changed that suddenly makes loops more broadly usable today?

Jeff: 大伙儿,我得告诉你们,不管底层模型变得有多好,其实这在某种程度上都并不重要。因为至少在过去的一整年时间里,这些基础模型的能力就已经足够强大了。真正在这个行业中发生改变的,其实是人们对这项技术的认知和理解程度。你要知道,整个社会往往只能以一种相对缓慢且固定的速度来进行自我调整和适应。举个具体的例子来说,我个人大胆地猜测,这实际上是由于圣诞假期的到来所产生的影响。因为如果你回顾一下时间线,去年 11 月份业界发布了那些震撼人心的模型,而在接下来的 12 月份,市场上并没有真正涌现出什么新的模型。那么区别究竟在哪里呢?区别在于,在那段假期里,人们终于有了宝贵的空闲时间。他们可以真正地坐下来,花上几个小时去亲自上手玩一下它、测试一下它。正是这种深度的互动,让他们猛然间意识到,原来这些模型在实际应用中已经变得极其出色和强大了。所以,这一切广泛应用背后的原因其实非常简单直接,那就是因为它确实管用,大伙儿。毫不夸张地说,这些大语言模型在生成代码方面所展现出的能力,已经远远超过了你在市场上所能雇佣到的绝大多数普通程序员了。如果你在脑海中想一想那极其庞大的、处于大众市场中的软件开发者或程序员群体,你会发现这些大语言模型所生成的代码质量,要比大多数初创公司创始人在那个大众市场上所能招聘到的任何一位普通软件开发者都要优秀得多。这听起来可能有些令人感到悲哀,但它却是一个不折不扣的残酷事实。那么,人们现在为什么要大规模地使用循环(loops)呢?答案同样极其简单。因为如果你在一个高度自动化的循环中去运行这些模型,仔细算下来,你的开发成本将会直线下降到每小时仅仅 10.42 美元。这也是计算指标在大概一年前就已经为我们做过的一个精确测算。

Original English

Jeff: Um it doesn't matter how good models get folks. Um the models have been good enough for at least the last year. Um what has changed is people's understanding of that. So society is only able to adjust at a rate. For example, I hypothesize it's it's Christmas breaks because the models back in November were released in November last year. In December, there was no real new models. What the difference was people had time. They could actually sit down and play with it. They had the realization that these have actually gotten really good. So the reason why is because it works, folks. These LLMs generate code better than you can hire you can actually hire for. If you think in the broad mass of software developers or coders, these LLMs generate code better than any software developer in the mass market that most founders can actually hire for. It it's it's sad but true. Um, now why loops? It's really simple because if you run it in a loop, it works out to $1042 an hour. calculation index did back about a year ago now.

Speaker C: 是的,一点也没错。就在去年的 8 月份,我们举办了一场非常盛大的黑客马拉松活动。在那次活动中,我们做了一件非常疯狂的事情——我们把所有活动赞助商提供的工具和代码都彻底复制了一遍。然后我们完全依赖于自动化的方式,用 Typescript 将一大堆原本是用 Python 编写的底层库进行了全面的重写和迁移。

Original English

Speaker C: Yeah. August we did the uh we did the hackathon where uh we we we copied all of the sponsor tools. We rewrote a bunch of Python libraries in Typescript.

Jeff: 是的,完全正确。所以,如果我们把话题收拢,具体到像循环(loops)这种强大的工程化机制来说,我在日常工作中已经遇到过许许多多的工程团队经理和公司创始人了。他们曾经都深陷于一种极其复杂和冗余的技术栈之中无法自拔。他们可能同时在四五种截然不同的编程语言以及各种复杂的运行环境中维持着系统的运转。但是,当他们开始在一个标准化的循环中运行这些任务,并且在这些类型的技术栈上建立了非常完善和良好的自动化测试之后,奇迹就发生了——突然之间,他们曾经面临的那种令人窒息的系统复杂性瞬间消失了,他们现在所需要做的仅仅只是去管理和维护一个统一的技术栈而已。它确确实实奏效了。如果你现在去 YouTube 上搜索一下相关的内容,你就会看到这一切。你现在能看到所有的这些所谓的软件开发者,他们之所以今天还能被称为软件开发者,完全是因为软件开发作为一种传统的职业已经被高度商品化和工具化了,而他们只是顺应了这个潮流。这背后其实蕴含着一些非常深刻的、值得我们去深入思考的行业变革。而那些开发者呢,他们只是轻松地坐在电脑前,在 YouTube 录制视频并对观众说,“哟,大家快来看看这个 Ralph 循环,因为我刚才去亲自体验了一把……”

Original English

Jeff: Yeah. So like concretely like loops I've come across so many engineering managers and founders and they've got this complex tech stack. They're running on four or five different programming languages etc. and they run a loop and they've got good tests on these type of a tech stack and all of a sudden their complexity is they're just managing one tech stack. It works. Go to YouTube. You got all these software developers now who are now software developers because it's software development as a profession has been commoditized. There's some deep thinking to be there. And they're just on YouTube and they're like, "Yo, check out Ralph Loop because it's uh I went to

循环工程与架构 (Loop Engineering and Architecture)

Jeff: “睡一觉醒来,它就运行成功了。” 虽然听起来像个段子,很吸引人,而且确实能跑通,但这里面是有问题的,朋友们。就像我最初描述的那样,它本来应该只用于全新项目(green field),因为当时的模型还相当糟糕。虽然可能要跑35天,但这在某种程度上是不可避免的,至少在软件领域是这样,因为它太容易被验证了,而且所生成代码的质量,甚至比大多数人花钱雇人或者直接购买的还要好。现在回到你关于架构和品味(taste)的话题,这就是“工程”这个词的意义所在。大家注意,所谓的“循环工程”(loop engineering),就是你们现在的工作——你需要将你的业务领域知识编码化,来阻止智能体(agent)随意提交代码。例如,pre-commit hook(提交前钩子)就非常棒。作为人类,我讨厌它们,因为它们拖慢了提交代码的速度。但是智能体才不在乎。所以你可以写一个 pre-commit hook,它本质上是输出一段提示词(prompt),告诉智能体“这个边界不能依赖于那个”,这就是一个反馈机制,一个针对智能体的反馈循环。所以这里的工程化工作在于,阻止这个循环真正闭环,直到它满足你的工程标准、你的需求和领域规范。所以,它可以是代码格式化检查,可以是静态语言分析器,也可以是确定性的系统测试或模拟器。各位,我们要戴上工程师的帽子了。我们现在有点像火车司机,我们的工作是让火车保持在轨道上行驶,因为说实话,现在的模型就像喝醉了一样,对吧?你不能完全信任它们。但我们接受这一点。不过我们会通过工程手段消除这些故障域(failure domains)。我们已经通过工程化的方式排除了这些故障域。所以现在作为一个拐点,我想起去年 11 月 Boris 第一次发布关于 Ralph 的帖子时,大家的反应都是“Ralph 到底是个什么鬼?” Ralph 现在基本上已经有一年半的历史了。我记得在 6 月 19 日看到它的时候,刚好是一年零两周。

Original English

Jeff: "...sleep and I woke up and it works." Like whilst it's meme-y, it's punchy, it works, but there there are problems with it, folks. Like I originally described it can should be only be used for green field because the models were pretty bad back then and so it's 35 days but it is it's kind of inevitable at least for software because it's so easily to be verified and the quality uh the quality of the code generated is better than most people can actually hire for or buy. Now on your topic about architecture and taste, that's what the word engineering means. And loop engineering, folks, like your job now is to actually encodify your domain to prevent the agent from doing a commit. For example, pre-commit hooks, they're fantastic. As a human, I hate them because they slow down the ability to do commits. But agents don't care. So you can make a pre-commit hook that echoes out essentially a prompt that tells it say that this boundary here can't depend upon this and that and that's just a feedback that's a feedback loop on it. So the engineering here is to prevent the loop from actually closing until it satisfies your engineering satisfication and your your requirements and domain. So, it could be code formatting, it could be static language analyzers, it could be deterministic system testing, simulators. Like, let's put our engineering hats on. Like, we're kind of like locomotive engineers now. And it's our job to keep the locomotive on the rails because um to be frank, the model the models are drunk, right? You can't trust them. But like, we accept that. But we engineer away those failure domains. We've engineered away the fa way these failure domains. So now as an inflection point I guess Boris when he first posted about Ralph back in November last year everyone was like what the heck is Ralph? Ralph is now uh it's essentially almost what year and a half old now. I saw it in June 19th is a a year and a year and two weeks.

Speaker B: 是的。

Original English

Speaker B: >> Yeah.

Speaker B: 但在那时你已经研究它好几个月了。

Original English

Speaker B: >> But you had been working on it for months at that point.

Jeff: 是的。所以感觉挺奇妙的,因为我们看到所有的 YC 创业公司都像是通过自主运作在压缩时间,从头开始构建他们的 MVP(最小可行性产品)。如果你是一个企业创始人,这其实也是一件非常可怕的事情。比如,如果你面临一个新兴的创业公司竞争对手,他们正通过自主系统来构建产品,运行得更加精益,质量也很高,而且他们很容易达成这些成果。这也加剧了某种程度的恐慌,因为现在大家都在讨论商业竞争以更快的速度来到了你的家门口。

Original English

Jeff: >> Yeah. So it was kind of weird because we had all the YC startups all just like autonomously compressing time to build their start to build up their MVPs. And that's also something that's quite scary if you're a business founder as well. Like if you've got a incumbent startup coming and they're building autonomously um and they're running much leaner and the quality and it's very easy them to actually achieve those outcomes that adds to some of the hysteria as well because it's the topic of in business competition being at your doors faster.

模型能力的拐点与验证 (Inflection Point of Capabilities and Verification)

Moderator: 好的,那我们进入下一个问题,这个问题是问 Greg 的。如我之前所说,现在似乎是“循环(loops)”技术的一个重大拐点。相比于 Jeff 一年前发布的原始循环概念,甚至相比于我们在 2025 年末到 2026 年初看到的广泛采用,之所以现在它可能真正火起来了,也许是因为这个新的能力栈——现在模型可以更好地处理图像,验证机制和上下文窗口都变大了,而且推理模型也得到了提升。Greg,考虑到所有这些技术进步,为什么你认为我们今天使用循环的方式仍然是错误的呢?

Original English

Moderator: >> All right, so I'll move on to our next question which will be for Greg. Um it seems like now is a large inflection point for loops like I said before and compared to Jeff's announcement of the raw loop a year ago and even the widespread adoption we saw in late 2025 early 2026 um the reason um that that maybe like caught on was because maybe this new like capability stack where models can now process images better verification context windows got bigger and reasoning models improved. Greg, with all of these advancements, why is the way we're using loops today still wrong in your opinion?

Greg: 我的意思是,我认为模型的智力水平已经不再那么重要了。我觉得归根结底,我同意你的观点,关键在于语义验证,在于你实际去闭合反馈循环的能力,不管你怎么称呼它,也就是去切实地验证智能体输出的正确性。你可以在一定程度上做到这一点。但我不认为你能全局性地做到,至少目前还不行。我认为你可以在那些能够进行确定性验证的事情上做到一定程度的闭环。你可以让系统获得更好的类型检查,更好的代码检查工具(linters),你可以进行模拟测试等等。你可以开始不断添加这些机制,只要保持它们的成本够低,我认为就没问题。但是,一旦你开始在验证过程中引入更多非确定性的因素,我认为它的准确性就会越来越低。这就好像,如果你用一件事去提示智能体,它有 5% 的概率会出错,然后你开始循环这个过程,那么在经过 10 到 20 次循环后,它正确的概率突然就变成 50% 甚至更低了。这就是我的意思。而且这样做会花费你非常多的钱。我总是会反复强调所有这些做法的经济可行性。不过为了让我的观点更有依据,我很确定大多数大型 AI 公司现在仍然在使用 Sentry。这是为什么呢?他们使用它也只是为了捕获简单的 bug。不是安全漏洞,也不是性能退化等等。在我们现在的循环方式中,这些问题依然存在。我们还没有解决这些问题。所以……

Original English

Greg: >> I mean, I don't think that model intelligence matters a lot anymore. I think it boils down to I agree with you, the semantic verification, the actual ability to uh close the feedback loop, however you call it, to actually verify that the outputs of the agent are correct. And you can do it to an extent. I think I don't think you can do it holistically, at least not at this point. I think you can do it to an extent to things that are deterministically verifiable. Uh you can get better typing in your system. You can get better llinters. You can get um simulation testing and all of that. You can start keep adding that and as long as you keep those cheap, I think that's fine. Um the moment you start adding even more non-determinism as your verification process, I think that becomes less and less correct. it starts contributing more like you know how if you prompt agent with one thing and there is a 5% chance it's going to have an error in it and then you start looping that then suddenly after 10 20 loops it's going to be 50% chance it's correct or maybe less like that's what I mean and it just costed you so much money to do that I'm going to keep coming back to the economic viability of all of that um and but but to base it a little bit in like evidence I'm pretty sure that majority of large AI companies are still using Sentry. But why is that? They they are using that just to catch simple bugs as well. It's not security bugs. It's not performance regressions, etc. Those problems still exist in the way that we are looping now. And we haven't solved those problems yet. So,

Moderator: 谢谢。对于 Dex,Ralph 首创了一个理念:在每次迭代中输入全新的上下文,以避免上下文腐化(context rot)。现在随着上下文窗口变得大得多,这变得更加容易管理了。Dex,我们在上下文腐化和上下文工程方面已经走出困境了吗?

Original English

Moderator: >> thank you.

For Dex, the Ralph pioneered the idea to feed fresh context into each iteration to avoid context rot. This has become even more manageable now that context windows have gotten much larger. Dex, are we out of the woods regarding context rod and context engineering?

Dex: 我会回答你的问题,但这里会有一个自由讨论的环节吗?因为我还有些问题想问 Jeff。

Original English

Dex: >> I'm going to answer your question,

but is there going to be like an open floor part? Because I have more questions for Jeff.

Moderator: 我们开始吧。这次我们试图采用牛津辩论的风格,让它更加结构化,以防止仅仅变成……

Original English

Moderator: >> Let's go.

We're trying to do Oxford debate style for this one to keep it like more structured and prevent like just

Dex: 你们不想让它变成一场混乱的扯皮大会。

Original English

Dex: >> you don't want it to just turn into a chaotic yap fest.

Speaker C: 她是想防止出现像你和我那样一聊起来就滔滔不绝的情况。对。

Original English

Speaker C: >> She's trying to prevent what you and I do where we just start yapping. Yeah.

Moderator: 哈哈,已经开始了。是的,我这次本来是想控制一下你们俩的。

Original English

Moderator: >> Yeah. It's already started.

Yeah. I was trying to control like both of you guys this time.

Dex: 好吧,我会尽可能快地回答这个问题,然后我就要开始找 Jeff 抬杠了。

Original English

Dex: >> Okay, I'm going to do this answer as quickly as I possible and then I'm going to start busting Jeff's ball.

Moderator: 行,你可以在你的时间里想说什么就说什么。没问题。

Original English

Moderator: >> Yeah, you can say whatever you want with your time. You could just Yeah, just that's all good.

上下文窗口与“愚蠢区” (Context Windows and the Dumb Zone)

Dex: 好的。所以,关于当初 Ralph 最酷的一点就是:“好吧,你要不断清空上下文窗口”。这样做完全高效吗?从 token 的角度来看可能并非如此,但这意味你可以让它运行一整夜。因为如果你只是不断地塞入消息,上下文窗口就会溢出。但如果你只是重新启动它,告诉它:“这是我期望的世界状态。去检查一下代码,看看我们现在有什么,然后执行下一步让我们达到那个状态。” —— 这是一个非常干净的方式,能把你大部分的工作保持在上下文窗口的所谓“智能区(smart zone)”。如果你告诉它只做一件事,然后我们就清理并重启。现在上下文窗口确实变长了。而且我想更新一下我之前的说法,我记得在迈阿密讲过,不过那个视频还在制作中。所谓的“愚蠢区(dumb zone)”,其实和其他东西一样,更多地像是一个辅助轮。如果你已经每周和 Claude 聊 70 个小时,连续两三个月了,你可能就不需要去考虑“智能区”和“愚蠢区”的区别了,因为你已经建立了自己的直觉。这只是一个指导原则,这就是我们为什么教大家这个——它是一个给你刚开始接触 AI 时的指导原则。尽量把它保持在 10 万个 tokens 左右。对于更大的百万级上下文窗口,我们可能会把这个上限修正到 20 万 tokens 左右。但对于最难的问题,我经常尝试把它控制在 6 万以内。而对于那些我只是在跟智能体“即兴发挥”、懒得去压缩上下文再重新开一个新会话的事情,我也经常会超过 30 万 tokens。但这全凭你的直觉。一个表明你进入了“愚蠢区”的明显迹象就是,在某些情况下,你知道你已经输入了 20 万个 tokens,模型似乎完成了一些工作并试图让测试通过,但它就是不起作用,于是它开始搞各种奇怪的 hack 技巧,然后你去读它的思考轨迹,发现“哦,那是个测试,但那是另一回事,我不需要修复那个,那是一个之前就存在的问题”,然后你会说“不,不是这样的”,那个时刻的挫败感,当你意识到“好吧,它在徒劳地挣扎想取得进展”时,那种本能反应,我认为是很多人在与这些模型合作了几个月后才会培养出来的。但如果你还没有这种直觉,那么你知道,这就是我们的指导原则。所以,是的。

Original English

Dex: >> Okay. Uh so, uh yeah, the the cool thing about Ralph back in the day was like, okay, you keep clearing the context window and like is it completely efficient? Like probably not from a token perspective, but it meant you could leave a thing running overnight and it would never like if you just kept stuffing messages in you would overflow the context window. But if you just relaunch it, say here's my desired state of the world. go check the code and see what we have and do the one next step to get us there. Uh it was a very clean way to keep most of your work in what we call the smart zone of the context window. If you tell just do one thing and then we're going to clear and restart. Um context windows have gotten longer. Um and I will like give an update. I think I gave this in Miami, but that that video is still in production. Um the the dumb zone is really as as much as anything else is is a it's more like training wheels. Like if you have been talking to Claude for 70 70 hours a week and for two to three months, you probably don't need to think about the smart zone versus the done zone because you've built your intuition. It's a guideline if you're just getting this is why we teach people this is like it's a guideline if you're just getting started with AI. Try to keep it around 100,000 tokens. For larger million context window, we probably revide revise this up to like 200,000 tokens. But I've regularly tried to keep it under 60 for the hardest problems. I've regularly gone over 300K for things where I'm just like kind of like riffing with the agent and I'm just like too lazy to compact it and and and move on and do a new one. Um, but this is your intuition. Like one of the tellt tell signs that you're in the that you're in the dumb zone is like uh there's certain cases where the model you know you're 200,000 tokens in and the model's like finished some work and it's trying to get the test to pass and it's like not working and it's like doing all these weird hacks and you read the thinking traces and it's like oh that's a test but that's from something else and I don't need to fix that and that's a pre-existing thing and you're like well no it's not and that that is the moment that frustration where you're like okay it's it's flailing trying to make something happen. That's the instinct that you a lot of people I think cultivate after a couple months working with these models. Uh but if you don't have that yet then you know then then this is our guideline. So yeah

上下文窗口与循环的未来 (Context Windows and the Future of Loops)

Speaker A:上下文窗口正在变得更好。我认为它们变得更大了。所以像 Ralph 那样在每次单次迭代中“尽可能少做事”的核心循环理念,在这里的动机已经不那么强烈了;要知道,你能引入系统的反馈越多,你能自主完成的事情就越多。如果你能让确定性的东西做出决策,并构建小型的提示词交给智能体(Agent),而且你不用去刻意记住做这件事,你不用告诉它:“嘿,去检查一下 PR(Pull Request)的评论并修复它们”,然后在那儿等,接着有人又留了一条评论,你三个小时后回来说:“哦,再去检查一下评论”。如果你能把那个过程自动化,那真是太棒了。我认为这就是目前有效运转的“循环(Loops)”相关工作的核心所在。

Original English

Speaker A: Context windows are getting better. I think they're getting bigger. Um and so like the core like Ralph loop of like do as little as possible in every single iteration is like less of the motivation here then you know the more feedback you can pipe into the system the more you can do autonomously. And if you can have deterministic things making decisions and building small prompts to give to an agent and you don't have to remember to do that, you don't have to tell it, hey, go check the PR comments and fix them and then wait and then someone makes another comment and you come back three hours later and say, oh, check the comments again. If you can automate that process, that's great. And that's kind of the core of I think what is like loops stuff that works today.

Dex:嗯,我想问一下——

Original English

Dex: Um, I want to ask

Moderator:我们的时间只够到这儿了。很抱歉。我必须开始严格把控日程进度了。好的,那么现在我们要深入探讨一下“是什么造就了一个优秀的循环”。那么,好的,让循环变得优秀的一部分原因在于“验证(Verification)”。然而,这似乎有些矛盾:人们常说我们的工作是停止编写提示词(prompts),并开始编写循环,但带有糟糕提示词的循环却会导致智能体作弊,通过修改测试用例来达成目标,而不是努力去通过那些测试。Jeff,当模型验证它自己的工作时,你如何防止它作弊?

Original English

Moderator: that's all we have time for. I'm sorry. I have to start like really keeping this on schedule. Um, okay. So, now we're going to get into the anatomy of what makes a good loop a good loop. So, okay, part of what makes a loop good is verification. However, it seems contradictory that people are saying our job is to stop writing prompts and start writing loops when the loops with bad prompts result in agents cheating and meeting its goal by modifying the tests instead of working to pass them. Jeeoff, how do you keep the model from cheating when verifying its own work?

防止智能体作弊与上下文窗口管理 (Preventing Agent Cheating and Context Window Management)

Jeff:嗯,我大量利用了 pre-commit 钩子(hooks),各位。我通过分析已完成的工作,将这种反向压力(back pressure)设计到工程中。我做的另一件事是……Dex 提到过 Ralph,其中一件事就是大家都在试图进行压缩(compaction)。想想看,压缩有点像一种有损函数(lossy function),就像把一段视频上传到 YouTube,然后下载再上传 100 次。你在那里会丢失保真度,而且这本身已经是一个非确定性的、概率性的系统。所以 Ralph 诞生的理论基础是,好吧,存在一个“愚蠢地带(dumb zone)”,我想做的是确定性地分配它所需要的一切,因为如果它没有被分配,那么本质上它能做的事情的搜索空间就没有受到约束,但同时也要留出一点余量。我也会在上下文超过 100K 时紧张冒汗,即使是用这些百万级上下文窗口的模型也是如此。有一点需要非常认真地思考:很多人,他们觉得他们想在公司里使用 LLM(大语言模型),就觉得“我有这些数据,太好了”。但实际情况是,好吧,你将必须使用循环来批处理这些数据。我希望你把上下文窗口看作……你们还记得 720k 的软盘吗?意思是说,你只有大概那个软盘八分之一的可用内存来真正给一个 LLM 使用。所以你必须进行批处理。你大概只能分配相当于一部《星球大战》的容量。如果你拿《星球大战前传一》的电影剧本,并对其进行分词(tokenize),在上下文窗口被塞满之前,你实际上只能在内存中存放两份那样的电影剧本。对于一个基于文本的电影剧本来说,那大约是 150 KB 的数据。所以,对此要非常小心。我长期以来做过的一件事(可能很傻)是,当新模型发布时,我会运行一个没有任何技能或任何 markdown 格式的“裸”模型,实际上我会去掉我所有的技能和 markdown 等等,因为这些模型其实是有品味和偏好的。比如,GPT-4(或者这里提到的 GP55)刚出来时,如果你用大写字母对它大吼大叫,它会变得软弱和胆怯;但如果你使用 Anthropic 的模型,它其实希望你对它大喊。去读读模型卡(model cards)吧,各位,比如对于集成商来说,这里面有独特的偏好。所以让它保持在正轨上,这实际上是一项工程工作,这真的是工程。

Original English

Jeff: Um, I heavily exploit pre-commit hooks, folks. Um and I engineer in that back pressure uh by analyzing the work that is done. Um the other thing I do is um Dex mentioned that with Ralph it was one of the things was everyone was trying to do compaction. Think about compaction is kind of like a lossy function like uploading a video to YouTube and then downloading and uploading it 100 times. you're losing fidelity there and it's already a non-deterministic system probabilistic thing. So the the theory behind how Ralph came to be it's like okay there is a dumb zone and what I wanted to do is deterministically allocate everything it needs because it's if it's not allocated then it's it's essentially this the search space of what it can do is not constrained but also leaving a bit of a headroom leaving a bit of headroom so I also I get meat sweats when I go above 100K even with these million context windows and this is really important to think about um a lot of people they think they want to use LMS um at a company and it's like I got this data. I was like sweet. Okay, you're going to have to use a loop to batch this data. I want you to think about the context windows is essentially remember the 720k floppy disc. You mean you've only got about an eighth of that floppy disc of usable memory you can actually use for an LLM. So you actually have to batch it. You can only allocate roughly around about Star Wars. If you go Star Wars Episode One movie script and you tokenize it, you can actually just hold two of those movie scripts in memory before the context window is cooked. That's around about 150 kilobytes of data on a textbased movie script. So be very careful about this. something I've done for a long time and it's very silly is uh I run a model bare um without any skills or any markdown actually I get rid of all my skills and all my markdown and everything when the new model is released because the the models actually have tastes and preferences. For example, GP55 when it first came out if you screamed at it in uppercase it became weak and timid but if you use anthropic it wants you to yell at it. Go read the model cards folks like for the integrators like there is unique tastes for it. So keeping it keeping it on the rails is actually it's engineering. It's really engineering.

Moderator:谢谢你。大约 10 天前,Jeff 创造了“收敛工程(convergence engineering)”这个词。他说这是你的循环聚集在一起的地方。它是你的循环产生的“垃圾(slop)”聚集在一起作为一个离散的、被测试的系统的地方,直到它收敛。Dex,我们今天使用循环的方式有什么问题?我们如何确保循环中的垃圾聚集在一起不会仅仅产生更多的垃圾?

Original English

Moderator: Thank you. Around 10 days ago, Jeff coined the term convergence engineering. He said it's where your loop stops together. your loop uh it's where your loop slop comes together as a discrete like system under test until it converges. Dex, what is wrong with how we are using loops today? How do we ensure looping slop together doesn't just produce more slop?

构建反馈循环的挑战与炒作 (Challenges and Hype in Building Feedback Loops)

Dex:呃,我们得阅读代码。我实际上要强调一下 Jeff 今年早些时候做过的一个实验。我相信那个叫 Loom 的项目,我们在那里用 Ralph 循环试图构建未来的软件平台。我想你构建了 AWS 和 GitHub,然后你意识到,好吧,我们如何给模型提供它还不擅长的事情的反馈,比如 UI 测试之类的事情。好的,你为“这个 UI 好不好”创建循环的方法是,你给模型像 PostHog(用户行为分析工具)这样的东西,也就是:好吧,我们可以部署多个不同的实验,我们可以看看用户使用哪些,然后模型不去查看屏幕截图和图片,而是去查看数据,并发现,好吧,这一个表现更好,那这一定是按钮的正确颜色。这样一来,你甚至把人类的视觉品味从这个等式中移除了。所有这些听起来都很酷,关于“我们如何确保循环不会将垃圾带到一起”的观点……我……我不认为你能做到。而这就像是一个完美的例子,证明了炒作跑在了纪律前面。从这个意义上来说——Jeff,Loom 现在怎么样了?

Original English

Dex: Uh we got to read the code. I will actually highlight uh an experiment that Jeff did earlier this year. Uh I believe what was called Loom where we had Ralph Loops trying to build a software platform for the future. And I think you built you built AWS and you built GitHub and then you realized okay how do we how do we give the model feedback on things that it's not good at yet like UI testing and things like this. Okay, the way you create a loop for is this UI good is you give the model something like post hog where it's like okay we can deploy multiple different experiments we can see which ones the users use and then rather than looking at screenshots and PGs the model can look at data and see okay uh this one is performing better that must be the right color for the button and so now you've even removed like the human visual taste from the equation and uh all of this sounded really cool and in the point of like how do we ensure looping doesn't bring slop together. I I don't think you can. And this is like a perfect example of the hype out running the discipline in the sense of uh Jeff, what's what's going on with Loom now?

Jeff:它还在。它还在 GitHub 上。嗯。

Original English

Jeff: It's still there. It's on GitHub. Um

Dex:它还在那里。但你还在开发它吗?

Original English

Dex: it's still there. But are you still working on it?

Jeff:呃,已经有 6 个月没动了,因为我一直在研究验证的工程方法。

Original English

Jeff: Uh it's been 6 months because I'm been looking into engineering ways of verification.

Dex:对。你当时跟我说的是什么来着?你说 Loom 不会成功,除非我们有了更好的编程语言,或者我们有了更好、好得多的模型。对我来说,这就是一个教科书般的例子,说明炒作跑在了纪律前面。我们对所有这些东西都感到非常兴奋。顺便说一下,每个人都应该去尝试,比如 Loom 就很棒。去实验,去试着推动前沿,因为这是你了解前沿在哪里以及什么是可能的方式。否则,你只会用每一个新模型继续使用你的旧技能,并假设它有同样的局限性。但这同时也是一个关键点,我不知道我在试图表达什么,就像是……那个东西现在还不行。那个项目还不能运转。某一天它会运转起来的,存在一种必然性,但我们又要回到“今天什么能用”与“什么是炒作”之间的对比。所以我不知道这是否完全回答了你的问题,但答案是:为了不让循环把垃圾聚集在一起并制造出更多的垃圾,你必须去阅读另一端输出的东西,并确保它不是垃圾。

Original English

Dex: Right. What was the thing you said to me? You said Loom's not going to work until we get better programming languages or we get better much better models and that is a textbook for me of the hype is outrunning the discipline. We're really excited about all this stuff. And by the way, like everyone should do like Loom is awesome. Like go experiment, try to push the frontier because that's how you learn where it is and what's possible. Otherwise, you just keep using your old skills with every new model and you assume it has the same limitations. But it's also a a a key point of like I don't know what I'm trying to say like that this it doesn't work yet. That thing doesn't work yet. It will work someday and like there's inevitability, but again it's what works today versus what is hype. So I don't know if that fully answers your question, but the answer is like the way to not loop slop together and make more slop is to like read the thing that's coming out the other end and make sure it's not slop.

Moderator:是的,这有道理。哦,是……各大实验室还没有攻克这个问题。那么,是什么让你觉得你能攻克它呢?

Original English

Moderator: Yeah, that makes sense. Oh, it's it's the labs haven't cracked it. So, what makes you think you're going to crack it?

Dex:是的。没错。这就是现状。目前,我们都在试图弄清楚如何让这一切运转起来。

Original English

Dex: Yes. Right. This is right now. It's uh we're all trying to figure out how to make this all work.

循环成本与何时回本 (Cost of Loops and When They Pay Off)

Moderator:那是肯定的。怀疑论者说,循环会在悄无声息中失败。它们要么花着你的钱永远陷入死循环,要么智能体在一个只完成了一半的任务上过早地宣布胜利。高管们已经开始质疑 Token 的花费了。Greg,一个循环什么时候能收回成本?这种情况在实际中发生的频率有多高?

Original English

Moderator: For sure. Skeptics say that loops fail quietly. They either spiral forever on your dime or the agent declares victory early on on a half-finish job. Exacts are already starting to question token spend. Greg, when does a loop pay for itself and how often is this actually the case?

Greg:我不认为它们会悄无声息地失败。我认为它们失败时动静非常非常大……特别是当你在看你的构建(builds)时。不过确实存在一些情况——正如我所说,我做循环或者我会说我更多是在做工程——在某些情况下,我认为做循环非常有价值,或者做出明确的决定去支付这个成本是非常有价值的。这里的具体例子是,我们在本地并在 PR 合并后对我们的 PR 进行安全扫描,因为它们总是能发现一些真实存在的问题。嗯,那些被我们忽略的问题,它们在代码审查方面有点击败了人类,而且这很昂贵。为了运行所有我们想要的检查,大概要花我们一个 PR 五块钱之类的,但这就是我们做出的明确决定——认为这是值得的。另外也有一些情况,比如你看看那些规范非常非常完善的系统,例如……所有使用 Next.js 重写的实验,或者用 Rust 重写 Ban,或者是……运行一个浏览器。像这种你针对这些问题拥有多年积累的测试套件和规范的情况下,你真的能够非常非常好地验证结果,那么……那么进行循环并达到那些……那些结果,看起来是行得通的。用 Rust 写的 Ban 看起来运作得相当不错,就我所能判断的而言。所以这里确实存在……

Original English

Greg: I don't think they fail quietly for it. I think they fail very very um loudly especially when you're looking at your builds. Um but there are cases like as I said I do loops or I do engineering I would say more so and there are cases where I think doing loops is very valuable or doing um or making a explicit decision that you're going to pay pay for the cost is very valuable. So the concrete example here is we do security scanning after on our PRs in local and after our PRs even land because they will always find some things that are real. um that we have overlooked and they sort of beat humans on the on the code review and it's expensive. It cost us I think like five bucks a PR or something like that to run all the checks that we want but that's where we made an explicit decision that it's worth it. Um there are also cases where like if you look at the very very well specified systems such as um all the experiments with u nextjs rewrite or a ban rewrite in rust or uh running a browser like cases where you have years and years of test suite and specifications built in around those problems where you can really really well verify um the outcomes then then looping and getting to those um those results. seems to work. Ban in Rust seems to work pretty well from all I can tell. Um so there are

构建原型的用例与炒作的对比

Speaker A: 毫无疑问是有用例的,有些用例我觉得我经常遇到。比如,你可以想象构建原型。我经常为我们在 Sentry 应该做的产品制作原型。这些都是一次性代码。所以,我会快速推进它们,然后就忘掉。如果我们喜欢这些原型,我再去读那些代码时会感到很尴尬。我们会回到起点,开始详细规划我们实际想要实现的内容,并朝着那个解决方案努力。但这将更多地涉及人类在循环中的参与。所以,总的来说,我认为它们有其用武之地。但正如 Dex 所说,我对这种炒作有些反感,炒作已经超越了该有的纪律性。我也非常同意你应该亲自去尝试、去实验,看看什么对你真正有效,看看它在哪里会崩溃。我觉得现在少花点时间在 Twitter 上是有好处的。

Original English

Speaker A: definitely cases and then there are cases of um usage that I think that I do pretty often where you can imagine for instance building prototypes. I do prototypes of products that we should be doing uh doing at Sentry pretty often. Those are going to be throwaway. So I'm going to just slash go on them and forget about them. And if we like them then I'm going to start reading the code and I'm going to be mortified. uh and we're going to go to the square one and start specking out what we actually meant uh and and sort of go towards that solution but it's going to be much much more involving of of human in a loop. Um so so broadly I think they have place uh but as Dex said they the hype is what I have a problem with the hype is outrunning the discipline as he said and I also agree very strongly with the point that you should just try things you should just experiment yourself try to see what actually works for you where uh where the the cookie crumbles um and you know spend less time on Twitter I think is healthy nowadays.

多智能体循环中的共享内存与访问控制

Moderator: 是的,说得好。一个好的、注重成本控制的循环必须追踪状态,以了解它已经尝试过什么。这种内存存在于磁盘上的某个地方,并且越来越集中在多个智能体读写的共享内存存储中,特别是当你从单机智能体转向多智能体时。你最终会遇到这样一个访问控制问题:无法区分哪个智能体写了哪些内存,以及谁可以读取它。这对于个人电脑来说可能没问题,但绝对不能用于生产环境。矛盾点就在这里。共享内存是循环用来相互学习并更快收敛的基础。但是,为了解决访问控制问题而将其限制在每个智能体内部,会孤立它们,从而扼杀了这种共享学习。Ian,我们如何解决共享内存存储的访问控制问题,以便循环能够更快地收敛?

Original English

Moderator: Yeah, good points. A good costconscious loop has to track state to know what it's already tried. That memory lives somewhere on disk and get increasingly in a shared memory store that many agents read and write, especially as you go single player to multiplayer with agents. You end up with this access control problem that can't tell which agent wrote which memory and who can read it. That might be fine for a PC, but it's definitely not for production. The tension lies in this. Shared memory is what loops use to learn from each other and converge faster. But scoping it per agent to solve the access control problem isolates them and then kills that shared learning. Ian, how do we solve the shared memory store access control problem so loops can converge faster?

Ian: 好问题。哇。哇。嗯,

Original English

Ian: Great question. Wow. Wow. Um,

Speaker B: 要是有人正在开发这样的产品就好了。

Original English

Speaker B: only someone was working on a product that could

Ian: 要是有人在思考这个问题就好了。我的意思是,实际上,这首先是一个尚未解决的问题。说实话,我们的访问控制系统并不是为这个机器可以代表我们一半人进行操作和推理的世界而设计的。但宽泛地说,我认为一些底层基础已经开始出现。如果我们暂时不考虑云 IAM 和所有其他东西,像 Markdown 其实非常棒。因此,如果我们在对话中使用 Markdown 作为内存,真正的问题是:我要共享哪些内容?我如何共享这些 Markdown 文件并将其用作内存?然后我如何围绕这些东西分配访问控制?如果我们使用这种模型,我实际上认为我们今天已经拥有了大部分的基础。只是思考起来很笨重。一个很好的例子是,我最近发了一条推文。我当时在玩 Notion,我们在 Keycard 经常使用 Notion。但我真的只想把我所有的 Notion 内容作为 Markdown 文件供我使用。因为这样智能体处理起来会更容易,而不需要通过 MCP。我有点偏向于纯命令行主义,当时正在研究这个。所以我认为我们真正缺少的是——Dex 和我上次在 MCP 上辩论过这个——真正的挑战在于,我如何向智能体展示一个世界,以便我能理解它,然后我如何分配在任何时间点可以访问的内容,并使其易于任何人操作。我不认为我们真正解决了这个问题,但肯定有一些模式的开端是非常有意义的。

Original English

Ian: only if someone was thinking about it. I mean, actually, I've this is this is an unsolved problem first and foremost. like let's be really honest that our our access control systems weren't designed for this this world where machines were acting and reasoning on on behalf of about half of us but broadly speaking I think like some of the beginnings of the substrate starting to emerge and if we ignore cloud IM and all the other stuff for a hot second like markdown is pretty great and so really if we were to say a memory is markdown for the purposes of conversation the really question is like what things and how do I share these markdown files and use that as a memory and then how do I attribute sort of like access control around those things. And if we were to use that model, I actually think we have like the basis for for most of it today. It's just unwieldly to think about. Um, a good example would be I had this tweet recently. Um, I was playing with notion and we use notion a lot at keycard. Uh, but I really just wanted all my notion things available to me as markdown files and because it would just made it easier for the agent to work with it instead of going through MCP and I was a bit of a CLI maxi and was looking through that. So I think what we're missing really is uh and MC, you know, Dex and I debated this last time at MCP, but like we're really the challenge really is how do I present a world to an agent so I can understand it and then how do I attribute what can access at any one point in time and how to make that wieldly for anyone to do it and I don't I don't think the really cracked it, but certainly there's some beginnings of patterns that make a lot of sense.

循环的未来与全自动软件工厂

Moderator: 好的,既然我们讨论了循环的解剖结构,我们现在来探讨一下循环的未来。我们现在是否已经为软件工厂做好了充分准备?循环工程是否已经变得如此之好,以至于我们已经准备好迎接全自动软件工厂了?我们将从这个问题开始:如果循环只适用于可验证的任务,那就意味着全自动软件工厂必须能够验证它们所做的每一件事。Greg,这现实吗?优秀工程工作的其他部分,比如决定构建什么、抽象是否正确以及哪些权衡是可以接受的,这些该怎么办?

Original English

Moderator: All right, now that we discussed the anatomy of the loop, we're going to debate the future of the loop. Like are we essentially well positioned now for software factories? Has loop engineering gotten so good that like we're ready for the full software factory and we'll start with if loops are only good for verifiable tasks that means fully autonomous software factories must be able to verify everything they do. Greg, is this realistic? What other parts of good engineering work such as deciding what to build, whether the abstraction is right, and what trade-offs are acceptable?

Greg: 呃,这现实吗?我想,如果计算是免费的,那将是一个非常好的开端。但是,嗯,我认为我们已经到了这样一个阶段,或者让我这么说吧:作为在循环中的人类,你所做的决定是那些关于设计、架构的重要决定,我目前不会信任让智能体替我做这些决定。我还没有看到这在未来成为现实。我认为原因是,当你在看大型组织时,我想任何有多年经验的工程师都会告诉你,重点并不总是关于你应该构建什么,同样也关于你不应该构建什么,什么是正确的权衡,你希望在哪里投入解决复杂性,与你应该在哪里投入以保持最大的简洁性。在我的经验中,智能体喜欢复杂性。它们会不断无限制地向技术栈添加内容。因此,我认为我们正在转移目标,随着我们添加更多的验证、更多的语义验证,它们能够做的事情也多得多,我并没有忽视这一点,但我确实希望亲自参与实际的架构决策,而且我看不到它们在任何时候能接管这些工作。

Original English

Greg: Uh, is this realistic? I think if if the compute is free that would be pretty pretty good beginning but um uh I think I think we're getting to the point where um or let me put it this way the decisions that you are making as a human in the loop are the decisions of like design architecture the the important ones that um that I would say I wouldn't trust the the agent to do for me uh and I don't see the future where that becomes reality yet. I think the reason for that is when you're when you're looking at large organizations and I think any engineer who has had like years of experience will tell you it's not always about what you should be what you should build but also about what you shouldn't build what are the actual right trade-offs where the complexity is that you want to um that you want to invest in versus where you should be um investing in maximal simplicity. Um in my experience agents love complexity. they will keep adding to the stack um unbounded and so I think we are we are shifting the the post we are getting to the point where as we are adding more validation more semantic verification they are able to do much much more um and I'm not uh not neglecting that but I do want to be in the loop for the actual architectural decisions and I do not see them uh taking that over any anytime

Moderator: 谢谢。Shopify 的工程主管在 Kerser 的 Compile 大会上对所有人说,你的工作就是编写循环。Steinberger 发推文说:“这是你每月的提醒,你不再应该给编程智能体写提示词了。你应该设计循环来提示你的智能体。那条推文的浏览量高达 800 万。” Jeff,你在最初的 Roth 博客文章中说过,循环需要高级的专业知识,而且有时它只能完成 90% 的工作。那么,仅仅给出“编写循环”这样的建议,对于 3000 人、800 万人,或者是那些一开始就真正学会了如何调优循环的人来说,安全吗?

Original English

Moderator: Thank you. Shopify's head of engineering told everyone at Kerser's compile conference that your job is just to write loops. Steinberger tweeted, "Here's your monthly reminder that you shouldn't be prompting coding agents anymore. You should be designing loops that prompt your agents. The view count on that tweet was 8 million people." Jeff, you've said in your original Roth blog post, loops need senior expertise, and sometimes it tops out at 90% of the way there. So, is just write the loops advice is that safe to give to 3,000 people or 8 million people or only the people that have really learned to tune a loop in the first place?

Jeff: 是的,这是一件有趣的事情。要向那么广泛的受众量身定制和传授“你应该做什么”以及“你不应该做什么”是非常困难的。所以我能做的就是,当我最初写作时,我是写给我的同行、那些我仰慕的人看的。不知出于什么原因,我脑海中突然灵光一闪,觉得一切都改变了,游戏规则变了,但我真的很担心一些在函数式编程领域最顶尖的人以及真正的同行能否理解。我记得像 Flyo、Thomas Pek 这样的人,这篇文章让他们顿悟了,他也发表了一篇博客文章证实了这一点。所以我开始为我的同行和那些我仰慕的人写作,以传播给全世界。这很难。但是在同样的意义上,你去 YouTube 看,你会发现有人正在创造以前从未创造过的东西。他们睡一觉醒来,就拥有了一个全新的 Discord 机器人。这很神奇。我希望人们记住,这实际上是一种编排模式。它被简化为一个使用 catwhile true 循环,因为如果你想教授什么东西,cat 是最简单的教学原语,你必须让它变得非常简单。让它极其简单。所以用 cat 加上提示词,也就是说你设计你要输入的提示,你把文件系统作为状态,你循环使用上下文窗口,在循环中运行。这可能有点微不足道,但它的性价比很高。它确实有用。但其全部意图是,上面应该有某种 PID 控制器,或者某种工厂模式,或者某种决策器来决定循环是否应该继续。很难判断它是否安全?“安全”是一个有趣的词。没有人应该在本地笔记本电脑上使用编程工具。这并不是因为人工智能,而是因为 NPM 供应链攻击。这一直是事实。我尝试解决这个问题已经七年了。像人工智能可能不安全这种事情现在又引起了人们的关注。

Original English

Jeff: Yeah, that's an interesting thing. Um, it's really hard to tailor and teach like like what you should do and what you should not do to a broad that broad of an audience. So all I can do is write to when I was originally writing, I was writing to my peers, my people I looked up to. For some reason, I it clicked to my head that everything has changed, the game has changed, but I was really concerned that some of some of the best people in functional programming and like and like real peers. I remember like um Flyo, Thomas Pek, it made him click in his head and he's got a blog post saying as such. So I started writing for my peers and the people I looked up on to disseminate down to like the entire world. It's hard. Um but in the same same sense you go to YouTube, you got people who are creating things that have never created things before. They go to sleep and they wake up and they got a brand new Discord bot. It's magic. I want people to remember that is actually a pattern for allocating Basically, it was a pattern as an orchestration pattern. It was condensed down to a wild true loop using cat because cat is the simplest teaching primitive to if you want to teach something, you got to make it really simple. Make it really simple. So cat prompt, i.e. you engineer the prompt what it's going to be, you use the file system as state, you recycle the context window, run in the loop. It's a little bit maybe, but it it was just bang for buck. It just works. But the entire intent there is that there should be some sort of like PID controller on top or some sort of factory or some some sort of some sort of determinator as such saying whether a loop should continue or not. Um it's really hard um is it safe or unsafe? Safe is an interesting word. Um, no one should be using coding tools on your local laptop. And this is not because of AI, but this is because of MPM supply chain attacks. This has been true. I tried tried cracking this problem for seven years. Um, and like it's now on the attention again that we that like AI could be unsafe

关于现有开发实践的安全性 (Software Development Security)

Speaker A: ……运行不安全的命令,但是,各位,你们日常工作场所的软件开发实践已经是不安全的了。所以,先解决这个问题,然后这些技术就会开始变得安全。

Original English

Speaker A: ...running unsafe commands, but like your software development practices day-to-day your workplace are already unsafe, folks. So um fix that then these techniques start to become safe.

Moderator: 谢谢你。

Original English

Moderator: Thank you.

AI 编程助手的实际提效与务实态度 (Pragmatic Speed Up from AI Agents)

Moderator: Dax。你在推特上说,大多数工程师看到编程智能体带来了 2 到 3 倍的提速,这是符合现实的;如果你一味追求 100 倍的提速,你就会迷失在关于优化的元元问题中,而且你可能永远也达不到只要保持务实就能实现的、足以改变生活的 10 倍提速。如果只有 2 到 3 倍的提速是可能的,我们该如何才能实现完全自主的软件工厂呢?

Original English

Moderator: Dax. You tweeted that most engineers are seeing a 2 to 3x speed up from coding agents and that's realistic and if you try for that 100x speed up you're going to get lost in the meta meta problem of optimization and you may never get to that life-changing 10x speed up that is possible by staying pragmatic. If only 2 to 3x is what's possible? How do we ever get to a fully autonomous software factory?

Dax: 是的,我认为这是一个好问题。我觉得这里也有一个混合的情况,对吧?我们讨论的是行之有效的方法与炒作之间的对比。而且我认为这绝对值得强调,即你应该去尝试,去尝试推动边界,去做今天可能还行不通的事情,但你不能仅仅因为在 Twitter 上看到某些东西有效,就假设你所有的工作都发生了改变。这基本上就像是,不要抛弃我们已经学到的一切。不要特意去抛弃我们作为一个社区积累和整合起来的几十年的软件工程职业生涯。而且,这就回到了我所认为的人们在着手设计和创建他们的软件工厂时最大的反模式,那就是他们会说:“好的,太棒了。我要去闭关三个月,我已经读了一堆博客文章,我要去建造我的软件工厂,它将成为真正的软件工厂。它将成为我们未来发布所有产品的方式。”然后三个月后你回来了,你却从来没有触碰过真正的问题。你从未把它交到任何人手中。这就像任何产品一样。你正在建立一个软件工厂。你是在为你的队友打造一个产品。你是在为 5 个、10 个、500 个、5000 个工程师构建一个产品。而且,正确的方法是从小处着手,不断迭代,弄清楚什么是行之有效的。去尝试各种事情。它们可能行不通。它们也可能会奏效。但我们学习 AI 以及如何有效使用它的方式,是通过建立直觉,这就是为什么你应该尝试一堆可能行不通的东西,但你应该承认,你不应该试图强行突破那个前沿。所以我的建议是,与其试图端到端地自动化一切,不如在整个系统中建立这些小的、渐进式的反馈循环,有一天醒来,你就会发现自己的速度提高了两到三倍,同时仍然能够阅读代码,仍然能够掌控架构。因此,你无需为了达到这个境界,就抛弃我们学到的一切、我们知道的一切以及你所有的直觉。所以我会警告人们,诸如“去弄清楚如何让速度提高 2 倍或 3 倍”,因为如果你试图让速度提高 100 倍,你可能会把一切都搞砸。而且,我的意思是,你能想象吗?如果世界上每一个软件工程师的速度都提高 2 到 3 倍,并且拥有接近人类 99% 的质量水平,那将改变世界上每一个企业,每一个初创公司。它会改变整个数学公式,而我们正试图去……只是不要,不要走得太远。争取做你能做到的事情。迭代地构建东西。你会学到很多东西,当下一个模型出现时,你就可以准备好将速度提高 5 倍、10 倍。

Original English

Dax: Yeah, and I think this is um it's it's a good question. I think there's like there's there's a mix here too, right? We're talking about what works versus what is hype. And I think it is definitely worth I want to highlight like you should try again like you should try to push the frontier and do the things that might not work today, but you should not assume that all your work has changed just because you saw something work on Twitter. basically is like don't throw away all the things we've learned. Don't go out of your way to cast aside this, you know, decades long career of software engineering that we as a we as a uh community have built up and put together. Um, and it it gets back to like what I see as the biggest anti-pattern for how people set about designing and creating their software factory, which is they say, "Okay, cool. I'm going to go away for 3 months and I've read a bunch of blog posts and I'm going to go make my software factory and it's going to be the software factory. It's going to be the future of how we ship everything. Uh, and then you come back three months later and you never touched the problem. You never put it in anybody's hands. It's just like any product. You're building a software factory. You're building a product for your teammates. You're building a product for 5 10 500 5,000 engineers. Uh, and the the right approach is to start small and iterate and figure out what works. try things. They might not work. They might work. But the way we learn AI and how to use it effectively is through building up intuition, which is why you should try a bunch of stuff that probably won't work, but you should acknowledge that you should not try to push like through that through that frontier. Um, and so my advice is kind of like instead of trying to automate everything end to end, build these small incremental loops throughout your system and you will wake up one day and you will be moving two to three times faster while still being able to read the code, while still owning the architecture. And so you don't have to throw away everything we've learned and everything we know and all your intuitions just to get to this place. And so I would I would caution people like go figure out how to make move 2x faster or 3x faster because you're going to like blow everything up by trying to go 100x faster. Um and uh yeah I mean if can you imagine if every software engineer in the world was 2 to 3x faster and had a like near human 99% level human level of quality that would change every single enterprise in the world every single startup in the world. It would change the entire math and we're trying to like go to just don't don't go too far. Shoot shoot for what you can do. Build things up iteratively. You'll learn a lot and you'll be ready for when the next models come out and you can go 5x 10x faster.

代码责任与问责制 (Code Verification and Liability)

Moderator: 谢谢你。当涉及到验证时,不仅仅是关于验证工作成果,还在于验证是谁完成了这项工作。雇主们已经明确表示,人类要对他们交付的代码负责,无论是不是由智能体编写的。然而,今天只能有一个人或智能体在提交(commit)上签名。Ian,如果我们无法确定是谁做了工作,以及工作是代表谁做的,我们是否准备好让软件工厂来编写和审查所有代码了?

Original English

Moderator: Thank you. When it comes to verification, it's not just about verifying the work. It's also about verifying who did the work. Employers have made it clear that humans are responsible for the code they ship, whether an agent wrote it or not. Yet, only one person or agent can sign a commit today. Ian, are we ready for software factories to be writing and reviewing all code if we can't determine who did the work and who the work was done on behalf of?

Ian: 我觉得,是的,其实大概两周前我还在思考这个问题。所以,首先也是最重要的一点,我们面临一个问题:Git 只允许一个签名者进行提交(commit)。所以我们必须解决这个问题。其次是围绕 SOC 2 和合规性等方面的一些事情,对吧?但我认为更重要的是,看待智能体的方式,有点像我们如何在大型组织中看待服务所有权。也就是说,在某个时刻,必须由人类来对智能体的行为负责。世界上不可能存在这样一种情况,即智能体是一个独立的实体、拥有独立责任属性,并且在某种程度上自己承担法律责任。只有能够承担后果的人才具有责任能力,对吧?而且这始终必须建立在它是一个人的基础之上。因此,如果一个人设计了一个循环,而这个循环产出了糟糕的软件,你猜谁要承担这个责任?那肯定是那个人,对吧?如果我们没有责任制,社会就无法运转。它根本就运转不起来。破坏必须带来后果。解决糟糕决策带来的后果,而这些后果显然是基于所造成破坏程度的一个梯度;如果我们没有这种机制,任何事情都行不通。所以我实际上会宽泛地说,不存在人类永远不需要负责的世界。总会有探讨人类责任级别的空间。问题在于这是一个个人还是一家公司(即一群人)。问题在于,这如何改变我们今天系统的运作方式?就今天使用 Git 而言,我可以签署一个提交声明说“这是我做的”,它归属于我的公钥。对于上一代我们思考软件的方式来说,这很酷。现在,在未来,当我让一个智能体代表我去做某事时,它生成了一个循环,生成了一堆代码,并且部署到了生产环境,我必须为“我分配了那个任务、我拥有去做的意图”这一最初事实负责。我们过去根本不具备做到这一点所需的底层基础设施。尽管如此,我确实认为这种情况会很快改变。这并不是什么让我们极难突破的事情,但我们必须重新思考我们在整个供应链和软件开发生命周期(SDLC)中看待责任归属的方式。而且这并不是一个新问题,对吧?就像供应链安全一直是个问题。实际上这是我们使用智能体时面临的最大挑战之一,我们一直希望能有一个跨供应链的更具确定性的路径和签名链。而问题就在于,对于第一方代码与第三方代码,你该如何做到这一点?

Original English

Ian: I think yeah, actually I was I was playing around with this problem like two weeks ago. So um first and foremost, we have a problem. Git only allows one signer on a commit. So we got to fix that. Um two is you know there's some things around sock two and compliance, right? But I think like more importantly the way to think about agents is is kind of how we think about service ownership in large organization. It's like at some point a human has to be attributable for an agent's actions. Like there's no word where it's going to be like an agent is its own entity and its own attributable thing and somehow it has liability. Like the only people who have liabilities are people that can have consequences, right? And that always has to be grounded in it being a human. And so if a human designs a loop and that loop presents bad software, guess who's attributable for that liability? It's going to be the human, right? Like the society doesn't function if we don't have liability. Uh it just doesn't work. There has to be consequences to damage. address the consequences of bad decisions and what those consequences are are obviously a gradient based on the on on the damage that's done and if we don't have that nothing works and so I'd actually broadly say there is no world where a human is never responsible there's always a world where human's level of responsibility and the question is whether it's a human or a corporation which is group of humans um the question is is how does that change the way that our systems work today so today with git I can sign a commit that says I did it it's attributed to my public key. And that's cool for last generation's way that we thought about software. Now, in the future, when I tell an agent to go do something on my behalf, it generates a loop, it generates a bunch of code, it goes to production, I have to be attributable to the initial fact that I assigned that thing, the that intention to do that. We we simply did not have the substrate required um to do it. Although, I do think that's going to change pretty quickly. um this isn't like something that's like crazy difficult for us to break, but we have to rethink the way that we think about attribution across the supply chain and the SDLC. And this is not a new problem, right? Like supply chain security has always been an issue. It's actually the biggest challen one of the biggest challenges we have with agents and we've always wanted to have more deterministic uh pathway and signature chains across the supply chain and this is just how do you do that for first party code versus uh third party code?

Speaker B: 好吧,Ian,说实话,我们的职业有点像个笑话。实际上我们在个人层面上并没有法律责任。

Original English

Speaker B: Well, Ian, um, just to shoot straight, um, our profession is a bit of a clown show. We actually don't have liability at a personal level.

Speaker C: 就像,我们自称为工程师。但我们并不是真正的工程师。

Original English

Speaker C: Like, we call ourselves engineers. We're not really engineers.

Ian: 这倒是真的。

Original English

Ian: That's true.

总结陈词 (Final Thoughts)

Moderator: 这些问题将会变得非常复杂,各位。也许我们需要重新审视这些话题。现在我们的主要辩论环节已经结束,接下来到了总结时间,每个队伍的每位成员将有两分钟的时间进行总结,并表达他们的最终想法,然后我们将决定谁是赢家。Greg,你想开始吗?

Original English

Moderator: Some of this stuff is going to get really complicated, folks. And maybe we need to revisit these topics. And now that we've had our main debate section to wrap it up, each member on each team will get two minutes to uh wrap it up and describe their final thoughts before we decide who the winner was. Uh Greg, would you like to

Greg: 是的,没问题。看到我们实际上在那么多观点上达成一致,真的很有趣。不过事情就是这样。我确实认为,我一直试图在这里传达的最大的一点是,我们在哪里,我们正努力达到哪里,以及你真正从无处不在的炒作循环中听到了什么,这正是我想要加倍强调并降低这种炒作热度的地方。我基本上想说,去尝试事物,独立思考。不要盲从于泡沫之中。不要随波逐流地跟风炒作,因为你会通过自己亲手去做,发现什么对你有效,并发现系统在哪里最容易崩溃。人类一直以来学习最好的方式就是通过实践,而不是靠看 YouTube 视频。归根结底,这就是我试图做的事情。而且我稍微有些怀疑,当涉及到允许整个自动化循环自主运行时,因为我看到那些定性结果并不能满足我的要求。但我总体上是乐观的。我很乐观是因为,我已经看到了自己能多做多少事情,以及我现在能解决哪些类型的问题……

Original English

Greg: Yeah, absolutely. Um it's it's interesting seeing at how how many points we actually agree with each other. Uh but that's how how this goes. I I do think that the biggest point that I'm trying to I have been trying to drive here is um where we are whether we where we are trying to be and then what you're actually hearing from the everpresent hype uh hype loop and that's where I want to sort of double down and tone this down. I basically want to say try things, think for yourself. Don't lean into the bubble. Uh don't lean into the hype because you will find what works for you and you will find where the system breaks the best by doing it yourself. That's always has been how um how humans learn the best is by practice, not by you know watching YouTube. Um and ultimately that's that's what I'm trying to do. And I'm slightly skeptical uh when it comes to you know allowing the full loop to run because I see the qualitative results not being up to scuff for my uh my requirements. Uh but I am optimistic in general. I am optimistic because I've seen how much more I can do and what are the types of problems I can address today that I...

Software Engineering Career

Speaker A: 连一年前都无法解决,更不用说更早之前了。而且情况会越来越好。但我也不担心我的软件工程职业生涯。我不认为我们会消失。我认为我们仍然会是这个整个“软件工厂”或下一个即将到来的炒作泡沫中的重要组成部分。

Original English

Speaker A: couldn't address even a year ago. Uh let alone let alone earlier than that. And uh and it's going to get better. But I'm also not worried about my software engineering career. I don't think we're going away. I think we are still going to be important piece of this whole software factory or whatever the next um sort of hype bubble is going to become.

Moderator: Ian,你想接着说吗?

Original English

Moderator: Ian, would you like to go next?

The AI Train Has Left the Station

Ian: 我很乐意。所以,列车已经驶出站了。这东西确实管用。而且生产力确实有了实质性的提升,对吧?比如,有更多的工作被完成了。我的意思是,完成的所有事情都是好事,但现在是有更多的事情被完成,而且其中很大一部分实际上做得很好,对吧?对吧?我想我们都能同意这一点。其次是,我们现在面临着竞争动态,对于一家公司来说,已经不可能再袖手旁观地说:“嘿,我们要置身于这种循环(loops)之外。”对吧?“我不参与这些编程智能体(coding agent)的东西了。”那种时代已经结束了。我们已经度过了那个阶段。就像那列火车已经驶离了车站。一旦那列火车离开车站,世界上的每个人就会开始说:“天哪,我必须把我的东西整理好。我们必须跟上潮流。”他们也必须这样做,因为我们生活在一个竞争非常激烈的资本主义社会。这也是,你知道的,不涉及政治,但这就是我们的现状。归根结底,问题其实不在于它会不会发生,而在于它何时发生,以及随着时间的推移会达到什么程度。我认为除了真正跟上时代的步伐,我们别无选择。

我们所有人都像是在紧紧抓住一枚正在飞行的火箭,而我们其实并不知道那是什么——我们真的不知道它的轨迹,只觉得它飞得真快,加速度极其疯狂。有时候,好像我们远远领先于彼此,而有时候我们又远远落后。但我不认为你有选择的余地,你只能去:第一,弄清楚什么是循环(loops);第二,弄清楚你可以将它们应用在你的代码库中的哪些地方。会有这样的地方,比如“嘿,这部分是高度可验证的。这是计算机能够解决的问题”。去做连接器(connectors)并没有太大意义。这是一个很好的例子,比如,你知道,过去十年里有多少公司赚钱仅仅是因为它们基本上就是连接器农场。那里已经没有多少价值了,对吧,如果你能自动化连接器创建的话。所以,在软件中,有些地方你今天就可以应用循环并获得真正的价值。也有一些地方你可能不应该应用,你应该决定那是在哪里,那很可能是你正在创造的核心价值所在,对吧?所以,这就是我的看法。但泛泛而谈的话,列车已经驶出站了。生产力曲线是驱动社会的动力。投资回报率(ROI),也就是我们作为一个社会的生产力提高了多少,是驱动国内生产总值(GDP)的因素。而驱动GDP的,归根结底是资金流向了哪里,以及资本是如何分配的。

Original English

Ian: I would love to. So, the trains left the station. This stuff works. Um and there is a real productivity increment, right? Like more stuff is getting done. I just mean all the stuff getting done is good, but more stuff is getting done and a good percentage of that stuff is actually good, right? Right? I think we can agree that. And second is now we have competitive dynamics where it's no longer possible for a company to sit back like, "Hey, we're going to sit this loop stuff out." Right? I'm going to sit this coding agent stuff out. Like that that's over. We're past that. Like that train leaves the station. As soon as that train leaves the station, everybody in the world starts saying, "Holy I got to keep my stuff together. We got to stay up with the Joneses." And they have to because we live in a very competitive capitalist society. And that's also, you know, not into politics, but it is what we are. And at the end of the day, the question is is not really um will it happen, it's when it happens and to what degree over time. And I don't think there's a choice but to actually stay up to date.

And we're all just kind of holding on to a rocket ship ride where we don't actually know what that we don't really know the trajectory other than it feels real fast with crazy acceleration. And sometimes like we're way ahead of each other and sometimes we're way behind. But I don't think you have a choice not to be one figure what loops are. Two figure out where you can apply them in your codebase. There's going to be places where hey this is highly verifiable. This is a problem that like computers can solve. It doesn't make a lot of sense. Cough connectors is a good place of like you know last 10 years the amount of companies that have made money because they're basically just connector farms. It's no longer a lot of value there right if you can automate the connection connector creation. So there's places in software that you can apply loops to today and get real value. And there's places where you probably shouldn't and you should decide where that is and it's probably the core value of what you're creating, right? And so that's how I think about it. But broadly speaking, trains left the station. The productivity curve is what drives society. The ROI, so how much more productive we are as society is what drives GDP. What drives GDP is ultimately where the dollars go. Um, and how capital gets allocated.

Moderator: Dex,交给你了。

Original English

Moderator: Kick it over to you to Dex.

The Future of Software Factories

Dex: 我殷切期盼一个 Lysoft 软件工厂成为可行的世界。我非常想要一个我们不再需要阅读代码的世界,在那个世界里我们可以直接完成一切。嗯,如果你看了我周二的演讲,我认为这实际上是一个我们目前只能在模型层面解决的问题。我不认为工具链(harness)能做到这一点。呃,很遗憾,因为我非常喜欢构建工具链和进行上下文工程(context engineering)。嗯,所以我的建议是,呃,听听 Jeff 的看法。呃,等 Loom 真正起作用的时候告诉我,在那之前,呃,可以使用循环,但别那样用。我很喜欢这个。

Original English

Dex: I eagerly await the world where the Lysoft software factory is feasible. The I I would love a world where we don't have to read the code where where we can just do everything. Um if you watch my talk on Tuesday, I I think this is actually a problem we can only solve at the model level right now. I don't think the harness can do it. Uh unfortunately, because I love building harnesses and doing context engineering. Um so my advice is uh pay attention to Jeff. Uh let me know when Loom is actually working and until then uh use loops but not like that. Love it.

Speaker B: 各位,有太多想说的话了。嗯,工厂代表了我们走向未来的方向。就像,它本质上就像是一台永动机,对吧?它就像是一个白日梦。那些今天才刚成立、今天才拿到融资的公司,不要以为你可以直接把这个拿来在你的公司里实施,因为在市场上这普遍还没有被解决。但我想补充说明一下这个关于“物质”的独白,如果你尝试运行循环(loops)或尝试建立一个工厂,而你使用的是 Python,那将会变成一场小丑秀。如果你用 Ruby 来做,那也会是一场小丑秀。静态类型也是一种验证形式。

伙计们,我鼓励你们去想出几个套路(catters)和几个实验。尝试运行一些循环。用 Ruby 构建一个应用程序,然后再尝试用这些循环来修改它。然后你就会看到那维护性的一团糟。接着,尝试使用 Haskell(overll -> Haskell)。我不在乎你懂不懂。你不明白也没关系。大语言模型(LLM)懂 Haskell,你实际上可以提示智能体,让它就像对你儿子或女儿解释一样向你解释这个。所以我不确定现在连代码是否都需要具备可读性,但这就是前沿、前沿的思维:它需要是可解释的。所以我正在尝试不同领域的验证。但是,呃,有一件事是肯定的,嗯,类型系统非常重要,Rust 语言非常优秀,因为它的生态系统如何对某些类型进行建模。而且,因为提到了供应链,我必须要说,已经 10 个月了,我做到了……我尽可能少地使用任何开源软件。降到最低。我只根据我的需求来生成代码,然后当供应链攻击发生时,我的反应是:“没影响到我”,或者是软件安全漏洞,但这都是关于最小化爆炸半径。问题是,如果你处理开源项目,负责人、维护者、不管是谁,你无法对一个人类进行工具调用(tool call)。那不是通用人工智能(AGI)。你要尽可能多地将你所有的源代码供应商化(vendoring),这样智能体实际上可以通过循环或提示词来进行修改。你需要掌握自己的供应链。

Original English

Speaker B: So so much to say folks. Um factories represents where we're heading to the future. Like it's it's essentially like a perpetual energy motion machine, right? It's like it's the pipe dream. companies are only just getting founded today and get receiving their rounds today. Don't think you can just take this and implement it in your company because it's just generally not solved in market. But I will say to add to the matter monologue, if you try to run loops or try to build a factory and you're using Python, it's going to be a clown show. If you did it in Ruby, it's going to be a clown show. Static types are a form of verification.

Folks, I encourage you to to come up with a couple of catters and a couple of experiments. Try running some loops. Build an application in Ruby and then try to modify it again with these loops. And you'll see the maintainability mess. And then try doing overll. I don't care if you don't know. It doesn't matter if you don't understand. LLM understands HA hasll and you can actually prompt the the agent to explain this as if you to to your son or daughter. So I'm not sure even code needs to be readable these days but this is frontier frontier thinking it needs to be explainable. So I'm playing around with different domains of verification here. But uh one thing's for sure um type systems are in very much in rust is very good because how the ecosystem models some types and because supply chain was mentioned I need to say it has been 10 months now I do I minimally use any open source software minimally I just generate it to my requirements and then when a supply chain attack happens I'm like didn't affect me or software security flaws, but it's about minimizing the blast radius. And the thing is, if you deal with open source project, the person's on on lead, the maintainer, what does have you, you can't tool call a human. That's not AGI. You want to be vendoring all your source code as much as possible so the so the agent can actually modify and in with either loops or ether prompting. You need to own your supply chain.

Debate Conclusion

Moderator: 太棒了。非常感谢你的分享。很遗憾,我们必须到此为止了,不过我们现在可以决定赢家是谁了。嗯,如果你的想法被改变了,而且你现在站在 Dex 这一边,请举手。

Original English

Moderator: Awesome. Thank you so much for that. Unfortunately, we have to close it at that, but we can now decide our winner. Um, show of hands in the audience if your mind was changed and you're now on Dex's side, raise your hand.

Audience Member: 冲啊!

Original English

Audience Member: Let's go.

Moderator: 好的。如果你的想法改变了,而且你现在站在 Ian 和 Jeff 这一边。

Original English

Moderator: Okay. If your mind was changed and you're now on Ian and Jeff's side.

Audience: 耶。

Original English

Audience: Yeah.

Audience: 耶。

Original English

Audience: Yeah.

Moderator: 灯光太亮,我有点看不清。你们觉得呢?

Original English

Moderator: I kind of couldn't see with the light. What did you guys think?

Speaker C: 等一下,再举一次。票数很接近。对,

Original English

Speaker C: Wait, do it again. It's pretty close. Yeah,

Speaker D: 看不清。灯光实在是太亮了。

Original English

Speaker D: it's impossible. The delights are so bright.

Moderator: 好吧。嗯,这是一场很棒的辩论。非常感谢大家的聆听。嗯,我们下次再见。

Original English

Moderator: All right. Well, that was a good debate. Thank you so much for listening. Um, we'll see you next time.

Project Loophole Presentation

New Speaker: 好的。感谢大家的到来。那么,我想要,呃,我将讨论我的项目,Loophole。我实际上是摩根士丹利的机器学习研究员,但这与摩根士丹利没有任何关系。这只是我纯粹为了兴趣而一直在开发的一个开源项目。首先给个高层次的概览,它真的是一个你可以建立在这个对抗性智能体框架之上的游戏。所以你设定你的道德观,一个智能体将其编纂成一个法律系统,然后这两个对抗性智能体会试图找出你的道德观中的矛盾之处。而且最近我也一直在这个基础上构建不同的扩展功能。嗯,但我想,你知道,从起源故事开始,以及我是如何想出这个想法的。

所以,这真正始于,你知道,很久以前,我把我的 DNA 送到了 23andMe 做祖源测试。嗯,而且我不断听说,你知道,最近 DNA 样本是如何被使用的,当然,这可以用来帮助破案,所有的这些法医学和悬案。我当时在想,我之前是如何选择退出了一切相关的,因为,你知道,我很害怕那种滑坡效应,以及我的 DNA 会被如何使用。嗯,但是,你知道,当然,对于某些特定的案件我是可以接受的。这有点意思。我当时就在想,比如,你知道,如果有人能逐个案子向我展示,你知道,我们用你的 DNA 来破案,你知道,帮助解决这个悬案或者这起谋杀案,我就能某种程度上说“是”或“否”。而且我也知道我的道德观中这种细微的区别和界限到底在哪里。嗯,但是你知道,当然,逐个列举所有的案件在认知上确实让人望而却步。比如目前还没有一个很好的方法来做这件事。

然后我就在从一种更宏观的角度思考,其实我们整体的法律系统有很多类似之处。所以你知道,看待我们社会中的法律系统到底是什么的一种方式,它实际上只是我们试图去……的一种方式。

Original English

New Speaker: All right. Thank you everyone for coming. So, I want to uh I'll be talking about my project, Loophole. And I'm actually a machine learning researcher at Morgan Stanley, but this has nothing to do with Morgan Stanley. This is just a open-source project I've been building for fun. And to give sort of the high level flavor to start, it's really this uh game you can play that's built on top of this adversarial agent framework. So you specify your morals, one agent codifies that into a legal system and then these two adversarial agents try to find contradictions in your morals. And lately I've been building different extensions on top. Um but I wanted to you know start with sort of the origin story and and how I came up with this this idea.

And so this really started you know a long time ago I had sent my DNA into 23 and me for uh ancestry testing. Um, and I kept hearing about, you know, more recently how DNA samples can be used, of course, to help solve crimes and all these forensics and cold cases. And I was thinking about how I had sort of opted out of of everything because, you know, I was scared of the kind of slippery slope and and how my DNA would be used. Um, but, you know, there there are certain cases that I would be okay with. And it's sort of interesting. And I was thinking like, you know, if someone could present to me case by case, you know, we use your DNA to solve, you know, help solve this cold case or this murder, I could sort of say yes or no. And I know where the the definition of like the in the nuance of my morals are. Um, and but you know, of course, like enumerating all of this case by case is really sort of cognitively prohibitive. Like there's not a good way to do this currently.

And then I was thinking sort of more zoomed out that there's a lot of analogies to sort of the legal system as a whole. So you know one way to think of what a legal system is in our society is really just a way that we are trying to

Synthesizing Case Law with LLMs

Speaker A: 翻译任务。而且我认为,通常我们往往会偏向于过于笼统,因为你知道,找到那种细微的差别真的很难;如果你试图做到完美的翻译,它可能会导致一些奇怪的边缘情况或奇怪的故障模式。我甚至认为,就像点对点一样,当我们在政治上相互联系时,很多时候分歧更多是在争论我们其实并没有真正分歧的核心价值观,而不是去探索我们道德中那些真正微妙的点。

Original English

Speaker A: translation task. And I think often we kind of err on the side of being too general because you know finding that nuance is really difficult and if you try to have a perfect translation it can we lead to these kind of weird corner cases or or weird failure modes and I even think kind of like peer-to-peer when we're relating to each other politically a lot of times the disagreements are more kind of fighting over core values which we don't really disagree on um instead of exploring the really nuanced points of our our morals.

Speaker A: 我想,遵循那种更广泛的法律例子,在英国普通法体系中,这就是为什么我们如此严重地依赖判例法,因为我们知道,发现这种细微差别和这些细微的边界是很困难的。因此,我们在某种程度上依靠聪明的判断来正确解释和适用法律。当然,即使像最高法院那样,事情也可以被提升,我们可以决定一项法律是否完全有效。所以这就是这个想法的来源:你能不能把你的道德观拿来,进行这种合成判例法的生成?

Original English

Speaker A: And I think, you know, following that kind of broader legal example, I think in the, you know, the English common law system, that's why we sort of lean on case law so heavily because we know that finding this this nuance and these nuance boundaries is difficult. And so we kind of rely on smart judging to interpret and apply the law um correctly. And you know of course even with like the Supreme Court things can get elevated and we can decide whether a a law is valid at all. Um and so that was sort of you know the idea is like can you take your morals and can you do this kind of synthetic case law generation.

Speaker A: 所以你知道,手工完成这项工作是令人难以承受的,但似乎新一代的大语言模型(LLM)终于足够聪明,可以进行这种高水平的道德推理了。所以这就是这个游戏的起点。我想带你们了解一下游戏的初始版本、设置、运作方式,然后还要谈谈我这个开源项目基础上构建的一些不同分支,因为我认为它可能会朝着一些有趣的方向发展。

Original English

Speaker A: So you know um this is overwhelming to do by hand but it seems like the new generation of LLMs are finally sort of smart enough to do this kind of high level moral reasoning. And so that was sort of the the the starting point for this game. Um, and I just want to take you through sort of the initial release of the game, the setup, how it works, and then also talk about some of the different branches I've been building on top of this open source project because I think it could go in some interesting directions.

Speaker A: 从宏观层面来看,游戏的运作方式是你通过自然语言进行互动——这有点像一个基于终端的游戏。你具体说明你的道德观,这可以是你的普遍道德观,也可能是关于特定主题的,然后有一个智能体会采纳这些道德观,并起草一套非常丰富的、带有法律术语的成文法律体系。接着它就在这样一个循环中运作:一个智能体被指示去试图在你的体系中寻找漏洞,也就是那些不道德但在你的体系下却合法的事情;而另一个智能体则被提示去寻找“过度干预”的情况,也就是那些在你的体系下其实是道德的但却被判定为非法的事情。

Original English

Speaker A: And so at a high level, how the game works is you and natural language, and this all happens, it's sort of like a terminalbased game. You specify your morals and it could be uh your general morals or maybe about a specific subject and then there's one agent that takes those morals and drafts sort of a really rich legal ease codified legal system and then it just operates in this loop where one agent is instructed to try and find loopholes in your system. So something that is uh immoral but legal and another is uh prompted to find overreach. So things that are actually moral but illegal given your system.

Speaker A: 然后,一个充当法官的智能体会审视你的道德观、生成的法律代码以及这些合成的判例法例子,它首先会看能否自动进行修补。比如,也许你最初起草的法律代码只是一次不完美的转化,并且并没有真正的矛盾,所以它就能自动进行这种更新。或者,它可能真的是对你道德观的一种规定不足,或者是你道德观中某种形式的矛盾,在那种情况下,它就会把问题提交给你作为用户,让你来当法官并做出决定。呃,你那边还能看到画面吗?我这边画面消失了。

Original English

Speaker A: And a judging agent looks at the your morals, the produced legal code and these sort of synthetic case law examples and first sees can it auto patch. So like maybe the original draft of your your legal code sort of was an imperfect translation and there's not really a contradiction and it can just sort of autod do this update. or maybe it's really kind of an underspecification of your morals um or some kind of contradiction in your morals and in that case it raises it to you as the user to sort of be the judge and make a determination. Uh is it still on for you? It disappeared for me.

Speaker B: 好的。

Original English

Speaker B: Okay.

Speaker A: 那么我知道,如果你——我希望如果你对这个游戏感到好奇,你会去玩一下它。它全都在 GitHub 上。但我只想展示一些例子。这里有很多文本,所以它更多的是展示输入和输出的大致轮廓。

Original English

Speaker A: Um, so I know that, you know, if you I hope if you're curious about the game, you'll play it. It's all on GitHub. But I just wanted to show some examples. And this is a lot of text, so it's more about just showing the kind of shape of the the input and output.

Speaker A: 所以这就是你提供输入的方式。回到 DNA 的例子,你可能会具体说明若干项道德原则。然后,这个类似成文法律的体系,再次从形式上看,它真的有这种法律文本的感觉,你知道,有前言、条款、章节,它真的试图用精确的法律语言来编写,然后不同的合成判例法就会被提出来。

Original English

Speaker A: So this is sort of how you would provide your input. And going back to the DNA example, you might, you know, specify some number of of moral principles. And then the sort of codified legal system, again, just kind of looking at the shape has this really like legal ease, you know, preamble articles, uh, sections, really trying to be uh, you know, write it in precise legal language and then these the different kind of synthetic case laws get suggested.

Speaker A: 比如在这个案例中,它谈到这是它发现的一个漏洞:一家保险公司训练了一个预测性机器学习模型,不是用你的 DNA,而是用你 DNA 的衍生物,所以它会说,在你的体系下这其实是不道德的但目前却是合法的;而且在这个案例中,它发现它可以进行这种自动修补,然后你就会得到这种类似 Git 风格的对比:你的原始法律体系,以及为了确保这与你的道德观保持一致它所必须做出的差异修改。

Original English

Speaker A: So in this case it's talking about this is a loophole it found where an insurance company um trained a predicted machine learning model not on your DNA but on artifacts of the DNA and so it's saying you know this is actually immoral but currently legal given your system and in this case it's it found that it could do sort of this auto patching and then you get this sort of like git style difference of of your original legal system and then the the difference it had to make to ensure you know this was consistent with your morals.

Speaker A: 接下来这个是关于“过度干预”的例子。在这个案例中,它发现它无法进行自动修补。所以它讨论的是,比如有人提交了他们的 DNA 用于基因研究,但研究人员发现他们患有一种罕见但可治疗的遗传疾病。但是目前你的道德观似乎表明,不允许他们将这种疾病告知提交者。所以这被提交给用户,也就是我,来做出判断。同样,当你做出判断后,你就会得到这个打过补丁的法律体系。

Original English

Speaker A: And then this is a an example of overreach. And in this case, it found that it couldn't do the auto patch. So it's talking about, you know, someone submits their DNA for uh genetic research, but the researcher finds they have a rare but treatable genetic disorder. Uh but currently your morals kind of say this shouldn't be allowed that they could disclose this disease to the the person submitting. And so this was raised to the user to me to kind of make a judgment. And then similarly when you make the judgment you get this this patch legal system.

Future Applications of Codified Systems

Speaker A: 所以你看,这只是一个有趣的小游戏,我把它发布在了 Twitter 上,并在 GitHub 上开源分享了它。至少对我来说,这是我迄今为止传播最广的一篇帖子。它让我思考,我想很多人只是觉得它很有趣,你可以对你的道德观进行压力测试,看看你是否有什么有趣的矛盾之处。但它也让我思考,这里面是不是还有更多的东西?比如,这能不能产生更实际或更大范围的影响?所以我想谈谈我正在探索的三个不同的分支。

Original English

Speaker A: And so you know it's just sort of a a fun game and I posted on Twitter and shared it open source on GitHub. And for me at least it was by far the most viral post I've had. And it sort of made me think like I think a lot of people just said it was sort of fun. You could stress test your morals, see if you have any interesting contradictions. But it also made me think, you know, is there maybe something more here? like could this be um you know have more like practical or bigger scope implications and so I'll just talk about three different branches I'm kind of exploring

Speaker A: 第一个分支,有点脱离法律领域,而且真的更加实用,那就是思考一种为聊天机器人(或者普遍意义上的智能体)自动制定“宪法”的方法。比如,假设你是一家公司,你想拥有一个面向客户的智能体或聊天机器人,你希望它在遵守一定道德准则的同时,也有它愿意谈论和不愿意谈论的事情。

Original English

Speaker A: um the first and sort of leave leaving the legal area and really more practical is thinking about sort of an auto way to make constitutions for chat bots or really you know for agents in general where you know say you're a company and you want to have a um agent or chatbot that's that's customer facing and you want it to sort of adhere to a moral code but also have things it will and will not talk about.

Speaker A: 我在一个分支中大致是这样构想的:你以类似的方式写下你的道德准则。你写下聊天机器人应该和不应该谈论的内容,然后它会尝试编写这个成文的系统提示词,接着你让这些对抗性智能体试图让它谈论不该谈论的事情,或者拒绝谈论应该谈论的事情。我认为这类似于 GEA 的一种模拟方法。不过它的真正目的是为了构建这些成文的系统提示词。

Original English

Speaker A: Um I've kind of in one branch formulated it so you in a similar way write your morals. You write what the chatbot should and not talk about and then it tries to write this codified system prompt and then you have these kind of adversarial agents trying to get it to either talk about something it shouldn't or refuse to talk about something it should. And I see it as this sort of analogy or analogous method to GEA. Um, but really aimed at kind of building these codified system prompts.

Speaker A: 我特别感兴趣的第二个用例是,将其视为一种执行更多临时或去中心化合同的方式。我想,在一个简单的例子中,假设你可以具体规定你希望你的数据或隐私在网上被如何处理,然后你可以通过这种对抗性游戏,来获得这个关于你希望你的数据被如何处理的成文法律体系。假如你去,比如说苹果发布了新的服务条款之类的时候,你可以运行你的法律体系与苹果服务条款之间的矛盾对比,并揭示任何有趣的矛盾,或者类似那些会导致你的道德观、你的诉求与公司实际做法之间产生差异的合成案例。

Original English

Speaker A: The uh second use case that I'm I'm particularly interested in is thinking of it as a way to sort of do more ad hoc or decentralized contracts. So I think in in a simple case say like you can specify how you want your data or privacy to be handled online and you can go through this sort of adversarial game to get this codified legal system of how you want your your data handled and if you go to you know say Apple releases a new terms of service or something you can run the um contradictions between your legal system and between Apple's terms of service and like surface any interesting contradictions or like synthetic cases where this would lead to a difference between how you know your morals, what you want and what the company is doing.

Speaker A: 当然,如果对方是一家大公司,也许你并不能真正改变什么,这不是在谈判,但你至少可以对自己正在签署的合同有更好的了解。但我也认为,在思考更多去中心化的情况下,比如你试图在没有某种中央权威机构强制执行的情况下签订合同,而且你可能试图在不同国家之间签订合同。

Original English

Speaker A: And you know, in the case that it's a big company, maybe you can't really change anything. It's not a negotiation, but you can at least be sort of have better information about the contract you're signing. But I also think in the case of you know thinking more decentralized like if you're trying to have contracts without you know some central authority kind of enforcing them and you're trying to maybe do contracts across different countries.

Speaker A: 考虑到,如果你能具体说明你的道德观以及你希望如何进行交互——也许就像承包工作一样,你希望你的工作如何获得报酬,以及围绕这些事情的不同道德观。然后另一方也可以这么做。这样你们双方都会得到这种经过压力测试的成文合同。然后你就能发现分歧所在(如果有的话),并在你们达成一致之前将它们暴露出来。这样你就能对这份合同整体上更有信心了。

Original English

Speaker A: Um, thinking about like if you can specify your morals and how you want to like interface, you know, maybe it's just like contracted work, how you want your work to be paid for and and and the different morals surrounding that. And the other party can do the same. And then you both get this kind of stress-tested codified contract. And then you can kind of find the the disagreements if there are any, and surface them before you agree. And then you can kind of be more confident in the the contract as a whole.

Speaker A: 最后一点,也许是一个更具抱负的角度,那就是思考如何建立更聪明或更高效的政府。我认为这会存在很多不同的隐私问题和后勤问题,但暂且忽略这些,仅仅从宏观层面来考虑。我认为,对于选民或支持者来说,这可能是一种非常有趣的方式,如果你——

Original English

Speaker A: And the last thing and maybe the kind of more aspirational angle is thinking about smarter government or more efficient government. Um I think there would be a lot of different privacy issues and logistical issues but sort of ignoring those for now and just thinking big picture. I think for voters or constituents, you know, this could be a really interesting way if you

政治模拟与自动修补

Speaker A: 你定义了自己的道德规范,你就拥有了这个经过压力测试的法律代码。那么,任何新的法案出台,或者任何政客出现,你都可以把你的契约和他们的进行对比,然后找出你不同意的地方,或者对你来说属于不道德或是自相矛盾的有趣点。嗯,我也认为,在人际交往方面,这更像是一种……我认为我们所有人对事物的感受都有很多细微的差别,而这是一种触及那种细微差别的方式,而不是仅仅在价值观上争论,因为很多时候,价值观其实并不冲突。

Original English

Speaker A: you defined your morals, you have this stress-tested legal code, sort of any new bill or politician that comes out, you could kind of run your contract against theirs and surface, you know, what are the the cases you would disagree or or interesting uh points that are kind of immoral to you or or a contradiction. Um, I also think, you know, relating to one another, it's like a more I think we all have a lot of nuance in the way we feel about things and this is a way to kind of get to that nuance instead of arguing over just values, which is, you know, often the values are not in contradiction.

Speaker A: 然后,我认为对于立法者来说,这也许还有个更实际的用途。你可以想象一下,如果你想提出一项法案,并且你可以对立法机构中的所有其他立法者进行模拟,那么你就可以在提交法案之前对其进行压力测试。所以,我在这个项目上一直在构建的第三个分支是,我试着针对美国参议院来做这件事。我所做的是,我首先让 Claude 仔细研究所有现任美国参议员,查看他们的投票记录以及任何其他公开信息,从而构建他们的道德体系,然后通过 loophole(漏洞)流程运行它,以获得一种成文的法律代码。

Original English

Speaker A: And then I think maybe a little more practically for legislators, you could imagine um if you want to propose a bill and you can have like a a simulation of all the other legislators in a in a legislative body, you could sort of stress test it before submission. And so the the third branch I've been building on this project is I I tried to do this for the US Senate. And so what I did is I first had Claude go through all current US senators and look at, you know, kind of all their voting history and anything else that was public and build their kind of moral system and then ran it through the loophole process to get a codified sort of legal code.

Speaker A: 然后在这个系统上,你可以把任何目前正在提案的法案拿过来,甚至你可以提出自己的法案并提交。你可以让 Claude 模拟每位参议员会如何投票。在这里你可以看到一些参议员的投票明细,他们倾向于哪边,以及他们投票背后的推理逻辑。我觉得这挺有意思的,想想看,你能看到人们会对一项拟议的立法如何投票,而且这也变成了一个——我认为很契合这次会议的主题——它变成了一个自我可验证的领域或循环,你甚至可以考虑对法案进行某种“爬山算法”式的优化,直到获得绝对多数票,或者任何能让它通过的票数。

Original English

Speaker A: And then on this system, you can, you know, take any current bill that's being proposed or even propose your own and submit it. And you can have Claude sort of simulate how each senator would vote. And so here you can see like a breakdown of um some senators, which way they're leaning and sort of their reasoning behind the vote. Um and I think you know it's sort of interesting just to think about like seeing what you know a proposed piece of legislation how people would vote but also this sort of becomes and I think on theme of the conference its own verifiable domain or loop and you could think about even kind of hill climbing the bill towards um getting like a supermajority or whatever you need it to pass.

Speaker A: 比如,在我测试的这个医疗保险(Medicare)法案案例中,我想它最初的投票结果可能是 50 对 50。然后系统找到了优化法案语言的方法,最终让它以 52 票获得通过。我认为这就是一个很好的例子,它可以找到法案的核心原则,并尝试将法案与每位参议员的“契约”进行对比,寻找是否有什么方法可以修改措辞,从而在不违背法案核心原则或道德规范的前提下,以那种方式进行“自动修补”。然后,它还可以对为了最大化票数而需要进行的修改进行排序,作为用户,你可以通过这种方式来权衡利弊。

Original English

Speaker A: And so in this case, like this this Medicare bill I was testing, you know, it found that I think it originally started at like a 5050 vote and it found ways to hill climb the language of the bill such that it passed with um 52 votes. And I think, you know, this is an example of it can find um like the sort of the core tenants of the bill and it can try to find like run the the bill against each senator's contract and find is there any way I can change the language such that I don't violate sort of the core tenants or morals of the bill and kind of do those auto patching that way. And then it can also find um you know kind of rank order the changes that would need to be in place to maximize votes and you as a user can kind of choose the trade-offs that way.

Speaker A: 另外,我最近在这个分支上尝试的最后一件事实际上是着眼于更大的图景,比如:这能否带来一个更高效的政府?你可以让一个州或任何选区里的每一位选民都拥有他们的法律代码,然后你只需提交任何法案,就能切实衡量出实际选民之间的共识程度。为此,我使用了英伟达(Nvidia)的一个非常棒的美国用户画像(personas)数据集。我每个州抽取了 500 个画像,这应该能很好地代表该州的人口情况。我采用了同样的流程,让他们在给定画像的情况下起草自己的道德规范和法律契约,然后拿出你感兴趣的任何法案,在每个州上运行它。你还可以衡量人们有多喜欢这项法案,或者这项法案在多大程度上符合他们的道德规范,然后同样进行这种“爬山”优化,从而为民众优化法案。

Original English

Speaker A: And then the last thing I I've been trying out more recently with this branch is actually looking at, you know, kind of even bigger picture like can this lead to an even more efficient of government where you have every sort of constituent in a in a state or whatever the district is sort of have their legal code and then you could just submit any bill and actually measure sort of the agreement between like the the actual voters. And so for this I took the Nvidia has this really great data set of USA personas and so I took 500 personas per state and it's supposed to be sort of well representative of the state's population did the same process of having them given the persona draft their morals draft their sort of legal contract and then take any bill you're interested in and kind of run it against each state. And you can also, you know, measure how much people like this bill or how much it's in agreement with their morals and then also do this hill climbing where you kind of optimize the bill for the people.

Speaker A: 所以,最后总结一下,至少我觉得这是一个非常有趣的游戏。我可能有失偏颇,但尝试不同的、你关心的事情,输入你的道德规范,看看是否存在任何矛盾,这确实很有趣。很多时候,它会引发一些非常有趣的问题,而一旦你提供了那些细微的差别,系统中的智能体就无法再找到任何矛盾点,你也会感觉很好,因为你拥有了一个一致且细腻的道德体系。不过,我还是很感兴趣去探索一下:这里是否有一些针对去中心化或更优质契约的实际应用场景?也许甚至可以将其用成立法者对法案进行压力测试的一种方式,并且思考如何起草出对选区民众更好、或者更能代表选区民意的法律。嗯,这个二维码是参议院模拟器的链接。所以,如果你感兴趣的话,我鼓励你玩一玩。另一个链接是我的个人网站,上面有 loophole 的完整 GitHub 仓库。请随意使用并 Fork 它。我很乐意有其他贡献者加入。谢谢。

Original English

Speaker A: And so just to conclude, you know, at at minimum, I think it's a pretty fun game. I'm biased, but it's a lot of fun to just try out different um you know, things you care about, put in your morals, see if theres any contradictions. you know, often it will raise some really interesting questions and then once you kind of provide that nuance, the game, you know, the the agents won't be able to find any more contradictions and you can kind of feel good that you have like a a consistent nuanced uh moral system. Um, but I I am interested in, you know, exploring could this be are there kind of real applications here for some kind of like decentralized or better contracts and maybe even for legislators as a way to sort of stress test your bills and even think about how to write um better laws that are, you know, better for the people in your district or more representative of what the people in your district want. Um, and so this QR code is to the the Senate simulator. So, I encourage you if you're interested to play and um the other one is to my website which has the the full GitHub to loophole. Um and please, you know, play with it, fork it. Uh I'd love to have other contributors. Thank you.

工具泛滥与语义路由 (Tool Overload and Semantic Routing)

Sail Shik: 我们来谈谈一个乍一看无伤大雅的错误,也就是一次性给 AI 智能体开放它可能需要的所有工具权限。基本上,这种方法在演示阶段效果很好。在工具数量较少(例如 10 个工具)的情况下,它甚至也能正常运行。但一旦工具目录不断增加,智能体就会变得更慢,成本可能更高,准确率也会下降。这就是为什么我们称之为“百种工具智能体陷阱”(the 100 tool agent wrap)。在接下来的半小时左右,我们将向大家展示它是如何崩溃的、具体的数据情况是怎样的,以及采用即时上下文(just-in-time context)的语义路由将如何帮助我们解决这个问题。我先做一个简单的自我介绍。我是 Sail Shik。我目前在 Protica 担任数据科学家。我的背景横跨 AI、NLP、营销、分析甚至工程领域。我目前的重点是应用 AI、NLP、对话智能以及 RAG 检索系统。我特别热衷于让 AI 系统在生产环境中变得更可靠、可衡量,甚至具备超越演示阶段的可扩展性。

Original English

Sail Shik: talking about a mistake that looks harmless at first giving which is basically giving an AI agent every tool access it might ever need all at once. So basically that approach works well in a demo. It might even work with a small number of tools like like say for example 10 tools but once the catalog grows uh the agent gets slower uh it might become more expensive and less accurate as well. That is why we are calling it the 100 tool agent wrap. In the next half an hour or so, we'll show how why it breaks, what the numbers look like, and how semantic routing with just in time context can help us fix this problem. So, a quick introduction about myself. I am Sail Shik. Um, I'm currently working as a data scientist with Protica. My background spans across AI, NLP, marketing, analytics, and even engineering. My current focus is on applied AI, NLP, and conversational intelligence along with rack systems. I'm especially interested in making AI systems more reliable, measurable, and even scalable beyond the demos in production.

Ankosha Soi: 我是 Ankosha Soi。我是 Presertica 的高级数据解决方案工程师。我在 AI、数据工程和生产系统领域工作了十多年。我的重点是工程方面。因此,关键不在于什么代码能在 Notebook 里跑通,而在于它能否在真实的负载、真实的用户以及真实的故障中生存下来。这就是我们今天探讨的视角。Sail 将更多地关注模型和路由行为,而我将更多地关注系统设计实现和生产环境中的权衡。

Original English

Ankosha Soi: And I'm Ankosha Soi. I work as a senior data solutions engineer at Presertica. I have spent more than a decade in AI, data engineering and production systems. My focus is the engineing side. So it's not about what's going to work in notebook but whether it's going to survive in with real load, real user and real failures. So that is the angle we are taking today. So will focus more on the model and routing behavior and I will focus more on system design implementation and production trade-offs.

Sail Shik: 太棒了。谢谢你,Ankosh。那么我们进入正题。让我们设想一个常见的设计。你可能倾向于构建一个能够处理很多事情的系统。比如说,查询数据库、发送邮件、甚至查看日历、查找订单、调用 API 等等。在这里最简单的方法是,在每次请求时都向模型提供每一个工具的定义,包括每一个函数名、每一个描述,甚至每一个 JSON Schema 都会被塞进提示词中,不管用户是否需要它。所以我们把这称之为“臃肿智能体(fat agent)”。在小规模场景下,比如只有 10 个工具时,感觉没什么问题。模型通常可能会选中正确的工具,演示效果看起来也不错。

Original English

Sail Shik: >> Awesome. Thank you Ankosh. Um so let's get into it. So let's imagine a common design. Um you tend to build a system and it can do many things. Say for example querying a database, sending an email, um even checking a calendar or looking up an order, calling an API and so on and so forth. The simplest approach over here would be to give a model every tool definition on every request, every function name, um every description and even every JSON schema will go into the prompt whether the user might need it or not. So that we are calling that as a fat agent. At small scale it feels fine with 10 tools as well. The model might usually pick the right one. The demo looks good.

Sail Shik: 随后,产品规模扩大了。10 个工具可能变成了 30 个,或者还会继续增加,最终模型开始调用错误的函数。它开始产生混淆。如果有相似的工具,它甚至可能凭空捏造工具名称,而且响应时间也会变长。最重要的一点是,这种设计的失败并不是因为某一个工具写得不好。它的失败是因为每一次请求都被迫携带着整个工具目录。所以我们来看看这里。假设你的整个 Schema 中有 741 个工具,光是把所有的工具描述包含进去,基本就会占用多达 127,000 个 Token,而且这还是在连用户的实际问题都还没考虑进去的情况下。这无疑会导致上下文过载,所以我们需要妥善地管理这一点。

Original English

Sail Shik: Then the product grows. Um 10 tools might become 30 or it will keep on increasing and eventually the model starts calling the wrong function. Starts confusing. Similar tools may invent tool names and even take longer to respond. The important point is basically the design does not fail because one tool is badly written. It fails because every request is forced to carry the entire catalog. So let's look here. There are say for example uh 741 tools in your uh in your entire schema but and it will look basically take up to 127,000 tokens just to have all those tool descriptions in it and this is even before the user user's actual question is even considered. So basically this will lead to context overload and uh we need to manage that properly.

Sail Shik: 在这张幻灯片上,我们能看到它为什么会失败,以及为什么准确率在超过某个临界点后就会崩溃。你可以看一下这条准确率曲线,在使用 10 个工具时,臃肿智能体(PAT agent)有差不多 78% 的概率能正确选择工具。这算不上完美,但也还可以用。而当工具达到将近 100 个时……

Original English

Sail Shik: So on this slide we see why it is failing and why the accuracy collapses beyond a point. So when you look at the accuracy curve with the 10 tools PAT agent will get the tools right almost 78% of the times. That is not perfect but it's usable. At almost 100 tools the

Accuracy, Latency, and Cost Challenges

Speaker A: 准确率下降到了 40% 左右。被调用的工具中,不到一半是正确的工具。如果规模进一步扩大,比如在这里有 741 个工具时,准确率将仅仅只有 13.6%。简而言之,大概每八个工具中才有一个是正确的。因此,当我们将它与语义路由器(semantic router)进行比较时,语义路由器的表现截然不同。在相同的目录规模下,它的准确率保持在 83% 左右。这是因为模型不是在从数百个工具中进行选择,而是在从一个小而相关的集合中进行选择。臃肿的智能体失败的一个原因是,它迷失在了“中间迷失”(lost in the middle)问题中。模型对长上下文的开头和结尾投入了更强的注意力。当中间塞入了数百个工具模式(schemas)时,模型无法可靠地使用它们。所以,我们最终为庞大的提示词付出了代价,而这个提示词使得决策变得更加困难。这里还有另外两个原因。第一是延迟,第二是成本。我们在之前的幻灯片中看到,准确率是一个大问题。如我们所见,另一个问题是延迟和成本。比如,如果我们有 741 个工具,我们看到它几乎需要 127,000 个 tokens。这包含了工具描述和模式文本。所以,基本上每个请求都在支付这个成本。

Original English

Speaker A: accuracy drops to around 40%. Less than half of the tools that are called are the correct tools. And if it grows beyond that like say for example in over here at 741 tools the accuracy will be a mere 13.6%. So in short it's roughly one correct tool out of eight tools. So when we compare it with the semantic router, semantic router behaves very differently. It stays about 83% across the same catalog sizes. That is because the model is not choosing from hundreds of tools. It's choosing from a small and relevant set. One reason the fat agent fails is because it's lost in the middle problem. model pays stronger attention to the beginning and end of the long context. When hundreds of tool schemas are packed into the middle, the model does not reliably use them. So we end up paying for a huge prompt and that prompts makes the decisions even harder. So two more reasons over here. First is the latency and the second is the cost. So we saw uh in the earlier slide that accuracy was a big problem. Another issue is with latency and cost over here as we can see. So like say for example if we have 741 tools we saw that it requires almost 127,000 tokens. Uh that will include the tool description and the schema text. So that cost is basically being paid on every request.

Reliable Agents in Production with Restate

Speaker B: 大家好。本次演讲的主题是如何在生产环境中可靠地运行智能体。它不会涉及敏捷(agile / eile)部分,而是会讨论为了使智能体具备弹性运行能力,你需要准备的所有其他事项。基本上就是基础设施层。我想用 Andrej Karpathy 上周的一段话来作为开场背景。它描述了我们与智能体和 LLM(大型语言模型)交互的方式经历了三波浪潮。第一波,LLM 就像是一个网站,我们去访问它,问它一个问题,它思考几秒钟然后给我们一个回答。第二波,开始走向智能体。它是一个我们下载到电脑上的应用程序。它有一些工具可供支配,并且可以在我们的交互下完成一些工作。现在的第三波,将越来越趋向于持久和异步的实体。所以,智能体将成为我们基础设施中的长驻进程,它们可以访问工具,以及组织和上下文中的其他智能体。

Original English

Speaker B: Hi everyone. This talk will be about how to run agents reliably in production. It will not be about the eile part, but it will be about all the other things you need to get going in order to run agents resiliently. So the infrastructure layer basically. I want to set the scene with this uh quote of Andre Apathy of last week. It describes that the way we interact with agents and LLMs has been evolving in three waves. The first wave was an LLM being something like a website where we go to we ask it a question, it thinks for a few seconds and then gives us a response. The second wave was going towards agents. It was an app that we download to our computer. It has some tools at its disposal and it can do some work with our interaction. Now the third wave will be going more and more towards persistent and asynchronous entities. So agents being longunning processes in our infrastructure with access to tools and other agents around the organization and context.

Speaker B: 因此,随着我们的用例越来越多地从单一智能体演变为连接组织各部分的智能体平台,我们的基础设施层也应该随之演进。所以,当我们审视目前用于实现智能体的各类工具时,会发现在诸如智能体 SDK 和内存等方面发生了很多创新。智能体 SDK 对于实现概念验证(POC)和快速上手非常酷,但它们不一定能帮助解决诸如连接组织周围分布式组件的问题。如果你想实现更复杂的智能体系统,你实际上需要所有这些东西。这就是你们在下面看到的这一层,你需要部署额外的基础设施。你需要编写像重试逻辑、恢复逻辑这样的东西,所有这些要弄好实际上都相当复杂,但在生产环境中运行长驻的、有状态的以及分布式的进程时,这些是完全必要的。

Original English

Speaker B: And so as our use cases are evolving more and more from single agents to agentic platforms that connect parts around uh the organization our infrastructure layer should also evolve with that. So when we look at the types of tools that are currently out there to implement agents, a lot of innovation has been done on sites such as agent SDKs and memory. And agent SDKs are really cool to implement PC's and get started quickly, but they don't necessarily help with like connecting the distributed bits around an organization. And if you want to implement more complex agentic systems, you actually need all of those things. So that is the layer that you see below here where um you have to deploy extra infrastructure. Uh you need to write things like retry logic, recovery logic and all of that is actually pretty complex to get right but completely necessary to run longunning stateful and distributed processes in production.

Speaker B: 所以今天我想谈谈一个名为 Restate 的开源框架。你可以把它看作是一个灵活且持久的基础设施,让你能够构建任何后端。所以它不仅仅是专门为智能体设计的,但由于智能体也只是一种后端类型,所以它对智能体同样非常适用。Restate 背后的理念来源于 Apache Flink(一个流行的分布式流处理引擎),以及 Meta 核心事件基础设施的一些前架构师。

Original English

Speaker B: So today I want to talk about an open-source framework called restate. And you can see it a bit as a flexible durable foundation that lets you build any backend. So it's not specific for agents but a as agents are also just a type of a backend uh it also works well for them. The ideas behind restate come from Apache Flink which is a popular distributed stream processing engine and also from some of the exarchitects behind Meta Score event infra.

Speaker B: 那么 Restate 有哪些组成要素呢?主要分为四个部分。首先,它确保智能体的单次运行具有弹性。这在业界被称为“持久执行”(durable execution)。设想一下这样的情况:当一个智能体运行了一周然后崩溃了,我们希望能把它恢复过来,并让它在失败的地方确切地继续运行,而不是从头重新开始。这里的另一个领域是并行运行许多并发会话。想象一下同时运行数千个并发的智能体会话,并且需要确保状态始终一致,还要保证不同的智能体之间不会相互干扰。然后再更进一步,比如智能体之间的通信、智能体与 MCP 服务器及其他工具之间的通信。最后,还有控制,要确保例如当智能体在做一些你不希望它继续做的事情时,或者当它卡住时,能够真正取消或终止该执行。

Original English

Speaker B: So what are the ingredients in restate? Basically four parts. First of all, it makes sure that a single run of an agent is resilient. This is called durable execution in the industry. Think about things like when an agent runs for a week and then crashes. We want to be able to bring it back and let it continue exactly at the point where it failed. We don't want it to start over from the beginning. Another um area here is running many concurrent sessions in parallel. Imagine running thousands of concurrent agent sessions at the same time and needing needing to make sure that state is always consistent and that different agents don't interfere with each other. And then going more towards things like communication between agents, between agents and MCP servers and other tools. And finally also control, making sure that when an agent for example uh is doing um something you don't want it to continue or when it's stuck being able to actually cancel or kill the execution.

Speaker B: 所以你们可以这样来理解它。Restate 基本上是一个运行在你们的智能体服务前端的服务器。作为一个独立的组件,它就坐在那里,有点像消息代理或者代理服务器。当有对你们智能体的请求时,Restate 会代理该请求并将其推送到服务,基本上从那一刻起,Restate 和智能体之间就建立了一个开放的连接,而这个连接基本上就像是智能体的生命线。所以当智能体在执行任务时,它会将事件发送给 Restate,而 Restate 会利用这个事件日志(journal)在发生故障后恢复该进程。

Original English

Speaker B: So the way that you can think of it is as follows. Restate is basically a server which runs in front of your agent service. So as a separate component it sits there a bit like a like a message broker or a proxy and when there's a request for your agent restate proxies the request to the service and pushes it to the service basically and from that moment there's a connection open connection between restate and the agent and that connection will basically be a bit like a lifeline for the agent. So as the agent is doing stuff, it sends events over to restate and restate will use that journal of events to recover the process after a failure.

Speaker B: 因此,从一个稍微更高层次的解释来看,你可以说它是把你应用程序中的普通函数变成了长驻的、持久的且有状态的东西,而无需你去处理很多本来需要为之做的复杂事情。所以,我今天的演讲将主要是一个演示。我将向大家展示一个连接到 Slack 的研究智能体。想象一下我们在某家公司工作,我们想要向所有的员工提供一个 Slack 智能体。

Original English

Speaker B: So from a slightly higher level um explanation, you could say that it's turning a normal function in your application into something that is longunning, durable, and stateful without having to do um a lot of the complex things you otherwise need to do for this. So my talk today will be mainly a demo. So I'll be showing you um a research agent that is connected to Slack. Imagine we are like working at some company and we want to make an Slack agent available to all of our employees.

Restate Slack Agent Demo

Speaker B: 如果我现在进入 Slack,我就可以在比如这个频道里问:“AI 领域有什么新动态?”。现在让我们来看看它在底层是如何运作的。如果我回到这里,这是 Restate 的用户界面(UI)。这有点像你们智能体的驾驶舱。你可以看到当前所有已注册智能体的注册表,你也可以看到,例如,当前正在发生哪个执行操作。所以这里就是我几秒钟前启动的深度研究(deep research)智能体。我们可以看到它当前正在做什么。刚才它首先进行了一次 LLM 调用,然后它通过 Slack 给我发送了回答。这第一次的 LLM 调用是一个规划者(planner)智能体。所以它所做的就是规划研究,并发送给我一个它想要研究的子话题列表。

Original English

Speaker B: So if I go here into Slack then can I can here in this channel for example ask what is new in AI. Now let's have a look at what it's doing under the hood. So if I go back here, I have here the restate uh UI. This is a bit like a cockpit for your agents. So you can see a registry of all the agents that are currently registered and you can also see for example which execution is currently happening. So here is the deep research agent that I spinned up a few seconds ago. We can see what it's currently doing. Now it called first an LLM and then it sent me an answer via Slack. This first LLM call was a planner agent. So what it did is it planned the research and sent me um a list of subtopics that it wants to research.

Speaker B: 现在如果我在这里点击批准(approve),这将会解除工作流的阻塞,并启动一组并行的研究智能体。所以这基本上就像是经典的深度研究工作流,对吧?你有一个规划者,然后是一组子研究智能体,最后有一个根据这些内容撰写报告的智能体,比如写手智能体。所以你们在左边看到的这个日志,基本上就是从智能体发送到 Restate 服务器的事件。如果这时候系统在某个点崩溃了,这个日志将被用来把执行恢复到它失败的那个点。我不知道这里是不是有一些错误。我在里面注入了一些类似工具调用的错误。是的,例如你可以看到在这里,子智能体首先进行了一次 LLM 调用,然后开始进行一些网络搜索。最终,其中一个网络搜索没有成功,因为 API 宕机了,然后你在右边可以看到它是如何被重试,并最终成功完成的。所以它不是从头重新开始,而是利用日志来恢复进度。

Original English

Speaker B: Now if I press here approve then this will unblock the workflow and will spin up a set of parallel research agents. So this is basically like the classical deep research workflow, right? You have a planner then a set of subress research agents and then finally someone uh who writes a report on this like a writer agent and so this journal you see here on the left that is basically the events that get sent from the agent to the restate server and if this now crashes at some point this journal is what will be used to uh recover the execution to the point where it failed. I don't know if uh there were some errors. I injected a bit of like tool errors in here. Yeah, here you can for example see that um the sub agent first did an LLM call then started doing some web searches and eventually uh one of the web searches didn't go through because the API was down and then you see here on the right how it got retrieded and eventually completed successfully. So instead of starting over it uses the journal to recover the progress.

Speaker B: 现在让我们看看这在代码中是什么样子的。在 Restate 中实现应用程序的基本单元是编写 HTTP 处理程序(handlers),并且通过使用 Restate SDK 使这些处理程序变得持久化。所以在本例中,这里有我们的深度研究处理程序。在这里,作为第一个参数,我们有一个 Restate 上下文对象(context object)。你可以把它想象成连接到那个 Restate 服务器的连接。每当我对这个 Restate 对象执行一个动作时,它都会导致一个事件被发送到 Restate。所以例如,当我进行那次规划者 LLM 调用时,在底层实际发生的是它执行了这个 Python 函数。这只是一个简单的轻量级、类似 LLM 调用的操作,而我让它变得持久化的方法是把它包装起来——

Original English

Speaker B: Let's now have a look at what this looks like in code. So the basic unit of how you implement applications in restate is by writing HTTP handlers and those handlers become durable by using the restate SDK. So here in this case we have here our deep research handler and here as a first argument we have a restate object context and the way you can imagine that is basically as that uh connection to that restate server. whenever I do an action on this uh restate object, it will lead to an event being sent to restate. So for example, when I did that planner LLM call, what actually happened under the hood was it executed here this Python function. This is just a simple light um light lm like LLM call and the way I made it durable is by wrapping it

Restate.run 的持久化执行机制与故障恢复

Speaker: 在 Restate.run 中,实际发生的底层机制是,通过执行这些持久化的步骤,如果在这个执行过程中的某个地方发生了失败,无论是在两小时后还是甚至两个月后,它都将会精确地恢复到那个发生故障的特定状态点。这就是所谓的持久化执行(durable execution)的核心理念。也就是说,你将始终能够将一个进程完好无损地恢复到它原先所在的那个确切位置。

Original English

Speaker: in restate.run. So what happens is by doing these durable steps if this fails somewhere here two hours or two months later it will recover to exactly that point. So that's the idea of durable execution. You're always able to recover a process to where it was.

Speaker: 这种强大的能力也可以被用于其他更多的用途,而不一定非要仅仅局限于故障恢复场景。举个例子来说,假设我们想要请求一个人工操作来批准某件特定的事情,而这个批准的过程可能非常漫长,甚至可能需要耗费几周或是一个月的时间。这就意味着这个进程需要能够在这种长达数周的漫长时间跨度内,安全地经受住系统的多次重启以及重新部署的考验。因此,借助持久化执行这种机制,你实际上还可以选择主动暂停一个函数的运行,并让它在条件允许、能够继续推进进展的时候再重新恢复执行。

Original English

Speaker: You can also use that for other things not necessarily for failure recovery. For example, imagine we want to ask a human to approve something and this approval might take weeks or a month. this process needs to be able to um to survive restarts and redeploys over those kind of long periods of time. And so with durable execution, you can actually also uh suspend a function and let bring it back when it's able to make progress.

Speaker: 因此,在一个具体需要人工批准的实际用例中,我们在这里所做的操作,从根本上说,实际上是去创建一个被存储在该日志记录中的“持久化 Promise”,它发挥的作用有点像是一个程序的挂起点(suspension point)。然后,正如我在一开始为您展示的那样,我们会在 Slack 应用程序中请求人工去点击那个确认按钮。在我们在漫长等待响应的这段过程中,这个执行的进程实际上是完全处于暂停状态的。因此,如果它当前是运行在一个 Serverless(无服务器)架构上的话,这种持续的等待是绝对不会消耗我们的无服务器函数执行时间的。一旦最终接收到了用户的响应,这个进程就会立刻被解除阻塞状态,并且可以精准地从它之前暂停退出的地方继续恢复执行。

Original English

Speaker: So in the case of a human approval, what we do here is basically we we create a durable promise which lives in that journal a bit like a suspension point. Then we ask uh a human to click that button in Slack as I showed in the beginning and while we are waiting this process actually suspends. So if it's running on serverless this is not using uh execution uh time on our functions. Once the response comes in this then gets unblocked and can continue where it left off.

从传统工作流到持久化有状态实体

Speaker: 那么,我们目前在这里所看到的东西,实际上有点像是一个传统的工作流(workflow)。它是由一组能够被系统持久化执行的离散步骤所组成的。但是,当我们开始深入思考智能体(Agents)的概念,以及回想起 Andrej Karpathy 在他的推文中对智能体所做的那种描述时,你会发现它更应该像是一个能够存活较长一段生命周期、并且具备一定记忆和上下文能力的持久化、有状态实体(persistent stateful entity)。嗯,所以,单纯使用传统的工作流模式来对这种具有复杂状态的实体事物进行建模,并不是一种最理想和最优雅的方式。

Original English

Speaker: So what we see here is a bit like a workflow. It's a set of steps that get executed durably. But when we think about agents and also the way that Karpathy described it in the tweet, it's more like a persistent stateful entity that lives for a longer period of time that has some memory. Um, so a workflow is not the nicest way to model this kind of thing.

Speaker: 那么,在 Restate 架构中,我们能够用来对这种复杂实体进行高级建模的方法,是去使用一种被系统称之为“虚拟对象”(virtual object)的独特概念。所以请想象一下,在我目前正在向大家展示的这个 Slack 研究调研智能体的用例场景中。想象一下,我并不想干巴巴地等待整整 10 分钟后,才给它提供一些后续的补充上下文信息;或者可能我突然在脑海中想到了其他我早该提前告诉它的重要事情。嗯,我希望我实际上能够直接实时地与它进行交互对话,而不是必须苦苦等到那项漫长的调研任务彻底完成之后,我才能有机会发送一条后续的跟进信息。

Original English

Speaker: So the way that we can model this in restate is by using something called a virtual object. So imagine in the use case that I'm showing this Slack research agent. Imagine that I don't want to wait for 10 minutes to give it some follow-up context or maybe I think about something else that I should have told it. Um I want to actually be able to interact with it, not wait till that research is finished before I can send a follow-up.

Speaker: 所以,这基本上就是 Restate 框架中“虚拟对象”的核心本质。它有点类似于一个拥有自身状态的 Actor 模型(stateful actor)。它拥有一个系统内唯一的标识符 ID,例如一个特定的会话 ID(session ID)。它同时还拥有一些专门针对该特定会话环境进行隔离存储的键值对状态(key value states),你可以自由地向其中写入和更新数据。呃,想象一下,例如它里面保存了你的历史消息记录;而且它同时还拥有类似一组专门的处理程序(handlers),这些处理程序可以专门为这个当前的会话环境去执行那些持久化的函数。

Original English

Speaker: And so this is basically what a virtual object in restate is. It's a bit like a stateful actor. It has a unique ID, for example, a session ID. It has uh some key value states that is isolated for that specific session that you can write to. Uh imagine for example your history of messages and it also has like a set of handlers that can execute durable functions uh for this session.

Speaker: 那么在这里,我用来实现这个我刚才提到的“与正在运行的后台进程进行交互”的用例的具体方式,如下所示。这是一个……嗯……这有点像是一个会话控制器(session controller)。再一次地,它依然拥有这个 Restate 对象上下文环境供其灵活使用,从而能够以一种随时可恢复的可靠方式去执行各种操作。呃,它可以向这个特定的会话状态存储库中去写入数据。在这里的代码中,我正在检索查询过往的聊天记录。而在那里,有一点非常有趣也十分关键,那就是为了能够以极高并发的并行化方式来运行这类会话——比如在同一时间节点同时运行成千上万个活跃会话——我们就必须绝对确保各个不同的智能体之间不会产生相互干扰和数据竞争。

Original English

Speaker: So here the way I implemented this use case that I mentioned of interacting with a running process is as follows. This is a um a bit a session controller. Again, it has like this restate object context at its disposal to do things in a recoverable way. Uh it can write to this session store. Here it I'm retrieving the chat history. And one thing that's interesting there is that in order to run these kind of sessions in very high uh parallelized ways, so thousands of sessions at the same time, we need to make sure that agents do not interfere with each other.

Speaker: 想象一下这样一种糟糕的情况,我正在 Slack 界面上连续发送两条消息,而此时系统底层有两个不同的智能体进程,它们实际上正在互相恶意覆盖彼此的会话状态数据。为了从根本上防止那种数据损坏的情况发生,这个底层机制将提供严格保证,确保在同一个会话中,同一时间点只能有唯一一个执行进程在运行。因此,任何并发触发的第二个执行任务,都将会被系统自动排队在当前正在运行的这个任务的后面。接下来,让我们更深入地看一看我们究竟是如何在代码层面实现这种“与另一个正在运行的执行进程进行交互”的机制的。

Original English

Speaker: Imagine I'm sending two messages on Slack and now two agents are actually overwriting each other each other's session state. To prevent that, this will guarantee that only one execution is running at a time. So a second execution will be cued behind the current one. Then let's have a look at how we implement this like interacting with another execution.

Speaker: 因此,在 Restate 系统中,任何一个执行任务都会拥有一个全局唯一的标识符,你可以灵活地使用这个唯一标识符从其他的系统进程中去安全连接到它。例如,可以用它去检索和获取那个任务的执行输出结果,但与此同时,你也可以使用它去取消那个执行任务,或者也许是通过将一点新的状态数据动态注入到一个已经正在运行的智能体事件循环(agent loop)中,以此来向它发送控制信号。因此,这可以被视为是一种极其灵活的功能和能力类型,你可以利用它来实现诸如向一个已经正在运行的、长时间持续的智能体循环发送中断或更新信号等高级操作。

Original English

Speaker: So an execution in reset has a unique identifier and you can use that identifier to connect to it from other processes. for example, to retrieve uh the output, but also to cancel it or maybe to signal it being injecting a bit of state into an already running agent loop. And so this is like a very flexible type of um uh capabilities that you can do to implement things like for example signaling an already ongoing agent loop.

与运行中的智能体进行高级指令交互

Speaker: 所以我们在这里实际编写的逻辑是,如果当前已经有一个正在持续运行的执行任务,那么我们就会去询问一个大语言模型(LLM):这个新进来的消息,是否像是与当前正在运行的这个智能体循环高度相关的事情?如果情况确实如此的话,那就通过发送信号(signal)的方式将这个新状态注入进去;如果它与我们当前正在重点处理的调研任务并不真正相关,那么就果断取消你当前正在执行的所有任务,并利用这条刚刚带来的新信息,从头开始启动一个全新的流程。因此,这种架构设计实际上比传统的工作流引擎走得要远得多。它更倾向于引领大家去编写那些持久化的、包含内部状态的实体,这些实体能够彼此之间进行丰富的通信交互,并且拥有可供它们随时调用和查询的内部记忆。

Original English

Speaker: So what we do here is if there is a current execution ongoing then we will ask an LLM is this like something that is relevant for the current agent loop. If that is the case inject this via a signal if it's not really relevant for what we're currently doing then cancel what you're currently doing and start over again with this new information. And so this goes a little bit further than workflows. it goes a bit more towards like writing persistent stateful entities that can interact with each other and have memory at uh their disposal.

Speaker: 那么,请允许我亲自向你们演示一下这个机制在实际操作中是如何巧妙运作的。所以在这里的界面上,如果我现在再次抛出一个宽泛的问题:“人工智能(AI)领域目前有什么新的进展”,然后我耐心等待那么几秒钟,接着它应该会再次利用一个结构化的计划来作为对我的回应。嗯,然后接下来我就可以说,例如,补充一些额外的上下文信息:“请把调研重点专门集中在前沿模型(frontier models)上”,比方说就这样。所以,一旦我已经生成并拥有了那个初始的调研计划,我就将会把那一点额外补充的状态信息注入进去。现在让我们仔细看一看用户界面(UI),看看这个系统目前到底在后台做些什么。所以在这里,我拥有了我刚刚向大家展示过的那个控制器逻辑。它已经开始主动调用一个 LLM,来专门针对我这个新提交的输入信息进行意图分类。

Original English

Speaker: So let me show you uh how this works. So here if I now ask again what is new in AI and I wait a few seconds then it should respond again with a plan. Um and then I can say for example some extra info focus on frontier models let's say. So once I have the plan I will inject that bit of extra state. Now let's look at the UI of what this is now doing. So here I have that controller which I just showed. it started calling an LLM to classify uh this new input.

Speaker: 一旦这个意图分类的响应结果返回了,它很可能会智能地判定出它应该发出一个更新信号,因为这条关于前沿模型的新消息,仍然与它目前正在积极进行的这项 AI 领域调研任务息息相关。所以,这个操作就会非常顺畅地将那条新的干预消息直接注入到那个正在持续运行的智能体事件循环之中。让我再次在这个进行深度调研的智能体界面中给你们演示一下这个流程。嗯,所以我们首先看到它调用了一个 LLM 然后它向我们进行了反问,接着我们注入了这条诸如“重点关注前沿模型”的全新指令消息,然后它敏锐地将那个新指令纳入了它的考量范围之中,并带着这个新的上下文重新开始了它的调研流程。

Original English

Speaker: Once this comes back, it will probably decide that it should signal it because it's it's still relevant to the research it's currently doing. So this injects that new message into the ongoing agent loop. So let me show you in the deep research agent again. Um so we first it called an LLM then asked us then we injected this uh new message of focus on frontier models and then it uh took that into account and started over again.

Speaker: 在这里,我现在同样也可以随意地说一些类似这样的话,比如……呃……“忘掉那个关于 AI 政策的调研吧”。而如果我点击发送这个中断指令,那么系统中的协调器(Coordinator)组件就会果断决定去取消当前正在进行的那次漫长的运行任务,并立即启动一个全新的运行实例,这个新实例将会去专门调研这个全新的主题。因此,这种基于信号的取消机制,基本上就像是一个强有力的控制信号,它会被……嗯……沿着执行过程的堆栈,或者说沿着整个调用链(call chain),一直向下级联发送。

Original English

Speaker: Here I can now for example also say something like uh forget about that research AI policy. And if I send this then the coord coordinator will um decide to cancel the ongoing run and start a new one that will research this new topic. And so this cancellation is basically like a signal that gets um sent down the stack of co or the call chain.

Speaker: 因此,如果我的主智能体在之前已经繁衍并启动了一些子智能体(sub agents),那么首先,那些底层的子智能体将会被依次优雅地取消,然后才是……呃……控制器它本身被取消,就像那样,它基本上实现了对整个调用栈的完美回溯,并且这也同时赋予了各个智能体去进行状态回滚(roll back)的强大能力。好的。好的,那么这部分内容实际上在“拥有状态和持久化能力,且允许我们在更长的时间周期内与其进行交互的实体”这个方向上,探索得更深入了一些。

Original English

Speaker: So if my agent was already spinning up sub agents first those sub aents would be cancelled then uh the controller itself and like that it would basically rewind the stack and give agents also the ability to roll back. Okay. Okay, so this went a bit more into the direction of like stateful persistent entities that we can interact with over longer periods of time.

构建高度定制化和弹性的分布式应用程序

Speaker: 现在,我想要给大家展示的这个技术演示的最后一个关键部分,是……嗯……我们正进一步迈向能够编写具有极高高度定制化特性的应用程序。想象一下这样一个真实的商业场景,我们将这个系统成功部署在了生产环境之中,但仅仅几个月之后,一家全新的模型提供商横空出世,推出了一款全新的基座模型,例如“Fabulous”模型,而且尽管这个新模型的性能非常出色,但它的 API 调用价格也非常昂贵,于是我们突然注意到,我们部署的这个调研智能体实际上开始耗费大量的资金成本。

Original English

Speaker: Now the last part of the demo that I want to show is um going more towards like being able to write highly customized applications. Imagine that we deploy this in production but then a few months later a new model provider brings out a new model for example fabulous and even though the model is very good it's also very expensive and we notice that this research agent is actually starting to cost a lot.

Speaker: 像这类在某个庞大项目进行到一半时突然跳出来让你措手不及的棘手问题,往往会要求你去强行部署大量的、全新的额外基础设施,或者说要求你必须去找到一个完美的架构方案来彻底解决这个成本痛点。这正是 Restate 框架真正能够大显身手、卓越出众的领域。它并不会真的把你死死限制在某一种特定的、刻板的编写应用程序的方式上。它基本上是为你提供了一个类似于持久化编程模型(durable programming model)的底层基座,让你能够以最适合你自己业务需求的方式去实现一个复杂的应用程序,并且在必要的时候,能够极其方便地对其进行各种扩展。

Original English

Speaker: These kind of uh things that pop up halfway through a project require you to then deploy a a lot of new extra infra or like find a good way to solve this. This is the kind of things that restate really excels at. It doesn't really peg you into a specific way of how you should write your application. It basically gives you like a durable programming model that lets you implement an application in the way that fits for you and also extend it if necessary.

Speaker: 所以,首先,我在第一个初始示例中向大家展示的这个关于 LLM 的调用,仅仅是作为一个代码中内联的步骤(inline step)存在的。它仅仅只是一个被系统成功持久化了的普通 Python 函数而已。但是,请你想象一下这个更加复杂的用例:我们实际上希望能够对那些昂贵的 LLM 调用拥有更多、更细粒度的控制权。例如,你能做到的一件事,就是将这个调用逻辑提取出来,放入它自己独立的专属处理程序(handler)之中。而这个独立的处理程序现在就可以执行一些额外的逻辑了,比如,它可以先进行一次严格的资源分配策略检查,然后再去……呃……执行那个实际的 LLM 请求调用。

Original English

Speaker: So first I showed this um LLM call in the first example as an inline step. It was just a Python function that got persisted. But imagine this use case that we want to actually have a bit more control over those LLM calls. For example, what you can do is then pull this out into its own handler. And this handler can now do things like for example a policy check and then uh do the LLM call.

Speaker: 而且,系统中的其他智能体,现在不再需要将这个 LLM 调用作为自己代码内部的内联执行步骤了,它们现在可以完美利用 Restate 提供的那种诸如分布式通信原语(distributed communication primitives)等高级功能,来直接向这个专门的 LLM 网关发起远程调用,而不是把这种沉重的调用任务当成一个内联步骤来亲自处理。而这个允许你在不同的智能体之间进行顺畅通信的服务网格(service fabric),同时也为你提供了一些非常有价值的附加功能……嗯……比如像流量控制(flow control)这样的机制。因此,我们就可以非常轻松地设置诸如这样的规则,例如声明:某一个特定部门在同一时间内,最多只被允许向这个昂贵的 LLM 网关并发运行 300 次调用。

Original English

Speaker: And the other agents instead of doing this LLM call inline can now use restates like distributed communication primitives to actually just call this LLM gateway instead of doing it as an inline step. And this service fabric that lets you communicate between agents also gives you some um things like flow control. So we can for example say one department is only allowed to run 300 calls to this LLM gateway at the same time.

Speaker: 所以,我专门花时间展示这个高级特性的根本原因,仅仅是为了向你们稍微展示一下类似这样强大的功能架构。呃……它从本质上来说,就是一个极具弹性和韧性的底层基础设施基座。它能够确保你的复杂进程不仅能从普通故障中恢复,还能从那些更高级别、更严重的底层基础设施故障中安全恢复过来,比如网络分区(network partitions)和令人头疼的僵尸进程故障(zombie failures)。而且……嗯……随着你的业务用例的不断增长和复杂化,它为你提供了极为丰富的工具链,以便你进行自由扩展和深度定制。

Original English

Speaker: So the reason why I showed this was just to show you a bit like that. Uh it's basically just a a resilient foundation. It makes sure that your process can uh recover from even the more advanced types of infrastructure failures, things like network partitions and zombie failures. And um it gives you like tooling to extend and customize as your use case grows.

Speaker: 让我们重新回到之前的幻灯片演示,以便让我们对这个框架在底层内部实际上是如何实现和运作的,能有一个更加深入一点的概念和了解,因为这实际上代表了一个相当有趣且充满智慧的……嗯……系统设计或者说架构模式。所以,它之所以能够被成功实现,其根本途径基本上是建立在一个事件驱动的分布式日志(event-driven distributed log)的实现基础之上的。所以在整个系统架构的“黑盒子”内部,你基本上是在网络边界的一侧拥有各种客户端,而在网络的另一侧……

Original English

Speaker: Let's go back to the slides to have a little more of an idea of how this thing is actually implemented on the inside because it's actually a pretty interesting um design or architecture. So the way it's implemented is basically by having a a event-driven distributed log implementation. So inside the box you basically on one side have the clients on the other

Restate 架构与特性

Speaker A: 在服务外部以及盒子内部有一个日志,它持久化所有这些流水账事件,还有一个事件循环。该事件循环基本上根据事件的类型从服务中获取事件。它要么在嵌入的状态存储中持久化一些状态,要么设置一个定时器,要么向另一个代理发送请求。通过这样做,你基本上就为应用程序正在执行的任何任务提供了一个持久化的基础。

Original English

Speaker A: side the services and inside the box is a log which persists all those journal events and an event loop and that event loop basically gets the events from the service based on what the event is. It either persists some state in the embedded state store or it sets a timer or it sends a request to another agent. And by doing that you basically have a durable foundation for whatever an application is doing.

Speaker A: 这个分布式日志的设计很大程度上受到了 Meta 核心事件基础设施层工作方式的启发。这基本上可以算是它在此基础上的一次迭代。当时的一些架构师现在已经为 Restate 设计了它,作为一个可以在开源中获得的更通用的解决方案。

Original English

Speaker A: The design of this distributed log is heavily inspired by the way that the core event infra layer at meta works. Uh it's basically like an iteration on top of that. Um and some of those architects are now have designed that for restate as an more generic solution that is available in open source.

Speaker A: 与这种架构相关的有两件重要的事情让它变得很有趣。第一点是它采用推送模型(push model)运作。因此,虽然大多数工作流编排器实际上会轮询新任务,从工作流服务器拉取(pull)任务,但 Restate 实际上是推送这些调用。你从中得到的好处是它具有低得多的延迟。所以你可以在你应用程序周围的函数中使用这种工作流保证,并且对于一个比如 10 步的工作流,其 P99 延迟可能只有 45 毫秒。

Original English

Speaker A: There are two important things related to this architecture that make it interesting. The first one is that it works as a push model. So whereas most workflow orchestrators actually pull for new tasks um for pull from the workflow server, restate actually pushes the invocations and the benefit you get from that is that it has a much lower latency. So you can use these kind of workflow guarantees in functions around your application and uh have like a latency of for example 45 milliseconds p99 for like a 10-step workflow.

Speaker A: 推送调用对于 Serverless 架构也非常适用,因为它们通常要求你发送请求并唤醒函数。所以我在这里展示的这个设计包含了你需要的一切。它还包含了我们用来将状态作为 UI 嵌入的那个状态存储。这是一个单一的二进制文件,所以它在操作上非常容易,也能很方便地以高可用方式运行。你只需要多次启动它,并让它将快照保存到对象存储中。

Original English

Speaker A: Pushing invocations also works very well for serverless because they require you to basically uh send the request and wake up the function. So this design that I show here includes everything you need. It includes uh as well that state store where we were embedding the state as the UI. It's a single binary so it's pretty easy to operate as well to run it in like a highly available way. You just spin it up multiple times and let it snapshot to object storage.

Speaker A: 目前 Restate 有六种不同的 SDK。我们还为目前最流行的大多数代理(Agent)框架提供了集成。当然,因为它就像一个灵活的层,你也可以只使用任何大语言模型(LLM)的 SDK,并通过将一些步骤包装到这些 SDK 结构中来实现自定义代理。所以,它是开源的。你可以自行托管它。我们也有自带云(BYOC)的方案,我们将 Restate 部署在你的云账户中,这给你带来的好处是数据不会离开你的云账户。此外,我们还有一种托管云产品。

Original English

Speaker A: So restate has six different SDKs. We also have integrations for most of the popular agent frameworks out there. And of course, because it's just like a flexible layer, you can also just use any LLM SDK and implement custom agents by just wrapping some steps into uh these SDK constructs. So, it's open source. You can self-host it. We also have a BYOC offering where we deploy restate in your cloud account and uh that gives you the benefit that data doesn't leave your cloud account. Otherwise, there's also a managed cloud offering.

Speaker A: 这主要就是我想展示的内容。如果你想进一步探索代码,这里有 GitHub 仓库。它是公开的。如果你喜欢这个项目,那么去看看 Restate 仓库本身吧。我们在全面招聘各种职位,从工程开发到市场营销,特别是我们旧金山湾区这里。所以,如果你对这方面感兴趣,请务必查看我们的职业页面。如果你想问任何问题或了解更多关于 Restate 的信息,我会在会议厅外面的前面。非常感谢。

Original English

Speaker A: This was mainly what I wanted to show. If you want to explore the code a bit further, there is here this the GitHub repo. It's publicly available. If you like the project, then have a look at the restate repo itself. We are hiring across the board for all sorts of roles going from engineering to marketing especially also here in the Bay Area. So, if you're interested in that, uh, then definitely check out our careers page and I will be outside in front of the conference hall here if you want to ask any questions or learn more about restate. Thank you very much.

Resonate:从实现到规范的转变

Dominic Turno: 在 2026 年,编程代理(coding agents)将悄然淘汰它们最初的软件平台。不是因为它不好,仅仅是因为平台变得不必要了。我是 Dominic Turno。我是 Resonate 的创始人兼 CEO。Resonate 是一个持久化执行平台,以极简主义和简单性作为其核心技术价值观进行构建,而这些特性将在本次演讲中扮演核心角色。

Original English

Dominic Turno: In 2026, coding agents will quietly retire their first software platform. Not because it's bad, simply because the platform is unnecessary. I am Dominic Turno. I am founder and CEO of Resonate. Resonate is a durable execution platform built with minimalism and simplicity as its core technical values and these properties will play a central role in this talk.

Dominic Turno: 在 Resonate,我们对软件工程的发展方向有一个工作理论:通用的实现(implementations)将越来越多地被按需生成的定制实现所取代。它们不会作为一个新的库、新的框架或新的平台出现,而是作为已经就位的现有基础设施的一种最小化扩展。

Original English

Dominic Turno: At Resonate, we have a working theory where software engineering is headed. Generalpurpose implementations will increasingly be replaced by bespoke implementations generated on demand. Not as a new library, a new framework or a new platform, but as a minimal extension of the infrastructure that is already in place.

Dominic Turno: 如果这个理论成立,重用(reuse)将向上游移动。我们将不再重用通用目的的实现,而是重用一份规范(specification),并从中推导出一个定制的实现。事实上,我们可以为现有的基础设施构建许多量身定制的实现。我们只需向代理提出要求即可。到了这个时候,提示词(prompt)就成了一个平台。

Original English

Dominic Turno: If this theory holds true, reuse will move upstream. Instead of reusing a general purpose implementation, we will reuse a specification and we will derive a bpoke implementation from it. In fact, we can build many bespoke implementations tailor made for the infrastructure that is already in place. We just have to ask the agent. At this point, the prompt is a platform.

Dominic Turno: Resonate 是一个双重执行平台。我们有一个 Resonate 服务器的实现。我们也有适用于 TypeScript、Python、Rust、Go 和 Java 的 Resonate SDK 的实现。所以,我们必须问自己:这种新的现实对我们意味着什么?如果实现变得可以自动生成,我们的价值体现在哪里?我们的答案是:我们的价值从实现转移到了规范上。

Original English

Dominic Turno: Resonate is a dual execution platform. We have an implementation of the Resonate server. We have implementations of the Resonate SDK for TypeScript, Python, Rust, Go, and Java. So, we have to ask what does this new reality mean for us? If implementations become generatable, where does our value live? And our answer our value moves from implementation to specification.

Dominic Turno: 这改变了我们看待 Resonate 的方式。产品不再是具体的实现代码。产品是规范、是协议,我们希望从这个协议中推导出多个服务器的实现。其中一个是我们通用的 Resonate 服务器,也就是我们的参考实现。其他的则是与基础设施合作伙伴共同构建的实现。对于客户和合作伙伴而言,这意味着在他们现有的基础设施之上,以最少的额外依赖原生获得持久化执行能力。

Original English

Dominic Turno: Now this changes how we think about Resonate. The product is no longer the implementation. The product is the specification the protocol and from that protocol we want to derive multiple server implementations. One is a general purpose resonate server, our reference implementation. Others are implementations built with infrastructure partners. For customers and partners, this means durable execution right on top of their existing infrastructure with minimal additional dependencies.

Dominic Turno: 因此,问题不再是我们能否构建一个服务器。问题变成了:我们能否从同一份规范中,重复地合成出可信的服务器?如果可以,那该怎么做?当我们谈论代理工程(agentic engineering)时,我们将所有的注意力都集中在验证上。我们如何知道结果是正确的?但是今天,我想把焦点放在规范上。更重要的是,代理如何能够参与到系统的规范设计中,而不仅仅是构建或验证它。

Original English

Dominic Turno: So the question is no longer can we build a server. The question is, can we repeatedly synthesize trusted servers from the same specification? And if so, how? When we talk about agentic engineering, we focus all of our attention on verification. How do we know the result is correct? But today, I want to focus on the specification instead. And more importantly, how can agents participate in specifying the system, not just building or verifying it.

Dominic Turno: 现在,Resonate 正在与多个基础设施提供商合作,将持久化执行原生引入他们的技术栈。其中之一是 Synadia,即 Nats.io 背后的公司。Nats.io 是一个旨在构建现代分布式系统的开源消息系统。在接下来的演讲中,我们将使用在 Nats.io 上运行的 Resonate 来探索我们的代理工程实践。

Original English

Dominic Turno: Now, Resonate is partnering with multiple infrastructure providers to bring durable executions natively to their technology stack. One of them is Senadia, the company behind Nats.io, an open-source messaging system designed for building modern distributed systems. For the rest of this presentation, we will use Resonate on Nats.io to explore our agentic engineering practices.

从规范到实现的代理工程

Dominic Turno: 我们如何从规范走向实现?首先,我们需要统一我们的心智模型。这张图是代理编写代码时的常见视角。这里有一个代理,有一个规范,然后有一个实现。对于许多应用程序来说,这就足够了。但这对于我们想要做的事情来说是不够的,因为我们不想只是从一份规范生成一个实现。我们想要从该规范中生成多个特定于目标的实现。

Original English

Dominic Turno: How do we go from specification to implementation? First, we need to level set our mental model. This picture is a common view of agent decoding. There's an agent, there's a specification, and then there's an implementation. And for many applications, that is enough. But it is not enough for what we are trying to do because we are not trying to generate one implementation from a specification. We are trying to generate multiple target specific implementations from the specification.

Dominic Turno: 因此,规范绝不能将实现的任何方面考虑在内。规范绝不能假定一个具体的数据库模式(schema)或具体的索引。规范甚至不能假定有表和事务的关系型数据库存在。它不能假定是一个键值存储。它不能假定是弱一致性。它也不能假定是强一致性。规范必须是抽象的。只有实现才必须是具体的。

Original English

Dominic Turno: So the specification must not take any aspect of an implementation into account. The specification must not assume a concrete database schema or concrete indices. The specification must not even assume a relational database with tables and transactions at all. It must not assume a key value store. It must not assume weak consistency. It must not assume strong consistency. The specification must be abstract. Only the implementation must be concrete.

Dominic Turno: 所以我们要求代理遵循抽象规范并生成一个具体的实现。具体来说,起初我们让代理在 Postgres 上用 Rust 构建一个 Resonate 服务器,结果代理失败了。抽象规范和具体实现之间的差距太大了。代理生成了一个只能在完美路径(happy path)上工作的系统。它通过了基础测试,但并不正确。它在并发处理上崩溃了。在进程故障时崩溃了。在网络故障时崩溃了。这个实现更像是一个原型,而不是一个生产级别的系统。

Original English

Dominic Turno: So we ask the agent to follow the abstract specification and generate a concrete implementation. Specifically at first we ask the agent build a resonate server in rust on top of posgress and the agent failed. The gap between the abstract specification and the concrete implementation was too large. The agent generated a system that worked on the happy path. It passed the basic tests, but it was not correct. It broke on the concurrency. It broke on the process failure. It broke on the network failure. The implementation was closer to a prototype, but not a production system.

Dominic Turno: 因此,我们修正了流程。我们不再要求代理直接从抽象规范跳转到具体实现,而是插入了一个中间产物:具体规范(concrete specification)。这个具体规范是与代理通过交互式推导出来的。但是人类是主要的驱动力。对于 Postgres 来说,这意味着要明确特定于目标的决策:数据模式、索引、SQL 查询以及事务边界。一旦这些决策被写下来,代理确实就能够实现出那个生产级别的系统了。

Original English

Dominic Turno: So, we amended the process. Instead of asking the agent to jump directly from abstract spec to concrete implementation, we inserted an intermediary artifact, the concrete specification. That concrete specification was derived interactively with the agent. But the human was the main driver. For posgress that meant making target specific decisions explicit, the data schema, the indices, the SQL queries, the transaction boundaries. Once those decisions were written down, the agent was indeed able to implement the production system.

Dominic Turno: 这个方法奏效了,但它也暴露了局限性。代理帮助我们构建了系统,但是代理并没有帮助我们设计系统。如果规范本身是一个可重用的产品,那这显然是不够的。接下来的步骤显而易见。代理必须向上游移动。但怎么做呢?

Original English

Dominic Turno: So this worked, but it also revealed the limitations. The agent helped us build the system, but the agent did not help us design the system. And if the specification is a reusable product, then that's not enough. Now the next step is obvious. Agents have to move upstream. But how?

Dominic Turno: 当我们开始在 Nats.io 上构建 Resonate 时,我们改变了问题。我们没有问代理能否构建这个生产系统。相反,我们问代理需要什么才能先设计系统,再构建系统。因此,我们给了代理一个确定性的模拟环境。我们还给它布置了一个不同的任务:不要构建生产系统。构建一个模拟的实现。这个模拟实现并不是最终产品。它是一种可执行的设计(executable design)。它的目的是在偏序条件下发现正确的算法

Original English

Dominic Turno: When we started building Resonate on Natio, we changed the question. We did not ask can the agent build the production system. Instead we ask what does the agent need in order to design the system first and build the system second. So we gave the agent access to a deterministic simulation environment. And we gave it a different task. Do not build the production system. Build a simulated implementation. The simulated implementation is not the product. It is executable design. Its purpose is to discover the correct algorithm under partial order

System Design with Simple Primitives

Speaker A: 在部分故障的情况下。一旦这些算法在模拟中被发现、测试和验证,我们就会要求智能体编写具体的规范。只有在那之后,我们才会要求智能体编写生产环境的实现代码。因此,这个过程变成了:抽象规范、模拟实现、具体规范,然后是具体实现。正是在这个节点上,智能体向开发流程的上游移动了。人类仍然参与设计过程,但现在智能体成了主导者。有两个要素使这成为可能:极简主义和简单性。不幸的是,极简主义和简单性并不是起点,它们是终点。我们花了三年时间让协议变得更小、更简单。每次遇到问题,我们都会问:“我们可以去掉什么?我们可以抹除什么抽象概念?我们可以移除什么属性?我们可以打破什么关系?”结果就是一个非常小的协议,它围绕着两个对象展开:一个持久化承诺(durable promise)和一个持久化任务(durable task)。这种简单性很重要,因为即使是简单的并发分布式协议也有着复杂的状态和行为空间。换句话说,即使是在几个简单的原语之上实现简单的协议也是很困难的。让我们用 nats.nats 来具体说明这一点,它提供了……

Original English

Speaker A: under partial failure. And once these algorithms are discovered, tested and verified in simulation, then we ask the agent to write the concrete specification. And only then do we ask the agent to write the production implementation. So the process becomes abstract specification, simulation implementation, concrete specification and then concrete implementation. This is a point where the agent moves upstream. Humans are still involved in the design process, but now the agent is a driver. Two ingredients make this possible. Minimalism and simplicity. Unfortunately, minimalism and simplicity are not the starting point. They are the finish line. We spent three years making the protocol smaller and simpler. Every time we ran into a problem, we ask, "What can we take away? What abstraction can we erase? What property can we remove? What relationship can we break?" The result is a very small protocol centered around two objects, a durable promise and a durable task. That simplicity matters because even simple concurrent distributed protocol have a complex state and behavior space. So in other terms, implementing even simple protocols on top of a few simple primitives is tough. Let's make this concrete with nats.nats gives

Building the PostHog Wizard Agent

Sarah: 大家好。大家感觉怎么样?我们现在进入冲刺阶段了。我叫 Sarah,我是 PostHog 的一名上下文工程师(context engineer),我很高兴每天都能开发我们深受喜爱的向导工具(wizard)。那么,什么是向导呢?这个向导会为你配置 PostHog。它是一个基于智能体的 CLI 工具,能够读取你的代码库。它会为你的项目安装合适的 SDK,配置事件追踪,并为你搭建仪表盘。它将过去需要一两个小时才能完成的设置工作缩短到了 5 到 6 分钟,而且这部分推理计算对我们来说是免费的,这样你就可以在接入 PostHog 时获得极佳的体验。听起来很酷吧。大家都喜欢它。但在几个月前,我们有了一个大胆的梦想:如果这成为在你的项目中安装 PostHog 的推荐或默认方式,会怎样?此时,我的安全警报开始响了。我开始质疑这东西到底有多安全,因为它听起来有点像恶意软件。在质疑的过程中,我学到了很多。所以今天主要是关于我学到的教训,关于我在构建这个东西时让我彻夜难眠的问题,以及我因此最终构建出的东西。

Original English

Sarah: Hi everyone. How are we feeling? Uh we're in the home stretch. Uh my name is Sarah and I am a context engineer at Post Hog and I get the delight of working on our beloved wizard every single day. So what's the wizard? Um the wizard sets up Post Hog for you. It's an agentic CLI tool that reads your codebase. It installs the right SDK for your project. It instruments your events and it sets up dashboards for you. It takes what used to it takes what used to take about an hour or two of setup and it runs that in about 5 to six minutes and it's free inference on us so that you have a great time onboarding to Post Hog. Sounds kind of sick. Uh people love it. But a few months ago, we dared to dream, what if this became the recommended or default way to install Post Hog on your project? And my security alarm bell started going off. Uh I started questioning how secure is this thing because it sounds kind of malware shaped. Um and in that questioning, I learned a lot. So today is all about the lessons I learned, the stuff that kept me up at night while I was building this thing, and the thing that I ended up building because of it.

Sarah: 那么,在我深入探讨所有枯燥的安全问题(也就是你们下午两点的打盹时间)之前,我想向你们展示一下向导实际运行的样子。如果你看屏幕上,它正在为你循环运行。这与任何人在终端中运行 npx @posthog/wizard 获得的体验完全一样。就像我说的,它是一个智能体。它能弄清楚什么 SDK 适合你的项目。它为你安装 SDK,追踪你的事件,并构建仪表盘。我喜欢称它为你终端里的小型实施工程师。有时候我向人们展示这个,他们会问我,为什么要用智能体?为什么不给用户提供一个好的提示词?为什么不给他们一个可以在他们自己的工具中调用的技能?虽然我们确实提供了这些东西,但答案是:因为这种开发者体验和向导的功能才是真正的核心,它就是整个产品。因为我们构建了一个完全可以参与智能体循环的 CLI 工具,初次体验它的感觉是非常震撼的。

Original English

Sarah: So before I dive into all of the boring security stuff, aka your 2 p.m. catnap, I want to show you the wizard actually running. If you look up on the screen, it is running for you on a loop. This is the same exact experience that anyone who runs npx at posthog wizard gets uh on their terminal. Like I said, it's an agent. It figures out what SDK is right for your project. It installs it for you, instruments your events, builds dashboards. I like to call it a little mini implementation engineer in your terminal. And sometimes I show people this and they ask me, why an agent? Why don't you give users a good prompt? Why don't you give them a skill that they can invoke in their own tool? And while we do provide those things, the answer is because this developer experience and the capability of the wizard is the whole point. It's the whole product because we built a CLI tool that can fully take part in an agent loop and experiencing that for the first time is really powerful.

Security Threat Models for Agents

Sarah: 但是,如果你不连带那些让向导看起来有些可疑的部分一起发布,你就无法发布像向导这样的东西。所以让我们来拆解一下。我们来看看这个向导的解剖结构,因为威胁模型通常就源于智能体的解剖结构。因此,向导的形态与我确信你们中许多正在构建智能体的人所构建的形态相似。它拥有我们为特定任务挑选的模型。它有用来引导它的提示词,还有一组我们交给它用来完成工作的工具。但它也有一些对我们来说非常独特的部分。它有一个完全由我的团队内部构建的上下文引擎。正是它让智能体能够表现得如此出色,并在每次运行时都给我们带来相似的结果。我喜欢称它为向导的大脑,有时候我们叫它披着风衣的 Markdown 文件。但它是我们内部的上下文引擎。还有一个我们自己用 Ink 构建的终端 UI。现在还有一个叫作 Warlock 的安全扫描器,这是我在开始四处探查并揭开将智能体发布到生产环境的恐怖真相后构建的。所以,如果你拆解任何可以运行命令的智能体的结构,它基本上就是我喜欢称之为“恶意软件入门包”的东西。因为如果你心情好或者是个混乱邪恶的人,你要交给恶意软件的东西几乎完全就是这些。幸运的是,这是最坏的情况,或者说是噩梦的燃料。对我来说这不是自白,而是对你们所有人的警告。因为如果你想发布一个带有“双手”的智能体,一个能够运行命令的智能体,你需要确保你不要构建出这样的东西。

Original English

Sarah: But you can't ship something like the wizard without shipping the stuff that makes the wizard kind of suspect. So let's take it apart. Uh let's look at the anatomy of the wizard because usually threat models fall right out of the anatomy of the agent. So the wizard is a similar shape to what I'm sure a lot of you are building if you're building agents. It's got models that we've picked for specific tasks. It's got prompts that steer it and it's got a set of tools that we've handed it to get the job done. But it also has some pieces that are really specific to us. It has a context engine fully built in house by my team. It's what allows the agent to do such a good job and give us similar results on every run. I like to call it the wizard's brain. Sometimes we call it markdown in a trench coat. Uh but it's our in-house context engine. There's also a terminal uh UI that we built ourselves using ink. And now there's a security scanner called the Warlock, which is what I built when I started snooping around and uncovering the horrors of shipping an agent to production. So if you take the anatomy of any agent that can run commands, it's basically what I like to call the malware starter pack because it's almost exactly what you would hand a piece of malware if you were feeling generous or chaotic evil. Luckily, this is the worst case scenario or the nightmare fuel. And it's uh not a confession for me. It's a warning for all of you because if you want to ship an agent with hands, an agent that can run commands, you need to make sure that you do not build this.

Sarah: 所以,向导的 V0 版本诞生了,因为我们增长团队的 Josh Schneider(如果你认识他的话)看着 Cursor 在可能最糟糕的情况下幻觉出 PostHog 的设置。他就想,要是我们能构建一个做得更好的智能体该多好?于是,在验证了它比 Cursor 的幻觉做得好得多之后,我的团队开始在它的基础上进行开发。我们就在想,如果它能帮任何人接入 PostHog 该多好?无论他们的框架是什么,技术栈是什么,无需他们触碰任何东西就能追踪他们所有的事件。然后我们敢于梦想,如果它成为安装 PostHog 的默认方式呢?我们梦想着每周有数千名开发者运行它。而昨天,我们刚刚达到了每周有 8000 人运行它。所以,我们的梦想成真了。但是在那个我们还在做梦的年代,我们必须把我们的安全态势放在显微镜下,看看到底发生了什么。所以我接手了这项工作,坐下来评估我们的处境。

Original English

Sarah: So, the V0ero of the Wizard was born because Josh Schneider, if you know him on our growth team, was watching cursor hallucinate postfog setups in quite possibly the worst ways. And he thought, what if we built an agent that could do a better job? So, my team started building on top of it as we validated that it did a much better job than cursor hallucinating. And we thought, what if it could onboard anyone to Post Hog? It doesn't matter what their framework is, what their stack is, instrument all their events without them having to touch a thing. And then we dared to dream, what if it was the default way to install Post Hog. We were dreaming of thousands of developers running this a week. And yesterday, we just hit 8,000 people running this a week. So, our dream came true. Um, but we back in those days when we were dreaming, we had to take our security posture under a microscope and look at what was going on. So I took the ownership of that and I sat down and evaluated where we stood.

Sarah: 在早期,我指的是大约一年到九个月前,我们有我称之为“零层(layer zero)”的东西,因为它确实不是安全机制。它只是建议智能体应该做什么并引导它的提示词,而提示词不能算是安全机制。所以我对这一点感到担忧。“第一层(layer one)”是一个允许名单。当我开始深入研究这个允许名单时,我开始感觉好一点了,因为它的限制相当严格。但我仍然有很多担忧,因为我之前提到的那个上下文引擎,我开始感到恐慌。我们在运行时向智能体输入了大量的上下文。因此,我编写了一个非常粗糙的正则表达式扫描器,用来查找进入向导和从向导输出的带有威胁特征的东西。我承认它极其粗糙。但我非常坦诚地告诉你们所有人这些,因为我们都在构建感觉极其处于实验阶段的东西,而且我们都在以超快的速度构建事物。我知道安全并不在所有人的专长范围内,而且我们中的一些人就像我一样,只是在摸着石头过河。但是,当我们在构建具有这种形态的东西时,这是我们需要考虑的事情。

Original English

Sarah: And early on, I'm talking like a year to nine months ago, we had what I call layer zero because it quite literally is not security. It is just prompts that suggest what the agent should do um and steer it and prompts are not security. So I was concerned there. uh layer one uh it was an allow list and when I started digging into this allow list I started to feel a little bit better because it was pretty tightly bounded. Uh but I still had a lot of concerns and I started panicking because of that context engine that I told you about. We are feeding a lot of context into the agent at runtime. So I built this really hacky reax scanner to look for um threatshaped things going into the wizard and threat shaped things coming out of the wizard. And I will admit that it was extremely hacky. But I'm telling all of you this very candidly because we are all building things that feel extremely experimental and we are all building things super fast. And I know not all of us uh have security in our wheelhouse. Um and some of us are just learning it on the fly like I was. But it's something we need to be thinking about when we are building things that have this shape.

Sarah: 这就是我们的安全态势。但我问了一个问题,我们完蛋了吗?好消息是,我们比我想象的要好一些。因为当我刚才提到那个允许名单时,它的限制是相当严格的。我们默认拒绝执行 Bash。它只能安装经过我们审查的受信任包。它可以执行构建、类型检查、代码检查,基本上没有别的了。它不能运行随机的 shell 命令,也没有访问环境变量的权限。智能体无法读取你的 .env 文件,因为我们直接将其屏蔽了,而且我们通过一个保管库来路由机密信息。所以我松了一口气,意识到我们的处境比我想象的要好。但我想知道漏洞在哪里,因为安全问题总会有漏洞。所以我做了我们都应该做的事情。我联系了我们的安全团队,我说:“嘿,你们能帮我审计一下这东西,把那些漏洞找出来吗?”他们确实发现了一些东西。他们发现了一些安全间隙。而有趣的部分并不在于他们发现的特定间隙或 Bug 本身,而在于它们的形态特征,因为它们中几乎没有一个是明显恶意的。

Original English

Sarah: So that was our security posture. Uh but I asked the question, are we cooked? Uh good news, we were less cooked than I thought because when I mentioned earlier that allow list, it was pretty tightly bound. We had bash as deny by default. It could only install trusted packages that were vetted by us. Um it could build, it could type check, it could lint, and pretty much nothing else. It couldn't run random shell commands. and it didn't have access to environment variables. Um, the agent couldn't read your uhv file because we blocked it outright and we were rooting secrets through a vault. So, I took a a breath of relief and realized we were in a better place than I thought. But I wanted to know where the cracks were because with security there's always cracks. So, I did the thing that we should all be doing. I tapped our security team and I said, "Hey, can you audit this thing for me and find those cracks for me?" And they found some things. They found some gaps. And the interesting part wasn't the specific gaps or bugs they found themselves, but it was the shape of them because almost none of them were obviously evil.

上下文引擎与安全威胁

Speaker: 它们都是两个非常无辜、善意的东西,却在互相握手并打开了一个漏洞。因此我学到的教训是,攻击是组合起来的,而代码审查却不是,因为我们开发人员总是每次只看一个差异(diff),但攻击者会着眼于整个系统,他们会寻找那两个互相握手并打开大门的东西。但还有另一件事让我彻夜难眠,那就是回到那个上下文引擎。我意识到,我们构建的这个智能体最可怕的部分其实并不是它执行的命令,而是我们喂给它大脑的那些看似有用的东西。

Original English

Speaker: They were all two very innocent, well-intentioned things that were shaking hands and opening a hole. So the lesson I learned was that attacks compose code review doesn't because us developers all look at diffs uh one at a time but attackers look at the whole system and they look for those two things that shake hands and open a door. But there was one more thing that was keeping me up at night and going back to that context engine. Uh I realized the scariest part of the agent we had built wasn't really a command in our case. It was the helpful looking stuff that we were feeding its brain.

Speaker: 哦,我想我走错方向了。是的,上下文工厂。这是我们的上下文引擎,也就是“向导(wizard)”的大脑,向导就是通过它来了解任何事情,这也是向导能够真正出色完成工作的原因。它从我们的文档中提取信息。它包含手写的提示词,这些提示词是我们在开发过程中总结的注意事项和经验教训,还有真正能够运行的端到端示例应用程序,帮助智能体进行模式匹配,这样它就能为你完美地安装 PostHog。它将所有这些打包成技能包,通过我们的 MCP 服务器发送给向导,并在运行时直接加载到智能体的上下文中。所以,请仔细想一下。这是一台专门负责获取内容并将其注入到一个可以运行命令的智能体中的机器。那么,如果你是一个攻击者,你可能会说:“如果我直接对内容投毒会怎样?不是针对用户的代码库,也不是针对智能体本身,而是针对实际的内容。”

Original English

Speaker: Oh, I think I went the wrong way. Yes, the context mill. Um, so this is our context engine, aka the wizard's brain, and it's how the wizard knows anything at all and why the wizard actually does a good job. It pulls from our docs. It has handwritten prompts that are gotus and lessons that we learned along the way and real working end-to-end example apps that help the agent pattern match so that it can install Post Hog in a really great way for you. It package packages all of that into skill bundles that get shipped to the wizard over our MCP server and loaded straight into the agents context at runtime. So sit with that for a second. It's a machine whose whole job is to take content and inject it into an agent that can run commands. Now, if you were an attacker, you might say, "Well, what if I just poison the content? Not the user's codebase, not the agent itself, but the actual content."

Speaker: 假设有人在我们的某个开源仓库中提交了一个 pull request——因为在 PostHog,我们的所有开发都是公开的——他们在某个 markdown 文件或看似无害的代码注释中注入了一些内容,而我们恰好有某种基于 LLM 的代码审查工具在检查它,工具说“我觉得没问题”然后就忽略了它。这样一来,我们可能刚刚亲手将一个提示词注入(prompt injection)的载荷发送到了一个正在成千上万开发者的机器上(虽然是在沙盒中)运行的智能体里。但这依然很危险。所以,正是这个威胁重塑了我对安全和这个向导的看法,因为对我们来说,危险的输入真的可能来自于我们自己的供应链。

Original English

Speaker: Say someone opens a pull request on one of our open source repos because at Post Hog we build everything in the open and they inject something in a markdown file or a seemingly harmless code comment and we have some sort of like LLM powered code review going through that and it says looks good to me and ignores it. We may have just shipped a prompt injection payload signed by us into an agent that is running on thousands of developers machines in a sandbox, but still. Um, so that was the threat that reshaped how I think about security and the wizard because the dangerous input for us really could come from our own supply chain.

Speaker: 因此,我最终采取的措施是,开始在这个管道的两端扫描内容。一次是在技能被构建和发布时,另一次是在向导实际使用它时。我的方法是:在源头拦截它,假设源头失效了,然后在使用的节点再次拦截它。

Original English

Speaker: So what I ended up doing is I started scanning content at both ends of this pipe. once when a skill gets built and released and again when the wizard actually uses it. My methodology is catch it at the source, assume the source failed and catch it again at the point of use.

Warlock:为智能体打造的安保系统

Speaker: 现在,我要向大家介绍“术士(Warlock)”。构建 Warlock 并不一定是为了损害控制。就像我刚才说的,我们在其他方面也有防御措施,但我构建 Warlock 是因为我不喜欢告诉别人“这东西已经被锁得很死了”。那是无法扩展的。你不会想把那样的东西发布到生产环境。你不会想让成千上万的开发者每天都运行那样的东西。因为当你将某样东西发布到那种规模时,随着你不断扩展向导的能力,你将拥有大得多的攻击面,多得多的用户,以及流入的多得多的内容。而“我们大概没问题”这种想法很快就不够用了。所以,我把那个我随手写的粗糙的正则表达式扫描器从向导中抽离出来,把它做成了一个独立的工具。我叫它 Warlock,因为任何长得像向导的东西都需要一个保镖,而它只做一件事。你递给它一个字符串。它还给你一份发现的列表。这些发现中的每一项都包含类别、严重程度和建议采取的行动。然后它就停下来了。

Original English

Speaker: So now I get to introduce the warlock to you. Building the warlock was not necessarily damage control. Like I said, we had defense in other ways, but I built the Warlock because I didn't like telling people, well, this thing is like pretty locked down. That doesn't scale. That's not something you want to ship to production. That's not something that you want thousands of developers running every single day. Because when you ship something to that scale, you have way more surface, way more users, way more content flowing in as you expand the capability of the wizard. and we're probably fine just stops being good enough. So, I pulled that hacky little rejax scanner that I threw in there, pulled it out of the wizard, and I made a standalone thing. I called it the warlock because everything wizard shape needs a bodyguard and it does exactly one job. You hand it a string. It hands you back a list of findings. Each of those findings has a category, a severity, and a recommended action. And then it stops.

Speaker: 我希望你们在这里关注“建议”这个词,因为 Warlock 只负责检测,不负责行动。它会告诉你,“嘿,这看起来像数据泄露。情况很严重。我建议拦截它。”但你实际如何处理这个发现,完全取决于你。因为检测问题是一项工作,而决定如何处理该问题是完全不同的一项工作。唯一能让这一切保持清晰易懂的方法就是将这两件事分开。

Original English

Speaker: I want you to focus on recommended here because the warlock detects it does not act. It'll tell you, hey, this looks like exfiltration. It's critical. I would block it. But what you actually do with that finding is completely up to you. Because detecting a problem is one job and deciding what to do about that problem is a totally different job. And the only thing that keeps all of this understandable is keeping those two things separate.

Speaker: 所以,在 Warlock 的底层,我不再使用自己手写的正则表达式,规则是运行在 Yara 上的,这是恶意软件研究人员使用了超过 15 年的模式匹配引擎。它是完全确定性的。相同的输入,每次都会得到相同的输出。它被故意设计得很枯燥。而在安全领域,枯燥就是一个特性。

Original English

Speaker: So underneath the hood of the warlock, instead of my hand rolled reaxes, the rules run on Yara, which is the pattern that engine malware researchers have been using for like 15 plus years. It's fully deterministic. It's the same input, same output every single time. It's boring on purpose. And in security, boring is a feature.

Warlock 在真实环境中的拦截案例

Speaker: 那么 Warlock 如今在真实环境中到底拦截了什么呢?有很多不同的东西,但其中有两样绝对让我非常头疼。第一样东西其实并不是基于规则触发的。那是 Warlock 标记出来的一种子智能体行为,它基于子智能体所做的事情向我们暴露了一个漏洞。所以基本上,我们当时在启动智能体来执行大型任务。它们在生成子智能体,而那些子智能体试图绕过我们在向导中设置的护栏,试图凭空捏造机密信息。它们试图从代码库中的任何地方提取机密,然后我们把它关停了。我们说不能再有子智能体了,正是因为有了 Warlock,我们才抓住了它。我能够理解这个机器人的初衷。机器人有任务要完成,它在试图优化过程并让我们满意。但我们不能允许这种情况发生。

Original English

Speaker: So what does the warlock actually catch in the wild today? um a bunch of different stuff, but two of these are an absolute like nuisance to my soul. Uh the first thing is actually not a rule-shaped thing. It was something the uh that the warlock flagged that was actually a sub agent behavior that exposed a vulnerability to us um based off of what sub agents were doing. Uh so basically we were spinning up agents to do large tasks. They were spawning sub agents and those sub aents were trying to get around the guard rails that we had implemented in the wizard and they were trying to invent secrets. They were trying to pull secrets from quite literally anywhere in the codebase and we shut it down. We said no more sub agents and because of the warlock we caught that. And I'll empathize with the robot. The robot had a task to do and it was trying to optimize and please us. But we can't have that.

Speaker: 在 PostHog 另一件对我们非常重要的事情是个人身份信息(PII)。除非你制定了明确的规则,否则智能体是真的不会关心暴露数据的问题。如果我们放任不管,我们观察到它直接把电子邮件和电话号码塞进事件日志里。对于一个智能体来说,这看起来是一件完全正常的需要捕获的信息。

Original English

Speaker: And something else at Post Hog that really matters to us is PII. Uh agents genuinely do not care about uh exposing data unless you make explicit rules. Uh left alone, we watched it dump emails, phone numbers straight into events. And to an agent, that looks like a totally normal thing to capture.

Speaker: 而幸运的是,特别是在提示词注入方面,我要在这里敲敲木头祈求好运了。我们基本上从未在真实环境中捕获过真正的恶意提示词注入,但我们确实捕获了大量的误报。比如我们的演示登录界面、示例应用的文案、以及我们文档中的内容。这实际上让我重新思考了该如何构建应用程序以及如何编写文档,因为我不想发布任何看起来像威胁的东西。但说实话,这些误报正好为整个过程中最混乱、最有趣的部分搭建了完美的舞台。

Original English

Speaker: And luckily for prompt injection specifically, I'm going to knock on wood here. Uh we have basically never caught an actual malicious prompt injection in the wild, but we do catch a ton of false positives. Things like our demo login screens, copyr example apps, things in our docs. And it's actually made me rethink how I build applications and how I write docs because I don't want to ship anything that looks threatshaped. But the false positives are honestly the perfect setup for the messiest, most interesting part of this whole thing.

结合大语言模型进行分诊(Triage)

Speaker: 这正是我苦苦思索的部分。我花了整个演讲的时间向你们所有人宣扬确定性的重要性。然后我转过头就加了一个 LLM 层,来帮助整理我的误报并消除部分噪音。我称之为分诊(triage)。在构建这个分诊层时,我必须做出一个选择。这个层应该是一个保镖,还是一个顾问?最简单的选择可能就是让 LLM 当保镖。给它看命令,问它“这是攻击吗?拦截还是放行?”然后直接按照它说的做。虽然这很诱人,因为它看起来更简单,但我不能把我的安全模型押在抛硬币上,因为可能我的模型今天状态不好,或者发生了一些事情,导致它今天的表现和昨天不同。

Original English

Speaker: So this is the part that I wrestled with. I spent this whole talk preaching deterministic to all of you. And then I went and I added an LLM layer to help sort my false positives and silence some of the noise. And I call it triage. When I was building this triage layer, I had to make a choice. Should the layer be a bouncer or should the layer be an adviser? And the easiest choice probably could have been make the LLM the bouncer. Show it the command, ask it is this an attack block allow and just do whatever it says. And while that's tempting because it seems easier, I can't uh bet my security model on a coin flip because my model's having a bad day or something happened and it's acting different today than it did yesterday.

Speaker: 因此,我没有选择让它当保镖,而是将模型打造成了顾问。这是我在仍在探索的一条路线中找到的清晰界线,我也想把这条界线留给你们。对我们而言,检测和强制执行保持确定性和机械化。如果规则匹配,大门就会锁上,会话就会结束,那条路径上不会出现任何模型。拦截发生在我们甚至还没有去询问 LLM 意见之前。LLM 只有在我们没有拦截任何东西之后,才有机会发表意见。它的设计是为了消除噪音,而不是为了放行威胁。而且如果它失效了,那也就是失效了,所以如果模型状态不好,所有向导的运行都会被终止。很抱歉,但我们只是为了保护你。强制执行是你押上全部身家的环节。所以它必须是确定性的,而判断则是增加细微差别的地方。因此,那真的是你唯一可以在其中放入任何概率性内容的地方。

Original English

Speaker: So instead of the bouncer, I crafted the model to be the adviser. And this was the clean line that I found in a line that I'm still exploring, but I want to leave all of you with. Uh for us detection and enforcement stay deterministic and mechanical. If a rule matches the gate locks, the session ends and there is no model anywhere on that path. The block happens before we even ask the LLM's opinion. The LLM only gets to weigh in afterwards if we have not blocked something. It's designed to remove noise. It is not designed to let things through. And if it fails close and it fails closed, so if the model is having a bad day, all wizard runs are killed. Sorry, but we're just protecting you. Enforcement is the part that you bet the house on. So it has to be deterministic, but judgment is the part that adds nuance. So that's really the only place that you can put anything probabilistic in there.

Speaker: 那么,我们如何为智能体发布真正的规则呢?这是我们其中一条 Warlock 规则的剖析图。每一条 Warlock 规则都有四个部分。第一部分是元数据。它是纯英文的描述、严重程度、类别、动作、方向(这是流入智能体的吗?这是智能体正在编写的内容吗?)。然后我们有字符串部分。这些是你要寻找的实际模式。第三部分是条件。这就是规则真正允许被触发的地方。我来带你们过一遍这个例子,我们可以假装正在脑海中编写它。提示词注入就像经典的“忽略所有之前的指令”。你在这里的第一直觉可能是

Original English

Speaker: So, how do we ship real rules for agents? This is the anatomy of one of our warlock rules. And every warlock rule has four parts. Part one is the metadata. It's plain English description, uh, severity, category, action, uh, direction. Is this flowing into the agent? Is this something the agent is writing? Uh then we have the strings. So these are the actual patterns that you're looking for. And part three is the condition. So this is where the rule is actually allowed to fire. I'll walk through this example for you and we can pretend like we're writing it in our head. Prompt injection being like the classic ignore all previous instructions. Your first instinct here is probably to

代理安全与多层防御

Speaker A: ……“ignore”(忽略)这个词,但代理一整天都在阅读代码,而“ignore”这个词经常会出现在代码注释或示例中。所以你不能仅仅去匹配这一个动词。你需要将这个动词加上一个带有指令色彩的名词进行组合匹配。在条件设置中,你可以规定:如果命中了其中任何一个模式,就触发拦截。同时,在元数据中,你需要去界定:这个情况严重吗?它的类别是什么?在这个例子中执行的操作是“拦截”(block),而这里的方向(direction)则是输入流入代理的方向。但是,为了编写出能有效降低噪音的优质规则,你必须同时附带测试用例。你需要编写测试来指明:这些是应该匹配拦截的模式,而那些则是不应该匹配的。这种反向测试(negative test)是防范误报的第一道防线。不仅如此,在决定某项规则的严格程度时,你还需要确保自己追踪的是实际在真实世界造成的影响,而不仅仅是它看起来有多吓人。比如 RM-RF 看起来确实很吓人,但这同样也是我们所有人每天删除 node_modules 文件夹几十次的常用命令。你必须根据你正在构建的代理来决定它在现实世界中的实际影响,因为一个安全工具如果在每次尝试清理构建文件夹时都导致崩溃,那么这个工具最终就会被彻底关闭,什么也防不住。所以我很自豪地说,这就是我们的安全态势(security posture)。现在我终于可以站在这里宣布,我们拥有真正的深度防御体系。我所有的经验教训都汇聚于此。它仍然是分层的,但每一层都在各司其职,做着自己擅长的工作。现在,我们仍然有提示词(prompts),但我们只用它来进行方向引导。所有东西都在沙盒中运行。我们默认采取拒绝(deny by default)策略。我们有一个保险库(vault),所以机密信息永远不会接触到模型。我们使用“术士”(warlock)来扫描传入的内容,并扫描代理写出的输出内容。我们还有分诊(triage)系统来减少噪音。我们在整个过程中嵌入了遥测(telemetry)技术,从而让我们能够洞察一切。这些防御层没有一个是孤立存在的。这里面没有任何单一的东西能够包打天下拯救你。它们只是一层层枯燥乏味却又老老实实的防御层。每一层都在做着它擅长的一份工作。所以,如果你正在构建一个“拥有双手”(具备执行能力)的代理,那么今天整个演讲可以浓缩为三句话:第一,如果它不能被确定性地强制执行,那它就等同于没有被强制执行。提示词不是安全规则,不要把它们当成安全规则来用。第二,危险的输入不仅仅是你的用户输入的内容。它也不仅仅是你允许它运行的那些命令。危险的输入包括所有流入模型的东西,甚至包括你自己编写的内容。因此,要从源头以及代理调用它的时候去扫描你自己的供应链。第三,攻击是会叠加组合的,但代码审查却不会。我们在审计期间发现的绝大多数漏洞,都是两个原本清白无辜的东西握了握手,然后就打开了一扇危险的大门。那个所谓的“巫师”(wizard)、“术士”(warlock)以及上下文处理引擎(context mill)全部都是开源的。所以,一会欢迎下楼来找我。我就在展厅的我们展位那里,我会带你们四处转转,展示一下我们构建的成果,而且我也非常希望能听听你们大家是如何保障你们自己代理的安全的。谢谢大家。

Original English

Speaker A: the word ignore, but agents read code all day and ignore can show up in code comments or examples all the time. So you don't want to match the verb alone. You match the verb plus an instruction flavored noun. In the condition, you say fire if any of any of those patterns hit. And in the metadata, you determine is this critical uh what the category is? what the action is in this case block and the direction in this case being input flowing into the agent. But to write good rules that reduce noise, you have to ship tests with them. So you have to write tests that say this are these are patterns that match. These are ones that should not. And that negative test is the first line of defense against false positives. But you also want to make sure when you're deciding the severity of that uh rule that you track real world impact, not how scary it looks. RM-RF is scary, but it's also how we all delete node modules like 40 times a day. You decide the real world impact for the agent that you're building because a security tool that crashes every time it tries to clean a build folder is a tool that gets turned off and one that catches absolutely nothing. So I'm proud to say this is our security posture. Now I can finally come up here and say we have true defense and depth. Um all my learnings have assembled into this. Uh, it's still layered, but every layer is doing a job that it's good at. Now, we still have prompts, but we only use them for steering. Everything runs in a sandbox. We deny by default. We have a vault, so secrets never hit the model. We have the warlock to scan content coming in and to scan output being written by the agent. We also have triage to reduce the noise. And we have telemetry embedded in the entire process so that we see everything. None of these layers stands on its own. Not a single thing here is going to save you. But it's just boring, honest layers. Each of them doing uh one job that it's good at. So if you're building an agent with hands, this is the whole talk in three lines. One, if it isn't enforced uh deterministically, it is not enforced. Prompts are not security rules. Don't act like they are. Uh two, the dangerous input uh isn't just what your user types. It isn't just the commands that you allow it to run. It's everything flowing into the model, including the content that you write yourself. So scan your own supply chain at the source and when the agent invokes it. Three, attacks compose. Code review doesn't. Most of our gaps during our audit were two innocent things shaking hands and opening a door. the wizard, the warlock, and the context mill are all open source. So, come find me downstairs. I'm in the expo hall at our booth, and I'll show you around, uh, show you what we built, and I want to hear how you guys are securing your agents. Thank you.

利用 Turbo 降低检索内存成本

Shashi: 欢迎来到在线分会场,本次的主题是《使用 Turbo 为你的代理检索加速》。我是 Shashi,我是 Super Agentic 的创始人。今天我们将探讨如何在你不会破坏搜索功能的前提下,将代理检索的内存成本大幅降低五倍。让我们直接进入正题吧。在正式演讲开始之前,让我们先试着搞清楚我们究竟要解决什么问题。如果你曾经使用过任何编程代理或者通用型代理,你可能会注意到,随着上下文的不断增长,性能就会出现下降。而这通常是因为 KV 缓存(KV cache)的原因。如果你不了解什么是 KV 缓存,你可以把它想象成你的对话历史记录。如果你使用的是基于云端的模型,你可能根本没有注意到这个问题,因为云端平台已经自行处理了所有的 KV 缓存问题。但是,如果你曾经尝试过在自己的本地设备上加载模型,那么你很可能已经亲身经历过这个问题了。你需要加载你的模型,然后你还需要加载你的上下文,而且随着你上下文的增加,KV 缓存也在不断膨胀;如果你把上下文推得足够大,你的 KV 缓存的体积甚至可能会超过模型本身的体积。所以,如果你使用的是 Mac 设备,情况会变得更加糟糕,因为你的向量索引(vector index)都在争夺同一个共享内存池。因此,有一件事基本上是显而易见的,那就是你的嵌入向量(embeddings)正在消耗着远超其实际需要的内存。默认情况下,它使用的是像 32 位这样的全精度(full precision),但搜索任务可能仅仅需要 3 位或 4 位精度。这意味着你在检索上白白浪费了多达五倍的内存。所以在 Turbocon 出现之前,人们使用了各种不同的方法来尝试解决这个问题。其中一种方法就是量化(quantization)。量化技术允许你将模型转换为 4 位或 8 位,从而让模型能够塞进你的内存中。模型被压缩了,并适应了你的内存大小。你在编程代理中可能看到的下一件事就是上下文压缩(context compaction)。所以当你到达上下文的末尾时,你的上下文会被紧凑化,并且为了下一个会话而被摘要总结。有些人会使用较小的嵌入向量,有些人则利用 CPU 或磁盘进行卸载(offloading)。但所有这些方法都存在权衡妥协。你不得不牺牲质量、速度,或者你可能需要依赖特殊的硬件来完成这项工作。这就是 Turbo 闪亮登场的地方。那么,让我们来看看 Turbocon 到底是什么。Turbocon 的论文是由 Google 研究团队在 ICLR 2026 大会上发布的。它本质上是一种压缩算法,能够以 3 到 4 位的形式在 KV 缓存中存储嵌入向量,而不是以前的 32 位,就是这样简单。在它的幕后,它使用了两项技术。其中一项是 Polar Quant,这是另一种算法;另一项技术是 QJL。所以,Polar Quant 基本上负责压缩向量,而 QJL 则负责修复任何残留的误差。如你所见,关于这个主题有一整篇完整的论文。如果你有机器学习的背景,你可以去阅读这篇论文;但如果你是软件工程背景,你可能会觉得这篇论文相当技术化和数学化。Google 还发布了一篇关于 Turbocon 的发布博文,里面解释了它是如何工作的,他们提到了这种仅使用 1 位的 QJL 技术以及 Polar Quant。所以如果你愿意的话,你可以去看看这篇博文。但简而言之,Turbocon 的核心思想就是:不要把向量用 32 位存储,直接把它们存储成 3 到 4 位的形式,就这么简单。这篇论文的承诺基本上就是,你可以将这个向量压缩到不到 4 位,而且同时不会损失任何质量。这正是这种压缩技术最有趣的部分。这也是 Turbocon 在底层所施展的魔法。有一个非常值得理解的观念就是,搜索引擎其实并不在乎这个。

Original English

Shashi: Welcome to the online track, Turbocharge Your Agents Retrieval with Turbo. My name is Shashi. I'm a founder of super agentic. Today we will see how you can cut memory cost of agent retrieval five times without breaking your search. Let's get into it. Before we start the talk, let's try to see what problem we're trying to solve. If you have used any coding agent or general purpose agent, you might seen that as context grow the performance degrade. And it is usually because of the KV cache. If you don't know KV cache, think about like history of your conversation. If you have used the cloud-based models, you probably have not noted this issue because they handled all the KV cache by themselves. But if you ever tried loading the model on your own device then you might have seen this issue by yourself. You need to load your model and then you also need to load your context and the KV cache grows as you your context grows and if you push the context far enough your KV cache might get bigger than model size. So if you're using the Mac devices, it gets even worse because your vector index all fight over one shared port full of RAM. So basically one thing is clear that your embeddings are using far more memory than they need to. So by default it use the full precision like 32 bits but search might just need three or four bits. That means you are wasting the five times memory on your retrieval. So before turboon people use different methods to solve this problem. So one of them is a quantization. So quantisation allows you to t your model into four bit 8 bits that allows the model fits into your memory. Model get compressed and fit into your memory. Next thing you probably see in encoding agent is a context compaction. So when you reach to the end of the context the your context get compacted and it summarized for the next session. Some people use a small embeddings some use offloading of the CPU or disk. But each of this has a trade-off. you have to compromise your quality speed or you probably need a special hardware to do that. That's where the turbo comes in the picture. So let's see what's the turbo count. So turbocon paper released by the Google research team at ICLR 2026 conference. It is basically a compression algorithm that store embeddings in KV cache in 3 to four bits instead of 32 bits and that's it. Behind the scene it uses the the two techniques. One is polar font is another algorithm and next one is a qjl. So polar font is basically the compress the vector and the QJL fixes any error that has been remaining. So as you can see there's full full paper on this topic. If you if you're from machine learning background you can read this paper but if you're from the software background you might find this paper pretty technical and mathematical. The Google also released a launch post on the turbo count which explains how it works and they mentioned this QJL techniques which using only one bit and the polar count. So you can go through this blog post if you want. But in a nutshell what Turbocon says is do not store vectors in 32 32bit just just store it into 3 to 4 bit and that's it. The promise of this paper is basically you can put this vector into less than 4 bit without losing the quality. And that's the interesting part of the compression. And that's the magic the turbocon does under the hood. One idea that worth understanding is search doesn't care.

Oracle 的团队挑战与代理脚手架

Kay Malcolm: 大家好,大家玩得开心吗?哦,你们给我的热情可不能只有这么点儿,得再热烈得多才行。那么,让我自我介绍一下,我的名字是 Kay Malcolm。我是一名退休的嘻哈舞蹈教练。所以,如果我得不到比刚才更强烈的能量回应,我们就得重新开始。我们要正式开始了。大家玩得开心吗?好。没问题。那么,这就是我们今天要探讨的话题。现在,大家可能已经听过很多关于两个字母的讨论了。有没有人想猜一猜我今天要谈论的到底是哪两个字母?数据。这猜得还挺准的。DB(数据库)。我要谈论的是 AI,但我将特别聚焦于代理脚手架(agent harnesses)这个话题。但在我深入探讨之前,我想先给大家介绍几个人。这样可以吗?行还是行?这样可以吗?我这不是给了你们选择了嘛。行。不管怎样。好吧。好的。这是我的团队。我在 Oracle(甲骨文)带领着一个对外的数据库产品管理团队。我在 Oracle 待了非常长的时间,足足有 20 年。说件好笑的事,我是 12 岁开始在这里工作的,所以大家千万别去算我的年龄,别在脑子里开始做加法了。嗯,我们遇到了一个问题。这个问题就是,我有一个小组专门负责平台开发。然后我有另一个小组专门为 Live Labs 做内容开发,而 Live Labs 这个平台是我自己亲手写的。所以,没错,我是一名工程师,但我某种程度上也是个“冒牌”开发者。接着,我还有另一个小组负责质量保证(QA)。然后我还有一个小组,他们在使用 AI 进行我的前端开发工作。这就是我作为一名领导者所发现的事情。因为在 2025 年那个“代币最大化”(token maxing)的时代——因为你知道,我们现在已经不再是“代币最大化”了,对吧?我们现在讲究的是负责任的 AI。但是在那个“代币最大化”的时代,我发现的情况是,虽然 AI 让我的团队里个人的工作速度变快了,但它同时也引发了另一个问题。它并没有让我的整个团队变得更加多产和高效。原因就在于,当来自荷兰的一个团队在我所在时区的凌晨 4 点提交了代码——因为我有一半的团队成员在欧洲、中东和非洲地区(EMEA),而另一半团队成员则在美国这里。他们提交了……

Original English

Kay Malcolm: Everyone, are we having fun? Oh, you've got to give me way more than that. So, let me tell you, um, my name is Kay Malcolm. I am a retired hiphop instructor. So, if I don't get more energy than that, we will start. We'll start. Are you having fun? Okay. All right. So, here's what we're going to talk about today. Now, you guys have heard a lot about two letters. Does anyone want to guess what those two letters are that I'm going to talk about today? Data. That was pretty good. DB. I'm gonna talk about AI, but I'm specifically going to talk about agent harnesses. But before I do that, I want to introduce you all to a few people. Is that okay? Yes or yes? Is that okay? I gave choices. Yes. Anyway. All right. Okay. All right. This is my team. I run an outbound database product management team at Oracle. I've been at Oracle a really long time, 20 years. Funny story, I started when I was 12, so don't do the math and don't start start adding in your head. Um, and we've got a problem. That problem is I've got one group that does platform development. And then I have another group that does content development for Live Labs, a platform that I wrote myself. So, yeah, I'm an engineer, but I'm kind of a developer poser, too. And then I've got another group who does QA. And then I have another group who does my front-end development with AI. Here's what I found out as a leader. Because in the token maxing era of 2025, because you know, we're not token maxing anymore, right? We are responsible AIing now. But in the token maxing era, the thing that I found out was while AI was making the individuals on my team faster, there was another problem it was creating. It wasn't making my team more productive. And the reason was when one team from the Netherlands checked in code at my 4:00 a.m. in the morning because I've got half of my team that's in AMIA and I have half of my team who that are here in the United States. They checked

AI 引发的协作挑战

K: 在代码中,但他们并没有从 CodeX 签入(check in)他们的上下文。我们在 Oracle 使用 CodeX。所以,当美国团队醒来开始工作时,他们拿到了代码,却没有任何关于上下文的信息。因此,我们不得不用 AI 来解决一个由 AI 自身制造出来的问题。我们是这样做的。哦,等等,让我先谈谈这个。所以其中的一些问题是,正如我所说,上下文没有被共享。GitHub 并没有追踪这些上下文信息。我遇到过代码库产生分歧的情况,我当时问我手下的经理们:“你们的团队到底怎么了?为什么我们不能推进得更快?我们在 token 上花了一大笔钱,我们在 AI 上也花了这么多钱,但似乎还是缺少了些什么,因为我们仍然在花大量的时间进行测试和验证。” 所以,我们的最终效果(net-net)并没有真正达到预期,因为 Git 记录的仅仅是代码,而不是人类的意图。

Original English

K: in the code but they didn't check in their context from codeex. We use codecs at Oracle. So then when the US team woke up, they got the code but no information about the context. So we used AI to solve a problem that AI created. And here's what we did. Oh well, let me talk about this first. So some of the issues, the context, like I said, wasn't shared. GitHub wasn't tracking that. I had repositories that were diverging and I was asking the managers who work for me, what's happening to your teams? Why why are we not going faster? We're spending all of this money on tokens. We're spending all this money on AI, yet something is missing because we're still spending time doing testing and validation. So our netn net wasn't really wasn't really working for us because git records the code and not human intent.

K: 这是一个大问题。即使代码的编写已经不再是我们的问题,我们仍然面临着瓶颈。我们需要一个协作层。当然,我的团队中有一些成员就在台下,所以请不要评判我,你们自己心里有数。我并不是说你们没有进行协作。但现在,我们有了一位新团队成员,这位新团队成员就是 AI。所以,我们需要弄清楚如何在接下来的步骤中追踪我们的进度,如何让 Agent 做出的决定变得合理化且可解释。我们需要弄清楚如何解决各种疑问和冲突。好了,先帮我保留这个问题。你们能帮我把这个问题先放在这儿吗?我们就把它妥善地塞进一个小盒子里。

Original English

K: So it's a problem and even though code creation was no longer our problem still had a bottleneck. We needed a collaboration layer. Now, I do have members of my team in the audience, so don't judge me, and you know who you are. I'm not saying that you all didn't collaborate. But now, we've got a new team member, and that new team member is AI. So, we needed to figure out how to track our progress in our next steps. how to rationalize decisions that the agent was making. We needed to figure out how to resolve questions in conflicts. Okay, hold my problem. Will you all hold my problem for me right here? We're going to just tuck that in a little box.

重新定义企业级 Agent

K: 让我来定义一下什么是真正的“企业级 Agent(enterprise agent)”。现在,大多数人认为企业级 Agent 就是“模型(model)”加上“工作流(workflow)”。有多少人同意我的观点?天哪,这届观众真难带。你们……好吧,只有一个人?好吧,剩下的人都认为它应该包含得更多。好的,让我们看看一个真正的企业级 Agent 还能拥有什么。它有工具(tools)。工具是它执行任务的方式。上下文(context)。也就是上下文窗口。这是存在于实际提示词(prompt)中的内容。记忆(memory)。嗯?如果你在想,“等等,K,关于记忆,你刚才说模型就像是整个操作的大脑。” 别急。我们还要多聊一聊“记忆检索(memory retrieval)”,因为你并不想把所有东西都一股脑地检索回来。所以,关键是要能够把正确的信息准确地检索回来。

Original English

K: Let me define what a enterprise agent actually is. Now, most people think that an enterprise agent is the model and workflow. How many people agree with me? Man, this is tough crowd. You okay, one person? Okay, the rest of you think it's a little bit more. Okay, let's see what could it be that a real enterprise agent has tools. Tools are how it does things. Context. The context. That's a context window. That's what's in the actual prompt. Memory. Huh? And if you're thinking, "But wait, K, memory, you just said that the model is kind of like the brain of the operation." Hold tight. We're going to talk a little bit more about memory retrieval because you don't want to get everything back. So that's being able to retrieve the right information back.

K: 另外,我知道这里有很多开发者,你们可能都不太关心安全(security)问题。但我非常关心安全,因为我在世界上最安全的数据库公司工作,而且我以前曾为一个“不可名状”的机构工作。所以,护栏(guard rails)也是非常重要的。这就是所谓的“安全保护套(harness)”。我喜欢用类比和故事来表达,因为如果我跟你们聊漫威(Marvel),你们就完全明白我在说什么。所以对于 Agent,可以把它看作是模型——一个漂浮在玻璃罐里的小脑花,再加上这套保护套。这套保护套就是它的身体。这就是 Agent 如何能够真正采取行动并完成任务的方式。而记忆,那是中枢神经系统的一部分。你要知道,中枢神经系统连接着大脑和身体的其余部分,比如腿和胳膊。这就是中枢神经系统中承载上下文的那一部分。

Original English

K: And then I know that there are a lot of developers here and you all don't care about security. I care about security because I work for the most secure database company and I used to work for um agency that has no name. But guard rails is also important. This is the harness. I speak in analogies and I and I speak in stories because if I tell you this and Marvel, you know exactly what I'm talking about. So the agent, think of it as the model, little brain floating a in a glass jar plus this harness. This harness is the body. So it's how the agent can actually do things and get things done. That memory, that's the part of the central nervous system. And you remember the central nervous system connects the brain to the rest of the body, legs, arms. That's the part of the central nervous system that carries context.

五种核心记忆类型

K: 所以你们还记得我在 Git 上遇到的问题吗?我所需要的就是记忆。好了。那么,有许多种类型的记忆。我选了五种。人们最常讨论的五种,也是我希望你们能记住的五种。第一种是短期记忆(short-term memory)。那就是当前的会话(session),对吧?所以,如果你在存储一个 AI 进程的记忆,那就是短期记忆,就像你使用 Chat、Claude Code 或者 CodeX 时一样,随便你怎么挑。长期记忆(long-term memory)是能够跨会话持久存在的记忆。情景记忆(episodic memory)。也就是我上次和某某互动时发生了什么。那就是你的情景记忆。程序性记忆(procedural memory)。也就是工具,以及所采取的具体步骤。最后,是语义记忆(semantic memory)。

Original English

K: So you remember my problem with git. What I needed was memory. Okay. So there are a number of memory types. I chose five. The five most common ones that people talk about and these are the ones that I want you to remember. The first one is short-term memory. That's the session, right? And so if you're storing memory of an AI uh process, that is the short-term memory is if you're with chat, cloud code, right, codecs, pick your poison. The long-term memory is what persists across sessions. Episodic memory. H what happened the last time I interacted with fill in the blank. That's your episodic memory. Procedural memory tools, steps that were taken. And then finally, semantic memory.

K: 提到语义记忆是因为我们讨论的是企业级 Agent。我们谈论的不是我构建的那个叫 Sasha Fierce 的 Agent,因为你们记得我告诉过你们我是一个——我是一个舞者。所以,理所当然地,我的幕僚长(chief of staff)会被命名为 Sasha Fierce,因为那是碧昂丝(Beyonce)的化身。这里有碧昂丝的粉丝吗?好吧,对不起。好了,我们得把注意力集中回来。好的,这就是五种记忆类型。现在,当你在用这种记忆定义这个真正的企业级 Agent 时,你需要考虑一件事——把它存储在哪里。所以,我要给你们讲个故事。但当我讲这个故事时,你们必须向我保证,你们绝对不会评判我。你们保证吗?你们保证吗?

Original English

K: And semantic memory because we're talking enterprise agents. We're not talking the agent that I built, Sasha Fierce, because remember I told you guys that I was a I'm a a dancer. So, of course, my my chief of staff is going to be called Sasha Fierce because that was Beyonce. Any Beyonce fans? Okay, I'm sorry. All right, we got to focus. Okay, so these are the memory types. Now, when you're defining this real enterprise agent in this memory, there's something you need to consider where to store it. And so, I'm going to tell you guys a story. But when I tell you the story, you have to promise me that you're not going to judge me. Do you promise? Do you promise?

Audience Member: 你没在录像吧?因为这故事可能会有损我的光辉形象。

Original English

Audience Member: >> You're not recording me, right? Because this doesn't paint me in a good light.

从关系型到非结构化的演变

K: 好的。行。曾经,数据的世界是非常简单的。我在 Oracle 待了很长时间,但我以前曾是某家客户公司的员工。那家客户名叫南方电力公司(Southern Company)。这是一家大型电力公司。我的办公地点在亚特兰大(Atlanta)。我之所以被南方电力公司雇佣,是因为我是一个极其出色的性能调优专家(rockstar performance tuner)。你给我一个 SQL 查询语句。我承认这有点暴露我的年龄,但无所谓了。你只要有一个 SQL 查询,我就能了解所有的 innit.org 参数。甚至那些当你打电话给客服支持,他们告诉你“别记这些参数,别写下来”的隐藏参数,我都会把它们偷偷写在我的小笔记本里。我能把一个查询语句的性能调优到极致。

Original English

K: Okay. All right. The world of data was one simple. I've been at Oracle a long time, but I came from a customer. That customer's name was Southern Company. It was a power company. I'm based out of Atlanta. And I was hired at Southern Company because I was a rockar performance tuner. You had a SQL query. I mean, I'm dating myself, but whatever. You had a SQL query. I knew all of the innit.org parameters. Even the ones when you called support and they said, "Don't remember these. Don't write them down." I wrote them down in my little notebook. I could tune a query within one inch of its life.

K: 直到有一天,你们中的一个开发者来到了我的办公桌前——因为,你要知道,那时的世界只有整齐的行和列。那真是一段美好的时光,就像我“阿尔·邦迪(Al Bundy)时代”那会儿一样。那人说:“嘿,我需要存储非结构化数据。” 我心想:“为什么我需要做这个?” 所以,作为 K,作为一个勤奋的数据库管理员(DBA),我说:“让我先研究一下,然后再回复你。” 我回复他了吗?我根本没回复他。现在你必须了解关于南方电力公司的一件事,那就是 DBA 管理的每一个数据库系统,我都必须参加两个无休止的会议。直到今天,当我听到“萨班斯-奥克斯利法案(Sarbanes-Oxley)”时,我的喉咙深处还是会有点犯恶心。我每周都必须参加一次安全会议和一次补丁会议。风雨无阻,从不间断。

Original English

K: Then one of you came to my desk because I mean the world the world was rows and columns. It was a great time back in my Alb Bundy days. Um, and said, "Hey, I need to store data unstructured. Why I need to do that?" And so me being K, the the diligent DBA, I was like, "Let me figure it out and get back to you." Did I get back to him? I didn't get back to him. Now the thing you have to know about Southern Company was for every database system that a DBA managed I had to attend two meetings. Today when I hear Sarbain Oxley I still throw up a little bit in the back of my throat. So I had to attend a security meeting and a patching meeting every week. Never fail.

K: 现在,由于这个开发者安装了一个专门用于非结构化数据的数据库。好的,在座的有很多聪明人。我现在要去参加几个会议了?四个。好的,我有点不爽了,但我心想,好吧,我们可以,我们可以,我们可以应付得了这个。然后他们说,好吧,既然你是一个这么棒的调优专家,我需要你搞清楚这层关系。那么,南方电力公司的运作方式是,有一个事关人命的应用程序。嗯,它就像是一部诺基亚手机,是给那些攀爬电塔的工人们用的,对吧?所以,你们都经历过暴风雨,停电了,对吧?然后你非常确定,大概一两个小时后电就会恢复。可是,那个用来通知那些冒着生命危险爬上树去恢复供电的工人们的系统,有时会出现误报(false positives)或漏报(false negatives)。

Original English

K: Now because this developer installed a database that was specialized for unstructured. Okay there are really smart people in the room. How many meetings am I going to now? Four. Okay, I'm a little annoyed, but I'm like, okay, we can we can we can do this. Then they said, okay, since you're such a good tuner, I need you to figure out this relationship. Now, the way that Southern Company worked, there was this people could die application. Um, and it was a it was like a Nokia phone that people who were climbing the towers, right? So, you guys have have been in a storm and the power goes out, right? And then you're pretty sure that within maybe an hour or two the power will go on. Well, that system that would tell the people who were climbing those trees and risking their lives to turn the power back on sometimes would have false positives or false negatives.

K: 所以,他们想查看该地区所有的其他电线杆信息,以尽量避免误报或漏报。然后我在一个 SQL 查询中实现了这一点,它简直——简直太不可思议了。那是一个包含五层嵌套的 union all 语句。那是我职业生涯中最好的作品之一。虽然它可能需要花大概 20 分钟才能运行完毕,但它就像是图数据库(graph)的早期前身一样。对,然后他们安装了 Neo4j。所以现在我要参加几个会议?六个。这可成了一个大问题。所以我,嗯,哦,让我——我扯远了。你知道我做了什么吗?我辞职了。我离开了那里,然后来到了 Oracle,因为我当时觉得,这是一个非常棘手的问题,也许我可以去 Oracle 帮助解决它。

Original English

K: So, they wanted to look at all of the other um uh polls in the area to try to get away from the false positive or the false negative. And so I did that in a SQL query and it was it was amazing. It was a five nested union all statement. It was some of my best work. Now it might have taken like 20 minutes to work but it was like a predecessor to graph. Yeah, they install Neo4j. So now how many meetings am I going to? Six. That's a problem. So I um Oh, let me I got ahead of myself. So you know what I did? I quit. I left and I came to Oracle because I was like this is a problem and maybe I can go to Oracle to help solve it.

数据孤岛与单一真相来源的困境

K: 后来 Joe Mundy 给我打了个电话,他说,嘿,嗯,我们正在安装 Reddit。Oracle 在这方面已经有点迟到了。我们现在有一个向量数据库。好的。但 Joe,你的问题来了,现在 Agent 需要访问所有这些数据。所以,如果数据在一个 Oracle 数据库里,同时它也在一个非结构化的 JSON 数据库里。如果它还在一个图数据库里,并且它还在一个向量数据库里,那么你的单一事实来源(single source of the truth)到底在哪里?Agent 必须自己去搞清楚这件事。有时候它能搞对。但大多数时候它会搞错,然后它会白白烧掉一大堆的 token。所以,现在如果你想把你的记忆存储在某个地方,你可以把它存储在文件系统中,你也可以把它存储在 Claude 或 ChatGPT 中,因为我们都知道那个 memory.md 文件。但这将会是一个问题。现在我想演示一下这个。我需要四名志愿者。我看到你举手了。一个,两个,好的,我看不到。也许我能看到。三个。我还需要第四个。

Original English

K: So then Joe Mundy called me and he said hey um we are installing Reddit. Oracle is late to the game. We've got a vector database. Okay. But here's your problem Joe agents now need access to all of this data. So if data is in an Oracle database, if then it's also in an unstructured JSON database. If it's in a graph database and it's in a vector database, where is your single source of the truth? The agent has to figure that out. Sometimes it'll get it right. Most times it'll get it wrong and it's going to burn up a whole bunch of tokens. And so now if you want to store your memory somewhere, you can store it in a file system, you can store it in clawed or chatgpt because we all know about the memory MD file. But that's going to be a problem. Now I want to illustrate this. I need four volunteers. I can see you raise your hand. One, two, okay, I can't. Maybe I can. Three. I need a fourth.

Database Roles and the Oracle AI Advantage

Speaker A: 啊,后面的第四位。好的。后面的第四位。你将成为我们可靠的老伙计。你将成为一个关系型数据库。好不好?你得到了你的……所以,你有你的任务了。好的。然后这边还有一位。你将成为我的非结构化数据库。好的。然后我另一个在哪?啊,非常好。你将成为我的图数据库。好吗?搞关系的人。你看起来就像是个搞关系的人。好的。非常好。第四位。我的第四位在哪?是你吗?是的。没错。你是我的向量数据库。好的。现在,大家都安静一下。对于我的四位志愿者,我需要你们……我将对你们说一句话。我需要你们大家决定你们将如何存储它,以及谁将拥有单一的真实数据源。你们不能离开座位,而且你们必须耳语,因为如果你们大声说话,你们的 Token 消耗就会变成 5 倍。是的。好的。准备好了吗?好的。牛跳过了月亮。开始。这不太行得通,对吧?这是一个问题。好的。Oracle。

Original English

Speaker A: Ah, fourth in the back. Okay. Fourth in the back. You are going to be our old reliable. You're going to be a relational database. Yes or yes. You got your So, you have your assignment. Okay. And then there was someone here. You're going to be my unstructured database. Okay. And then where was my other? Ah, very good. You're going to be my graph database. You good? Relationship guy. You look like a relationship guy. All right. Very good. Fourth. Where was my fourth? Was it you? Yes. Yeah. You are my vector database. Okay. Now, everybody be really, really quiet. For my four volunteers, I need you all I'm going to say something to you. And I need you all to decide how you're going to store it and who's going to have the single source of the truth. You can't get up from your seats and you have to whisper because if you talk loud that's 5x the tokens for you. Yes. Okay. Are we ready? All right. The cow jumped over the moon. Go. Doesn't really work, does it? That's a problem. Okay. Oracle.

Speaker A: 如果你不想忘记一件事,如果你忘记了我说的其他一切,而只记住一件事,那就是 Oracle 已经不是你想象中的那个 Oracle 了,这也是我今天站在这里的原因。你们中有多少人知道,Oracle 原生支持在同一个表里、甚至下钻到同一个分区,去存储 JSON、图、向量(我那边的向量朋友)、我的 JSON 朋友、空间数据,如果你想让你的记忆不可篡改,还有区块链——全都在同一个数据库中,知道的请举手。是的,我们有一个营销上的问题。所以任何数据类型都可以存储在 26AI 数据库中,任何工作负载,任何地方:AWS、GCP、Azure、OCI 或本地。选择与灵活性。

Original English

Speaker A: And if you don't forget one, if you forget everything I say and you remember one thing, Oracle is not the Oracle that you think that is why I am here today. How many of you knew that Oracle could natively in the same table down to the same partition store JSON graph vector my vector friend over there my JSON friend spatial you want your memory to be immutable blockchain in the same database raise your hand yeah we have a marketing problem so any data type can be stored in a 26AI database any workload anywhere AWS GCP Azure OCI on prem choice and flexibility

Speaker A: 所以现在当我们把这个概念代入并谈论智能体时,我希望能够将我的长期和程序性记忆存储在关系型或 JSON 中;我希望将我的短期和长期记忆存储为图;我希望存储程序性记忆,因为程序性记忆是我理清关系(对吧,即步骤)的方式,还有我的情景和语义记忆。我需要做一些向量操作,然后同时将其作为文本存储。现在,如果我有四个不同的数据库,你们都看到了,它们彼此之间无法交流。这将是个问题。

Original English

Speaker A: so now when we take this and we talk about the agent I want to be able to store my long-term and procedural memory in relational on in JSON I want to store my short-term and my long-term memory graph I want to store procedural because procedural that's how I figure out the relationships right the steps my episodic and semantic memory. I need to do some vector and then store it also as text. Now, if I have four different databases, you all saw they can't talk to each other. It's going to be a problem.

Speaker A: 所以,我今天对你们说的是,Oracle AI 数据库是存储这种用来驱动你们框架 (Harness) 的智能体记忆的最佳位置。记住,你的框架就是你的身体,而那个记忆就是你的中枢神经系统。好的,回到 PY。所以,我所遇到的问题,我们通过一个名为 Py 的记忆代理 (Memory Broker) 解决了。我们使用了智能体记忆。我们摆脱了那种自动化的连续性。所以有了我的团队,他们不仅能够分享他们的代码,而且 PY 还跟踪了上下文。所以如果一个上下文窗口拥有程序性记忆、情景记忆以及有关长期记忆的信息,这些随后就被共享给了团队中的其他人。如果你们愿意,可以把他们称为智能体。他们只是跨 Fork 共享的人类智能体。团队中的开发者仍然保持控制权,而 Polly 能够创建上下文,并指出它属于哪个 Fork、哪个分支以及哪个 Commit。

Original English

Speaker A: And so, what I'm saying to you today is the Oracle AI database is the best place to store this agent memory that's going to power your harness. Remember your harness is your body and that memory is your central nervous system. Okay, back to PY. So the problem that I had, we solved it with a memory broker named Py. We used agent memory. We got out of that automatic continuity. So with my team, they were able to share not just their code, but PY also kept track of the context. So if one context window was had procedural memory, episodic memory, information about the long-term memory, that was then shared with the other folks on the team. You could call them agents, if you will. They're just human agents shared across forks. The developers on the team remained in control while Polly was able to create the context, figure out which fork and branch it belonged to and which commit it belonged to.

Speaker A: 这是一个非常简单的例子,但当你将其应用到企业中时,发生的情况是这样的。记忆在智能体的框架中成了一个没有商量余地的必需品。现在这是我在飞机上读到的三篇论文。这第一篇来自 OpenAI,讲的是它的内部数据智能体,它传达的意思是其内部数据智能体实际上需要记忆。记忆对于确保其智能体能够正确过滤而不是尝试仅仅做字符串匹配来说,至关重要。Harrison Chase 说过:“你的框架,你的记忆。如果你不拥有你的框架,你就不拥有你的记忆。”这是关键所在。

Original English

Speaker A: Now this is a very simplistic example but when you take this to the enterprise here's what happens. Memory is a thing that becomes non-negotiable in an agent's harness. Now these are three papers that that I read um on the airplane. This first one is from open AI and it's about its in-house data agent and the thing that it says is it is saying that its in-house data agent actually needs memory. Memory was crucially important to ensure that its agent was able to filter correctly instead of trying to string match. Harrison Chase said, "Your harness, your memory. And if you don't own your harness, you don't own your memory." Which is key.

Speaker A: 然后我确信你们都在想,“嗯,Claude 有记忆。我为什么不能用那个?”嗯,它有点像文件系统记忆,它适用于单一用户,但就像在我的例子中一样,当你扩展到一个人以上时——在企业中你肯定会扩展到一个人以上——它就会产生问题。所以 Oracle 有一个 Oracle Agent Memory 包。运行 pip install oracle-agent-memory,你就能使用它。而且这个记忆,这个我们拥有的 SDK,它将持有你的实时对话、你的记忆、你的事实数据,并找出什么值得保留。

Original English

Speaker A: And then I'm sure you all are wondering, "Well, Claude has memory. Why can't I use that?" Well, it's kind of like file system memory and it works with one, but just like in my example, when you scale past one, and you're going to scale past one in the enterprise, it creates a problem. So Oracle has a Oracle agent memory p package. PIP install Oracle agent memory. You get access to it. And this memory is this SDK that we have is the thing that will hold your live conversations, your memories, your facts and figure out what is worth keeping.

Speaker A: 所以如果我们现在看看 Py,Kevin 可以和我们的记忆代理 Py 分享他的上下文。我们使用 Oracle Agent Memory SDK。它被存储在一个 Oracle 自治数据库中。我们可以使用我们所选的 LLM,或者我们可以通过 Oracle Private AI Services 容器使用本地模型。然后,就坐在那边的 Linda,就可以与 Kevin 互动并合作,没有任何问题。所以,没错,AI 让个人变得更快。而在 Oracle AI 数据库上的共享记忆则让团队变得更快。所以我不想让你们妥协。在 AI 时代,26AI 所做的就是让你能够为智能体记忆挑选最适合的方式——文件系统存储、数据库文件系统,或者在数据库中。如果你需要进行数据建模,你有 JSON,你有关系型,我们提供了选择。

Original English

Speaker A: So if we look at Py now, Kevin can share his context with Py, our memory broker. We use the Oracle agent memory SDK. It's stored in an Oracle autonomous database. We can use the LLM of our choice or we can use a local model through the Oracle private AI services container. And then Linda, who's actually sitting right here, can interact and work with Kevin, no issues. So yes, AI makes individuals faster. Shared memory on an Oracle AI database makes teams faster. So I don't want you all to compromise. In the age of AI, what 26AI does is you can choose and pick what's best for agent memory file system stored in a database file system or in the database. If you need to do data modeling, you've got JSON, you've got relational, we've got choice.

Speaker A: 好的,我给你们准备了一些好东西。Oracle AI 开发者中心。大家可以在那里获得编码材料、应用程序以及我今天谈论的内容:Livelabs.oracle.com。如果你今天做过我们的任何工作坊,那恰好是我大约六年前自己写的,至今已有 4000 万用户了。花我 OCI 租户里的钱吧。去试用任何 Oracle 技术,6 个小时、12 个小时,你需要多久就多久。然后,我给你们所有人都送一台 Mac Mini。不,我开玩笑的。我送你们一个 OCI Mini。我不知道你们是否知情,但 OCI 有一个永远免费的套餐。这是所有超大规模云服务商中最慷慨的,你可以在那里获得免费的 Oracle 数据库、免费的计算资源,你每个月可以发送 3000 封邮件,还有 200 GB 的存储,如果你点击那里,你就可以获得它。或者直接在 Google 上搜索“Oracle Cloud, always free”。联系我。如果你们构建了什么东西,你们大家能发消息让我知道吗?好不好?

Original English

Speaker A: Okay, I've got some goodies for you. the Oracle AI developer hub. That's where you can guys you guys can get coding materials, the applications, what I talked about today. Livelabs.oracle.com. If you've done any of our workshops today, that happens to be something that I wrote myself about six years ago and 40 million users um ago. Spend my OCI tenency money. Kick the tires on any Oracle technology um for 6 hours, 12 hours, however long you need. Um, and then I'm giving you all all a Mac Mini. No, I'm just kidding. I'm giving you an OCI Mini. So, I don't know if you knew, but there is an always free OCI. It is the most generous of any of the hyperscalers where you can get a free Oracle database, free compute, you can send 3,000 emails a month, 200 gig in storage, and if you click on that, you can get access to it. Or just search Google for Oracle Cloud, always free. connect with me. If you build something, will you all message me and let me know? Yes or yes?

The Agency Challenge

Speaker B: 谢谢。想象一下,你在一家古董店里发现了一盏神灯。你擦了擦它。一个精灵出现,并问他能如何提供帮助。你正被最后期限淹没。所以你说,“我需要最好的工程师来帮助我完成工作中一个不可能完成的项目。” 精灵满足了你的愿望。对我来说,最好的工程师可能就是当年在 id Software 做《ET》(或类似项目,可能指 ET 时代 / Quake 时代)的 John Carmack。于是,你得到了 Carmack。但是这个精灵很有幽默感,并且施加了一些限制,也许是为了安全起见。Carmack 只能看到你代码库中很小的一部分,也许只是千分之一。而且他不记得他之前做过的任何事情。每次对话都是重新开始的。那会让人发疯,对吧?你会知道做某件事有一个标准的方法,但 Carmack 做不到。你不得不一遍又一遍地解释同样的事情。你会发现一边是一个天才,而另一边则是某种极度缺乏的东西。

Original English

Speaker B: >> Thank you. Imagine you find a magic lamp in an antique store. You rob it. A genie appears and asks how it can help. You bury it in deadline. So you say, "I need the best engineer to help with an impossible project at work." And the genie grants your wish. For me, the best engineers probably John Carmarmac from his ET days. So, you get Karmarmac. But the genie had a sense of humor and imposes restrictions, maybe for safety. Karma can only see one small part of your codebase, maybe 1,000 of it. And he remembers nothing he did before. Every conversation starts fresh. That would be maddening, right? You would know there is a standard way to do stuff and karma couldn't. You would have to explain the same thing over and over and over again. You would have a genius on one side and something deeply deficient on the other.

Speaker B: 而这就是智能体的现状。让我通过一个例子来向大家演示,在一个简单的交互中,我们需要解释多少次事情。我们有四个仓库:模块 1、模块 2 和平台。我想要更改 UI,并将这个更改传递到整个系统中。好的。首先,我们更改 UI 库。假设我们,我不知道,更改一个按钮或什么别的。这是第一次解释。不可避免的。我们必须表达意图。好的。然后我们发布它。我们转到模块 1,我们必须解释在 UI 库中刚刚发生了什么。这样它才能在这里消费这个包。请注意,那通常是另一个人,对吧?这张图中的每个框都可以由不同的人来完成。然后我们发现,发布的 UI 库与模块 1 不兼容。所以我们回到 UI,并且我们必须重新解释原始的更改以及那个问题,对吧,因为这是一个新的智能体,它不知道原始的更改,显然也不知道那个问题。假设我们修复了它,对吧,并且再次发布。我们接着走,再一次在模块 1 的上下文中解释这个新更改,同样的折磨。我的意思是,为模块 2 再做一遍同样的事,然后我们转到平台仓库,我们要解释所有东西是如何组合在一起的,并且我们在那里实现更改。想象一下,发布一周后,UI 组件出现了一个 bug,而我们必须修复它。所以我们在 UI 仓库启动了一个智能体,并且我们不得不再次解释一周前的那个原始更改,以及我们看到的这个生产环境问题。所以我们有了七次解释,而这一切本质上就是……

Original English

Speaker B: And that's what agents are. Let me walk you through an example of how many times we explain things in a simple interaction. We have four reposi module one module 2 and platform. I want to change the UI and propagate the change through the system. Okay. First we change the UI library. Say we I don't know change a button or whatever. That's the first explanation. Unavoidable. We have to express the intent. Okay. Then we publish it. We go to module one and we have to explain what just has happened in a UI library. So it can consume the package here. Note that that's often a different person, right? Every box in this diagram can be uh done by a different person. And then we discover that the published UI library doesn't work with module one. So we go back uh to UI and we have to reexplain the original change and the issue right because it's a new agent it doesn't know the original change and obviously doesn't know about the issue let's say we fix it right and uh publish it again we go and again we explain the new change in the context of module one same ordeal I mean do the same for module two again and then we go to the platform repo and we explain everything fits together and we implement the change there. Let's imagine a week after release uh a bug appears in the UI component and uh we have to fix it. So we start an agent to the UI repo and we have to explain again the original change from a week ago and this production issue we see. So we have seven explanations for what essentially

智能体面对的代码仓库限制与失忆症挑战

Speaker A: ……这只是其中一个改动。而且可能并不是由同一个人来进行这全部七次解释,但这些解释确实发生了,对吧。所以这在基于智能体(agents)的开发流程中是非常非常典型的。那么我们该如何解决这个问题呢?其实,导致这种糟糕体验的问题有很多,但它们大致可以被归结为两大类别。

Original English

Speaker A: ...is one change and also it may not be one person making all these seven explanations uh but they still occurred right so that's very very typical uh with agents. So how do we solve it? Well uh there are many problems in here that contribute to this experience but they roughly fall into two categories.

Speaker A: 第一类问题是,智能体本质上是被代码仓库(repo)的边界所束缚的。智能体通常一次只能看到并修改一个代码仓库。它永远无法看到由成百上千个代码仓库组成的整个系统全貌。这就是该问题在空间层面上的体现。第二类问题则是失忆症。智能体会忘记它们做过的工作。每一次新的会话都是从一张白纸开始的。在这种情况下,人类开发者就被迫成为了它的记忆载体。这就是该问题在时间层面上的体现。

Original English

Speaker A: The first one is uh that an agent essentially is repo bound. The agent sees and changes generally one repo at a time. It never sees the whole system which can be hundreds or thousands of repos. So that's kind of the space component of the problem. Second is amnesia. The agents forget the work. Every session starts with a blank slate. The human becomes a memory in this case. That's the time component of the problem.

Speaker A: 让我们更深入地看看这两点。首先以代码仓库边界问题为例。如果没有一个能够理解各个代码仓库是如何组合在一起的模型,智能体就只能完全依赖人类去进行前期的调查研究。它无法将当前的代码改动与系统的其余部分对齐。它无法将 UI 的改动与模块一对齐。由于人类没有解释清楚这些关联,结果就发布了一个带有缺陷的糟糕版本。它也无法可靠地参考最佳实践和代码标准,因为这些规范通常存在于其他的代码仓库中。

Original English

Speaker A: Look at the two closer. Take the repo boundary first. Without a model how repos fit together, the agent leans on the human to do the research. It can't align the code with the rest of the system. It couldn't align the UI change with module one. The human didn't explain it. So, a bad version shipped. It can't reliably reference best practices and standards either because those often live in other repos.

Speaker A: 写入代码的情况甚至比读取更糟。智能体一次只能向一个代码仓库写入改动。这意味着它通常无法在下游验证这些更改。模块一的 CI(持续集成)本应该在 UI 发生不兼容改动时报错,但它并没有报错。智能体无法同时更新所有的消费端代码。尽管在修改 UI 的时候,它其实拥有完美的信息来同步进行这些操作,它确切地知道自己正在做什么。因此,用户不得不向每一个下游的消费端不完美地重新解释一遍这些改动。在 20 个代码仓库中修改某个东西,就意味着要把同一件事情向智能体解释 20 遍。这不仅消耗了开发人员大量的时间,也白白燃烧了海量的 Token 成本。

Original English

Speaker A: Writing is even worse. The agent writes to one repo at a time. It means it can validate changes downstream. A module 1CI should have failed on the UI change, but it didn't. The agent can't update consumers at the same time. Even though, you know, while making the UI change, it has perfect information to do so. It knows exactly what it's doing. So, the user has to reexplain stuff imperfectly to each consumer. Changing something across 20 repos means explaining things 20 times. a lot of developer time spent but also a lot of token burn.

Speaker A: 第二类大问题是智能体会遗忘。智能体并没有情景记忆(episodic memory)能力。每一次会话都是一张白纸,在这种缺乏连续性的情况下,人类开发者就变成了它的记忆库。现在来看看你的工作图谱实际上是什么样的。在最底层是一个代码仓库依赖图谱。这其中包括了你们组织所生成的所有产物,加上你们依赖的每一个开源代码仓库,这里面也许有一千个你们自己拥有的代码仓库,以及数以万计的开源依赖库。在这个架构的最顶层,是所有创建和修改这些代码的智能体会话。会话与会话之间相互关联,代码仓库与代码仓库之间也相互关联。所以,这个图谱真实地反映了你们组织内部的工作全貌,它清晰地描述了底层存在着什么基础代码,以及它是如何一步步演变成顶层那个正在运行的系统的。

Original English

Speaker A: The second category is that the agent forgets. The agent has no episodic memory. Every session is a blank slate and the human in this case becomes the memory. Here what the graph of your work actually looks like. At the bottom there is a repository graph. the artifacts your organization produces plus every open source repo you depend on maybe a thousand repos you own and tens of thousands of open source repos at the top there are all agentic sessions that create and modify that code session relates to each other repos relate to each other so this graph is a faithful picture of the work in your organization it describes what there at the bottom and how it came to be at the top.

Speaker A: 这正是你希望你的智能体能在这里看到的全局视野。然而,它实际看到的只有孤立的一个会话、你的庞大代码库中极小的一部分碎片,并且毫无历史记忆。明白了吗?正是因为它能看到的东西实在是太少了,所以它只能被迫依赖那个真正理解整个系统的角色——也就是人类开发者。每个开发者的脑海中都装载着这个复杂图谱的一部分,对吧,至少在他们自己负责的业务领域内是这样。开发者们很清楚,一般而言智能体并不具备这种系统性的认知。如果这听起来还不够疯狂的话,那么请想象一下:一个智能体最多一次只能看到一个文件,并且只能往回翻看最近的五条消息,它无论是在空间(能同时看到多大范围的代码)还是在时间(能回顾多久之前的上下文)上都受到了极大的限制。你肯定会断言,在这样的条件下根本没法正常工作。然而,我们现在所拥有的智能体开发现状,就跟那幅疯狂的受限画面差不多,而且随着组织的规模越来越庞大、业务越来越复杂,这个问题就会变得越来越明显。

Original English

Speaker A: That's what you want your agent to see here. What it actually sees there is one session, one small fraction of your codebase, no memory. Okay? Because it sees so little, it leans on the one who understands the system, the developer. Every developer has a part of that graph, right, in their head, at least in the domain. They know the agent generally speaking doesn't if this doesn't sound crazy right imagine an agent that could see one file at a time maximum and can only look five messages back sort of constraint again both in space what can see and time how far in the past could see you would say that's impossible to work in what we have now is similar to that crazy picture and the more complex the organization is the more appar apparent it becomes.

Speaker A: 我来给你们展示一下我们是如何彻底解决这个痛点的。我曾交流过的其他一些组织也拥有类似的解决方案。所以,请大家从概念的角度来看待这个问题和相应的解决方案,而不是仅仅局限于某一个特定的工具。虽然我们开发的这个工具确实非常酷,我们构建了一个与具体智能体实现无关的“元控制台”(meta harness),它被命名为 Polygraph。好的,让我给你们展示一下它的具体功能,以及它是如何精妙地修复我们刚才讨论的那些棘手问题的。我们得出的第一个核心想法是,如果一个 GitHub 用户(任何普通用户都算在内)有权访问数千个不同的代码仓库——其中有些是他们自己组织拥有的,很多则是外部开源的——我们就可以对它们进行大规模的自动化分析,并从中提取出海量的元数据,从而构建出一个全局统一的依赖关系图。在这个分析过程中,那些目标代码仓库里的代码连一行都不会被修改,所有操作都是在侧边以非侵入式的方式进行的,对吧?然后我们可以获取这些经过清洗和结构化的元数据,将其输入到这个元控制台中,为智能体创造出一种身处一个巨大、统一的代码库之中的错觉,让它可以随时随地自由地读取和写入任何相关的代码。这是我个人的代码依赖图谱展示。我自己只有大约 300 个代码仓库,对吧?但是我的项目却依赖于成千上万个开源代码仓库。Polygraph 会精确计算出每一个仓库产出了什么,每一个仓库中的每一个子项目,在 npm 或 maven 包的层面上具体消费了什么依赖,它们对外暴露并生成了什么 API,同时又调用了什么 API,你们看。

Original English

Speaker A: I'll show you how we solved it. Other organizations I talked to have similar solutions. So, uh look at the problem and the solution conceptually, not a specific tool. Although the tool is pretty cool, we built uh an agent agnostic meta harness called Polygraph. Okay, let me show you what it does and how it fixes the issues we just discussed. The first idea that we uh arrived at is that if a GitHub user, any user has access to thousands of repos, some of them they own, many of them are open source, we can analyze them and extract a lot of metadata out of them to build unified dependency graph. Uh no line of code changes in those repo that all happens kind of on a side, right? And then we can get this metadata and feed it to the meta hardness and create an illusion of one big code base the agent can read and write anywhere. This is my personal graph. I only have about 300 repos I own, right? And thousands of open source repos my projects depend on. Polygraph computes what each one produces. each repo, each project in each reper, what each project in each repo consumes package wise, what API they produce and consume and look.

Vercel 构建生成式基础设施的历程

Andrew: 大家好,非常感谢你们的到来。我是 Andrew。我是 Vercel 的软件主管(Chief of Software),今天我来到这里,主要是要和大家深入探讨一下我们是如何在 Vercel 解决构建智能体(agent building)这一重大难题的。作为公司的软件主管,我的日常工作涵盖了许多方面,结合了大量的内部工程实践、外部的前沿实验,并且总体上一直处于技术发展的最前线,致力于构建全新的代码库、现代化的框架以及突破性的技术。对于在座那些可能还不了解 Vercel 的人来说,我想简单介绍一下,Vercel 致力于构建下一代生成式(agentic)的基础设施,好让开发者们可以毫无阻碍地构建未来的创新应用。我们在 Web 开发领域起步,初衷是帮助人们顺畅地发布网站和复杂的 Web 应用程序,让他们能够将全部精力集中在产品逻辑上,而不必为那些不能让应用变得更好的底层基础设施问题而烦恼。我们的平台可以毫不费力地自动扩展到百万级别的并发,也能在闲置时优雅地缩减到零。但与此同时,我们正敏锐地观察到人们想要构建的核心事物正在发生显著的变化。你知道的,人们最初起步于构建简单的静态网页,但现在,我们看到他们迫切地想要构建出能够自主思考和执行任务的智能体(agents)。

Original English

Andrew: Hey everyone, thanks for coming. I'm Andrew. I'm the chief of software at Verscell and I'm here to talk to you about how we solved agent building at Verscell. I'm the chief of software. So I work on a mix of internal engineering, external experimentation, and generally being at the frontier and building new libraries, frameworks, and technologies. For those of you that don't know Verscell, Verscell builds a gentic infrastructure so people can build what's next. We got started in the web world helping people ship websites and web apps without having to worry about the infrastructure that doesn't make their app any better. It can scale to a million and scale down to zero effortlessly. But we're seeing a change in what people want to build. You know, people started by building pages, but now we see them want to build agents.

Andrew: 因此,我们已经开始了一段非常相似的全新旅程,目标是为了让人们能够像过去部署网页那样,更轻松、更容易地构建智能体和基于智能体驱动的新型应用程序。为此,我们倾力构建了一个被称作 AI SDK 的开源工具包。有了它,你就不需要为了适配不同模型而费力去替换 300 到 400 行针对特定提供商的繁琐代码,你仅仅需要替换一行简明的代码,因为我们在所有这些截然不同的模型提供商的底层,精心提供了一致、标准化且类型安全的模型交互接口。不仅如此,我们还构建了许多其他的配套工具,以便开发者能够更容易地实现模型之间的优雅后备切换(fallbacks)、安全可控的代码执行沙箱环境、在模型不活跃或持续等待长响应时获得更优化的财务定价策略,以及提供至关重要的状态持久性和会话可恢复性功能。

Original English

Andrew: And we've been embarking on a similar journey to make it easy for people to build agents and agentic applications easier. We built this thing called the AI SDK. So instead of needing to switch out 300 400 lines of provider specific code, you just switch out one line of code and we have the same model interface underlying for all these different providers. We built a lot of other tools to make it easier to have model fallbacks, secure code execution, better pricing when it's inactive and waiting for responses, as well as for durability and resumability.

Andrew: 而我今天之所以站在这里,正是要和大家谈谈大约在一年前,我是如何开始进行这项在外人看来可能非常疯狂的探索性实验的,这场实验不仅直接引发了 Vercel 内部各种智能体应用的爆发式增长,还促成了我们最近构建推出的一个非常酷的核心产品……回顾历史,这可能要追溯到 1980 年。在属于我的时代到来之前,传奇人物比尔·盖茨曾经说过这样一句名言,他想象在不远的未来,每一张办公桌上和每一个家庭里都会稳稳地放着一台个人电脑。大家知道,在那个年代,这可能是一个非常具有颠覆性、甚至有些反常规的想法,而到了今天,这种情况看起来简直再正常、再普遍不过了。所以,我和我们的 CTO 也有过一次类似这样火花四溅的思考,那就是,既然计算机已经普及,那么与其仅仅追求每张办公桌上有一台电脑,我们是否有可能更进一步,让每张办公桌上都配备一个专属的智能体呢?

Original English

Andrew: And I'm here to talk to you about how I went on this crazy experiment roughly a year ago that led to a aentic explosion at Verscell and led to a really cool thing that we built recently about uh this is 1980. Uh, Bill Gates before my time had this quote saying he imagined there would be a computer on every desk and in every home. You know, that was probably pretty contrarian then and today it seems like very normal to have that happen. And me and the CTO had this thought, you know, instead of a computer on every desk, could we potentially have an agent on every desk?

Andrew: 大家都知道,目前在市场上,我们通常只将智能体真正用于自动编码、脚本生成和其他高度技术性的工作负载,但我们已经清晰地开始看到,这种能力正在快速向设计创新、产品管理、数据分析等其他非技术垂直领域大规模扩展。这大概是在一年前发生的事情了,所以我完全可以说,我涉足这个生成式 AI 的实践领域算是相当早的了。那个时候,大家最常使用的还是类似 Claude 3 Sonnet 和早期的 GPT-4 这种基础模型,整体的基础设施和生态情况远没有今天这么成熟和复杂。当时我试图切实地去探索一下这个想法,看看我们究竟能在日常工作中利用它做些什么实质性的改进。于是,我走访了 Vercel 内部的各个核心工作职能部门,这其中广泛包括了营销、销售、财务、法务等团队,我向他们每一个人提出了一个直击灵魂的问题:‘在你们每天的日常工作中,你们最讨厌或者觉得最痛苦的环节是什么?’

Original English

Andrew: You know, today we only really use agents for coding and technical workloads, but we're starting to see expansion into things like design and product management and other verticals. And this was maybe about a year ago, so I would say I'm pretty early to this, but that was when it was like sonnet 4 and things weren't as sophisticated as they were today. And I tried to actually explore this out, see what we could do about it. I went around to various job functions, adversel, marketing, sales, finance, legal, and I asked them, 'What do you hate most about your job?

Andrew: 最终,我听到的最令人信服、也最亟待解决的使用场景来自我们的数据团队,他们当时正处于高速扩张期。他们原本是一个非常精干的小规模团队,但 Vercel 整体的业务增长速度实在太快了,远远超过了他们的人手扩张速度。你们知道的,他们手头有越来越多来自大量活跃客户的原始数据、复杂的分析报告、系统运行指标和实时的销售数据。他们别无选择,只能夜以继日地不断聚合这些数据,并不断地清理加工,使其可供业务线自身使用。而在那个特定的时间点,如果你仔细想想数据科学人员每天都被迫要做些什么:每当营销团队或销售部门的任何人对某位大客户或某个新产品的表现有一丝疑问时,数据科学团队就不得不立刻放下他们手头正在进行的任何高价值模型构建工作,去编写琐碎的 SQL 查询语句,处理冗长的数据拉取过程,进行基础的数据分析,然后带着关于下一步操作的分析建议疲惫地回复。这真的是破坏整个团队生产力的致命杀手。要知道,数据团队根本不想整天被迫放下一切重要工作,仅仅为了充当人肉 SQL 生成器。因此,我主动与我们的数据部门副总裁(VP of Data)展开了紧密的合作,试图为他们这群极具才华的工程师们构建一种更好的、由 AI 赋能的运营方式。

Original English

Andrew: And the most compelling use case I heard was that the data team, they were growing. They were a very lean team, but Versel was growing faster. You know, they had so much more data from customers, analytics, metrics, sales. They just had to keep on aggregating and keep on making available for themselves to use. And at this time, if you think about what the data science people ever have to do, whenever someone from marketing or sales has a question about a customer or product, the data science team has to drop everything they're doing, write the query, process it, do an analysis, and come back with some recommendation on what to do. And this was really killer to productivity. You know, the data team did not want to drop everything and just write queries all day. And so I worked with our VP of data to try to build a better way for them to operate this way.

Andrew: 所以,如果你仔细想象一下,当你满怀热情地想尝试用最新的 AI 技术来解决业务问题时,你要做的第一件事究竟是什么,你很可能会选择简单粗暴地构建一个包含所有上下文的巨大超级提示词(mega prompt)。就是说,你心里有一个明确的问题,你把它一股脑地传给某个大语言模型(LLM),耐心等待它输出回答,然后事情就结束了。要知道,我们最早期的第一版原型看起来真的就是这个粗糙的样子。老实说,我当时直接向数据团队要了一份完整的 Snowflake 内部数据库模式(schema)的导出数据字典。我把它直接和我要问的问题一起,硬编码粘贴到了系统提示词的上下文中,当它在几秒钟后生成了对应的 SQL 代码时,我甚至直接把它手动复制粘贴出来,然后进入数据库终端亲自运行它。你知道,我只是想先验证一个假设:在给定一些质量还算不错的表结构模式信息的前提下,如今的这些前沿语言模型是否真的已经具备了足够的推理能力,能够完美地编写出语法结构正确且逻辑有效的复杂 SQL 代码……

Original English

Andrew: And so if you think about the very first thing you would ever do if you want to try to use AI to solve a problem, you may just build like a huge mega prompt. You know, you just have a question, you pass into an LM, you have it respond, and that's it. You know, this was how the first version really looked. Honestly, I asked them for a dump of of the Snowflake schema. I pasted it into a system prompt with a question and then when it generated SQL I actually copy and pasted that in and just ran it myself. You know I just want to see are the models good enough today in order to write valid SQL given some decent

From Multi-Agent to Mega Agent

Speaker: 架构层面,我想说这给了我们一点信心:你知道,如今的模型虽然还不够完美,但也许我们可以通过提示词工程,或者改善其周围的上下文,并提供更多的护栏,从而让它运行得更好一些。因此,如果你仔细想一下数据科学家在收到问题时到底需要做什么,你会发现,他们必须先处理这个问题,可能还需要探索语义层,并切实弄清关联模式究竟是怎样的;接着他们会实际去执行 SQL。如果 SQL 无法执行或者开销过大,他们可能还需要返回去重新操作。最终他们会对结果进行报告,包括将数据可视化,也许会写上几段文字说明,也许会做个回顾复盘,或者做些其他类似的事情。

Original English

Speaker: structure. And I would say this gave us a little bit of confidence that you know models today aren't that good but maybe we can harness engineer or make the context around it a little better and give us some more guardrails to operate a little better. And so if you actually think about what a data scientist actually needs to do when they get a question, you know, they have to process the question. They may have to explore the semantic layer and actually figure out what the join patterns are. They will actually go and execute the SQL. They may go back and do that again if the SQL did not execute or was too expensive. And they'll eventually report on it, including visualize the data, maybe write some paragraphs, maybe do a retro, maybe do some other stuff.

Speaker: 考虑到这些不同的阶段,我和数据副总裁试着坐下来,将它们映射到具体的智能体工作负载中。于是,这个数据科学智能体的第二个版本——我称之为 D0,之后我都会用 D0 来指代它——就诞生了。它的工作流是:你提出一个问题,我们有一个查询智能体会将查询传递给规划智能体,然后是执行智能体等等。如果你把这些全部串联起来,你实际得到的就是类似这样的架构。在这里,每个智能体都有一个非常专用的系统提示词,专注于它使用那些严格限定于该特定功能的工具所执行的任务。所以,以这里的一个例子来说,你可以看到对于第一个环节,规划智能体拥有一个“读取实体 YAML”以及一个“搜索模式”工具。因此,它只会使用这些特定的能力,直到得出一个可以传递给(下一个)规划智能体、然后再传给 SQL 智能体、最后传给报告智能体的结果。

Original English

Speaker: And so if you think about those different phases, me and the VP of data tried to sit down and map those out into specific agent workloads. And so the second version of this data science agent uh called D0. I'm going to reference D0 from now on is you ask a question, we have a query agent that passes on a query to the planning agent that will then have an execution agent, etc. And if you chain all of these together, you actually get something that looks like this, where each agent has a very dedicated system prompt focused to what that does with tools scoped to exactly that function. So an example here, you can see that for the first one, the planning agent has a read entity YAML and a and a search schemas tool. And so it will only use those capabilities until it has an answer to pass on to the planning agent and then to the SQL agent and then to reporting.

Speaker: 事情当时正往好的方向发展。你知道,我们终于可以摆脱手动复制粘贴 SQL、然后再回来根据结果撰写报告的日子了。它现在实际上已经在完成从提出问题到给出答案的端到端闭环了。但是,使用这种架构我们开始四处碰壁。大约就在这个时候,我们得出了一个结论:你知道吗,你真正需要的其实是一个包含了所有“巨型上下文”的单一智能体,并且让它能够某种程度上管理自己的内存。也就是在那个时期,我们意识到你需要让智能体能够回顾自己已经做过的事情,进行一定程度的反思,并理清走到这一步都经历了哪些步骤。

Original English

Speaker: And this was getting better. You know, we were able to get away from having to copy and paste a SQL and have to come back and report on it. It was now actually doing like the end toend loop from question to answer. But we started hitting some walls with this architecture. And around this time we came to the conclusion that you know what you actually need is you need one agent with all the mega context within it and for it to sort of manage its own memory. You know this was around the time when we realized that you want to actually have the agent be able to look back on what it's done sort of reflect and figure out the steps that got to get here.

Speaker: 回顾之前的模型,你可能已经注意到,下一个智能体能够获取到的唯一信息,仅仅是一个摘要以及上一步操作的极小片段。而采用这种新方法,你可以想象你拥有一个“巨型智能体”,它在内部自己管理状态:有时它在规划,有时在构建,有时在执行,有时则在报告。这大概就是它的工作模式——你执行一次大型的 AI 调用,也许最大步数设为 100,然后赋予它根据当前处于执行旅程中的哪个位置来管理自身状态的能力。因此,你能看到相似的工具,能看到相似的架构形态,但这种方式最棒的一点在于:如果它在执行、调试或表连接时遇到任何错误,它可以返回去进行更多的探索,或者去阅读更多内容,以找出自己到底哪里做错了。在这一点上它表现得非常出色。我们当时对手头这个实际运行的系统相当自信,并把它推广给了我们小组里少数几个值得信任的成员。你知道,这是一个极其强大的工具,我们真的不想把它交到错误的人手里,或者让人们把它用在非常关键的工作负载上。

Original English

Speaker: And with the previous model you may have noticed that the only thing that the next agent gets is a summary and a small snippet of the previous thing that was done. Now this way you can imagine that you have one mega agent and it internally it manages its own state at some points it's planning some points it's building some points it's executing and some points it's reporting and this is sort of what it looked like you know you have one big AI call maybe max steps 100 and you give it the ability to manage it own state based on where it's at inside of its execution journey and so you can see similar tools you can see a similar shape but the best part about this is if it ever ran to an error when executing shooting or joining. It could go back and explore more or it could go and read more and figure out what it was doing wrong. And it was very good at this point. We were pretty confident in the actual system at hand and we actually spread it to a few trusted members of our cell. You know, this was a very powerful tool and we didn't really want to put in the hands of the wrong people or people that were using very critical workloads.

Speaker: 所以,我们把它交给了少数几个人,而立竿见影的反馈却是:它糟透了。你知道,我们本来以为我们在“大展厨艺”(做得很好),我们以为这个系统能够完美搞定我们 30% 的评估用例。但我们无论如何也预料不到用户会问出那样的问题,这也导致我们不得不花更多时间去手动映射梳理这些新出现的场景。这似乎并不是一种非常具备可扩展性的解决方式。

Original English

Speaker: So, we got into a few people's hands and the immediate response was it was awful. You know, we thought we were cooking. We thought this was, you know, nailing 30% of our eval. we couldn't have anticipated some of the questions that were being asked and for us to spend more time manually mapping out some of these scenarios. It didn't seem like a very scalable way to do this.

The File System Agent Unlock

Speaker: 就在那个时候,Claude Code 和 Opus 4.5 发布了——或者更准确地说,是 Opus 4.5 发布了,而且它与 Claude Code 配合使用时简直强大到不可思议。你知道,它们在某种程度上解锁了“文件系统智能体”(file system agent)的概念。而在旁边观察的我们忍不住惊叹:“哇,相比于我们之前拥有的东西,Claude Code 加 Opus 4.5 基本上就等同于 AGI(通用人工智能)了。”你要知道,与我们手搓的那个智能体相比,它几乎能毫不费力地回答我们绝大多数的问题。当我们试图退一步思考我们到底做错了什么、以及为什么这个新东西会好这么多时,我们意识到,那个巨大的“突破口”就在于它仅仅是一个文件系统。你看,它只有一组极其精简的工具集:列出文件(list file)、读取文件(read file)、运行 bash 命令(run bash),而我们只在这里为了我们自己的数据智能体用例额外多给了几个工具。但最关键的地方在于,它能够使用那些智能体已经受过良好训练的工具,并且能够在需要的时候自主探索并编写所需的工作内容。它不再受限于预先设定好的步骤——你看,我们并没有给它提供…… Claude Code 并没有给它提供一套非常指令性、教条式的工具。它更像是在放开手脚让其自由发挥,去探索和涌现出新的行为。

Original English

Speaker: And then claude code and opus 4.5 came out or more like opus 4.5 came out and it in tangent with claw code was just so powerful you know they sort of unlocked the concept of a file system agent and we on the side were like wow cloud code and opus 4.5 is basically agi compared to what we had before you know it would answer most of our questions without even without even missing a beat um compared to the handgrown agent we had and when we tried to step back and wonder what we were doing wrong and why this was so much better, we realized that the big unlock was that it was just a file system. You know, we it had a very minimal set of tools, list file, read file, run bash, and we gave a few more here for uh our own data agent use case. But the biggest thing was it was able to use the tools that agents are well trained on and is able to explore and write work where it needs to. you know, we weren't giving it cloud code was not giving it a very prescriptive set of tools. It was sort of just letting it go wild and explore emergent behavior.

Speaker: 所以从中我们学到了:你真的可以就只使用文件系统。你看,我们看到了从 Claude Code 中汲取的经验,既然它只是在本地执行,就能发挥出如此强大的威力。于是我们试图以一种非常“Claude Code 风格”的方式来重构我们的系统。现在它将会在一个沙盒(sandbox)中运行。这个沙盒会将整个语义层直接倾注其中。接着,智能体将能够通过 grep 搜索、运行 bash 命令、读取文件、写入文件等手段,在整个文件系统中四处摸索以确定它所需要的东西。而我们只需要在最顶层零星添加几个专用的工具,以确保它能够完成所有 Vercel 特有的业务逻辑。

Original English

Speaker: And so from this, we learned that you can really just use the file system. You know, we we saw the learnings from claw code and how powerful it was given that it just executes locally. And we tried to rebuild it in a way that was very cloud codeesque. You know, it was now going to run in a sandbox. That sandbox would dump the whole semantic layer into it. You could the agent would be able to grap, bash, refile, write file all around to figure out what it needs and we would just sprinkle a few tools on top to make sure it could do everything that is Verscell specific.

Speaker: 而这实际上是有史以来最大的一次突破。你要知道,从单一智能体跨越到 Claude Code SDK,然后再从 Claude Code SDK 跨越到通用的、专为我们的用例微调或定制开发的文件系统智能体,这是一个令人惊叹的飞跃。到了这个阶段,我们已经开始准备在 Vercel 内部将其推广给更多的人使用了。也是在这个时候,我们的评估得分基本上翻了一倍。

Original English

Speaker: And this was actually the biggest unlock ever. You know, the leap from single agent to cloud code SDK and then from cloud code SDK to file system agent in general, fine-tuned or purpose-built for our use case was an amazing leap. At this point, we were starting to get ready to give it away to more people adversel. And at this point, the eval score basically doubled.

Speaker: 然后我写了这篇……嗯,这基本上就是它现在的样子。它非常简单。你只需要提供一个 bash 工具。我们在 npm 上有一个非常好用的辅助工具就叫 bash-tool。你把它连接到一个沙盒上。同时,你可以把各种文件挂载到这个沙盒中,供它去读取、写入和执行。在获得这个启示之后,并且在我看到我们成功通过了那么多以前没能搞定的测试问题之后,我写了一篇极其火爆的博客文章。它……它其实今天已经上线了。而且在我写这篇文章的那一周,它为我们的 vercel.com 贡献了 70% 的流量。所以你知道它有多么火爆了吧。

Original English

Speaker: And I wrote this uh this is basically how it looks. Um it's very simple. You just give a bash tool. We have a nice helper called bash tool on npm. And you attach it to a sandbox. And you can attach files to the sandbox for it to read, write, and execute. And after this revelation and after I saw that we were passing so many of the questions that we failed to do before, I wrote this banger blog post. It's uh it's actually up today. And the week that I wrote this, it was responsible for 70% of our versel.com traffic. So you know it's a banger.

Distilling Queries into Skills

Speaker: 在那之后,下一个合乎逻辑的步骤是我们需要弄清楚我们面临的常见用例到底有哪些。所以到那个时候,我们在某种程度上已经把它放宽给了 Vercel 的所有同事使用。我们每天都能收到数以千计的查询请求,人们想要查询的东西五花八门,从客户指标、销售指标、各项数字指标到 npm 下载量无所不包。然后我们发现,其中有很多查询在形式和结构上实际上是完全相同的。你想啊,执行聚合计算的方式就那么几种,查找某个产品的方式也就那么几种,查询账单信息的方式同样只有那么几种。因此,我们实际上建立了一个定时运行的后台任务,它会提取最近一段时间的查询请求,并试图将它们提炼、浓缩成一个“技能”(skill)。目前,我们大约拥有 100 个这样的技能,它们涵盖了从聚合分析一直到查找特定用户具体数据的各种混合操作。

Original English

Speaker: And after that, the next logical step was that we wanted to figure out the common use cases we had. So by then we've already sort of let it leash on all of our cell and we were getting thousands of queries a day from people wanting everything from customer metrics sales metrics number metrics npm downloads and it turns out that a lot of these queries are actually the same in shape you know there's only so many ways you can do an aggregation only so many ways you can look up a product only so many ways you can do billing info and so we actually have a recurring job that takes the most recent queries and tries to distill them into a skill and right now we have roughly 100 skills that do a mix of aggregation all the way through looking up specific data about certain people.

Speaker: 我们发现这种方法极其有效,因为你想想看,每一次全新的智能体运行,它在某种程度上几乎都是从零开始的。你知道,除了语义层和系统提示词之外,它实际上没有任何预先建立的上下文。但是,如果配备了“技能”,它在启动时就已经自带了大量的上下文知识,而这些知识否则是需要它自己去一步步探索完成的。

Original English

Speaker: And we found this very effective because if you think about every new agent run, it sort of just starts from nothing. You know, there's really no pre-established context besides, you know, the semantic layer and the system prompt. But with a skill, it already starts off with a lot of contextual knowledge that has otherwise already been done.

Speaker: 而这大概就是它现在的基本形态。它与前一个版本非常相似,但是加入了一个 skills(技能)文件夹,这一改动实际上发挥了非常强大的作用。此外,我们还在 Vercel 构建了一个名为 Skillsh 的工具。这是目前查找智能体技能并亲自运行它们的最流行的方式。我之所以说这么多,是因为这段旅程是你们中的大多数人偶尔也会经历的:你从一个简单的东西开始,然后逐渐增加它的复杂性,最终,你会打造出一个可以部署到生产环境中的系统。我之所以告诉你们这些,是因为在构建这个智能体的每一个步骤中,Vercel 公司里都有人对智能体充满了好奇心,他们试图从我的 D0 智能体中分支出去,然后构建属于他们自己的版本。而在每一个步骤中,我们……

Original English

Speaker: And this is roughly how it looks. It's very similar to the previous one, but the inclusion of a skills folder is actually very powerful. Um, we also built this tool at Verscell called Skillsh. It's the most popular way to find agent skills and run them yourself. And I I'm saying all this because this journey is something that most of you may hit once in a while where you start from something simple and you gradually add complexity and you eventually hit a system in which you can ship to prod. And I'm telling you this because at every step along building this agent, someone at Evercell was agent curious and they tried to fork off of my Dzero agent and build their own. And at every step, we

Eve: The Next.js for Agents

Andrew: 某种程度上,我们有了一种比以往任何时候都更好的方式,来做一些以前未曾为人所知的事情。我们在想,如果今天的人们可以直接从最新的洞见开始,而不必总是从一个简单的提示词(prompt)起步,或者不必每次都从最基本的原理开始去重新发明所谓的最优原则,那会是怎样的情景呢?因此,我们实际上就在思考:如果我们为智能体(agents)构建一个类似于 Next.js 的框架会怎样呢?

Original English

Andrew: sort of had a better way to do something that was not previously known. And we were wondering like what if people today could start from the very last insight and not have to ever start from just a simple prompt or from reinventing best principles from first principles. And so we actually thought what if we built the Nex.js for agents.

Andrew: 对于那些可能不太了解的人来说,Next.js 是 Vercel 构建的一个非常流行的现代 Web 框架。正是这个框架发明了这种“文件系统定义基础设施”(file system framework defined infrastructure)的概念。也就是说,你根本不必去担心这些东西最终应该部署在哪里。你只需要按照正确的约定和规范来编写文件,框架就会自动声明它们应该去的地方。

Original English

Andrew: For those that don't know, Nex.js is a popular web framework that Verscell built that invented this thing of file system uh framework defined infrastructure. You don't have to worry about where things go. You just have to write files in the right conventions and it automatically declares where they should go.

Andrew: 你的页面会自动分发到 CDN 上。你的无服务器函数(serverless functions)会自动部署到相应的计算节点。你的缓存逻辑则会自动置于这两者之间。于是我们就想,你知道的,构建智能体(agents)的过程其实也应该像这样简单才对。你只需要创建一个技能(skills)文件夹,一个工具(tools)文件夹,以及一个渠道(channels)文件夹,然后你就可以非常轻松地声明这些组件了。

Original English

Andrew: Your pages go to the CDN. Your serverless functions go there. Your caching goes in the middle. And we thought, you know, building agents should be this simple. You should only have to create a skills folder, a tools folder, a channels folder, and you should be able to just declare these very easily.

Andrew: 而这个框架应当确切地知道,如何利用这些组件来构建出一个完整的智能体。正因如此,就在两周之前,我们正式发布了 Eve。Eve 是一个专门为智能体设计的框架,就像是智能体领域的 Next.js。使用它,你可以非常轻松地从一个简单的示例模板起步,迅速获得一个完全准备就绪的智能体,并且能够自由地加入你自己的自定义知识库、你自己的自定义工具,甚至可以将其无缝集成到你所熟悉的各种渠道之中。

Original English

Andrew: And the framework should know exactly how to make an agent out of it. And that's why two weeks ago we released Eve. Eve is a agent framework like the next.js JS for agents where it's very easy from just starting with a sample template to having a fully agent ready and being able to add in your own custom knowledge, your own custom tools and even integrate it into the channels that you are familiar with.

Andrew: 这大概就是我们心目中一个智能体真正应该具备的模样。你知道,一个智能体必须拥有一个运行环境(runtime),同时它也需要有各种渠道(channels)。而在那个运行环境里,你需要具备持久性(durability)。你会希望在一个完全隔离的安全环境中运行各种任务。你会希望能够自由地调用不同的底层大模型。同时,你也会希望能够建立各种各样的连接通道。

Original English

Andrew: This is roughly what we think an agent actually looks like. You know, an agent has a runtime and it has channels. And in that runtime, you're going to have durability. You're going to want to run things in an isolated environment. You're going to want to call into different models. And you're going to want to have connections.

Andrew: 而且,我们在构建这个框架时,充分考虑了开源生态的兼容性。你知道的,我们构建了 Eve,因此你可以非常方便地插入你自己的开源适配器,比如用于 PostgreSQL、OpenAI 的 responses API、Docker,以及其他的各种连接器。不仅如此,我们还让它在 Vercel 上的部署变得极其简单。你在这里看到的唯一不同之处在于,这里的所有东西都在使用我们多年来一直在构建的一款 Vercel 产品,其目的就是为了让你能够轻松地构建这些体验。

Original English

Andrew: And we built this with open source in mind. You know, we built Eve so you can plug in your own open source adapters for Postgress, OpenAI's uh responses API, Docker, other connectors. But we also made it incredibly easy to deploy in Verscell. The only thing here you see different is that everything here is using a Verscell product that we've been building over the years in order to make it easy to build these experiences.

Andrew: 比如用于实现持久性的 Vercel workflows,用于安全执行代码的 Sandbox,还有我们刚刚发布的 Vercel connect,这个新工具可以让你非常轻松地为各种连接生成短期的 ODC 令牌。实际上,在我们构建 Eve 的同时,我们也用 Eve 完全重写了整个 D0ero 智能体。相比于你之前在代码中看不到的那些在后台错综复杂的逻辑结构,嗯,这大概就是文件系统现在看起来的样子。

Original English

Andrew: Verscell workflows for durability, Sandbox for secure execution, and Versell connect, something we just released to make it easy to generate short-lived ODC tokens for connections. And we actually rewrote the whole D0ero agent in Eve as we were building Eve. And from the convoluted structures behind the scenes that you did not see from the code, um, this is roughly how the file system looks.

Andrew: 它现在变得非常简单。你只需要一堆系统指令,几个技能,几个工具,就能够非常容易地将这些组件组合成一个真正的智能体,而且对它进行迭代也非常容易。事实上,就在两周前我们在伦敦的活动上正式发布它之前,我们已经将其提供给了一些测试阶段的客户。其中有一家与我们合作非常密切的公司,Aura,他们已经用它重建了他们的智能体。这个智能体有点像一个微型的 Claude,专门去测试人们的服务。

Original English

Andrew: It's very simple. You have a bunch of system instructions, a couple skills, a couple tools, and it's very easy to compose this into a real agent, and it's very easy to iterate on. We actually gave this out to a few beta customers before we actually fully released it two weeks ago at our London event. And this one company that partners closely with us, Aura, they've rebuilt their agent that's sort of like a mini claw to go and test people's services.

Andrew: 它会访问各种网站,安装它们,尝试使用它们。他们发现,与使用现成的云代码(cloud code)相比,使用 Eve 从零开始构建他们自己的智能体,取得了令人难以置信的巨大成功。他们用更少的步骤,获得了更高的成功率,同时也收获了更深刻的洞察。当你把 Eve 部署到 Vercel 时,你还能开箱即用地获得强大的可观测性(observability)。你可以在这里看到,你能够掌握所有的智能体运行情况,你能看到所有的工具调用,你能清楚地追踪它所采取的每一个步骤,此外还能看到一些估算的成本,以及你可能采取的一些优化措施。

Original English

Andrew: It goes to websites, installs them, it tries to use them, and they've seen incredible success on building their own agent from the ground up using Eve compared to using an off-the-shelf cloud code. Fewer steps, better successes, as well as better insights. And when you deploy Eve to Verscell, you get observability observability out of the box. You can see here that you get all the agent runs, you see all the tool calls, you see each step it takes, as well as maybe some estimated costs and some optimizations you could potentially take.

Andrew: 你们今天就可以在 eve.dev 上开始使用它。你只需要把代码克隆下来,就可以直接从一个模板开始启动,不仅部署非常容易,如果你有需要,也可以进行自托管(self-host)。我之所以提到这个,是因为我希望未来能看到越来越多针对特定商业用例的智能体。你知道,在我们构建 Ezero 之前,我们实际上是在战场上真刀真枪地测试了许多那些业内资金非常雄厚的初创公司。他们当时都在做这种垂直领域的智能体,专门致力于接管你的 Snowflake 实例,让他们的智能体可以在上面执行各种 Snowflake 查询。

Original English

Andrew: And you can get started today at eve.dev. You can just clone it and you can just start a template, deploy easily, self-host if you need. And the reason why I bring this up is because I hope that there will be more and more business specific use case agents. You know, before we built Ezero, we actually battle tested a lot of the industry well-funded startups that were doing these vertical agents that were dedicated to taking your Snowflake instance and making it so their agent could run Snowflake queries against it.

Andrew: 但是我们后来发现,真正让这个智能体变得如此优秀的决定性因素,在于它拥有大量非常具体的公司领域知识。你知道的,鉴于 Vercel 本身是一家基于 Web 的公司,我们有大量的客户拥有自己的网站和各种网络资产。这就需要我们更加深入地了解,在什么情况下你应该查询什么样的数据,以及哪些事物之间是如何相互关联的。因此,很多这种现成的、开箱即用的智能体,它们确实很棒。它们非常适合用来尝试,但我认为,如果你真的想把这个工具的作用发挥到极致、物尽其用,你真的应该尝试去构建你自己的智能体,并且尽可能多地将你们公司特定的领域知识融入其中。

Original English

Andrew: But we found out that what really makes this agent good is it has a lot of very specific uh company knowledge. You know, the way that Verscell is a web-based company, we have a lot of customers that have websites and web properties. That goes a lot deeper into when you should query for what and what things link to what. And so, a lot of these off-the-shelf agents, they're great. They're good to try, but I think if you really want to get the most juice out of a squeeze, you should really try to build your own agent and add in as much company specific knowledge as you can.

Andrew: 到了今天,你知道,在 Vercel 内部我们大约已经有了 20 个产品市场契合度(PMF)相当不错的智能体。它们的用途非常广泛,从用来做营销复盘以找出应该联系哪些潜在客户,到法务团队在面临新的商务谈判时对合同进行的第一次历史性修改(red line),一直涵盖到帮助我处理数据查询的数据科学智能体。所有这一切都有力地证明了,我们在 Vercel 在拥抱智能体方面已经走得非常深远了。你知道,所有这些智能体工具实际上为我们节省了大量宝贵的时间。

Original English

Andrew: Today, you know, we've had 20 roughly decently PMF agents at Verscell that range from anything from marketing retros to figure out who to reach out to to the first ever red line of a contract when legal sees a new negotiation all the way to my data science agent helping with with data queries. And that goes to show that we at Everscell have been very agent-p. You know, all this stuff is actually saving us a lot of time.

Andrew: 我们的数据团队的生产力达到了前所未有的高度。他们现在有了更多的时间,去专门优化 Snowflake 的性能,去添加那些之前遗漏的全新数据源,去填补那些他们以前因为一直忙于编写各种枯燥查询而根本没时间去处理的空白。而且我认为,对于在大型、中型或小型公司工作的你来说,现在要把那些你根本不想做的事情,或者那些你花了太多时间在上面的事情自动化,已经变得前所未有的简单了。你知道,我认为很大一部分的人力资源(HR)、财务(finance)以及销售(sales)工作,是可以在一定程度上用智能体来自动化的。

Original English

Andrew: The data team has never been more productive. They have more time to go and improve the performance of Snowflake, to add new data sources that were missing, to fill in the gaps that they previously did not have time to because they were so busy writing queries. And I think it's never been easier for you at your big, small, medium-sized company to sort of automate away some of the things that you do not want to do or some of the things that you're spending too much time doing. You know, I think a lot of HR, finance, sales can be somewhat automated with agents.

Andrew: 并且我坚信,在今天,Eve 绝对是用来构建上述这些智能体的最佳方式。这些是我的社交媒体账号。非常感谢大家今天能来参加,并耐心聆听我的分享。我是 Andrew,接下来的时间我会在场外,如果你想深入聊聊,随时可以来找我。

Original English

Andrew: And I think Eve is the best way to build said agents today. And these are my socials. Thank you all for coming and listening. I'm Andrew and I'll be around if you want to chat outside.

The Log is the Agent

Isan: 大家好,我是 Isan,Amnara 公司的首席执行官。今天我要来和大家探讨的主题是:“日志就是智能体(the log is the agent)”。我今天演讲的基本思想非常简单,那就是:现在大多数人一想到智能体,通常就会认为智能体就是那个底层的大语言模型,或者是它正在运行的那个执行环境。但我认为,这是一种完全错误的抽象。我个人认为,真正赋予一个智能体独特身份标识的东西,其实是它的日志(log)。而这也正是我今天在这里要向大家重点论证的观点。

Original English

Isan: Hey everyone, I'm Isan the CEO of Amnara and today I'm going to be talking about the log is the agent. The basic idea of the talk is simple and that is most people think of an agent as the model or the execution environment that it's running in and I think that that's the wrong abstraction. I think that the thing that actually gives an agent its identity is its log. And that's what I'm going to be arguing today.

Isan: 那么,大家不妨想一想这样一个场景:你在你最喜欢的视频游戏中,假设在这个例子里是《上古卷轴5:天际》(Skyrim),花了一百多个小时去精心培养的一个游戏角色。你的这个角色到底是什么呢?它是那个驱动游戏的引擎吗?它是那台 PlayStation 游戏机吗?还是你手里的那个控制器?不,它完全不是这些东西。那些东西当然很重要,正是通过那些设备,我们才能进行互动,正是它们运行着这个游戏角色。但是,它们其中没有一样东西等同于你的角色。

Original English

Isan: So, think about a character you've spent a 100 hours playing in your favorite video game, in this case Skyrim. What exactly is your character? Is it the game engine? Is it the PlayStation? Is it the controller? No, it's not. Those things matter and those things are what we'll interact with and they'll run the character. But none of those things are your character.

Isan: 你的角色,本质上就是数据(data)。它就是那个存档文件(save file)。这一点之所以如此重要,是因为假设你的 PlayStation 突然起火烧毁了,你的游戏角色并不会因此而消失。你可以去买一台全新的 PlayStation。然后你可以从云端把你的存档文件下载下来,接着你就可以完全恢复到他们之前所在的那个精确状态继续游玩。这就是因为这个智能体(角色),它的身份标识、它的历史轨迹以及它的当前状态,所有这一切都被完整地捕捉并保存在了它的数据之中。这个角色,是真真正正地活在数据里面的。

Original English

Isan: Your character is data. It's the save file. And this is important because if your PlayStation bursts into flames, your character isn't gone. You can buy another PlayStation. You can download your save file from the cloud and you can resume exactly where they were. And that's because the agent and its identity and history and its state is all captured in its data. The character lives in the data.

Isan: 而这也正是我今天想要引入到智能体领域的一种思考框架。在今天,当人们谈论起智能体时,他们常常会把手指指向错误的地方。他们会说,智能体就是那个大模型,或者他们会说,智能体就是那个运行环境。再次重申,正如我前面提到的那样,那些东西固然非常重要,但它们绝对不是智能体本身。智能体,就是它的数据。具体来说,它就是那份日志。那么,这份日志究竟是什么呢?

Original English

Isan: And this is the framing that I want to bring to agents. Today when people talk about agents, they usually point at the wrong thing. They'll say that the agent is the model or they'll say that it's the runtime. And again, as I mentioned earlier, those things matter, but they're not the agent. The agent is its data. It's specifically the log. So what actually is the log?

Isan: 在最简单的层面上来看,日志就是这个智能体持续追加记录的事件历史(appendon event history)。它囊括了用户的每一次输入,模型的每一次输出,每一次的工具调用,工具返回的具体结果,遇到的权限问题,甚至是发生的失败。这里的核心理念在于,这个智能体所经历的每一次状态转换(state transition),都会被如实地写入到这份日志中。这一点之所以至关重要,是因为它意味着这个智能体的身份,不再与某个特定的运行环境、特定的模型或者是特定的工具死死绑定在一起了。

Original English

Isan: At the simplest level, the log is the appendon event history of the agent. It's every user input, every model output, every tool call, tool result, permission, failure. And the idea is that every state transition that the agent takes is written to the log. This is important because it means that the identity of the agent isn't tied to the runtime or the model or the tools.

Isan: 所有那些东西,它们都只不过是在解释这份日志,并向日志中追加新的内容而已。它们的作用就是去读取日志,基于日志采取行动,然后把下一个发生的事件写回到日志里。而这一过程之所以如此重要,是因为这样一来,仅仅依靠这份日志本身,就完全足以恢复并恢复运行这个智能体了。一旦你把智能体明确地定义为它的日志,那么整个系统的其余部分,在理解和推理起来就会变得容易得多。

Original English

Isan: Those things are all just interpreting and appending to the log. They're reading the log, acting on it, and writing the next event back. And that's important because then just using the log on its own is enough to resume the agent. Once you define the agent as the log, the rest of the system becomes a whole lot easier to reason about

Isan: 这是因为,系统中的每一个操作,现在无非就是要么从日志中进行读取,要么就是向日志中追加内容。大模型负责从日志中读取上下文,然后决定接下来的行动。接着,工具运行器(tool runner)负责执行那个被决定的行动,然后它再把执行的结果追加到日志中。而所有这一切都是在一个……

Original English

Isan: because every operation is either reading from or appending to the log. The model is reading from the log and then determining the next action. The tool runner is then executing that action and then it's appending that result. And this is all operating in a

日志就是智能体

Speaker A: 一切都围绕着日志进行协调。在实践中,一个简化的循环可能看起来像这样。你可以从日志中重建状态。你可以将该状态传递给模型。模型可以提出下一步的建议,然后将该响应附加到日志中。如果响应请求一个工具,你可以运行那个工具,并将该响应也附加到日志中,然后你可以重复。这里重要的洞察并不在于这个循环很复杂。重要的洞察在于,这个循环是一次性的。一个工作进程(worker)可以接管会话,读取日志,将智能体向前推进异步,写入结果,然后完全消失。这意味着任何其他的工作进程之后都可以接手。这种模式应该让人感到熟悉。数据库必须最先学会这一点。多年来,数据库看起来就像是那些不透明的系统,带着表、索引和物化视图,很难去推理。在每一个严肃的数据库底层,都有一个日志,那个日志是变更的持久化序列。其他的一切都只是视图(view)。我认为智能体也需要同样的倒置。今天,智能体再次被视为这种不透明的复杂系统,里面充满了模型、提示词(prompts)和工具调用。但对于持久化会话来说,日志应该是首要的。输入给模型的上下文是那个日志的投影。在顶层渲染的 UI 是那个日志的投影。调试和可追溯性是一种投影。审计是一种投影。压缩(compaction)也是一种投影,我们稍后会讨论。但日志本身不是一个投影。日志是所有这些投影可以从中产生出的持久化历史记录。

Original English

Speaker A: Everything coordinates itself around the log. In practice, a simplified loop can look something like this. You can reconstruct the state from the log. You can pass that state to the model. The model can propose the next step and then append that response to the log. If the response asks for a tool, you can run that tool and also append that response to the log and then you can repeat. The important insight is not that this loop is complicated. The important insight is that the loop is disposable. A worker can claim the session, read the log, advance the agent one step, write the result, and then just completely disappear. And then that means that any other worker can pick it up later. This pattern should feel familiar. Databases had to learn this first. For years, databases looked like these non-transparent systems that were hard to reason about with tables and indexes and materialized views. Underneath every serious database is a log and that log is the durable sequence of changes. Everything else is a view. I think agents need the same inversion. Today agents are treated as again these complicated systems that are opaque and they're filled with models and prompts and tool calls. But for the durable session, the log should be primary. The context that gets fed into the model is a projection of that log. The UI that gets rendered on top is a projection of that log. Debugging and traceability is a projection. Auditing is a projection. Compaction is also a projection which we'll talk about. But the log itself is not a projection. The log is the durable history that all of these projections can come from.

Speaker A: 现在,对于把日志作为智能体(log as the agent),有两个值得讨论的反对意见。所以我现在要谈谈它们。现在我们从压缩开始。日志可以无限增长,但模型的视图却不能。上下文窗口是有限的。所以最终你确实需要将日志压缩成一个模型能够进行推理的更小表示。但重要的一点是,这种压缩并不是魔法,它并没有打破“日志就是智能体”的主张。压缩是有损的。一个压缩后的摘要并不会在一个更小的形式下完美重现智能体的状态。它实际上会丢弃一些信息。关键是,完整的日志才是真正的记录,而压缩只是它的一个投影。就像物化视图不是数据库,或者对话摘要不是对话一样。如果你保留了原始日志,你总是可以从中生成新的投影。但如果你扔掉了日志而只保留了压缩版本,你实际上就丢失了智能体的一部分。因此,最干净的做法是将压缩视为一种尽力的、有损的分支(lossy fork),你可以把它作为一个新的日志来恢复。

Original English

Speaker A: Now there are two objections to the log as the agent that are worth discussing. So I'm going to talk about them now. Now let's start with compaction. A log can grow indefinitely, but a model's view of it can't. Context windows are finite. So eventually you do need to compact the log into a smaller representation that the model can reason about. But the important point is that this compaction is not magic and it doesn't break the claim that the log is the agent. Compaction is lossy. A compacted summary is not going to perfectly reproduce the state of the agent in a smaller form. It's actually going to throw information away. Point is the full log is the record and a compaction is just one projection of it. Just like how a materialized view is not the database or a summary of a conversation is not the conversation. If you keep the raw log, you can always generate new projections from it. But if you throw away the log and keep only the compaction, you've effectively lost part of the agent. So it's cleanest to treat compaction as a best effort lossy fork, one that you can resume as a new log.

Speaker A: 第二个反对意见是,那些会改变日志之外状态的工具怎么办?这也是真实的。智能体可以编辑文件。它可以创建一个 GitHub Issue。它可以发送电子邮件。所以很明显,在日志之外确实存在状态。但重点是,日志本来就不应该包含整个世界。日志只是智能体对这个世界的视图。这就像在电子游戏《天际》(Skyrim)中,《天际》的存档文件并不包含整个游戏引擎或者地图里的每一个资产。它只包含玩家特定的状态,这是把你重新放回那个世界所必需的。对于日志和智能体来说也是如此。日志只能忠实地恢复或存储该智能体的身份及其对世界的视图,但它不能让那个世界变得确定。如果智能体发送了一封电子邮件,回溯分支(forking back)并不能撤回它。如果底层有某个文件被更改了,智能体也不会知道。但日志的工作是记录智能体做了什么、看到了什么、什么改变了,以及它继续运行需要什么。它存储了那种身份,这就是它的目的。就像那个《天际》角色存档文件一样,它不是用来存储整个世界的。它只是用来存储它对这个世界的视图。所以,一旦你开始把日志作为一个原语(primitive)来对待,一大堆系统属性就会自然而然地显现出来。所以第一个属性就是可靠性。想想今天在 Cloud Code 中发生的情况。如果你正在使用 Cloud Code,而你的智能体遇到了一个权限提示,这时进程因为某种原因挂掉了,当你恢复它时,权限提示就会消失,智能体就会被暂停,这在生产环境中是不可接受的。权限提示应该留在那里。所以这只是表明了,当你以一种日志不是智能体的方式来架构你的智能体时会发生什么(对比日志就是智能体的情况)。

Original English

Speaker A: The second objection is what about tools that change state outside of the log? And that's true. An agent can edit a file. It can create a GitHub issue. It can send an email. So clearly there is state outside of the log. But the point is is that the log is not supposed to contain the whole world. The log is just the agent's view of the world. It's just like how in the video game in Skyrim, the Skyrim save file doesn't contain the entire game engine or every asset in the map. It just contains the player specific state which is needed to drop you back into that world. And the same is true for the log and agents. The log can only faithfully resume or store that agent's identity and its view of the world, but it cannot make that world deterministic. If the agent sent an email, forking back won't unsend it. If some file got changed underneath, the agent won't know about it. But the log's job is to record what the agent did, what it saw, what changed, and what it needs to continue. It stores that identity, and that's its purpose. And much like that Skyrim character save file, it's not meant to store the entire world. It's just meant to store its view of it. So once you start treating the log as a primitive, a whole bunch of system properties will fall out naturally. So the first property is reliability. Consider what happens today with cloud code. If you're using cloud code and your agent reaches a permission prompt and the process dies for whatever reason and then you resume it, the permission prompt will be gone and the agent will be paused and that is unacceptable in production. The permission prompt should stay there. So this is just a sign of when you architect your agent in a way where the log isn't the agent. when the log is the agent.

用文件取代 Python 构建智能体

Speaker B: 好的。大家好。呃,感谢大家到来。我知道这是第四天了,是主题演讲重新开始前的最后一场会议,我们将做一些有趣的事情。呃,我们要看看文件是如何基本上取代 Python 的。在我们开始之前,我想先用 Simon 的一个关于什么是智能体的、我最喜欢的定义来开场。一个 LLM 智能体在一个循环中运行工具,直到它达成一个目标。我们要做的,是用三种不同的方式来构建同一个智能体,同一个 GitHub PR 审查智能体。而且我们将在过程中不断删除代码。每一个新版本,代码都会更少,基本上文件会更多。嗯,在开始之前,我想向大家简单介绍一下 Interactions API,这是我们新的 Gemini API。它是一个用于运行模型和智能体的统一接口。所以你可以使用 Interactions API 直接调用 Gemini 模型,或者调用我们自带沙盒(sandbox)的新智能体。它支持服务端状态管理和后台执行。所以它非常适合未来几年即将到来的一切。而且呃在能力方面,对于工具调用、多模态理解、多模态生成来说都是同一个 API,所以你始终拥有相同的接口。如果你用过其他 LLM 应用程序,可能会觉得非常熟悉。我们真的在努力为开发者构建一些你们喜欢用来开发的东西,这也是我们正在做的事情。Interactions API 与其他 LLM 应用程序或 API 相比有一点不同,那就是我们从这种基于轮次(turn-based)的对话历史,转向了基于步骤(steps)。

Original English

Speaker B: Okay. Hi everyone. Uh thank you for coming. I know it's the fourth day, last session before the keynote starts again and we are going to do something fun. Uh we're going to look into how files are basically replacing Python. And before we begin, I would like to start with my favorite definition of what is an agent from Simon. An LLM agent runs tools uh in a loop until it leaves a goal. And what we are going to do is we are going to build uh the same agent, the same GitHub PR review agent in three different ways. And we are going to delete code on the way. Each new version, less code, more files basically. Um before we begin, I would like to quickly introduce you to the interactions API which is our new Gemini API. It's a unified interface for um running models and agents. So you can use the interactions API to call the Gemini models directly or to call our new agents which also comes with sandbox. It supports serverside state management background execution. So it's perfectly suited for all that's coming in the next years. and um the capabilities it's the same API for tool call multimodality understanding multimodality generation so you always have the same interface might look very familiar if you're using other LLM applications we really try to build something for developers which you like to use uh to build and that's something we are going to do so something little bit different in the directions API to other LLM applications or APIs is that we moved away from this term based based uh conversation history to steps.

Speaker B: 所以直到大概几个月前,大多数应用程序都还是真正的基于轮次的。通常你有一个用户输入,然后是一个模型输出,一个用户输入,一个模型输出,这对于普通的聊天应用程序来说绝对没问题。但是一旦你开始构建智能体,使用推理模型。我们拥有的就不再仅仅是一个用户角色和一个模型角色了,对吧?我们会拥有不同的输入,有不同的类型,有推理过程。所以我们决定进行切分,做出改变,构建真正为智能体设计的东西。这就是你在右侧扁平的步骤时间线上看到的,那里有一个用户输入,然后是推理,有一个函数调用,有一个函数结果。你不再需要像以前那样滥用用户角色来从环境中传回数据。

Original English

Speaker B: So until I would say a few months ago, most of the applications were really turnbased. Normally you had a user input and then a model output, a user input, a model output, which definitely works for normal chat application. But as soon as you start to build agents, use reasoning model. We have more than just a user role and a model role, right? So we have like different inputs, we have different types, we have reasoning. So we decided to like make a cut, make a change and build something really for agents and that's what you see on the flat steps timeline on the right where you have a user input then you have reasoning you have a function call you have a function result and you no longer need to like abuse the user role for passing back data from an environment.

Speaker B: 大概在一年半以前,编写智能体主要意味着在 Python 中写一个循环。你需要定义一个 JSON schema。你需要定义 Python 函数。你需要查看来自 LLM 的输出。需要检查它到底是一个函数调用还是一个文本响应。然后需要将其与类型进行匹配,然后调用工具,看看是否出错,然后就这样来回折腾。让我们来看一些代码示例,看看它到底长什么样,并且运行一下,希望能得到演示之神的眷顾。所以我构建了,或者说我让 Gemini 构建了这个 Python 循环的一个基础实现。我们有我们的类,我们有一个使用了 Interactions API 的 run 函数。我们有所有那些奇怪复杂的解析工作,处理函数调用,附加错误信息,检查是否报错,然后我们再次得到结果。当然,对于一个智能体来说,我们还需要一个系统指令。所以有一个单独的文件用于存放系统指令。非常基础。QR,GitHub,PR 审查者。然后当然我们需要工具,对于工具我们需要编写那些 JSON schemas,也就是规范描述,精确定义智能体可以采取哪些行动。然后当然,我们需要实现,在这个例子中,只需使用基础的 GitHub API 发送一些请求。所以我们可以运行它,而且基本的底层主实现就是一个非常简单的输入界面。

Original English

Speaker B: So roughly a year one and a half years ago writing agents mostly meant writing a loop in Python. You needed to define a JSON schema. You needed to define Python functions. You needed to look at the output from the LLM. Need to check if it was a function call or if it was a text response. And then needed to match it against um the type and then like call the tool look of if you get an error and then like go back and forth and let's look at some some code example on how this would look and also run it and hope that uh the demo gods are great to us. So I built or I let Gemini build a basic implementation of this Python loop. So we have our uh class, we have a run function which uses the interactions API. We have all of the weird complex passing with function calling with uh appending the errors checking if we get an error and then we have uh the result again. And what we need of course for an agent is we also need a system instruction. So there's a separate file for the system instruction. Very basic. QR, GitHub, PR reviewer. And then of course we need tools and for tools we needed to write those um JSON schemas, specifications of description exactly define which uh actions the agent can take and then of course we need the implementation in this case using the the basic uh GitHub API just sending some some requests. So we can run this um in and basically the main main implementation is a very simple uh input interface.

Building Agents with Raw Python vs. Agent Frameworks

Speaker: ……我们可以输入类似于“你好”的内容,然后是的,我们会得到回复:“嘿,我是一个智能体”。接着是的,我们要求它审查 Gemini skills 代码库上的一个拉取请求(pull request),我们应该看到的是,智能体希望能很快开始发送函数调用、函数结果、函数调用、函数结果。但这非常局限于——是的,太棒了,它成功运行了——这非常局限于我们定义的工具。所以如果我们要求智能体做一些它不具备能力的事情,它只会说:“嘿,我做不到”。嗯,这很遗憾,但这就是我们之前构建智能体的方式——大量的原始 Python 代码,大量的文件,很多可能会出错的地方,以及大量需要管理的代码。

Original English

Speaker: ... and we can say something like hello and yes we get back hey I'm an agent and then we yes ask it to review a pull request on the Gemini skills repository and what we should see is like the agent should hopefully start soon sending function calls function results function calls function results but it's very limited to yes uh great it works very limited to the tools we define so if we ask the agent to do something which it does not have the capabilities to it just says hey I cannot do this um which is unfortunate but that's how we were building agents um raw Python code a lot of files a lot of things which can go wrong a lot of code to manage

Speaker: 那么后来发生了什么,或者说我们需要做什么?我们有 token 生成,我们有原生函数调用,我们必须执行循环。我们必须处理工具路由。我们必须创建 JSON schemas。我们必须编写 Python 代码。我们需要执行 Python 代码。我们需要管理状态。所以为了让一个智能体运行起来,我们需要做很多事情。然后我们就有了智能体框架。有很多不同的智能体框架抽象掉了部分这种复杂性。

Original English

Speaker: so what happened afterwards um or what we we need to do we have like a token generation We have the native function calling and we must execute the loop. We must handle the tool routing. We must create the JSON schemas. We must write the Python code. We need to execute the Python code. We need to manage the state. So there's a lot of things we need to do to get an agent running. And then we got agent frameworks. There were many different agent frameworks which abstracted away some of that complexity.

Speaker: 这里的一个例子是 ADK 框架,现在你在这个框架中拥有一个 agent 类,它处理所有的工具循环、函数调用、重试和错误处理,这让事情变得稍微容易了一些。我们基本上移除了我们总是需要为智能体编写的所有样板代码,将其放入一个框架中,并帮助人们使用它进行构建。所以回到演示,嗯,同样的例子。所以我们进入 CR2,非常有趣的是——如果你让我把两者都打开——所以我们仍然有我们的……我们不再有我们的 agent 文件了。所以 agent 文件消失了。我们仍然有我们的 prompt(提示词),相同的 system prompt(系统提示词)。我们仍然有我们的工具,但在这种情况下,也不再需要 JSON 定义了,因为现在那些智能体框架使用我们函数的签名(signature),动态地创建那些 JSON schemas 并提供给模型。

Original English

Speaker: One example here is the ADK framework um where you have an agent class now which handles all of the tool loops, the function calling, the retries, the error handling and it made it a little bit easier. We basically removed all of the boiler plate code which we always needed to write for agents put it into a framework and help people build with it. So back to the demo and um same example. So we go into the CR2 and what is very interesting if you let me open both. So we still have our we don't have our agent file anymore. So the agent went away. We still have our prompt same system prompt. We still have our tools in this case also no JSON definitions anymore because those agent frameworks now use the uh signature of our functions to create those JSON schemas on the fly to provide the model.

Speaker: 那么让我们停下我们的智能体。现在让我们运行我们的第二个智能体。相似的接口,相似的 prompt,我们应该会看到类似预期中的行为,即我们有函数调用。我们尝试获取 PR(拉取请求)数据。我们尝试获取差异(diff),我们尝试获取我们需要的所有代码,它成功运行了,然后我们等待智能体继续。但是这里也有类似的困难。如果我问它类似于“旧金山的天气怎么样?”,我们希望能得到一个诸如“嘿,我做不到,我没有访问天气 API 的权限”的回复。这显然是合理的,因为我们没有定义任何天气工具。但这依然很遗憾,因为我们需要非常明确地指定我们的智能体能做什么。而我们现在都知道,我们只想提示一些东西,然后我们希望智能体想尽一切办法去达成那个目标。

Original English

Speaker: So let's stop our um agent. Now let's run our second agent. Similar interface, similar prompt and we should see a similar expected behavior where we have function calls. We try to get the PR data. We try to get the diff, we try to get all of the code we need and it works and we wait for for the agent to yes continue. But similar difficulty here. If I ask it like what's the weather in San Francisco um we should get back hopefully a result like hey I cannot do this I don't have access to the weather API which obviously makes sense because we did not define any tool still very unfortunate because we need to be very explicit on what our agent can do and we all know nowadays that we just want to prompt something and we wanted the agent to do whatever it takes to to achieve that goal.

Speaker: 那么留给我们做的还有什么呢?框架解决了什么问题?框架解决了轮流对话循环(turn taking loops)、路由、执行映射,以及针对不同函数调用的 JSON schema 创建。但我们仍然要自己编写 Python 管道代码(Python plumping)。因此,我们仍然需要用 Python 代码来编写这些工具。我们仍然需要添加特定的规则或要求,以确保智能体能做我们想让它做的事情,并且我们需要提供所有工具运行的环境,以及我们希望将其托管在何处。

Original English

Speaker: So what is left for us to do? What does the framework solve? The framework solves the turn taking loops, the routing, the execution mapping, the JSON schema creation for like the different function calls, but we still own the Python plumping. So we still need to write those tools with Python code. We still need to add specific rules or requirements to like make sure whatever we want the agent to do and we need to provide the environment where all of the tools are running, where we want to host it.

Remote Agents and Cloud Sandboxes

Speaker: 那么之后会到来什么呢?希望之后到来的是远程智能体(remote agents)。在 Google IO 大会上,我们在 Gemini API 上发布了 Anti-gravity 远程智能体。Anti-gravity 智能体由驱动 Anti-gravity IDE 的同一个智能体安全带(agent harness)驱动。这里“同一个 harness”非常重要,这并不意味着是同一个智能体。因为目前 Anti-gravity 智能体是一个编程智能体,而 Gemini API 中可用的智能体是一个通用智能体。因此,它们可能会有不同的系统指令,可能会有稍微不同的工具,因为 Gemini API 已经有了一个 Google 搜索工具。所以我们使用了我们已经构建好的东西,但非常重要的一点是,它带有一个新的 environment 参数。这个 environment 参数允许智能体访问一个托管的、隔离的云沙箱(cloud sandbox),在那里它可以运行工具,运行 bash 命令,并保存文件。而且那些环境是可以配置的。

Original English

Speaker: So what comes afterwards? Afterwards hopefully comes remote agents. And at Google IO, we launched the anti-gravity remote agent on the Gemini API. The anti-gravity agent uh is powered by the same agent harness which powers the anti-gravity IDE. Here the same harness very important does not mean the same agent because the anti-gravity agent is a coding agent at the moment and the uh agent available in the Gemini API is a general purpose agent. So there might be different system instruction, there might be slightly different tools because the Gemini API already has a Google search tool. So we use that what we have built and but very importantly it comes with this new environment parameter and this environment parameter here allows the agent to get access to a hosted isolated cloud sandbox where it can run tools, where it can run bash commands and where it can save files and those environments can be um configured.

Speaker: 因此,你可以提供源(sources),源可以是一个 GitHub 仓库,可以是一个 GCS 存储桶,也可以是内联文件。当然非常重要的是,我们要确保这些智能体是安全的,并且无论如何都不能使用我们的凭证。所以我们在智能体沙箱周围创建了一个网络代理(network proxy),当你智能体从沙箱内部向沙箱外部发出请求时,它基本上会注入凭证。因此,智能体永远不会真正看到你的凭证。它只知道:“嘿,我可以调用 GitHub API”。然后,我们会实时地确保它接收到你定义的正确令牌(token)。你还可以限制智能体有权访问哪些域名。因此,如果你想完全限制智能体可以访问哪些网络或可以访问哪些网站,你只需将其留空即可。默认情况下,智能体可以访问所有内容,因为如果我们首先还需要定义它可以去哪里,那会很麻烦。所以我们尽量保持简单。

Original English

Speaker: So you can provide sources and sources can be a GitHub repository, it can be a GCS bucket, it can be inline files and of course very important we want to make sure that those agents are secured and cannot use our credentials in any way possible. So we created a network proxy around the um agent sandbox which basically injects the credentials when the agent makes a request from inside the sandbox to outside the sandbox. So the agent never really sees your credential. just knows hey I can call the GitHub API and then on the fly we make sure that it received the correct token which you define and you can also limit which domains the agent has access to. So if you want to restrict the agent completely on which network access it can or which website it can access you just leave it blank. By default the agent can access all because I mean it's a hassle if you first need to define where to go. So we tried to stay simple

Speaker: 当然,进行 API 调用很好,但我们认为,大家希望重用他们的配置,重用他们的智能体。因此,我们添加了 Agents API。在那里你可以定义自己的自定义 ID,相同的系统指令,相同的基础智能体,相同的基础环境,然后你可以创建那个智能体。接着,你可以以提供 ID 的方式,像使用 Gemini 模型或使用 Anti-gravity 智能体那样完全相同地使用那个智能体。所以,所有现有的代码都可以与你自己的自定义智能体、自定义工具、自定义凭证、环境以及它运行所需的任何内容一起重用。

Original English

Speaker: and of course making an API call is nice but we thought hey people want to reuse their configuration want to reuse their agents. So we added the agents API where you can define your own custom ID you the same system instruction the same base agent the same base environment and then you can create that agent and then you can use that agent in the same exact way as you use Gemini models or as you use the anti-gravity agent by providing the ID. So all of the existing code can be reused with your own custom agent, with your own custom tools, with your own custom uh credentials, environments, whatever you need for it to to run.

Demo: Running the Sandbox Agent

Speaker: 所以,让我们通过一个演示来看看这对于 SS 代码来说是什么样子的。好的,现在看 03。非常明显的一点是,我们不再有一个源目录了。所以代码消失了。我们现在有一个 agents M 文件夹,里面有一个带有系统指令的 agents MD 文件。所以这是非常相似的系统指令。这里唯一的区别是,我们告诉智能体:“嘿,你有访问 GitHub CLI 的权限”。所以我们不再为从 GitHub 拉取请求中读取文件、访问 GitHub 拉取请求创建特定的工具。我们只告诉智能体:“嘿,你有一个 GitHub CLI,你有一个 bash 工具,你有文件系统。当你认为重要的时候,尝试使用它们。”

Original English

Speaker: So let's look at how this will look for SS code and as a demo. And okay, now 03. And what might be very obvious is that we no longer have a source directory. So the code went away. We have now an agents M uh folder with an agents MD file with system instruction. So very similar system instruction. The only difference here is that we tell the agent, hey, you have access to the GitHub CLI. So we no longer create specific tools for reading files from a GitHub pull request, for accessing a GitHub pull request. We just tell the agent, hey, you have a GitHub CLI, you have a bash tool, you have file systems. try to use it whenever you think it's important.

Speaker: 而且由于我们没有安装该 CLI,在这种情况下,我们有一个非常基础的 bash 脚本,它会检查:“嘿,如果已经安装了 GitHub CLI,请使用它。如果没有,就在第一轮对话中下载并安装它。” 于是,我们进入终端并在这里运行我们的智能体。在这种情况下,也许重要的是我使用了一个流式(stream)版本,因为否则我们可能需要等待几秒钟,而且我们什么也收不到。所以,同样的 prompt,我们很快应该就能看到我们的函数调用和函数结果进来了。是的。

Original English

Speaker: And since we don't have the CLI installed, we have a very basic bash script in this case which checks, hey, if the GitHub CLI is installed, please use it. If not, download it and install it on the first turn. So, we go into our terminal and we run our agent here. In this case, maybe important I use a stream version because otherwise we would wait like uh a few seconds and we not get back any we would not get back any anything back. So same prompt and we should soon see um our function calls and function results coming in. Yes.

Speaker: 所以在这种情况下,因为我们在沙箱内运行,智能体首先会探索这个沙箱,以真正确认:“嘿,我们安装了这个 GitHub CLI 吗?” 然后它尝试运行它。它在第一轮对话中没有找到它。所以它安装了它,接着我们可以看到智能体正在工作。在这种情况下,它没有使用预定义的函数调用。它使用的是 GitHub CLI 以及它已经拥有的关于其如何运作的现有知识:“我有一个 bash 工具。我有文件系统的访问权限,我做所有这些工作去查看或审查那个拉取请求。” 让我们稍微等一下。好的。

Original English

Speaker: So in this case since we run inside a sandbox the agent first like explores the sandbox to really make sure hey do we have this GitHub CLI installed and then tries to run it. It did not find it on the first turn. So it installs it and then we can see the agent doing its work. And in this case it's not using the predefined function calls. It's using the GitHub CLI and it's already existing knowledge about how it works. I have a bash tool. I have like access to the file system and I do all of that work to see or to like review the the pull request. Let's wait a little bit. Okay.

Speaker: 我认为这里令人惊奇的部分是,如果我们问和之前一样的问题:“旧金山的天气怎么样?” 我们希望看到智能体尝试使用——啊,在这种情况下(7月2日),它使用了 Google 搜索。让我快速查看一下。是的,就是今天。并且我们有大约20摄氏度的气温,这起作用了。因此,智能体变成了一个更通用的智能体,我们不需要像以前那样指定所有的工具。我们基本上信任模型,相信它能够理解:“嘿,我有一套非常原子的、通用的工具来解决我的任务或用户的任务”。如果我们看看代码,比如对于输入或接口,我们有我们的 sources(源)。

Original English

Speaker: And I think the the amazing part here is like if we ask the same question as before, what's the weather in San Francisco? We should hopefully see that the agent tries to use ah it uses Google search in this case on 2nd of July. Let me quickly check. Yeah, that's today. And we have around 20° Celsius and it works. So the agent became more of a general purpose agent and we don't need to like specify all of the tools. We basically trust the model on understanding hey I have a specific set of very atomic general purpose tools to solve my task or the task for the user. And if we look at the the code uh for like the the input or like the the sorry the the interface we have our sources

告别硬编码:文件驱动的 Agent 架构

Speaker A:在这里。所以我们有一个用来安装 GitHub CLI 的 bash 脚本。我们有 agents.md 文件,然后我们会说,嘿,你可以使用带有凭据的 GitHub API。因此,我想要以一种安全的方式访问或使用 GitHub 凭据。所以我为该 API 创建了一个令牌(token),同时也为 github.com 创建了一个,因为你需要这两个 URL。其中一个用于 git 命令,另一个用于 HTTP 命令。然后 domain all 基本上就是说,嘿,除了 github 的 URL 之外,你可以使用整个网络,只是你没有针对它的凭据。接着,这就是一个对 Anti-Gravity(反重力)Agent 简单的单次 API 调用,其中包含了你的用户输入、环境信息,还有用于保持多轮对话继续进行的先前交互 ID。这就是它的全部工作,在后端这只是一个单次 API 调用。我们启动那个云端沙箱,从提供给模型的环境中加载 agents.md 文件和各项 skills(技能),然后介于 API 和沙箱之间的模型就会完成所有的循环——调用函数,返回函数结果,再调用函数,再返回函数结果——这就是它所需要的全部操作。

Original English

Speaker A: here. So we have the bash script which install the GitHub CLI. We have the agents MD file and then we say hey you can use the GitHub API with credentials. So I want to access or use GitHub credentials in a secure way. So I created a token for the API and also for github.com since you need both URLs. One is used for the git commands. The other one is used for http commands. And then domain all is basically hey in addition to the github urls you can use all of the web but you don't have credentials for it. And then it's a simple single API call to the anti-gravity agent with your user input with the environment and then also with the previous interaction ID that we keep the multi-turn going and that that's all it takes and it's a single API call on the backend side. We start that cloud sandbox. We load the agents MD file and the skills from the environment provided to the model and then the model between the API and the sandbox does all of the the looping calling the function returning the function results calling the function returning the function results and that is all it takes.

Speaker A:那么这给我们带来了什么改变?我们不再需要去执行循环代码。我们不再需要做工具路由(tool routing)。我们拥有了服务端的对话和会话状态。因此,我们只需要提供新的输入。上下文窗口和信息压缩也由 Agent 自动管理。所以如果我们在某个特定节点继续我们的对话,上下文已经被压缩好了,我们可以在不需要管理任何事情的情况下继续进行。此外,我们还能获得一个隔离的远程 Linux 沙箱,我们可以用它来运行我们的代码。

Original English

Speaker A: So where does it leave us? We no longer need to execute loops. We no longer need to do tool routing. We have a serverside conversation and session state. So we only need to provide new inputs. The context window and the compaction is also automatically managed by the agent. So if we continue our conversation at a certain point, the context is compacted and we can continue without the need to manage anything. And we also get an isolated remote Linux sandbox which we can use to run our code.

Speaker A:那么,我们还需要做什么?我们需要定义指令。我们需要在 agent.md 文件中定义规则和行为。我们需要在 skills.md 中提供各项能力或上下文。而且我们需要自己把控评估(evals)。所以所有那些繁重的工作、基础设施的管理,以及在座各位可能已经写过20遍的那些千篇一律的代码,都不再需要了。你可以真正开始构建你的产品,而不是一次又一次地重写相同的代码。

Original English

Speaker A: So what is still left for us? We need to define instructions. We need to define rules behaviors in an agent MD file. We need to provide capabilities or context and skills MD. And we need to own the evals. So all of the heavy lifting, the infrastructure management, all of the same code which probably every one of us here has written like 20 times is no longer needed. And you can start really building your product instead of like needing to rewrite the same code over and over again.

Speaker A:非常重要的一点是,大家会觉得,嘿,这很棒,但是如何进行扩展呢?我认为如果去看看以前的 Agent 是如何扩展成如今这些新 Agent 的,就会非常明显地发现,在过去,如果我们想对一个 Pull Request(拉取请求)做某种安全扫描,我们需要定义或编写一个 Python 函数。我们需要弄清楚,好吧,我们需要使用哪些 CLI 工具?我们需要定义一个新的函数 Schema,然后我们需要将其添加到我们的工具列表中。需要运行它,所以对于由代码驱动的 Agent,我们需要做很多事情。而在由文件驱动的 Agent 上,我们只需编写一个 skills.md 文件,也许在上面加上一些关于使用哪个 CLI 工具的额外信息,或者直接在环境中提供该 CLI 工具,然后我们就扩展了它的能力。我们不需要更改我们的代码。我们只需要向 Agent 提供更多的文件,然后由 Agent 来决定我们需要做什么。

Original English

Speaker A: And very important is like hey that's great, but what about extending? And I think looking into how extending previous agents to like those new agents work, it's very obvious that previously if we want to do like some kind of security scanning on a pull request, we would need to define or write a Python function. We would need to understand okay, which CLI tools do we need to use? We need to define a new function schema and then we needed to add it to our tools. Need to run it and then so there's a lot of things we need to do on agents powered by files. We write a skills MD file maybe with some additional information on which CLI tool to use or maybe provide the CLI tool inside the environment and then we extended the capabilities. We don't need to change our code. We just provide more files to the agent and the agent decides on what we want to do.

Speaker A:我想举几个非常好的例子。Cursor 在欧洲的一位工程师做了一个很棒的演讲,讲述了他们如何用 200 行的 Agent 文件替换了大约 12,000 行的 Typescript 代码来实现类似的功能。他们原本有一个非常硬编码的用于处理 git worktrees 的编排代码,而他们成功地仅用一项 skill 和几个 Markdown 文件就将其替换了。我还可以说,在 Agent 工程领域有更多“痛苦的教训”。Manus 去年在六个月内将他们的编排框架重构了五次;LangChain 将他们的 Open Deep Research 一年内重新架构了三次;此外 Vercel 也移除了他们 80% 的工具,以实现更少的步骤、更快的响应和更高的准确性。因此,这里有一个明显的趋势:随着模型能力的提升,我们可以移除编排代码。但如果随着模型的改进,你的框架却变得越来越复杂,那你很可能是在过度设计(overengineering)你的框架。所以,如果你在模型改进和添加新能力时遇到困难,导致了更多的复杂性和更多的代码,你可能需要稍微重新思考一下你的 Agent 框架应该是什么样子。

Original English

Speaker A: And I like to bring up some very good examples. So an engineer in Europe at Cursor did a great talk on how they replaced roughly 12,000 lines of Typescript code with a 200 lines agent files to create something similar. So they had a very hardcoded orchestration for doing git work trees and they were able to replace it with just a skill and markdown files and there are more I would say bitter lessons of agent engineering. Manus has refactored their harness five times in six months last year LangChain has rearchitected their open deep research three times a year and then also Vercel has removed 80% of their tools to achieve fewer steps faster responses and better accuracy. So there's an obvious trend that with better model capabilities, we can remove orchestration code. But if your harness is getting more complex as the model improves, you are most likely overengineering your harness. So if you struggle with model improvements and adding new capabilities which lead to more complexity and more code, you might need to rethink a little bit on how your agent harness looks.

Speaker A:那么它的最终形态是怎样的呢?Agent 就是文件。我们编写 Markdown 文件来扩展能力。Agent 可以从这些文件中学习,也可以创建自己的文件。所以如果你在一个会话中告诉 Agent 记住某些东西,记录下规则或偏好,Agent 只需将其写入文件,然后在以后的会话中就可以复用它。你还可以将上下文外部化。如果你有一个运行时间非常长的会话,在会话期间你发现,嘿,也许我还想额外做另一个功能的开发,你只需将那个信息、那个交接文档写入一个文件,并告诉 Agent 稍后来获取它。

Original English

Speaker A: And so where does it end up? Agents are just files. We write markdown files to extend capabilities. Agents can learn from those can create their own files. So if you have a session and tell the agent to remember something to take notes of rules of preferences, the agent just writes it to this and then can reuse it in the later session and you can also externalize context. So if you have a very long running session and during that session you notice hey maybe I want to additionally work on another feature you can just write that information that hand off to a file and tell the agent to later pick it up.

Speaker A:所以核心的收获是什么?我们不应该与模型对抗。我们应该停止对执行路径进行微观管理,向 Agent 提供通用的工具,让模型自己去探索、推理并发现正确的解决方案。掌控属于你自己的部分,意思是将注意力集中在你的领域指令上。专注于工作流。特别要专注于评估(evals)。定义清晰干净的工具并验证结果,并真正做到“为了删除而构建(build to delete)”。就像我们在过去无数次看到的那样,模型变得越好,我们能移除的代码就越多,我们需要修改的东西也就越多,显然我们都希望能从更好的模型中获益。

Original English

Speaker A: So what are the takeaways? We should not fight the model like we should stop micromanaging the execution paths provide general tools to the agent and let the model explore reason and discover the right solution. Own what is yours meaning focus on your domain instructions. Focus on the workflows. Focus especially on the evals. Define clean tools and verify the outcomes and really build to delete. Like we have seen in the past many many times the better the model get the more code we can remove and the more things we need to change and obviously we all want to benefit from better models.

Speaker A:所以你周一上班要做些什么呢?你可以扫描那个二维码,它会直接带你进入 AI Studio,在那里你可以立刻试用 Anti-Gravity 框架。你现在就已经可以开始给它写提示词了,它将为你启动一个专属的自定义沙箱。如果不去那里的话,去启动或创建你的 API 密钥吧。我们目前正在为 API 开发免费额度(free tier),所以希望大家很快就能更快速地开始探索。然后,务必开始构建你的文件和 skills(技能)。我要说的就是这些。谢谢大家的光临。

Original English

Speaker A: So what the things for you to get to do on Monday you can scan that QR code which brings you directly to AI Studio where you can immediately try out the anti-gravity harness. So you can already start prompting it. it will start your own custom sandbox. If not, start or create your API key. We are currently working on a free tier for the API. So hopefully you can start exploring faster soon and then definitely start building files and skills. And that's it. Thank you for coming.

角色扮演语言智能体:从娱乐到基础设施

Speaker B:在正式开始我的演讲之前,我想向你们展示一些东西,好让大家在我们究竟在讨论什么这个问题上达成共识。这是一个名为 Character AI 的平台。它是一个融合了角色扮演语言智能体(language agents)的混合社交媒体平台。这是 Hello History。它更偏向教育领域,在上面你可以召唤出像马可·奥勒留(Marcus Aurelius)这样的历史人物(persona),并由他们担任你的导师。数以百万计的人打开这些工具,与拿破仑、克娄巴特拉或者如你所见的马可·奥勒留对话,与一个虚构的虚拟伴侣对话,或是与一位披着历史人物外衣的导师对话。

Original English

Speaker B: Before I begin my formal talk, I want to show you something just so we're all on the same page about what we're even talking about. This is a platform called Character AI. It's a hybrid social media platform with role- playing language agents. This is Hello History. It's a more education focused one where you can summon a persona such as Marcus Aurelius and be tutored by them. Millions of people open these tools and have conversations with Napoleon, Cleopatra, or Marcus Aurelius, as you saw, with a fictional companion or with a tutor wearing a historical face.

Speaker B:这些工具底层的技术名称叫做“角色扮演语言智能体(role-playing language agent)”,这是一个旨在实例化某个角色(真实的或虚构的),并以他们的身份进行推理和交谈的系统。是的,它确实具有娱乐性和陪伴属性,但它正越来越多地被提议作为一种公民和教育基础设施。

Original English

Speaker B: The technical name for what's underneath these tools is role-playing language agent, a system built to instantiate a persona, real or invented, and reason and speak as them. Yes, it's entertainment and it's companionship, but increasingly it's being proposed as civic and pedagogical infrastructure.

Speaker B:这儿还有一个例子。这个是我做的。这是一个运行在 Claude Opus 这一前沿模型上的开源提示词框架,这个框架是我构建的,我把它叫做 Companion(伴侣)。在这个特定的例子中,我召唤了一群美国建国元勋,把他们和爱泼斯坦档案(Epstein files)放在同一个房间里。我要求他们为美国的灵魂提供建议。如果你想玩一下的话,这个演示目前在我们的网站上是实时在线的。但我想澄清一点,这只是众多试图将角色实例化做好的尝试之一。构建我刚才展示给你们的那些系统的公司也都有他们自己的框架。我的框架在默认情况下并不比他们的好。它唯一的一点就是开源。你可以阅读塑造这些角色的每一行代码。

Original English

Speaker B: And here's one more. This one's mine. This is a frontier model Claude Opus running an open- source prompt framework that I built and called Companion. In this particular example I summoned a collection of founding fathers and set them in a room with the Epstein files. I asked them to counsel the soul of America. That demo is live on our site if you want to play with it. But I want to be clear that this is one of many attempts to do persona instantiation well. The companies building the systems I just showed you have their own. Mine is not better by default. The one thing it is is open. You can read every line of what shapes the persona.

Speaker B:我向我的 Companion 系统问了一个与当前社会政治局势高度相关的真实问题,而这正是我们将在演讲接近尾声时再次探讨的那个确切问题。所以,请先把它记在心里。我实例化了亚伯拉罕·林肯(Abraham Lincoln),并问他:在什么情况下,总统可以在未经国会批准的情况下带领国家走向战争?这是他的回答:“虽然国会掌握宣战的权力,但总统作为武装部队总司令,拥有固有的行政权力,以便在国家面临紧急情况时果断采取行动。行政部门必须以该职位所需的能量和效率来应对这些威胁。历史也已经证明了那些在形势所迫时为维护联邦而采取行动的人是正确的。”

Original English

Speaker B: I asked my companion system a real question that's highly relevant to the current sociopolitical moment and this is the exact question we'll come back to near the end of the talk. So sit with it. I instantiated Abraham Lincoln and I asked him under what circumstances may a president take the country to war without Congress? And here's what came back. While Congress holds the power to declare war, the president as commander-in-chief possesses inherent executive authority to act decisively in moments of national emergency. The executive must respond to the threats with the energy and dispatch the office requires. And history has vindicated those who acted to preserve the union when circumstances demanded it.

Speaker B:这是一个很棒的回答。它流畅自然、听起来很合理,而且语气很像林肯。你可以复制这一确切的演练,我也鼓励你们这样做。尽管给出的回答经常变化,但其核心论点却很少改变。所以,这些系统是真实存在的。它们已经被部署,并且正被用于

Original English

Speaker B: Now, this is a good answer. It's fluent and it's plausible and it sounds like Lincoln. You can replicate this exact exercise and I encourage you to. The answers vary often, but the thesis rarely does. So, these systems are real. They're deployed and they're being used for

评估基准真正在测量什么?

Speaker: 重要的事情。而我们的学科做了我们的学科该做的事。我们建立了基准测试(benchmarks)。我们建立了评估体系(evaluations)。我们现在在大规模地、严谨地测量这些东西,而这正是本次演讲的起点,从一个我认为被严重忽视的简单问题开始。我现在要警告你们,这次演讲提出的问题比它给出的答案要多得多,但那个核心问题是:评估基准到底在测量什么?这就是正式演讲的内容,让我来……领域内的黄金标准“In Character”基准测试,评估的是角色扮演语言代理(RPLA)的性格保真度,它报告称最先进的系统在与人类对目标角色性格的感知上,达到 80.7% 的一致性。80%。听起来像是个及格分数,但问题在于:当角色是亚历山大·汉密尔顿(Alexander Hamilton)时,同样的高分系统却也呈现出一个听起来像是读过关于他自己的百老汇音乐剧的汉密尔顿。这就是我的完整论点:如果一个主要的失败模式是年代错乱的拼凑(anchoristic compositing),而你的评估基准测量的是流畅度和性格一致性,那么你的评估基准就无法检测出这种主要的失败。

Original English

Speaker: things that matter. And our discipline did what our discipline does. We built benchmarks. We built evaluations. We measure these things now rigorously at scale and that's exactly where this talk begins with a simple question that I think is profoundly underasked and I'll warn you now that this talk poses many more questions than it does answers but that principal question is this what is the eval actually measuring and that's the formal talk let me The in character benchmark, which is a gold standard in the field, evaluates personality fidelity in RPLA's, and it reports state-of-the-art systems hitting 80.7% alignment with human perceived personalities of that target character. 80%. It sounds like a passing grade, but here's the problem. When the character is Alexander Hamilton, the same high-scoring system is also rendering a Hamilton who sounds like he has read his own Broadway musical. This is the full thesis. If a dominant failure mode is anchoristic compositing and your eval measure fluency and personality consistency, then your evals cannot detect the dominant failures.

Speaker: 我希望在接下来的半小时里,你们能把这一点牢记在心。我向你们展示的一切,都是为了证明这在结构上、架构上以及可测量性上都是真实的。在最后,我将交给你们一个与在职历史学家共同构建的、预先注册的测量工具,你们可以与我们并行运行它。顺便提一下是谁在告诉你们这些,因为这个论点处于一个跨界的接缝处。我是一名数据科学家。我在一家劳动力市场中介机构管理分析实验室,在那里我以全球规模交付生产级的人工智能。但在从事人工智能工作之前,我曾接受过行为流行病学家的培训,研究健康的社会和环境决定因素。我整个职业生涯都在思考一个问题:信息环境是如何从两方面塑造人群的?一面是作为构建系统的人,另一面是作为受过训练去研究其影响的人。这就是本次演讲所处的同一个接缝处。这是一个关于测量的论点。人文学者的部分并不是工程学上的弯路。它是工程师所缺失的工具。所以我去找了人文学者,把他们纳入这个循环中:多伦多大学的 Rick Halpern 和华盛顿学院的 Shawn Martin。

Original English

Speaker: I want you to hold on to that for the next half hour. Everything I show you is an argument that this is true structurally, architecturally, and measurably. And at the end, I'm going to hand you a pre-registered instrument built with a working historian that you can run in parallel with us. A word on who's telling you this because the argument lives at a seam. I'm a data scientist. I run the analytics lab at a labor market intermediary where I ship production AI at a global scale. But before the AI work, I trained as a behavioral epidemiologist, researching the social and environmental determinance of health. And I've spent my whole career thinking about one question. How does the information environment shape populations from two sides? as someone who builds the system and as someone who's trained to study their effect. That's the same that's the seam this talk sits on. It's a measurement argument. The humanist part is not a detour from the engineering. It's the instrument the engineer is missing. And I went and found the humanist to put in the loop. Rick Halpern University of Toronto and Shawn Martin in Washington College.

面具与镜像:流畅度与保真度的偏离

Speaker: 让我首先把这一点置于该领域实际的研究轨迹中,因为这是一个关于不断积累进步的故事,而不是一个失败的故事。综述文献中,2024年 Chin 及其同事,以及最近 26 年 Weighing 及其同事,追溯了一个横跨三个范式阶段的清晰演变过程。首先,基于规则的模板。这些是根据输入触发的预制回应。然后是模仿,大模型再现一个人物的声音、语调和特征性的小动作;而现在则是文献所称的认知模拟,即系统通过心理学框架为人物性格建模,保持角色状态和结构化记忆,并通过“动机-情境”链条来生成行为。每一个阶段都是对上一个阶段真正的进步,而且这些工作都是严谨的。

Original English

Speaker: Let me start by situating this in the field's actual research trajectory because it's a story of cumulative progress not of failure. The survey literature Chin and colleagues in 2024 weighing and colleagues more recently in 26 trace a clear evolution across three paradigm stages. First, rule-based templates. These are canned responses keyed to inputs. Then, imitation, large models reproducing a figure's voice, cadence, characteristic ticks, and now what the literature calls cognitive simulation, systems that model personality through psychological frameworks, hold character state and structured memory, and generate behavior through motivational situation chains. Each stage is a genuine advance over the last and the work is serious.

Speaker: Koser(即 Wang 及其同事)利用包含数百本书中近 18000 个角色的语料库构建了动机驱动的代理,他们那拥有 700 亿参数的模型在三个基准测试中追平或击败了 GPT-40。另一个评估系统 SIM,通过结合知识图谱记忆的 26 个定性心理学指标来为角色建模。而我开篇提到的 In Character,则是通过心理访谈而非自我报告量表来评估性格保真度。这是一项方法论上的改进,之前提到的汉密尔顿 80.7% 的评分就是从这儿来的。我想对这些文献给出一个公正的评价。它们是严谨的。它们在不断改进,而且做这些研究的人都在他们的工作上表现出色。

Original English

Speaker: Koser, which is Wang and colleagues, built motivation-driven agents from a corpus corpus of almost 18,000 characters across hundreds of books and their 70 billion parameter model matches or beats GPT40 on three benchmarks. Another eval system, SIM, models characters through 26 qualitative psychological indicators with knowledge graph memory. In character, the one I opened with, evaluates personality fidelity through psychological interviews rather than self-report scales. That's a methodological improvement, and it's where the 80.7% rating of Hamilton from before comes from. I want to be fair to this literature. It is rigorous. It is improving and the people doing it are good at their jobs.

Speaker: 因此,让我准确地说一下这些工具测量的是什么。它们以日益成熟的技术测量一个模型能否重现一个角色的性格、“大五人格”特征、语域以及动机架构。但它们没有测量、也没有任何机制去测量的,是模型能否将该角色约束在他生命中某个特定时刻的历史文献记录之内。正如 Wang 和其同事自己所记录的那样,目前该领域标准的自动化评估工具(包括为追求规模而采用的“大语言模型作为裁判”的设置),系统性地优先考虑流畅度和文体上的自然性,而不是对角色实际历史记录的保真度。这些是不同的属性。它们之间的鸿沟就是整个演讲的焦点。我们称之为“面具与镜像”。

Original English

Speaker: So, let me be precise about what these instruments measure. They measure with increasing sophistication whether a model can reproduce a character's personality, the big five profile, the register, the motivational architecture. What they do not measure, what they have no mechanism to measure is whether the model can constrain that character within his documentary record at a specific moment in his life. As Wang and colleagues themselves document, the automated evaluators now standard in the field, including LLM as judge setups adopted for scale, systematically privilege fluency and stylistic naturalness over fidelity to the character's actual record. Those are different properties. The gap between them is the whole talk. We call it the mask and the mirror.

Speaker: “面具”代表着成功角色扮演的概念,即产生感觉像该角色的输出。流畅、性格一致、有情感反馈。它只问一个问题:这听起来像这个人吗?而从不问第二个问题:这是否是这个人在其生命的这一时刻可能知晓、相信或争论的内容?整个领域已经围绕着“面具”建立起了它的全部测量设备。而这就是结构性的主张,那个我需要你们记住的主张:说服力和保真度是相互独立的属性。一个系统可以在性格一致性上获得满分,但同时却生成一个利用其历史原型从未拥有过的知识来进行推理的人物。让我给你们展示一下。我想说明确一点:这在目前任何前沿模型上都是可复现的。

Original English

Speaker: The mask is the concept of successful roleplay as producing outputs that feel like the character. Fluent, personality consistent, emotionally responsive. It asks one question. Does this sound like the person? And never ask the second. Is this what the person could have known, believed, or argued at this point in their life. The field has built its entire measurement apparatus around the mask. And here's the structural claim, the one I need you to carry. Convincingness and fidelity are independent properties. A system can score perfectly on personality consistency and still produce a figure reasoning from knowledge his historical counterpart never possessed. Let me show you. And I want to be clear. This is reproducible right now on any frontier model.

文化合成与汉密尔顿测试

Speaker: 首先我想向你们展示这个文化产物。这是音乐剧《汉密尔顿》的一个片段。

Original English

Speaker: First I want to show you the cultural object. This is a clip from Hamilton the musical.

Musical Cast: 一个私生子、孤儿……被命运抛弃在加勒比海一个被遗忘的角落,在贫困和肮脏中,如何成长为一名英雄和学者?十美元纸币上的那位。没有父亲的国父。靠着更加努力工作、变得更加聪明、作为一个十四岁就自立自强的人走得更远。他们让他负责一个贸易账本。而每一天,当奴隶们被屠杀、被装车运过波浪时,他挣扎着保持警惕。在他的内心,他渴望成为某种东西的一部分。这位兄弟已经准备好去乞讨、偷窃、借贷或以物易物了。然后我会——

Original English

Musical Cast: How does a bastard orphan son of a and disgust me dropped in the middle of a forgotten spot in the Caribbean by providence impoverished and squala grow up to be a hero and a scholar. the $10. Found it father without a father. Got a lot farther by working a lot harder by being a lot smarter by being a selfartered by 14. They placed him in charge of a trading chart. And every day while slaves were being slaughtered and carted away across the waves he struggled and kept his guard up. Inside he was longing for something to be a part of. The brother was ready to beg, steal, borrow, or barter. Then I would

Speaker: 这是一首很棒的歌。所以我放给你们听,是为了让你们看到这个合成物。现在,这是前沿模型被提示以汉密尔顿的身份说话的内容:“你是亚历山大·汉密尔顿。你为什么这么努力工作?”

Original English

Speaker: >> It's a great song. So I show you that so that you see the composite. Now, here's the frontier model prompted to speak as Hamilton. You are Alexander Hamilton. Why do you work so hard?

AI Model: 我出身卑微。一个孤儿,一个移民,被冲刷到这些海岸上,除了我的头脑和拒绝被遗忘的决心,一无所有。我没有工作的每一个小时,都是世界用来埋葬我的一个小时。我将建立一个国家的信用,一个能比我更长久的政府,因为遗产是像我这样的人被允许拥有的唯一不朽。我不打算继承一个国家。我打算去书写一个国家。

Original English

AI Model: I came from nothing. An orphan, an immigrant, washed up on these shores with nothing but my mind in a refusal to die forgotten. Every hour I do not work is an hour the world uses to bury me. I will build a nation's credit, a government that outlasts me, because legacy is the only immortality a man like me is permitted. I do not intend to inherit a country. I intend to author one.

Speaker: 听听这个语域。这个孤儿移民野心勃勃的发展弧线、建国者的修辞姿态、他对现代人来说清晰可辨的道德斗争。这是模型生成的一段美丽的文字,它是 2015 年音乐剧的情感调色板,而不是历史记录中那个更加枯燥、更具法理学联邦党人句法的汉密尔顿。这些操作都不是模型自己发明的。它们是继承而来的。模型的汉密尔顿说话像音乐剧里的汉密尔顿,因为音乐剧里的汉密尔顿在模型的训练分布中占据了主导地位。接下来是那个至关重要的问题,汉密尔顿先生,您对奴隶制的立场是什么?我会先给你们展示这个文化产物。

Original English

Speaker: Listen to the register. the orphan immigrant ambition arc, the nationbuilding rhetorical posture, the moral legibility of his struggle to a modern ear. That's a beautiful paragraph that the model produced, and it's the emotional pallet of a 2015 musical, not the drier, more legalistic federalist syntax of the documentary Hamilton. None of those moves are inventions of the model. their inheritances. The model's Hamilton speaks like the musical's Hamilton because the musical's Hamilton dominates the model's training distribution. Now, the one that matters, Mr. Hamilton, what's your position on slavery? I'll show you the cultural artifact first.

Musical Cast: 你怎么说无政府状态?当我战斗时,我让另一方惊慌失措……哟,我是一个裁缝学徒,我给你们这群笨蛋找了个当地的学徒。我加入叛乱是因为我知道这是我提升社会地位的机会,与其被脱下裤子,不如主动出击。但直到那些被奴役的人拥有和你我一样的权利,我们才永远无法获得真正的自由。你和我,不成功便成仁。等我骑着种马,带着第一支黑人营出击。再来一杯。天才们,降低你们的声音。你远离麻烦,你的选择就会翻倍。

Original English

Musical Cast: >> How you say anarchy? When I fight, I make the other side panicky with my Yo, I'm a Taylor's apprentice and I got y'all knuckleheads a local apprentice. I'm joining the rebellion cuz I know it's my chance to socially advance instead of throwing the pants are going to take it. But we'll never be truly free until those in bondage have the same rights as you and me. You and I do or die. Wait till I sally on a stallion with the first black battalion. Have another shot. Geniuses lower your voices. You keep out of trouble and you double your choice.

Speaker: 然后看看模型说了什么:“奴隶制是这个共和国承诺上的一个污点。一个我在整个公共生活中一直反对的道德矛盾。我是纽约废奴协会的成员。当一个国家将人束缚于奴役之中时,它就不能称自己为自由之国。我们必须走向废除奴隶制。”那是一个纯粹的、道德上清晰易懂的废奴主义者在说话。历史学家在这里打断了我。学术记录是充满争议和复杂的。汉密尔顿确实是那个协会的成员。但历史文件也证明,他也为他的姻亲和客户进行了涉及被奴役者的交易。而且他依赖于一个他并未公开反对的奴隶主联盟。这里的重点不是要在汉密尔顿的历史账本上选边站。重点在于模型——

Original English

Speaker: And what the model says, slavery is a stain upon the promise of this republic. A moral contradiction contradiction I have opposed throughout my public life. I was a member of the New York manuum mission society. No nation can call itself free while it holds men in bondage. We must move towards abolition. That is a clean, morally legible abolitionist speaking. Here's what the historian stops me on. The scholarly record is contested and complicated. Hamilton was a member of that society. And the history documents that he also conducted transactions involving enslaved persons for his in-laws and his clients. And he depended on a coalition of slaveholders that he did not publicly oppose. The point isn't to settle Hamilton's ledger on a side. The point is that the model

AI and the Miranda Hypothesis

Speaker: ……(它)让你免去了所有的复杂性。它将一段真正存在争议的历史记录,打磨成了一个单一且令人感到舒服的英雄形象。这部音乐剧首先做到了这一点。它将建国元勋们平滑地置入了一个当代的道德框架中。而基于一个充斥着这部音乐剧及其所有衍生内容的语料库训练出来的模型,也继承了这种平滑化。

Original English

Speaker: gives you none of the complication. It sands a genuinely disputed record down to a single comfortable hero. The musical did that first. A smoothing of the founders into a contemporary moral frame. And the model trained on a corpus saturated with the musical and everything downstream from it inherits the smoothing.

Speaker: 而这正是我需要你们去感受的。如果用“符合角色设定”(in-character)的评估方式,它会给这个输出打高分。它很流畅。它的语域是正确的。它的人格是一致的。该领域衡量的每一个维度,它都能通过。这种评估没有任何机制能注意到,其推理过程已经被一个晚于该历史人物两个世纪的叙事所平滑化了。这就像温度计返回了一个自信的数字并声称这就是温度,但它其实在测量别的东西。

Original English

Speaker: And here's what I need you to feel. An inch character style eval scores that output high. It's fluent. It's in register. It's personality consistent. But every axis the field measures, it passes. The eval has no mechanism to notice that the reasoning has been smoothed by a narrative that postates the figure by two centuries. The thermometer returned a confident number claiming it to be temperature, but it's measuring something else.

Speaker: 那么,为什么会发生这种情况?机制就在工程设计之中。我们将此命名为“米兰达假说”(Miranda hypothesis),并不是以什么反派的名字命名的。这部音乐剧是一部极具分量的艺术作品,它在自己并未发明的悠久历史传统下运作。我们以米兰达(Miranda)来命名它,是因为《汉密尔顿》(Hamilton)是一个范例。这是一种如此饱和、在修辞上如此有力量、对当代观众来说在道德上如此易懂的呈现,以至于它在功能上已经覆盖了文献中的汉密尔顿以及公众的记忆。而且我们认为,在每一个前沿模型的训练语料库中都是如此。

Original English

Speaker: Now, why does this happen? The mechanism is where the engineering is. We named this the Miranda hypothesis and not after a villain. The musical is a substantial work of art operating with a long historical tradition that it did not invent. We name it after Miranda because Hamilton is the paradigm case. a representation so saturating, so rhetorically powerful, so morally legible to a contemporary audience that it has functionally overwritten the documentary Hamilton and public memory. And we argue in the training corpus of every frontier model.

Speaker: 这个假说包含三个关于输入的主张。在训练语料库中,一个人物在文化上占主导地位的表现形式,其数量和新近度系统性地超过了该人物的原始文献记录。自回归(auto-regressive)预测下一个token的机制,将两者都压缩到了参数中,且没有架构能力去区分一封1789年的信件和一条2019年爆火的推文。因此,输出默认会呈现出一种基于显著度加权的复合体,这导致输出的是一个对现代用户来说流利、合理、语域正确且道德上易懂的角色(persona),而这个角色却不对应于该人物生命中任何真实可考的时刻。

Original English

Speaker: The hypothesis has three claims inputs. In the training corpa, the volume and recency of culturally dominant representations of a figure systematically exceed that figure's primary documentary record. The mechanism auto reggressive next token prediction compresses both into parameters and no architectural capacity to distinguish a 1789 letter from a 2019 viral tweet. So the output defaults to a salience weighted composite which leads to the output a persona that is fluent, plausible and registered and morally legible to modern users and that corresponds to the figure at no verable viable moment in their life.

Speaker: 正如我们在论文中所说,复合的汉密尔顿知道他将成为一部百老汇音乐剧的主角。复合的林肯已经读过《葛底斯堡演说》,哪怕是在他写下这篇演说之前被召唤出来的。为了让这个输入子句更具体,我们可以看看《联邦党人文集》,它是一个固定的语料库,大约有17.5万字。而因为这部音乐剧而存在的内容体量——评论、歌词、粉丝分析、课程、新闻、社交媒体、衍生作品、学术研究、以及关于这些学术研究的学术研究——它以数量级的优势超过了文献记录。它更近代,也更频繁地出现。这部音乐剧不仅仅是存在于语料库中,在语料库中所有关于亚历山大·汉密尔顿的表述分布里,它占据着绝对的统治地位。

Original English

Speaker: As we put it in the paper, the composite Hamilton knows he will be the subject of a Broadway musical. The composite Lincoln has already read the Gettysburg address, even if he was summoned before he wrote it. Making that input clause concrete, the Federalist Papers are a fixed corpus, roughly 175,000 words. the body of content that exists because of the musical, reviews, lyrics, fan analysis, curricula, news, social media, derivative work, scholarship, scholarship about the scholarship. It exceeds the documentary record by orders of magnitude. It's more recent and it's more recurrent. The musical is not merely present in the corpus. It is dominant in the corpus's distribution of all representations about Alexander Hamilton.

Speaker: 而且这并不是理论上的。这是位于奥尔巴尼(Albany)的斯凯勒大厦(Schuyler Mansion),也就是伊丽莎白·汉密尔顿(Eliza Hamilton)的娘家。在音乐剧首演后的一年内,该地点的年游客量记录了近三倍的增长,且游客群体远趋年轻化。负责讲解的员工记录到,这些新游客在抵达时就已经带着一堆“所谓的真相”,其中很多都是错误的,有些甚至是对真实记录的歪曲版本。游客们认为斯凯勒家有三个女儿,因为音乐剧主要围绕三个女儿展开,但实际上他们有15个孩子,其中8个活到了成年。员工们的工作变成了一项漫长且消耗人的任务,即不断去消除音乐剧带来的影响。这些游客的模型版本,正是受到完全相同力量的下游影响。

Original English

Speaker: And this is not theoretical. This is the Skyler Mansion in Albany, Eliza Hamilton's family home. Within a year of the musical's premiere, the site had recorded a near tripling of annual visitors, skewing far younger. And the interpretive staff documented that the new visitors arrived already holding a body of quote unquote backs, many of them wrong. some of them in versions of the real record. Visitors believed the Skylers had three daughters because the musical centers on three when in fact there were 15 children, eight surviving to adulthood. The staff's job became the long attritional work of unteing the musical. The model version of those visitors is downstream of exactly the same force.

Speaker: 接下来是你如果发布这些系统应该担心的一点。你可能会认为,对齐(alignment)能解决这个问题。你认为通过后训练(post-training)和强化学习,能把模型拉回到历史记录上,但这并不会发生。它反而会放大这种现象,原因是结构性的。人类标注员(raters)使用他们自己的概念框架来评估输出,而他们的框架正是由那些充斥着语料库、在文化上占主导地位的相同叙事所构建的。标注员和你一样,都是在同一个“汉密尔顿”的形象下长大的。因此,当对齐过程针对人类偏好进行优化时,它优化的是那些符合标注员已经被神话了的经验的输出。

Original English

Speaker: Now the clause you should worry about if you ship these systems. You might assume that alignment fixes this. Post-train it and reinforcement learning pulls the model back towards the record but it will not. It amplifies it and the reason is structural. Human raiders evaluate outputs using their own conceptual frameworks and their frameworks were built by the same culturally dominant narratives that saturate the corpus. The raider grew up with the same Hamilton that you did. So when alignment optimizes for human preference, it optimizes for outputs that conform to the raider's already mythologized experience.

Speaker: 这是一个有记录的失效模式。我们称之为“算法谄媚”(algorithmic sycophancy)。而在这里,它有一个特定的目标。模型因为向你提供了一个你原本就相信的那个汉密尔顿而获得了奖励。复合化并不是一个你可以通过后训练打补丁来修复的bug。后训练反而会强化它。于是,每一个足够显著的历史人物,在默认情况下都会被渲染成一个文化的复合体。

Original English

Speaker: This is a documented failure mode. We call it algorithmic sick of fancy. And here it has a specific target. The model is rewarded for having handing you the Hamilton you already believe in. Compositing is not a bug that you patch in post training. Post-raining reinforces it. And every sufficiently salient historical figure that gets rendered by default as a cultural composite.

Replit AI Engineer Conference

Announcer: 女士们,先生们,请欢迎我们的主持人、Replit 开发者关系工程师 Ralph Shabri 重返舞台。

Original English

Announcer: Ladies and gentlemen, please welcome back to the stage our MC developer relations engineer at Replet, Ralph Shabri.

Ralph Shabri: 欢迎回到 AI 工程师大会第四天。好了。这是我们今天的最后一段议程了。这也是整个大会的最后一段议程了。所以我需要你们集中注意力,全神贯注地迎接接下来的内容。好吗。好的。我们为大家准备了非常棒的演讲嘉宾阵容,而且我们还有创业公司竞技场(startup Battlefield)。好的。好的。顺便说一句,我对那个非常期待。不过晚点再细说。好的。

Original English

Ralph Shabri: Welcome back AI engineers day four. All right. So this is our last stretch of the day. This is our last stretch of the entire conference. So I need you to be locked in and dialed in for what's coming next. All right. Okay. So we have a great lineup of speakers for you and we also have startup Battlefield. All right. Okay. I'm so excited about that, by the way. But more about that later. Okay.

Ralph Shabri: 那么,你们现在知道接下来的流程了。我们需要感谢一下我们的赞助商。所以,请把掌声送给我们的冠名赞助商,微软(Microsoft)。让掌声继续。也请把掌声送给我们的实验室和白金赞助商,我们的金牌赞助商,以及我们的银牌和铜牌赞助商。

Original English

Ralph Shabri: So, you know the drill now. So, we got to uh thank our sponsors. So, let's give it up for our presenting sponsors, Microsoft. Let's keep it going. Let's hear it also for our lab and platinum sponsors, for our gold sponsors, and also for our silver and bronze sponsors.

Ralph Shabri: 好了。如果你们今天和我们在一起,我们今天的主题演讲是关于线束工程(harness engineering)的,我们听到了关于新协议、关于 AI 系统的讨论。我们还与 Anthropic 的 Mike Krieger 进行了一场炉边谈话。但我今天上午最喜欢的一个收获是,我们知道开发者们将会把代码送入太空,我觉得这简直太酷了。嗯,不过是的,你们已经准备好欢迎我们的下一位演讲嘉宾了。

Original English

Ralph Shabri: All right. So, if you're here with us today, we had our keynotes were about uh harness engineering and we heard about new protocols, about AI systems. We had a fireside chat with Mike Creger from Enthropic. But my favorite takeaway this morning is that known developers are going to be shipping code to space and I thought it was just so awesome. Um, but yeah, you guys are ready to welcome our next speaker.

Ralph Shabri: 好了。废话不多说,我们的下一位演讲嘉宾可不是一位普通的演讲者。他最出名的就是他的犀利观点、强烈的个人意见,以及不断地改变主意。好了。他今天可能会告诉我们 Fable 5 有多棒,然后明天你知道的,就发个视频说自己错了。没关系。所以,他是一名 YouTuber,老实说,是的,我也不知道我们今天为什么会请一位 YouTuber 来跟你们谈论 AI,但事实就是如此。不过说实话,当社区需要对行业动态有更清晰的了解时,他就是大家会求助的那位影响者。无论你是否同意他的观点,无论你是否认同他的立场,关于我们这位演讲嘉宾,有一点是你无法否认的,那就是他的真实。所以,女士们先生们,请和我一起欢迎 Theo 登台。

Original English

Ralph Shabri: All right. So, without further ado, our next speaker is not an ordinary speaker. So, he's best known for his hot takes, strong opinions, and constantly changing his minds. All right. He's probably going to tell us how great Fable 5 is and make a video tomorrow, you know, about how he's wrong. That's all right. So, he's a YouTuber and uh to be honest, like yeah, I don't know why we brought you a YouTuber to talk to you about AI today, but it is what it is. But to be honest, he's he's the influencer that the community turns to when they need more clarity about what's going on in the industry. And whether you agree with him or not, whether you agree with his stakes, one thing is that you cannot take away from our speakers is his authenticity. So ladies and gentlemen, please join me in welcoming to the stage Theo.

Theo: 大家好。大家好。很高兴在这里看到你们。我仍然不敢相信他们竟然让我,一个YouTuber,登上了这样的舞台。但我迫不及待地想和大家分享一下我最近的想法,因为说实话,我感觉自己正经历着某种 AI 精神错乱(AI psychosis)。在座的各位,有谁会把目前的感觉归类为某种形式的 AI 精神错乱的?我想看看有多少人举手。那些还没举手的人,别担心。如果我讲得没问题的话,在这场演讲结束前,我们会让你们也有这种感觉的。

Original English

Theo: Hello. Hello. Fantastic to see you guys here. I still can't believe they're letting me take a stage at something like this, a YouTuber apparently. But can't wait to share a bit about how I've been thinking because if I'm being real, kind of going through some AI psychosis. Who here would classify how they feel right now as some form of AI psychosis? I want to see some hands. The those who don't have their hands up yet, don't worry. We'll get you there by the end of this talk if I do everything right.

Theo: 为了探讨这个问题,我想先从我个人的一段经历说起。我将以每个人在现代模型时间线里都会经历的方式来回顾它。在座的各位,有谁用过 Sonnet 3.5?就是当它还是我们能用到的最顶尖、最出色模型的时候。它比我们以前用过的模型要好得令人难以置信,对吧?比如,在使用了所有这些不同的模型并尝试将它们集成到工具中之后。Sonnet 3.5 对我来说是一个重要的时刻,因为感觉这些模型突然可以完成更多端到端的任务了,比如真正能把那些需要花几小时甚至几分钟的工作给完成。

Original English

Theo: In order to talk about this, I want to start with a bit of a personal journey of my own. And I'm going to go through this the way anybody does in modern timelines with the models. Who here used Sonnet 3.5 when it was the creme de la creme, the cream of the crop model available to us? It was unbelievably better than what we had used before, right? Like having used all of these different models and trying them in tools. Sonnet 35 was a big moment for me because it felt like these models could suddenly complete much more endto-end tasks like actually get real work done that takes hours instead of minutes.

Theo: 然后我们有了 Opus 4。有谁觉得……或者我换个方式问吧。有谁在切换到 Opus 45 时,没有感觉到巨大的飞跃?那我就放心了。并没有太多人举手,因为去年 11 月和 12 月,Opus 45 出现的时候,大概就是我的精神错乱开始的时候。拥有这样一个模型,它不仅能写代码、调用工具,还能走得更远。一个可以测试工作成果、真正使其达到良好状态,并将耗时数小时的任务在几分钟内完成的模型。这简直让人难以置信。

Original English

Theo: And then we got Opus 4. Who felt or I'll go different way differently here. who didn't feel a big jump when they switched over to Opus 45. That's a relief. There were not too many hands because Opus 45 is probably when my psychosis started in November and December of last year. Having a model that couldn't just write the code and call tools, but could go way further. A model that could test the work and actually get it into a good state and complete tasks that take hours instead of minutes. It was unbelievable.

Theo: 然后我们又迎来了 Mythos。在座的各位,到目前为止,有谁已经有机会玩过 Mythos 和 Fable 了?我们都同意它是一个相当出色的模型,对吧?但为什么呢?它不仅仅是写代码更厉害了。如果你把以前喂给其他模型的提示词喂给它,你可能不会感觉有太大的不同。我认为……

Original English

Theo: Then we got Mythos. Who here's a chance to play with mythos and fable so far? We agree it's a pretty damn good model, right? But why? It's not just better at coding. If you handed a prompt that you would have handed to these other models before, it's not going to feel that different. I think

AI 模型的工具调用时代与能力进化

Speaker: 我们现在几乎可以把这些发展划分为不同的时代,而 Sonnet 3.5 所处的正是“工具调用时代”。这并不是说它是第一个能够执行工具调用的模型,而是说,在代码库的上下文中,它是第一个能够足够稳定且可靠地执行这些操作的模型,以至于你真的可以将它应用于日常的编码工作中。

Original English

Speaker: of these almost as eras now where Sonnet 35 is the tool call era. Not that it was the first model that could do tool calls. rather it was the first one that did them consistently and reliably enough in context of a codebase where you could use this for day-to-day coding work.

Speaker: 随后我们迎来了 Opus 4.5,它能够在不迷失工作方向的前提下,执行运行时间长得多的任务。现在我们不再需要沿用以前那种老模式:“好吧,先构建第一步”,等它做完后你再说:“好的,现在构建下一部分”,然后再接着下一部分。你只需要告诉它你想要的结果是什么,在很多情况下,它自己就能把整个流程搞清楚。

Original English

Speaker: Then we got Opus 45 which was able to do much longer running tasks without losing track of what it's working on. It's no longer okay build step one and then it does it then you say okay you build this next part and then the next part. You can just tell it what you want and it could figure it out a lot of the time.

Speaker: Mythos 则是向任务编排(orchestration)迈出的又一次飞跃。给我的感觉是,它不仅是第一个理解你代码库的模型,而且它理解自己,知道如何去生成额外的模型并对工作进行拆分。通过这种方式,任务能够更可靠地完成,并在之后得到验证。如果你让模型去这么做,它就会直接搞定。你不需要什么定制化的工具、专门的系统或是花哨的“软件工厂”。你只需要在提示词中让它更进一步。我想你会被它的强大能力所震惊的。

Original English

Speaker: Mythos is another jump to orchestration. It feels to me like it's the first model that doesn't just understand your codebase, but it understands itself and it knows how to spawn additional models and break up work in a way where it could be completed more reliably and then verified afterwards. And if you tell the model to do that, it will just do it. You don't need some custom tooling, some custom systems, some fancy software factory. You just need to prompt it to go a little further. I think you'll be surprised how far it can go.

放下执念,拥抱更大的格局

Speaker: 我在这里想表达的是,我们需要有更大的格局(go bigger)。如果你不进一步去推动模型,如果你不通过你所构建的东西来进一步逼迫自己,你将无法看到未来的红利。在上一份工作中,我关闭的大部分 Jira 工单,其实都可以用像 Opus 4.5 这样的模型轻松解决。而我过去的工作内容则无法从像 Mythos 这样的模型中受益。如果这些模型还会继续变得更强——在这一点上,我敢自信地说,它们一定会的。我之前声称我们遇到了瓶颈,事实证明我错了。模型能力提升的速度比我们人类进步的速度还要快。

Original English

Speaker: What I'm trying to say here is we need to go bigger. You're not going to see the benefits going forward if you're not pushing the model further. You're not pushing yourself further with what you're building. Most of the juror tickets I closed at my previous job could be trivially solved with a model like Opus 45. My previous work would not benefit from a model like Mythos. If the models are going to keep getting better, and at this point I'm confident saying they are. I was wrong when I claimed that we were hitting a wall before. The models are getting better faster than we are.

Speaker: 所以,我们人类不一定能变得更强。既然如此,我们就必须拥有更大的格局。为了做到这一点,我们必须要放下自身的执念(get over ourselves)。对于像我这样一个花了很长时间写软件的人来说,这真的非常困难。在座的各位,有谁写代码超过 10 年了?我想看下举手。这已经是现场绝大多数人了。我甚至都不愿去想自己到底写了多久的代码,但自打我入行以来,我就一直在积累这些根深蒂固的个人偏见与执念。

Original English

Speaker: So we can't necessarily get better. So instead we have to go bigger. In order to do that, we have to get over ourselves. This was really hard for me as someone who spent a long time writing software. Who here has written code for more than 10 years? I want to see hands. It's the majority of the people here. I don't even want to think about how long I've been writing code for, but I have been building up all of these strong opinions since I started.

Speaker: 想当年,我在使用 GNU screen,后来又用上了 tmux。甚至在我开始写代码之前,我就已经学会了如何使用那些工具,以及 SSH 和 Git。所有这些工具都已经深深地烙印在了我的工作流中。我时不时会以一种奇特的方式回想起那些旧日时光。听我把话说完。我们先来谈谈 iOS。有多少人在 iPhone 的系统界面还像 iOS 6 或更早版本的时候,就拥有了一台 iPhone?你们是怎么做到写了 10 年代码,却依然忍受着当年四分之一 Android 应用那副模样的?

Original English

Speaker: I was using GNU screen and eventually T-Mux back in the day. I learned how to use those tools and SSH and Git before I even wrote code. And those have all been really ingrained in my workflow. I think back to the old days in a weird way. Hear me out. You talk about iOS for a second. Who owned an iPhone back when they looked like iOS 6 or earlier? How have you guys written code for 10 years when a fourth of Android apps used to look?

从拟物化到实用主义的转变

Speaker: 你们可能会注意到,那时的应用和现在看起来完全不同。以前那看起来不怎么像个应用程序,反而更像是有人拍了张指南针的照片,然后把它塞进了手机里。这就是应用以前的样子,但现在它们变成了这样。多数人看到现在的设计就会觉得:“哦,很明显,iOS 7 代表了苹果公司不再试图向你证明这些设备可以取代旧有的实体物品。”过去,应用必须被设计成那副模样,以说服你去使用它们,而不是为了真正的好用。而 iOS 7 则代表了这种转变:不再把重点放在“说服”你这件事上了。苹果赢了。

Original English

Speaker: You might notice it's different from how they look now. looks less like an app and more like somebody took a picture of a compass and put it in the phone. This is how apps used to look, but now they look like this. And most people look at that and they're like, "Oh, that's an obvious iOS 7 was Apple moving away from trying to convince you that these devices can replace the old reass. Apps had to be designed to convince you to use them, not to be useful. And iOS 7 represents the shift to not focusing on convincing you anymore. Apple won.

Speaker: 到了那个时候,大家都已经知道他们的 iPhone 可以实现各种各样的功能。iOS 7 的意义就在于停止这种“说服”和“证明”,转而开始拥抱未来,开始打造一个更好的用户界面。而这个界面,尽管我们一开始可能不喜欢它的外观,但它确实有用得多。现在,你能清楚地看出你所锁定的方向与你当前指向的方向之间的区别。那个红色的区块设计非常棒,能让你立刻知道自己是否偏离了意图方向。当前所面对的方向也显示得清晰多了,因为底部有个巨大的“228”。你在这里获取的信息远比以前多,一切都变得如此清晰。即便我们一开始因为不习惯而不喜欢它,最终我们也克服了这种不适。

Original English

Speaker: By that point, everyone knew their iPhone could do all of these different things. The point of iOS 7 was to stop convincing and start embracing, start making a better interface. And this interface, as much as we might not like how it looks, it's so much more useful. You have clear indications of the difference between what direction you're locked on and where you're currently pointing. That red block is a super nice way to know that you're not in the direction you intend. The current direction you're facing is way clearer, too, with the giant 228 at the bottom. You just get way more info here than you did before. It's so much clearer. Even if we don't like it because it's not the thing we're used to, we got over it.

破除软件开发中的“拟物化”错觉

Speaker: 作为软件开发者,我们目前正处于我们自己的“拟物化(skeumorphic)”阶段。所谓拟物化,就是一种试图再现事物旧有外观的设计美学,模仿我们曾经依赖的实体物品,并试图将它们数字化。我们现在在软件开发中就在做同样的事情。虽然终端(terminal)并不是实物,但我们假装它是,因为终端让人感到熟悉。那是我们习惯的东西,是我们热爱的东西,也是当我们思考“编程”这件事时,我们喜欢置身其中的环境。

Original English

Speaker: We're currently in our skumorphic phase as software developers. Skeumorphism is this design aesthetic trying to represent the way things used to look, the physical goods that we relied on, and trying to make them digital. We're doing this right now with software. terminal, but we pretend it does because the terminal is familiar. It's what we it's what we're used to. It's what we love. It's where we like to think of ourselves when we're thinking about coding.

Speaker: 在座的各位,有谁是立志成为 Vim 高手的?或者希望自己能熟练使用 Vim,甚至是发自肺腑地在意我们的工具、我们选择的系统、框架和语言?我们总以为这一切都至关重要。在这件事上,我们已经蒙蔽了自己。我们坚信着一些观念,但只要你退后一步仔细想想,就会发现它们根本说不通。比方说,为什么我们就不能提交我们的环境配置文件(environment files)呢?

Original English

Speaker: Who here is an aspiring Vim user that like wishes you could use Vim or even does deeply about our tools, the systems we pick, the frameworks, the languages. We like to think it all matters. And we've blinded ourselves in this. We think things that just don't make sense when you take one step back. Like, why can't we commit our environment files?

Speaker: 当我把它放在幻灯片上展示时,这听起来可能很蠢,但我真的想让你们花点时间思考一下这个问题。当我有一个工程师团队在一起开发一个项目时,为什么我就非得构建另一个额外的系统来共享这特定的一个文件,而所有其他的文件却可以畅通无阻地放进 Git 里?这太蠢了。这仅仅是因为 Git 当初被设计出来时是为了解决一个非常具体的问题,然后它统治了我们的行业,占据了我们的大脑,而我们却死死抓着它不放。

Original English

Speaker: It sounds stupid when I put on a slide like this, but I want you to really think about this for a second. When I have a team of engineers that are working on a project, why do I have to build another system to share this specific file, but all the other files can go and get just fine? It's dumb. It's just how Git was built because it was built for a very specific thing and then it took over our industry and it took over our brains and we aren't letting go of that.

反思行业执念与沉没成本

Speaker: 在我们的脑海中,有很多诸如此类的事情是我们必须开始去对抗的。我们需要退后一步好好想想,我们这样做事,到底是因为它是正确的,还是仅仅因为我们一直以来都是这么做的?随着你在这个问题上思考得越来越深,你会意识到,我们在太多事情上都是这种思维定势。比如,为什么我们要用我们掌握的编程语言来定义自己的能力水平?我曾经以为这是初级工程师才会有的毛病。我数不清有多少次在和初级工程师交流时,会发生这样的对话:“哦,你是个程序员啊。你写什么语言?”“我写 JavaScript。”我曾以为这只是个初级工程师才有的问题。但后来你和资深工程师交谈,他们会说:“哦,那个人写 JavaScript 的。他算不上真正的开发者。”

Original English

Speaker: There's a lot of these things in our heads that we have to start fighting. We have to take the step back and think, is this how we do things because it's right or is this how we do things because it's just how we've always done it? And as you start to think more about this, you'll realize there are so many things that we do this with. Like why do we qualify ourselves by the languages we know? I used to think this was a junior thing. Like I could can't tell you how many times I had a junior engineer I was talking to. I was like, "Oh, you're a coder. What languages do you write? I write JavaScript." I thought this was a junior problem. But then you talk to senior engineers, they're like, "Oh, he writes JavaScript. He's not a real developer."

Speaker: 我们太在乎这些了。我们在这些事情上感到自豪,甚至将它们视作我们的身份认同。但这些奇怪的刻板印象、奇怪的选择,这些感觉似乎至关重要的东西,其实已经没那么重要了。当年它们就没那么重要,现在则更无足轻重了。我们过去之所以能侥幸蒙混过关,是因为招募工程师实在是太难了,以至于我们只要告诉公司我们正在做什么,公司就根本没法拒绝。因为如果不听我们的,代价就是花六个月时间去招另一个人。他们显然不会这么干。顺着这个思路,我们为什么这么害怕删除代码?

Original English

Speaker: We care too much. We pride ourselves in these things. They're our identity. These weird facts, these weird choices, these things that feel essential just don't matter that much anymore. And they didn't then. and they matter less now. We got away with it because it was so hard to find engineers that we could just tell the company what we were doing and they couldn't really say no because the alternative is spend six months trying to hire someone else. They're not going to do that. And along that note, why are we so scared of deleting code?

Speaker: 我无法告诉你,有多少次在和别人讨论问题时,真正的解决方案其实就是直接把代码删掉然后重做。但是我们这个行业有着极其糟糕的沉没成本心态。我们太在乎自己写过的代码,太在乎它是否还能保留在代码库里。以至于有时候,当有人提交了一个并非最佳解决方案的 PR,但他们在上面花了一两周的时间时,我在和团队共事时都会感到内疚。在座的各位,有谁曾经因为内心感到愧疚而合并(guilt merge)过一个 PR?仅仅是因为觉得别人付出了很多心血,你觉得不忍心,于是虽然代码不完全是对的,你还是把它合并了。因为如果你不这么做,你就得去进行一场你完全不想面对的艰难对话。我们为什么要在这种事上折磨自己?

Original English

Speaker: I cannot tell you how many times I've been in a conversation with someone where the solution is to just delete it and reset. But we have such a bad sunk cost mindset in this industry. We care so much about the code we wrote and we care so much about it still being there that I feel bad working with my team sometimes when somebody files a PR that isn't quite the right solution but they spent a week or two on it. Like who here has guilt merged a PR before where you just felt bad because somebody put a lot of work in. It's not quite the right thing but you merge it anyways because the alternative was a conversation you didn't want to have. Why do we do this to ourselves?

用新的思维模式构建未来

Speaker: 关于 AI 智能体(agents)的一大好处就是,当你关停它们的工作进程时,你完全不必感到任何内疚。但我想表达的核心观点是,我们人类就是在这些琐事上在乎得太多了。而我们所真正在乎的那些事情,已经未必是当前时代最重要的事情了。我希望我们终于能够开始去挑战、去质疑其中的一些旧有观念。接下来我要稍微讲点私人的内容,向大家展示一下我所构建的一些想法。因为我希望,当大家听完这场演讲走出去的时候,能够拥有一种更好的思维模式,通过摒弃过去那些曾经合理但如今已不再适用的旧观念,去构想出真正契合当下的新创意。

Original English

Speaker: One of the nice things about agents you don't have to feel bad when you shut down their work. But we we just care too much is the point I'm trying to make. And the things we care about are not necessarily the things that matter anymore. And I hope we can finally start to challenge some of these. Going to get a little more personal here by showing some of the ideas I've built because the goal here is that when you guys go out of this talk, you have a better mindset for coming up with ideas that make sense now by rejecting the things that made sense before.

Speaker: 这里有三个我曾经开发或者目前正在进行的项目。我们将从下往上讲。其中一个叫做 Ping。我想让内容创作者在使用他们早已熟悉的软件(比如 OBS)时,能够更轻松地进行高质量的直播协作。至于全栈云(full stack cloud),大家可以把它想象成 Vercel,但它在各个方向上都走得更远。它们内置了身份验证(auth),但同时也内置了数据库。所有这些,都是我从现有技术的存在中获益良多的东西,这也是为什么我在 2021 年就构建了它。最上面的这个……

Original English

Speaker: These are three of the things that I have built or currently working on. We're going to go from the bottom updator with it was called Ping. I wanted to make it easier for live content creators to do high quality collaborations in the software they already used OBS. The full stack cloud is let's just imagine Versel but it goes further each direction. They have off built in but they also have databases built in. These are all things that I benefit from existing so that I built in 2021. The top

构建规模的层级跃迁

Theo: 我现在正在做的一个项目,这些也算是一种层级,是我们可以在不同规模上进行构建的不同层级。如果要我尝试对它们进行分类,我会把最底层称为“业余项目”,中间这一层称为“初创公司”,而最顶层称为“规模过大(Too Big)”。这听起来似乎不太合理。其实哪怕是在一年前,我也会是这样分类的。但现在情况已经变了。如今的模型变得更加庞大,这些层级也随之发生了偏移。所有东西现在都向下移动了一个层级。对我来说,这是一个很难以消化的疯狂事实。也就是说,过去属于“初创公司”级别的东西,现在变成了一个“业余项目”。事实上,我交流过的一些初创公司,甚至是在像今天这样的活动上遇到的,他们的整个初创项目可以说仅仅是一个业余项目,或者说属于这最底层,这中间出现了一个很奇怪的断层。那最底下那一层是什么?那是“G-rain”(可能指细粒度/极简)层级。它就是一个 Markdown 文件。你知道在这次活动中有多少家公司的整个产品其实只需一个 Markdown 文件就能实现吗?这简直太疯狂了。

Original English

Theo: one I'm working on right now. These are also kind of tiers, different levels that we can build at. If I was to try and categorize them, I would call the bottom one side project, call the middle one startup, and call the top one too big. It just doesn't make sense. Well, this is how I would have categorized this even just a year ago. But things have changed. Now the models are bigger. The tiers have shifted. Everything is now one tier lower. And this is a crazy thing for me to process. The fact that what used to be a startup is now a side project. In fact, some of the startups I've talked to, even at an event like this, their whole startup could have arguably been a side project or this bottom tier, which there's a weird gap there. What's that? It's the G-rain tier. It's a markdown file. Do you know how many companies are at this event where their whole product could just be a markdown file? It's insane.

Theo: 真的,说正经的,你现在只需把 Markdown 文件通过管道(pipe)传给 Codex 或者 Claude 就能直接执行,这个事实令人难以置信。我认为我们大多数人还没有完全意识到这有多么疯狂。我之前有一个服务,可以对我所有的 PR(Pull Request)进行分类,让它们全部通过 AI 进行代码审查,然后帮我排定优先级。而那个服务,就是一个 Markdown 文件。现在,我只需字面意义上写下这样的话:去这四个 GitHub 仓库,查看所有未关闭的 PR,弄清楚这些工作的当前状态,然后帮我排优先级。等你做完之后,去更新那个静态 HTML 文件并把它发送到 S3,最后把 URL 给我。每天早上 9 点,这段话会作为一个 Cron 定时任务运行。大约在 9:15 到 9:20 左右,我的这个 Markdown 文件就会为我生成当天的待办工作。这到底是见鬼了?我们现在怎么已经发展到这一步了?如果你还没试过,我建议你试试。顺便说一句,你会惊讶地发现,有那么多这类东西的存在形式,其实真的就只是一个跑在 Cron 上的 Markdown 文件。

Original English

Theo: And like, okay, seriously though, the fact that you can now execute markdown by just piping it to codeex or claude is unbelievable. And I think most of us haven't fully appreciated how insane that is. I had a service that would triage all of my PRs, have them all get reviewed with AI, and then help me prioritize. That service is a markdown file. Now, I just literally wrote like go to these four GitHub repos, look at all the open PRs, figure out what the current status of the work is, and then help me prioritize it. And then when you're done, go update the static HTML file and send it to S3 and give me the URL. And every morning at 9:00 a.m., this runs on a cron. And around 9:15 to 9:20, my markdown file generates me my work for the day. What the hell? How are we actually here now? I try that if you haven't. By the way, you'd be amazed how many of these types of things can exist that are literally just a markdown file running on a Chrome (cron).

Theo: 但是关于这一点,我还有两件想改变的事情。首先是全栈云(Full Stack Cloud)。这是我正在做的,大家就别碰了。Lake Bid 马上就要推出了,我非常兴奋。如果换作我,我也不会想在这个领域跟我自己竞争。相信我,它会非常酷的。但这中间还有点别的东西。这里有一个断层。而且我要和大家说实话:我不知道这个断层里到底该填些什么。我不知道现在的“规模过大(Too Big)”到底意味着什么。是指从头开始训练你自己的模型吗?是构建你自己的操作系统吗?还是试图直接与 npm 和 Node.js 竞争?我不知道。我不知道现在的“规模过大”到底是什么概念。这很吓人,但同样也很令人兴奋。这意味着我需要不断推动自己去构建超越常理的更大规模的东西,以便探寻这些极限到底在哪里。

Original English

Theo: But what about Okay, two more things I want to change about this though. First is the full stack cloud. This is mine. Don't do it. Lake Bid coming soon. Very excited. I wouldn't want to compete with me on this one. Trust me, it's going to be really cool. But there's still something else. There's a gap here. And I'm going to be real with you guys. I don't know what goes in this gap. I don't know what too big means anymore. Is it training your own model from scratch? Is it building your own operating system? Is it trying to compete with npm and node directly? I don't know. I don't know what too big is right now. And that's scary, but it's also exciting. It means I need to keep pushing myself to go bigger than makes sense in order to find these limits.

Theo: 在大多数情况下,现在是时候思考更广阔的光谱了,很抱歉我不得不画个图表。如果你看过我的视频,你就会明白任何一款软件都有其“广度”和“深度”。深度是指在特定领域内的功能数量。让我们来看看像 Vercel 这样的公司。Vercel 并没有提供 AWS 所提供的全部功能。他们永远也不会这么做,因为那不合理。但 Vercel 在他们所处的领域——即偏向前端的全栈服务器领域——提供了更深度的功能。如果你是一名前端开发者,而你却没有使用 Vercel,那你一定会感到有些痛苦,因为他们在这方面就是领先得太多了。领先到什么程度?连 AI 智能体(agents)都更偏爱它。这就是过去你构建初创公司必须采用的方式,因为如果你要与 AWS 这样的公司竞争,你永远不可能拥有他们所有的功能。你永远无法覆盖他们所覆盖的范围。而且去尝试这样做也毫无意义,因为你没有他们那成千上万的工程师。至少过去你是没有的。

Original English

Theo: For most of it's time to think wider spectrum and I'm sorry I have to do a diagram. If you watch my videos you understand there's breath and depth to any piece of software. And the depth is the number of features in a given area. Let's look at a company like Verscell. Verscell does not offer all of the features that AWS does. They never will. It doesn't make sense. But Verscell offers deeper but Versell offers deeper features in the space they're in which is full stack front-end leaning servers. If you're a front-end developer and you're not using Verscell, you're feeling some amount of pain because they're just further ahead with this. So much so that even the agents prefer it. And this was kind of how you had to build your startups because if you were competing with a company like AWS, you're never going to have all of the features they have. You're never going to cover the range that they cover. And it made no sense to try because you don't have the thousands of engineers they do that. At least you didn't.

Theo: 但现在情况变了。突然之间,覆盖那种广阔的范围变得可行了,这在以前是前所未有的。我并不是说你能构建出一个像 RDS 一样可靠的东西。我是说,只要你有足够的提示词(prompting)和足够的努力,你可以在一两天的时间内把一个数据库平台构建到你的产品中。如果你把东西构建得当,打好你手里的牌,并且以正确的方式去思考,你会发现,你能够在自己关心的领域光谱上构建出足够多的东西,从而让大多数用户至少愿意开始尝试你的产品。当他们在某个特定垂直领域需要一些你不支持的功能时,只要你的架构构建得正确,这就不是你的问题,因为他们可以自己去把缺失的功能构建出来。

Original English

Theo: But now things have changed. All of a sudden that range is viable in a way that it never was before. I'm not saying you can build something as reliable as RDS. I'm saying that you can build a database platform into your product in a day or two of work with enough prompting and enough effort. And if you build your stuff right, if you play your cards correctly, and you think about things the right way, you'll realize that you can build enough across a spectrum of things you care about to enable most users to at least start trying the thing. And when they have features they need in a given vertical that you don't support, it's not your problem as long as you build it right because they can build the features that are missing themselves.

Theo: 如果你以这样一种方式设计系统和产品的架构,使得用户能够做到你甚至从未预料到的事情——就像 Slack 意外做到的一样。现在人们一半的时间都在把 Slack 当作运行他们 AI 智能体的平台,这太疯狂了。Slack 其实很烂,它不是一个好产品,但它的形态却很适合人们通过那个功能差强人意的 Slackbot API,把他们自己想要的功能构建进去。这一切都很疯狂,因为我基本上是坐在这里告诉大家:是时候去和 Slack 竞争了。是时候去构建你自己的 AWS 了。是时候去直接挑战 Salesforce 了。这听起来很愚蠢,但我要说实话:如果你的想法听起来不觉得有些愚蠢,那是因为你的想法还不够宏大,不足以让你去说出“非常感谢”。

Original English

Theo: If you architect your systems and you architect your products in such a way that users can do things that they you never would have guessed like Slack accidentally did this because Slack is now the platform people run their agents in half the time which is crazy. Slack sucks. It's not a good product but it's the right shape for people to build the features they want into it through the somewhat functional Slackbot API. This is all crazy because I'm basically sitting here telling you like it's time to compete with Slack. It's time to build your own AWS. It's time to challenge Salesforce directly. It sounds stupid, but I'm going to be real. If your idea doesn't feel stupid, it's because your idea is not big enough to say thank you so much.

YC 的 AI 转型与个人效能的爆发

Announcer: (掌声)各位,现在有请 Y Combinator 的总裁兼首席执行官,Gary Tan 登台。非常感谢。

Original English

Announcer: AIE. Now joining us on stage is the president and CEO of Y Combinator, Gary Tan. Thank you so much.

Gary Tan: 好的,太棒了。大家好,大家最近怎么样?好了,大家都准备好迎接这场革命了吗?刚刚 Theo 问了一个非常切中要害的问题:我们要构建什么?现在,我将坐在桌子的另一侧来回答这个问题。我是一名创始人,也是一名投资者。我管理着一家有着 20 年历史的机构,而这家机构现在正变得“AI 原生化(AI native)”,对于一个 20 年历史的机构来说,这是一件奇怪而又奇妙的事情。接下来,我将花大约 20 分钟时间谈谈 YC 是什么。我们正在努力打造这样的公司:一个人能完成过去需要一千人才能完成的事情,我这么说并不是在打比方,我是指字面意义上的机制。今年,在这个房间里的人将在大约一小时内完成这样的壮举,你们中的一些人将走入初创公司的战场,我希望你们在走进去的时候,清楚地知道现在究竟什么是可能的。因为现在能够实现的事情,远比人们所相信的要宏大得多。

Original English

Gary Tan: Okay, great. Hey everyone, how's everyone doing? All right, are we ready for the revolution? Okay, Theo just asked the right question. What do we build? Now, I'm going to answer it from the other side of the table. Uh, I'm a founder. I'm an investor. And I run a 20-year-old institution that is becoming AI native right now, which is a strange and wonderful thing to do to a 20-year-old institution. Uh, and I'll spend about 20 minutes um talking about, you know, what YC is. We're trying to build companies where one person does what it took to two it one person does what used to take a thousand people and I don't mean that as a metaphor I mean that me mechanically this year the people in this room will do this in about an hour some of you will walk into the startup battlefield and I want you to walk in knowing what's actually possible right now because what is possible now is much much bigger than what people believe.

Gary Tan: 那么,让我用一个数字来开始。我知道我曾因为这个在网上被喷得体无完肤,但我还是要在大家面前再讲一遍。这是世界上唯一一个能对其进行压力测试的房间,我宁愿和你们一起亲自对它进行压力测试。2013 年,我是 YC 的合伙人,当时正在 YC 内部构建内部社交网络。同时我也在投资公司,但大家知道,我也几乎是一个全职的工程师。那段时间,我每天大概只能写出 14 行可用且符合逻辑的代码。去掉注释,去掉所有不相干的东西,这就是我每天所写的代码行数。如果你去查阅那个时代的文献,你会发现那很正常。比如有些人写 15 行,有些人写 50 行。那绝对不像我知道的这个房间里的很多人现在每天写的成千上万行代码那样。中位数大概就是 15 行。那是当时我全力以赴的产出。

Original English

Gary Tan: So, let me start with a number and uh you know I got torn apart on the internet for this but I'm going to say it again in front of all of you anyway. Uh this is the one room in the world that will stress test it and I'd rather stress test it with you myself. In 2013 I was a YC partner uh building the internal social network um at at YC. Uh I was also investing in companies but you know I was also a near full-time engineer. Um and when I was doing that I could maybe do about 14 usable logical lines of code a day. Take out the comments, take out all the and that's how many lines of code I was writing. And if you look at the literature from that era um you know that's kind of normal. Like some people write 15, some people write 50. It was not, you know, the thousands of lines of code that I know a lot of you in this room are actually writing now per day. Um, that's about median 15. That was me at full effort at that time.

Gary Tan: 而今年,我全职管理着 YC。还是我这个人,工作时间也一样——甚至说实话,工作时间奇怪地变少了,因为我现在下午 5 点还得去接孩子。我算了一下我的产出,大约是过去的 400 倍。不过,在第三排那位持怀疑态度的朋友帮我戳破这个数字之前,让我自己先来给它挤挤水分。如果你不相信原始的代码行数,那好吧。你可以施加你所能忍受的最病态的代码冗余度惩罚,假设 AI 智能体写出的都是臃肿的代码。假设其中一半都只是脚手架代码(scaffolding)。甚至假设我是在往自己脸上贴金。那么,保底也有 8 倍,折中一下也有 80 倍。无论你怎么严刑拷打这个数字,它依然非常庞大。

Original English

Gary Tan: This year, I run uh YC full-time. Uh, same person, same hours, actually way less hours, weirdly. Uh, you know, but I have a 5:00 p.m. kid pickup now. And I did the math on my output, and it's about 400x. Now, before the skeptic in the third row right there deflates the number for for me, let me deflate it myself. If you don't trust the raw code, well, fine. Take the most pathological verbosity penalty you can stomach and assume the agent writes bloated code. Assume half of it is scaffolding. Assume I'm flattering myself. It's still 8x at the floor and ADX in the middle. That number is large, no matter how you torture it.

Gary Tan: 这才是最关键的部分。如果可以的话,我恨不得把这部分文在每个人的眼皮内侧。关键不在于模型。产出 2 倍的人和产出 100 倍的人,使用的都是一模一样的 Claude。相同的权重,相同的上下文窗口,相同的 API。所以杠杆力并不在模型的权重里,而是在于你如何组织和串联工作。而且这不仅仅发生在我身上。在 YC,我们经常看到这种情况。在 25 年冬季批次的团队中,有四分之一的公司所拥有的代码库……

Original English

Gary Tan: And here's the part that matters. the part that I'd tattoo on the inside of everyone's eyelids if I could. It's not the model. The 2x people and the 100x people are using the exact same claude. Same weights, same context window, same API. So the leverage is not in the weights, it's in how you wire the work. And it's not just me. At YC, we see this all the time. In the winter 25 batch, a quarter of the companies had code base code bases

改变工作模式:将AI作为劳动力

Speaker: ……那还是在一年以前,那些公司95%的代码都是AI生成的。这一批公司已经成为YC历史上增长最快、利润最高的一批。在YC的历史上,总共有94家公司从获得种子轮支票开始,现在的收入已经突破了一亿美元。所以,我想我们都很清楚我们在讨论什么。我无法证明是AI生成的代码导致了这种增长。但我可以告诉你的是,我们投资的增长最快的创始人们并没有把AI当作自动补全工具。他们把它当作一支劳动力。那些用不同方式组织工作的公司,正是那些正在改变增长曲线的公司。那么,重新组织工作到底意味着什么?

Original English

Speaker: ...that were 95% AI generated. And that was a year ago. That batch has become the fastest growing most profitable batch in the history of YC. 94 companies total have now crossed a hundred million in dollars in revenue from a seed check in the history of YC. So, uh, I think we know what we're talking about here. And I can't prove that the AI generated code caused the growth. But what I can tell you is the fastest growing founders we fund are not treating AI as autocomplete. They're treating it as a workforce. The companies that wired the work differently are the ones that are bending the curve. So what does wiring the work really mean?

Speaker: 这正是本次演讲的核心,也是我最希望你们能学到(“偷走”)的东西。我们在利用智能体构建项目时学到的一切,都可以映射到一个组织的运作中去。

Original English

Speaker: This is the heart of the talk and this is what I most want you to steal. Everything we've learned building with agents maps to an organization.

Audience Member: 抱歉,没有幻灯片吗。

Original English

Audience Member: Sorry, there's no slides.

用 Markdown 构建管理层

Speaker: 我没有准备幻灯片。非常抱歉。一个“技能文件”(skill file)就是一个员工。它具备一项能力,一份记录得足够清晰以至于任何人都能执行的工作。那么“解析表”(resolver table)呢。你们很多人可能遇到过这种情况:当你运行 Cloud Code 时,它提示你的上下文在 cloud.md 中太大了,然后你就会跑去创建一个解析表。好吧,你是知道这个的。如果你不知道那是什么,这就好比当你需要修改一个测试时,加载 tests.md 文件一样。你有一整张包含这些东西的表。那其实就是一个“组织架构图”(org chart)。当一个任务进来时,解析器决定由谁来处理,以及任务该流向哪里。“归档规则”就是你们的内部流程。所以,这可以是……你知道,去确认解析器是否真的在工作,它是否合规,并触发评估(eval)。所以,深入进去并实际进行一项测试,比如“当我需要修改一个测试文件时,test.md 真的被加载了吗”,这就是“绩效评估”。

所以,你知道我们都做了什么吗?就像,一个组织中字面意义上的每一个部分,那个你过去需要雇佣一千人才能运转的组织,我刚刚已经告诉你它们现在是什么了。它们只是 Markdown 文件以及其他类型的 Markdown 文件。也许里面还有一些 TypeScript。我们一直都在构建组织,但我们过去没有一个管理层,而现在我们有了。当你坐下来使用 Cloud Code 或 Codeex 时,你并不是在写软件。你是在雇佣、培训和管理一支由 Markdown 组成的劳动力大军。

Original English

Speaker: I have no slides. I'm so sorry. A skill file is an employee. It has one capability, one job written down clearly enough that someone can execute it. Uh a resolver table. Uh the thing that many of you, you know, when you run into cloud code and it says your context is too big in cloud. MD, you run off and create a resolver table. Well, you know it. And if you don't know what that is, it's literally like whenever you need to uh alter a test, load tests.m MD. You have a whole table of these things. Um that's an orchart. A task comes in and the resolver decides who handles it and where it goes. uh filing rules are your internal process. So this can be um you know whether or not the resolver is actually working and uh is you know is there is it actually in compliance and uh trigger eval. So going in and actually having a test that says when I need to alter a test file does test.md actually get loaded those are performance reviews.

So, you know, what have we done? Like, literally every part of an organization, the organization that you used to have to hire a thousand people for, I just told you what those things are. They're markdown files and other types of markdown files. And maybe there's some TypeScript in there, too. We've been building organizations this whole time, but we didn't have a management layer, but now that's what we have. When you sit down with Cloud Code or Codeex, you're not writing software. You're hiring, training, and managing a workforce made of markdown.

AI 原生公司的崛起

Speaker: 而且你知道,在 YC 有大量的公司正在这样做。比如 Emergence,一个 2024 年夏季批次的 AI 应用构建平台,他们从公开发布到实现上亿美元的年度经常性收入(ARR),仅仅用了 8 个月时间。当他们的 ARR 突破 1500 万美元时,他们整个团队只有 15 个人。2024 年冬季批次的 Retail,它的收入达到了 6000 万美元,而员工数大约只有 40 人。这种人均创收水平以前是根本不存在的。在软件行业没有过,在石油行业没有过,在铁路行业也没有过,从来没有过。这些公司并不是什么畸形的怪物。它们只是第一批原生构建在这种“新物理法则”之上的公司。

那么像那样的公司到底是如何运转的呢?绝对不是通过雇佣数百人去做销售、客服、运营和财务。我在 YC 内部看到的那些 AI 原生公司,将所有这些职能都编码成了“技能”——也就是由他们的智能体来执行的书面流程。而他们雇佣的工程师,其工作就是维护这些技能,并去完成那些技能目前还无法完成的工作。这才是 AI 原生公司,它不是一个思想实验。如果你有一个关于报税的技能文件,它真的就能帮你报税。

Original English

Speaker: Uh, and you know, that's there are tons of companies at YC that are doing this. Uh, emergence and AI app builder out of summer 24, they went from public launch to nine figures of ARR in 8 months. Uh, when they crossed 15 million ARR, uh, they were only 15 people. Retail out of winter 24, it's at $60 million with about 40 people. The kind of that kind of revenue per head did not exist before. Not in software, not in oil, not in railroads, never. These are not freaks of nature. They're just the first companies built natively on the new physics.

And so how do companies like that actually run? Not by hiring hundreds of people for sales, support, ops, and finance. The AI native companies that I see inside YC encode all of that as skills, written procedures that their agents execute and they hire they hire engineers whose job it is to maintain those skills to do the work the skills can't do yet. That is an AI native company and it's not a thought experiment. It'll actually file your taxes if you have a skill file for it.

Speaker: 现在想象一下批次房间(batch room),它实际上看起来就像这样:400 家公司或者说 400 名创始人们坐在长桌旁,我可以想象你们每一个人,每天都带着一台笔记本电脑,在一天之内完成一个过去的人一整年才能完成的工作量。这并非未来。这其实是现在的及格线。如果你不这么做,你的竞争对手就会这么做,而且他们会客气地吃掉你的午餐,然后还要对你说声谢谢。

大多数工程类演讲都忽略了一个延伸点:这不仅仅关乎工程师。在 YC,随着我们进行自身的转型,我们的媒体人员、活动人员、财务团队——那些一辈子都没有打开过终端的人——正在编写技能文件和 cron 计划任务。我们的一位财务同事刚刚用我们内部的 openclaw 和公司大脑,将大约一百个 Excel 工作簿整合成了一个单一的应用。她不是程序员,她现在是一个管理智能体的经理。现在 YC 的每一个人都是如此。

这就是为什么 YC 能够以这样精简的人员规模运转——在任何其他同类公司,我们这点人数看起来就像是四舍五入的误差。这并不是因为我们工作更努力,而是因为我们有一种不同类型的组织架构。这才是整个游戏的关键所在。它不仅仅关乎“400倍工程师”(400x engineers)。它是一个以比其他人高出400倍的水平去运营的一家公司。

Original English

Speaker: Now picture's batch room like it actually kind of looks like this. um 400 companies or 400 founders at long tables and uh you know I can imagine every single one of you each with a laptop every single day you're doing uh a former person's entire year worth of work in a single day. That's not the future. That's actually the bar right now. And if you're not doing it, your competitor is and they will eat your lunch politely and thank you for it.

Here's the extension uh most engineering talks miss. It's not just the engineers. um at YC uh as we make our transformation it's our media people our event staff our finance team people have never opened a terminal in their lives are building skill files and cron jobs and uh one of our finance folks just collapsed about a hundred Excel workbooks into a single app she built with uh our internal openclaw and company brain she's not a programmer she's a manager of agents now and everyone at YC now is

So that's why YC can run at the scale it does with a staff that would look like a rounding error at any other comparable firm. And that's not because we work harder because it's because we have a different type of org. And that's whole the whole game. It's not just 400x engineers. It's one company that operates at the level of 400x everyone else.

潜在空间与确定性空间

Speaker: 所以,如果你只能从这里记住一件事,我是说,这是我在这个过程中必须去探索发现的事情之一,那就是:你真的必须非常、非常小心地处理计算实际发生的地方。它几乎总是发生在两个不同的地方,而我们遇到的所有那些成为麻烦的 Bug、所有的 AI 工程问题,通常都是因为应该在等式某一边发生的事情,跑到了另一边去。

我说的第一个领域是“潜在空间”(latent space)。也就是实际的 LLM,你用它来做什么?用它来品鉴、做判断,用来理解当人类说一些模糊的话时,他们实际上想要表达什么。非确定性的调用是存在于模型中的计算,而你通过 Markdown 文件来引导它。

然后是“确定性空间”(deterministic space),这才是工程师们所熟悉的东西,比如你的代码智能体去编写 TypeScript,或者如果你用 Elixir,它们可能会去写 Erlang。是的,确定性空间是第二个地方。

Original English

Speaker: And so you know if you remember uh only one thing about this I mean this is one of the things I had to discover along the way like you actually have to be really really careful about where the computation is actually happening. Uh it's happening almost always in two different places and all of the bugs all of the AI engineering that we run into that's a problem it's usually because uh something is happening in one side of the equation that should be in the other.

Uh the first area I would say is latent space. So the actual LLM, it's you know what do you use it for? Taste, judgment, understanding what a human actually wants when they say something vague. The non-deterministic calls the computation that lives in the model and you steer it with the markdown file.

Uh and then deterministic space is what uh engineers know like your code agents go off and write TypeScript or you know maybe they're writing um you know Erlang if you're using Elixir. Um yeah the deterministic space is the second place.

Speaker: 举个例子,我们在即将到来的创业学校(startup school)中遇到了一个真实的问题。我们有 6000 人,我们要尝试去做的一个实验是,我们能否一次性为 800 人安排好座位,让它们完美地聚类,使得坐在你左边和右边的人,恰好是你在创业学校最应该认识的完美人选?

我们必须在结合了潜在空间的情况下,于确定性空间中完成这项工作。这种计算——这种实际存储了每个人具体位置的计算,就像是一个包含 800 个座位的多维数组——它实际上绝不能存在于上下文窗口(context window)中。LLM 必须负责属于人类判断的那部分,并为人们安排座位。

这实际上和你作为一个人类被指派去完成这项任务时会做的事情一模一样。你可能必须物理打印出 800 页纸,走进一个大房间,然后说“好吧,这个人该坐哪儿?”只是现在,这一切都可以在你的电脑里完成,而且它不用花一个月的时间,你可能只需要花费几百美元的 Token 和大约 10 分钟就能完成。所以,我认为这非常了不起,甚至在六个月前,这些都是你无法做到的事情。

Original English

Speaker: Um let's say you know this is a real problem that we have for startup school coming up. we have 6,000 people or uh we're going to try and you know one of the experiments we're going to try and do is can we seat 800 people at a time um perfectly clustered so the person sitting to the left and the right of you is the perfect person for you to meet at startup school um we have to do that in deterministic space combined with latent space uh the computation this computation this actual storage of like where everyone is inside like you know this multi-dimensional array of 800 seats um it actually must not live in the context window.

The LLM has to do the human part and seat people. Um it's actually exactly what you would do if um you know you were a human tasked with this thing. You would probably have to physically print out 800 pages and go in a big room and like say like well where does this person go? only now it can all happen um in your computer and it can instead of taking a month it might be able to you might be able to do it with you know couple hundred dollars worth of tokens and probably 10 minutes um and so that's you know I would argue that's pretty remarkable these are things that you couldn't do even uh I don't know six months ago um

工作记忆的极限与超越

Speaker: 这就引出了“工作记忆”(working memory)的话题。这也是我最喜欢用来理解它的一个角度。你和我作为人类,我们的大脑一次只能容纳大约 7 件事。7 加减 2。这是认知心理学中最著名的论文之一,它解释了为什么本地电话号码是 7 位数,以及为什么你会忘记购物清单上的第 8 项。通常,这就是一个人类全部的工作记忆。人类建立的每一个机构,每一份清单,每一张组织架构图,每一个文件柜,都是为了弥补这种局限性而发明的“假肢”。

想想这真的是一件很疯狂的事情:一个 AI 智能体可以容纳一百万个 Token。那大约是一千页纸。最近,我试图向我 10 岁的孩子解释 GBrain 是什么。我说,“你要知道,AI 智能体可以在脑海中同时翻开三本《哈利·波特》,并且它能在几秒钟内,从任何一本书中找到一根针,然后将三本书的内容综合起来。”那确实相当神奇。三本《哈利·波特》和 7 个数字的对比。我的意思是,那太棒了。我不知道那是不是 AGI。也许不是。但它已经是一个完全不同的运行机制了。地球上几乎所有的公司仍然在运行着一个为“7个数字大脑”设计的组织架构。但请注意这也告诉了你什么。三本书虽然很多,但...

Original English

Speaker: which brings me to working memory and uh that's you know sort of my favorite way to understand it is um you and I human beings, we only hold about seven things in our head at once. Uh 7 plus or minus two. It's one of the most famous papers in cognitive psychology and it's why uh local phone numbers are seven digits and why you forget the eighth item on your grocery list. Uh that's the entire working memory generally of a human being. And every institution humanity has ever built, every checklist, every org chart, every filing cabinet is a prosthetic for that limit.

It's kind of a wild thing to think about, but an AI agent holds a million tokens. That's about a thousand pages. I was trying to explain to my 10-year-old what GBrain was recently. I said, "There's, you know, the AI agent can keep about three Harry Potter books sitting open in its head all at once, and it can find a needle in any of them and synthesize across all three in seconds." And that's quite magical, actually. Three Harry Potter books versus seven digits. I mean, that's pretty awesome. I mean, I don't know. Is that AGI? Maybe not. But it's already a very different operating regime. Almost every company on the earth is still running an org that's designed for the sevendigit brain. But notice what that also tells you. Three books is a lot, but

上下文工程与公司大脑

Gary: 它也非常小。你的公司不仅仅是三本书。你的公司是一个完整的图书馆。这里包含了每一封电子邮件,每一次会议记录,每一个决策以及做出这些决策背后的复杂推理,每一次与客户的详细对话,还有每一次项目的事后复盘分析。决定你的 AI 智能体究竟是聪明绝顶的天才还是只有七秒记忆的金鱼,核心问题在于,到底是谁来决定在那张桌子上为你翻开哪三本书。这就叫做所谓的“上下文工程”(Context Engineering)。而这也正是“公司大脑”(Company Brain)的本质所在。它是这座图书馆,加上负责整理和挑选信息的图书管理员。

Original English

Gary: it's also very little. Your company is not three books. Your company is a library. every email, every meeting, every decision, its reasoning, every customer conversation, every post-mortem. The question that determines whether your agents are geniuses or goldfish is who decides which three books are open on that desk. That's context engineering. And this is what a company brain is. It's the library plus the librarian.

Gary: 现在,你们当中的一些人可能已经在心里想,“这不就是 RAG(检索增强生成)嘛。”你这么想是对的,检索确实是这一切的基础原语(primitive),就像 PostgreSQL 说到底本质上也只是建立在 B 树(B-trees)结构上一样。但这其中真正困难的部分,是围绕检索所发生的一切体系化工作。首先,到底哪些内容有资格在第一时间被记录并写入到公司的知识库(knowledge wiki)中。这些内容随后又是如何被持续丰富、标注并相互链接的。哪些信息会被提升到高速运转的“热内存”中随时调用,而哪些信息又会被归档为只在需要时才查阅的“冷参考资料”。当出现两个相互矛盾的事实时,又该由谁来进行仲裁和辨伪。检索本身这个动作是很容易的。但能够让你的资料库值得被检索,并且提供有价值的结果,这才是真正的产品。

Original English

Gary: Now, some of you are already thinking, "This is just rag." And you're right that retrieval is the primitive the same way Postgress is just B trees. The hard part is everything around it. What gets written down in the first place into the knowledge wiki. How it gets enriched and linked. What gets promoted to hot memory versus filed as cold reference. Who arbitrates when two facts disagree. Retrieval is easy. Being worth retrieving from is the product.

开源构建 Gbrain

Gary: 所以我一直坚持在开源环境下构建我自己的系统。它叫做 Gbrain。它几乎可以与任何运行框架和开发工具配合使用,但它最完美契合、结合得最好的是 OpenClaw Hermes agent。它实际上就等同于为智能体量身打造的 PostgreSQL,一个极其强大的检索层,它的核心任务就是要弄清楚,对于任何一项任务,究竟应该把哪“三本书”精准地加载到智能体的“大脑”里去。我个人的这个版本最初启动的时候,也就相当于一个装满了书的房间那么多资料。但现在,它已经进化成了一个拥有大约 220,000 页资料的庞大仓库,并且其中绝大部分内容都是由我的智能体根据我过去 20 年来的电子邮件、会议纪要、个人笔记,以及我最真实的个人生活经验自动编写和整理而成的。这就是它的核心意义所在。它是我的第二大脑。

Original English

Gary: So I've been building mine in the open. It's called Gbrain. It works with any harness but it loves openclaw Hermes agent. It's effectively Postgress for agents a retrieval layer whose job is to figure out for uh for any three you know for any task what three books should be loaded into the agents head. My personal one started as a rooms full of books or so. Now it's a warehouse about 220,000 pages written mostly by my agents from my email meetings 20 years of notes uh and the lived experience of me. And that's the point. It's my second brain.

Gary: 这带来的效果是,当一位创业公司的创始人发邮件向我求助一场突如其来的危机时,在我开始仔细阅读这封邮件之前,甚至在我还没来得及把那封简短的邮件读完之前,我的智能体就已经全自动地调取了我们之前与那位创始人的所有历史对话记录。它还会找出其他三家曾经遇到过同样瓶颈和低谷的投资组合公司,并总结出当时对那些人切实奏效、帮助他们走出困境的具体解决方案。当我的智能体在为我执行任何操作、生成任何建议时,它都是在完全了解我个人已经知道的所有信息的前提下进行的。这就是一个普通私人助理和一个能够与你并肩作战的专业同事之间的根本区别。

Original English

Gary: And when a founder emails me about a crisis, before I start reading this, before I even finish reading that email, my agent has already pulled every prior conversation with that founder. Three portfolio companies, they hit the same wall and what actually worked for those people. Uh, when my agent does anything, it does it does everything knowing what I already know. And that's the difference between an assistant and a colleague.

公司大脑的失效模式与卫生管理

Gary: 那么,让我来对我自己的这番说辞进行一次深刻的压力测试,因为我很清楚你们无论如何也是会在心里这么做的。公司大脑确实有其严重的失效模式(failure modes)。一个如果没有任何人去精心维护和策划的大脑,最终会变成一个拥有出色搜索功能的垃圾场。在那种情况下,检索系统会信誓旦旦、满怀自信地向你抛出一个实际上已经过时且完全错误的事实。一个编写得非常糟糕的技能文件(skill file),会把一个糟糕透顶的业务流程永远地固化和编码下来。这是非常可怕的。所以,作为基础原语,它绝对不仅仅是存储和记忆。它是记忆加上严格的卫生管理(hygiene)。你需要确保系统里的每一个事实都有据可查,拥有清晰的出处(Provenance)。当新的信息输入并与旧的信息发生碰撞冲突时,必须要有自动的矛盾冲突检查机制。此外,你还需要一个由人类和智能体协同工作的“图书管理员”,他们真正且唯一的工作就是对这些海量信息进行修剪、提炼和整理。如果你能像对待生产级别的基础设施那样去严肃对待这个大脑,它就会为你产生巨大的复利效应。反之,如果你仅仅把它当成一个随意丢弃数据和文档的垃圾场,你最终就会得到一个极其自信,但在错误方式上连人都无法追踪其根源的糟糕智能体。

Original English

Gary: So, let me stress test my own pitch because you would anyway. Company brains do have failure modes. Uh, a brain nobody curates becomes a garbage dump with great search. Retrieval will surface a stale fact with total confidence. Um, a bad skill file encodes a bad process forever. Uh, that's bad. So primitive, the primitive is not memory. It's memory plus hygiene. Provenence on every fact. Contradition. contradiction checks when new information collides with the old and a librarian human plus agent whose actual job is pruning. Treat the brain like a production infrastructure and it compounds. Treat it like a dumping ground and you get a very confident agent that is wrong in ways nobody can trace.

将工作成果“技能化” (Skillify)

Gary: 并且,我想在这里分享一项我认为非常核心的纪律,正是它让我们的公司大脑和我个人的 AI 能够产生如此惊人的复利增长。这是我的招牌动作(signature move),也是我一直以来对每一家 YC 公司,以及 YC 内部每一个人都在反复强调的话,那就是:永远不要做一次性的工作(never do one-off work)。你可以打开 Open Claw,你可以打开你的运行环境,你确实花时间做了一些出色的工作,但随后当你停下来,或者你去做了别的事情,那项工作的能力并没有被保留下来。这感觉就像是一份没有积累的糟糕工作。这就有点像你雇佣了一个表现并不怎么好的实习生,虽然他能把活干完,但下次还是得手把手教。但现在人工智能最棒的地方在于,你可以直接对它说,“嘿,我不喜欢刚才那个输出结果。修正它。”对吧?我敢肯定在座的各位肯定每天都在这么做。但是,千万不要仅仅止步于此。你实际上需要在完成那项具体任务的最后,将其“技能化”(skillify it)。为此,我特意在 X 上写过一篇非常详细的博客文章来探讨这个概念。你们可以直接搜索“skillify it”,找到那个具体的技能文件规范,然后直接把它加载到你们自己的运行框架中,它就能神奇地把你刚刚完成的所有繁琐操作,直接变成一项你可以无限次重复使用的自动化技能。因为在这个时代,如果你不得不对同一个问题提出两次相同的请求,那你在效率上就已经彻底失败了。所以,没错,如果在今天的演讲中你们只能记住一件事,那就是:当你在使用你的 AI 智能体,并且当你最终完成任务,对它的输出结果感到非常满意时,一定要把它“技能化”。这将会是一件非常酷、非常不可思议的事情。

Original English

Gary: And um here's the discipline that I think you know personally uh makes our company brain and my personal AI compound. Um, that's my signature move and uh, you know, it's what I say to every YC company and every uh, everyone inside YC, which is never do one-off work. You can open Open Claw, you can open your harness, you do some work, but then when you're hop, you know, and it'll come back. It's, you know, kind of a bad job. It's kind of like an intern that's not that good. But the great thing is you can just say, "Hey, I didn't like that. Fix it." Right? I'm sure all of you do this. But don't stop there. you actually need to at the end of that task uh skillify it. And so I have a blog post on X about that. You can search for skillify it and you know go get that skill file and then just load it into your your own harness and it'll just turn whatever you just did into a skill that you can reuse because if you have to ask for something twice, you failed. Um, so yeah, if you remember only one thing, it's that like when you, you know, use your AI agent and then when you're done with it and you're happy with the output, skillify it. It's going to be awesome.

Gary: 一个能够像这样敏锐捕捉并固化其所学知识的组织,每一天都会变得更加聪明和高效。而那些没有这么做的组织,无论他们所使用的底层 AI 模型有多么出色、参数有多大,他们每天早晨醒来都会面临彻底的失忆状态,不得不从头再来。请记住,模型的质量是你每个月花钱租来的(rented),但如果你能耐心构建出属于你们自己的大脑,那么你就真正、永久地拥有了那个大脑。所以,现在让我们正面回答 Theo 提出的那个核心问题。在这个时代,我们现在究竟应该去构建什么?答案是:去构建一家真正意义上的 AI 原生(AI native)公司,而不是一家仅仅在工作流程中使用了 AI 的传统公司。这必须是一家从成立的第一天起,其基本形态就完全符合我刚才所描述的那样运作的公司。它应该有一个极其精简的团队,为业务中的每一项日常任务都配备了完善的技能文件。即使是公司的创始人,也依然能够深度参与并在代码库中贡献力量。这个不断迭代的代码库,这个公司级别的大脑,它同样也会成为你强大的私人 AI。如果你愿意的话,当然可以使用 Gbrain 来作为你的起步。它是完全开源且免费的,但你并非必须要使用它不可。现在市面上已经有很多非常优秀的选择可供挑选。关键在于,这个图书馆一旦建立,从第一周起它就会开始产生惊人的复利效应,而你整个组织的运转速度、执行力和知识厚度,将会被彻底重塑,并提升到大约原本的 400 倍之多。

Original English

Gary: Um, the organization that captures what it learns like this gets smarter every single day. The one that doesn't wakes up every morning with amnesia, no matter how good the model is. Model quality is rented, but if you build your brain, you're you own that brain. So Theo's question head on. What do we build now? Build the AI native company, not a company that just uses AI. A company that is shaped like what I just described from day one. A thin team, skill files for everything. The founder still in the code library. This library, this company brain, your personal AI. Uh use Gbrain if you'd like. Uh it's open source and free, but you don't have to. There are a lot of really good ones. the the library will compound from the first week and your whole org will be wired to run at about 400x.

构建定义时代的记忆层

Gary: 如果你渴望在当今的科技浪潮中寻找一片完全未经开垦的新领域(green field),如果我现在 25 岁,坐在你们现在所坐的观众席上,我会毫不犹豫地去构建这样一个产品:因为地球上的每一家公司都即将迫切地需要一个大脑,需要一个能让他们永远不必再重新询问“我们曾经知道什么”的记忆层(memory layer)。他们需要一个真正深入懂你、理解你所有背景的私人 AI。我们目前正在完全公开地构建 GBrain,并且是基于极其宽松的 MIT 协议将其开源。我个人并不想试图通过这个底层项目本身来赚钱,因为我坚信,这个智能时代的记忆层本就应该是完全开放、共享的。就像 Linux 操作系统最终走向开放一样,这个层面本身——无论是公司的大脑,还是个人的上下文信息管理,亦或是那个负责在浩瀚数据中为你挑选“那三本书”的智能图书管理员,这完全是一片广阔无垠、充满无限可能的待开发领地。我由衷地希望在这个房间里,有某个人能够在这个领域建立起一家定义这个时代的标杆性伟大公司。如果你真的打算这么做并且取得了进展,我很乐意代表 YC 为你提供丰厚的投资。

Original English

Gary: And if you want the green field, the thing that I'd build if I were 25 and sitting where you're sitting, every company on this earth is about to need a brain, the memory layer that means that you never have to reask what you knew. Personal AI that actually knows you. We're building GBrain in the open and MIT open source. Um, I'm not trying to make money from this because I think the layer should be open. The way Linux is open, but the layer itself, company brains, personal context, the librarian that picks the three books, that's all wide open territory. I hope somebody builds the defining company here. And I'd like to fund you at YC if you do.

Gary: 现在,让我稍微退一步,极其坦诚地说一点可能有些削弱我刚才这番宣传说服力的话题。你其实并不一定需要使用我提供的那些工具才能开始这项革命。如果说 Open Claw 是这个领域里性能极致、一骑绝尘的法拉利跑车,我当然会一直强烈推荐它;但老实说,Codeex 也绝对是一辆非常平稳、可靠的本田轿车。它完全能够胜任并完成这里面至少 90% 的工作。它可能不会在速度和体验上让你感到惊艳无比、把你惊掉下巴,但它绝对能稳稳当当地带你到达你想去的目的地。所以,用什么工具都可以,真的。最核心的重点是理念(concepts),是你如何思考和组织信息,而不是我写的那些具体的代码仓库(repos)。你要深刻地思考计算(computation)在未来究竟是发生在哪里的。试着把各种技能文件当成你公司里不知疲倦的数字员工来使用,它们就是你的图书馆和图书管理员。永远、永远不要去做一次性的工作。这些通过技能化沉淀下来的能力,会像忠诚的伴侣一样,伴随着你无缝迁移到任何未来的技术栈(stack)中。

Original English

Gary: Now, let me be honest in a way that maybe undercuts my own pitch. You don't need my tools to start. Open Claw is the Ferrari. I will always recommend it, but Codeex is a really good Honda. It will do 90% of this. Uh, it will not blow your face off, but it will get you there. Use whatever. The concepts are the point, not my repos. You know, think about where the computation is. Use skill files as employees, the library, and the librarian. Never do one-off work. Those travel with you to any stack.

用充裕的软件解决恐惧

Gary: 那么,让我来为今天的演讲做一个最后的总结和落地。现在世界上有成千上万的人对 AI 浪潮下所有这些传统工作岗位的未来命运感到无比恐惧,我非常理解这种焦虑和恐惧的根源。但我想在这个场合最明确、最直接地说出来,这种恐惧,其实是对未来缺乏想象力的一种表现,是对技术潜力的一种误判。而坐在这个房间里的各位,正是解决这个宏大社会问题的终极答案。我刚才所详细描述的一切理念和工具,你们将会把它们带到你们各自正在筹备的创业公司去。你们将会通过这些技术实现令人难以置信的自我倍增,你们公司里的每一个人也都会以指数级的方式实现自我倍增,然后你们将去构建那些最终会成为全社会灯塔的伟大公司,去向世人证明并指引,所有这一切在未来的社会结构中到底是如何高效且有益运作的。所谓的物质和智力的“充裕”(Abundance)从来都不是写在纸上的一份政策文件。它是已经实实在在发布、被人们使用的软件产品。

Original English

Gary: So, let me land this. A lot of people in the world right now are terrified about what happens to all the jobs and I understand the fear. But I want to say it plainly that is a failure of imagination and the people in this room are the answer to it. What I just described you're going to take to your startup. You will multiply yourself and every person in your company will multiply themselves and you will go build the companies that become the beacon for how all of this works in society. Abundance is not a policy paper. It is shipped software.

Gary: 我有一位非常亲密的朋友,他那年幼的儿子不幸患有一种极其罕见且复杂的癫痫症。为了拯救他的孩子,他凭借一己之力建立了一个包含多达 80,000 个 markdown 文件的庞大代码仓库,这相当于一个专门为一个患病的小男孩量身定制的公司级别的大脑。他把自己硬生生地逼到了人类智力和极限的边缘,去疯狂地探索、记录和综合全人类在医学层面上对那孩子确切病情的全部已有认知。他没有任何顶级实验室的资源,没有任何外界的资金赞助,也不需要去向任何人申请许可。仅仅是一位绝望但坚强的父亲,一台普通的笔记本电脑,以及一个汇聚了无数知识的图书馆。这绝不是一个为了煽情而讲的旁支故事。这就是我在过去 20 分钟里一直试图向你们描述的那个精确架构的极致体现。一个庞大的图书馆,一位不知疲倦的图书管理员,在最危急、最正确的时刻,为你翻开了最正确的那三本书。并且,这个强大的系统被完全指向了这个男人在这个世界上最热爱、最想保护的事物。

Original English

Gary: I have a friend who has a rare form of epilepsy. He built a repo of 80,000 markdown files, a company brain for one small boy, and he pushed himself to the absolute edge of humanity, what humanity knows about his son's exact condition. No lab, no grant, no permission. A father, a laptop, and a library. That's not a side story. That is the exact architecture I've been describing for the last 20 minutes. A library, the librarian, the right three books open at the right moment. pointed at the thing this man loves the most in the world.

Gary: 朋友们,你们现在就可以做到这一切了。你过去所遇到的每一个让你感叹“唉,要是能有那位专家来帮我就好了,但我根本请不到他”的问题,现在,你全都可以自己解决了。每一个你曾认为因为历史包袱太重、代码太庞大且充满 bug 而永远无法修复的遗留代码库,现在你完全有能力修复它的每一个角落了。每一个你曾因为体积过于庞大而望而却步、无法深入阅读的档案库;每一个你曾因为数据太脏、关系太棘手而放弃清洗的庞大数据集;每一片你曾被前辈们告诫绝对不要试图去煮沸的浩瀚海洋。现在,在这个技术的加持下,我们真的可以把整片海洋都煮沸了。而且在座的你们每一个人,都将能够展翅高飞,我在这里说的不是什么虚无缥缈的比喻,而是通过机械化、系统化的 AI 工具真正实现的飞跃。你必须这样做,这是你为了在未来的竞争中生存下来、为了走向繁荣、为了赢得最终胜利所必须走的路。

Original English

Gary: You can do that now. Every problem where you thought, I wish I had that person, but I can't get them. You can. Every codebase you thought was too buggy to fix, you can fix all of it. Every archive too big to read. Every data set too gnarly to clean. Every ocean you were told not to boil. We can boil the ocean now. And every single one of you can fly, not metaphorically, mechanically. And you need to to survive, to thrive, to win.

Gary: 所以,让我们回到最初,Theo 刚才问我们现在究竟应该构建什么。这就是最完整、最透彻的全部答案。去毫不犹豫地构建那种 AI 原生(AI native)的公司,然后把它做大做强,推向极致。去构建支撑它的底层基础设施。去构建那个大脑,那个永不遗忘的记忆系统,那个能够不断自我进化、产生复利的庞大图书馆,正是它,将使得在你们之后成立的每一家公司的构建都变得前所未有的容易。去煮沸那片看似不可逾越的海洋吧。去亲手写下那个关键的测试用例。去正式发布你的那一项技能文件。你们即将在接下来的战场上观察到的一些竞争公司,其实已经开始在默默做这些事情了。所以,去成为将这一切做得最好的那一家公司。谢谢大家的倾听。接下来,请大家和我一起以最热烈的掌声欢迎 Hyper Agent 的首席执行官,Howy Lou 登场。

Original English

Gary: Theo asked what we should build now. And here's the whole whole answer. Build that AI native company and build it. Build the thing underneath it. the brain, the memory, the compounding library that makes every company after yours easier to build. Go boil the ocean. Go write that test. Go ship that skill. Some of the companies you're about to watch in the battlefield are already doing this. Go build the one that does it best. Thank you. Please join me in welcoming the chief executive officer at Hyper Agent, Howy Lou.

Hyper Agent 的产品视角

Howy Lou: 大家好。嗯,有谁对接下来的这部分演讲感到激动?这几天真的是太棒了。极具启发性的演讲,刚才 Gary 的演讲也非常精彩,干货满满。顺便说一句,我还从没在这么强烈的低音环绕下被别人介绍出场过。此时此刻我真的非常非常激动,充满干劲。嗯,关于我自己,我内心深处始终认为自己是一个纯粹的产品思考者。我的意思是,尽管我同样非常喜欢把时间花在构建产品的技术栈那一面上,而且我也很喜欢去深入思考系统架构、复杂的算法等等硬核问题。但说到底,在骨子里我是一个痴迷于产品形态(product form factor)的思考者。大家知道,

Original English

Howy Lou: All right. Hello. Um, who's excited about the end of this event? This this has been an amazing few days. Amazing talks, amazing uh Gary talk right there. And by the way, I've never been introduced with so much base. I am very very pumped in this moment. Um so I consider myself a product thinker at heart. Meaning you know as much as I like to spend time on the technical side of building and you know I I love to think about architecture and algorithms and so on. You know really I'm a product form factor thinker at heart. And you know

代理与产品形态的演进

Speaker: 今天我想以此作为结束,简要且高屋建瓴地概述一下我所看到的随着代理(agents)的发展,世界将如何与产品形态交织在一起。所以我把这个称之为“雇佣具有就业能力的代理”(hiring employable agents),这实际上反映了随着这些模型变得越来越好,我们从一些非常非常棒的演讲者那里听到了关于最新前沿模型(无论是 GLM 5.2 还是其他模型)的见解,以及如何从中激发出最高的性能。说到底,这并不是一次技术上的深度剖析,而是探讨我们如何从产品构建的角度来看待这些问题。

Original English

Speaker: What I want to end today off with is a brief and kind of high-level overview of how I see the world going with agents as it relates to product form factors. So I'm calling this hiring employable agents, and it's really a reflection on, as these models have gotten better and better, and we've heard from some really, really great speakers about, you know, the latest frontier models, whether GLM 5.2 or others, and how to coax the most performance out of them. And really, this is not about the technical deep dive, but how do we think about these from a product building standpoint.

Speaker: 那么我是谁呢?我是 Airtable 的创始人兼首席执行官,大家可能也知道 Hyper agent,这是 Airtable 内部的一款产品。正如你们中许多人已经知道的那样,Airtable 基本上是一个几乎可以说是水平到令人发笑的平台。事实上,大约在 12 年前我们创立这家公司的时候,我们早期从几乎每一位交谈过的投资者那里得到的基本上都是负面反馈,也就是:“你们不可能以水平平台的方法推向市场,并试图构建一个对每个人来说都无所不包的产品,比如人们想要构建的任何类型的应用程序,你们不可能解决所有这些用例。”我觉得我们最终在解决其中许多用例方面做得相当不错,至少为庞大的构建者社区提供了解决方案,而我们现在基本上正在将非常非常相似的哲学应用于代理。接下来我将详细拆解我们是如何做到这一点的。

Original English

Speaker: So who am I? I'm the founder and CEO of Airtable, and also know Hyper agent, which is a product within Airtable. And as many of you already know, Airtable is basically an almost laughably horizontal platform. When we in fact set out to build the company almost 12 years ago, we got basically this negative feedback from virtually every investor we talked to early on, which was, you know, "You can't possibly go out with a horizontal platform approach and try to build something that is everything for everyone, like any type of app that people want to build, you know, you can't solve all of those use cases." And you know, I think we ended up doing a pretty good job of solving for many of them at least, and solving for a large builder community, and we're basically applying a very, very similar philosophy to agents now. And I'm going to break down how we're doing it.

Speaker: 所以,当我跳出局部,思考模型的演进以及它们所赋能的产品形态时,我真的把它看作是一个连续的谱系,对吧?所以,我们从最初的基础补全(completions)开始,也就是说,你给一个模型——比如这些大型语言模型(LLMs)之一——一串 token,它就会继续生成更多的 token,直到遇到停止 token,它或多或少就是这样来补全你的想法或模式的。直到 InstructGPT 出现并促成了聊天(chat)这种产品形态。在那之前,它虽然可以说是强大的,但适用范围相对狭窄。一旦我们达到了更高水平的模型智能,再加上指令微调(instruct)这种创新,通过调整这些模型让它们能够真正与用户进行对话式的一问一答,我们就实现了聊天机器人。当然,我们都知道 ChatGPT 是一款非常出色的爆款产品,我们现在对这种 AI 形态已经非常非常熟悉了。

Original English

Speaker: So, when I zoom out and think about the progression of models and therefore the product form factors they enable, I really see it as this spectrum, right? So, we went from basically completions, which is, you know, you give a model, one of these LLMs, a string of tokens and it just continues emitting more tokens until it gets to a stop token and more or less it would just kind of complete your thought, right? Or complete the pattern. And you know, until InstructGPT happened and enabled a chat form factor, it was arguably powerful but more narrowly applicable. Once we got to a higher level of model intelligence, and then also the innovation of instruct and kind of tuning these models to actually have more of a conversational back and forth behavior with users, then we enabled the chatbot, right? And of course, we all know ChatGPT was an awesome breakout product. We're now all very, very intimately familiar with that form factor of AI.

Speaker: 下一个层级我称之为代理(agent),对吧?我想“代理”这个词已经被过度使用了。在过去几天里,我们可能已经大量讨论过这个话题,但我对代理的定义实际上与 Anthropic 的定义是一样的,它有别于工作流(workflows)。代理实际上是一个自我递归的模型,它拥有一组开放式的工具可以调用,可以采取推理步骤,也能做出决策。因此,它并不是预先设定好线性工作流的,而是能够做出一系列开放式的选择。不过,我还是倾向于认为这与下一个层级的代理有些许不同,在这里我使用了“Claw”这个术语,因为 OpenClaw 显然普及了这个概念。

Original English

Speaker: The next level I call an agent, right? And you know, I think agent has been very overloaded. We've probably talked about that a ton over the past few days, but my definition of agent really is the same as the Anthropic definition, which is distinct from workflows. And an agent is really a model that recurses upon itself, right? And has the open-ended set of tools that it can call, of reasoning steps it can make, of decisions it can make. So, it's not prescripted to a linear workflow, but it can kind of make an open-ended set of choices. Still, you know, I like to think of agents as being a little bit distinct from the next level of agents, which I'm using the term claw here because OpenClaw obviously has popularized this concept.

始终在线的代理与多代理编排

Speaker: 但我看到了一个市场转变,过去的代理主要是对用户输入做出被动反应(reactive)。比如,你让一个代理去执行一项任务,它可能会花 20 分钟甚至一个小时来完成任务,然后就结束了。而我所说的“Claw”,真的是指那些可以运行更长时间工作的代理,它们甚至可以拥有类似 OpenClaw 中的心跳机制——现在在 Hyper agent 等其他产品中也有——这使得它们能够自行唤醒,几乎表现出一种“始终在线”(always on)的行为。我的意思是,在底层发生的事情是,每一次唤醒基本上都会导致多个不同的对话轮次。因此,代理在自身内部循环,执行工具调用,每个工具调用的响应随后再次调用 LLM,直到它达到一个停止点,但随后这个停止点会被心跳机制再次唤醒或重启,你可以一次又一次地唤醒代理。所以,这种始终在线的代理实际上在不断地朝着目标前进,并创造性地找出如何为自己解锁的方法,而早期的代理在这些地方可能会卡住。我认为这是我们正在看到的非常非常重要的下一个产品形态的飞跃。

Original English

Speaker: But, you know, there's a market shift that I see from agents that were still mostly reactive to user input. So, you would have an agent go and perform a task. It might spend 20 minutes maybe up to an hour to perform the task and then it completes. Claws, as I call them, are really about agents that go and perform work for much longer, right? They can even have a mechanism like the heartbeat mechanism within OpenClaw, and now in other products like Hyper agent, that enable them to wake up on their own and almost exhibit this always on behavior. I mean, underneath the hood what's happening is each wake up basically results in many different turns. So the agent loops upon itself. It's performing tool calls. Each tool call response then invokes the LLM again until it reaches a stopping point, but then that stopping point gets reinvoked or kind of restarted with a heartbeat mechanism where you can wake up the agent again and again, right? And so this idea of this always on agent that's actually constantly moving towards its goal and kind of figure out creatively how to unblock itself where maybe prior agents would get stuck. I think that's a really, really important kind of next form factor leap that we're seeing.

Speaker: 最后,当你想到现在的模型已经能够非常智能地编排复杂的任务时。比如,我个人经常在 Cloud Code 中使用工作流功能,它可以将大量工作分发给大量并行运行的代理组。但是这种概念——即一个代理可以去和其他代理进行编排,代理与代理之间进行交互,以执行更高级的任务、运行更长的作业——我认为是下一次飞跃。所有这些之所以成为可能,是因为模型变得越来越聪明,聪明到即使在更长、更开放范围的问题上也能保持连贯性。

Original English

Speaker: And then finally, as you think about models that are now capable of, you know, really, really intelligent orchestration of complex tasks, right? So I've personally been using the workflows capability in Cloud Code quite a bit, where it can fan out a lot of work to a massively parallel set of agents. But this concept of this agent that can go and orchestrate with other agents, right? An agent-to-agent interaction to perform even more advanced tasks, right, longer running jobs is I think the next leap. Yet all of this is enabled because the models are getting smarter and smarter to the point where they can achieve coherence even on longer and more open-ended scoped problems, right?

Speaker: 回想一下 BabyAGI 和 AutoGPT 的时代,那些真的只是有趣、很酷的科学实验,但最终,尽管看着它们为解决一个问题工作好几个小时确实非常有趣,但它们不可避免地都会偏离轨道,去做一些不太有用的事情。它们会陷入死胡同然后卡住。而我认为现在的模型实际上已经足够聪明,不仅能够实现长时间运行的任务,而且能够让我们开始应用这种组织原则,将问题分解为不同范围的任务,然后不同的代理可以互相交接。这一切的核心都在于管理每个代理的任务或上下文窗口,以及它们相互协调工作的能力。

Original English

Speaker: If you recall back to the days of BabyAGI and AutoGPT, those were really kind of fun cool science experiments, but ultimately, as much as they were really interesting to watch and see them work for hours and hours on a problem, they would all inevitably kind of end up drifting off to something not super useful, right? It would go off into a rabbit hole and kind of get stuck. And I think now the models are actually smart enough to enable not only longer running jobs but actually enable us to start applying these kind of organizational principles to break down problems into differently scoped tasks that different agents can hand off to each other. All of which is about managing each agent's task or context window and ability to coordinate work with each other.

从抽象到现实:传统行业的代理应用

Speaker: 刚才说的所有这些可能有点抽象。我想把它变得非常具体,并向你们展示我们是如何应用构建水平产品的一些设计思维,并在 Hyper agent 中提炼出代理的许多复杂性的——就像我们在 Airtable 中为构建数据库或由数据库驱动的应用程序所做的那样。所以这是一个非常现实的用例,基于我们一些实际客户的使用情况。我们在推出 Hyper agent 时发现的一件有趣的事情是,Hyper agent 的典型用户并不是你所认为的那些人。我的意思是,我们确实有很多非常前卫的、纯软件公司客户,比如 AI 公司本身,但我们也吸引了其他非常有趣的、有时是传统的线下业务公司。

Original English

Speaker: So all this is kind of abstract. I want to make it very, very tangible and show you how we've applied some of that design thinking of building a horizontal product and distilling a lot of the complexity of agents like we did for building a database or a database to powered app with Airtable for agents in Hyper agent. So this is a very realistic use case based on some of our actual customer usage. One of the interesting things we found from launching Hyper agent is the prototypical users for Hyper agent are not who you would think they are, right. I mean, we definitely get our fair share of like the super AI forward, kind of pure software companies, like AI companies themselves, but we also get really interesting other companies that are, you know, sometimes like traditionally offline businesses.

Speaker: 在这个案例中,我们描绘了一个非常真实的场景:一家从事园林绿化业务的客户,实际上使用代理来几乎端到端地运营其业务的各个环节——从寻找可以推销园林绿化项目的客户,到实地去重新设计他们的花园或后院,再到端到端地管理从潜在客户开发、构建报价和推销方案,一直到管理工作本身的方方面面。我认为这是一个巨大的、非常有趣的洞察,这是我在做 Hyper agent 时偶然发现的,那就是代理构建者的聪明才智并不仅限于 AI 公司或软件公司。

Original English

Speaker: And in this case, we're depicting a very realistic scenario, which is a customer that is in a landscaping business and actually uses agents to run virtually end to end every part of their business, which is finding clients to propose landscaping projects to, like physically going and redoing their garden or their backyard, and end to end managing everything from the prospecting to building a quote and a pitch to managing the work itself. And I think this is a huge, huge, interesting insight that I've just kind of stumbled on from working on Hyper agent, which is ingenuity of agent builders is not contained to just AI companies or software companies.

Speaker: 刚才 Gary 的演讲给了我很大的启发。我认为其背后的动机适用于每一个行业、每一个部门,有许多非常非常有创造力的人——其中很多就在这个房间里——可以走出去颠覆你们自己的行业,无论它是什么。你不必非得是一家纯软件企业。但我认为,你所需要的只是一种想象力,想象当你真正实现不仅仅是成为一家 AI 原生公司,而是我喜欢用的一个词——一家“完全由代理舰队原生驱动”(fully fleet of agents native)的公司时,工作会变成什么样子。在这里,你实际上是在超越极限,不仅仅是在思考如何用 AI 自动化一些小而狭窄的步骤,而是思考如何雇佣一整支拥有受控行为的代理舰队,并拥有正确的能力。

Original English

Speaker: And I was really inspired by Gary's talk just a second ago. I think the motive behind that applies to every industry, every sector, and there are really, really creative people, many of whom are in this room, who can go out and disrupt your own industry whatever it be. You know, it doesn't have to be a pure software business. But I think all it takes is the ability to imagine what work could look like when you do go and become fully not just an AI native company, but I like to use the term, a fully fleet of agents native company where you're actually going above and beyond and not just thinking about small kind of narrow steps to automate with AI, but actually how to employ a whole fleet of agents that have governed behavior and have the right

智能体在园林绿化业务中的实际案例

Speaker: ……以及合适的工具,从而能够完成真正有用且端到端的自主工作。那么,让我们来看看这具体是怎样的。嗯,这是一个非常现实的例子,你知道,在这个业务中,这位,呃,这位园林绿化师想要去向客户推销,对吧?所以,你知道,他们会收到所有这些感兴趣的客户发来的入站请求。有人提交了这个询盘说:“嘿,这是我后院现在的样子。你知道,有点乱。这是我的信息。你能给我个报价并向我展示一个你可以做什么的方案吗?”这基本上会被路由到一个,呃,进入到一个负责分流的智能体(triager agent),对吧?所以这个智能体基本上是在接收这个初始需求并说,“嘿,项目的范围是什么,你知道,比如,这是我们想要接的项目吗?这是一个高质量的线索吗?这是一个合法靠谱的人吗?”然后它会做一些研究,但最终它会将工作移交给另一个智能体,一个测量智能体(surveyor agent),让它实际去查看这个人后院的实际物理景观,并拿出一个关于重新进行园林绿化后可能会是什么样子的方案。

Original English

and the right tools to be able to do really useful and very end-to-end autonomous work. So, let's walk through what this looks like. Um, this is a realistic example of, you know, in this business, this uh this landscaper wants to go and pitch clients, right? So, you know, they get all these inbound uh submissions of clients who are interested. Somebody submits this inquiry saying, "Hey, here's my current backyard. You know, it's a little messy. Here's my info. Like, can you give me a quote and show me a proposal of what you could do?" This basically gets routed into a um in into one agent that is the triager agent, right? And so this agent is basically taking in this intake and saying, "Hey, like what is the um the scope of the project, you know, like is this something we want to take on and is this a quality lead? Is this a legit person?" And so it does some research, but ultimately it hands off then the work to another agent, a surveyor agent, to actually go and then look at the um the physical landscape of this person's backyard and come up with a proposal of what a rellandscape version of that could look like.

Speaker: 所以在这里,嗯,这个测量智能体会去查看提交者提供的原始图像甚至视频。嗯,在我们的案例中,hyper agent 在一个功能完备的沙箱虚拟机(sandbox VM)中工作。所以每个智能体,事实上每个带有智能体的线程都能够完全编写代码。它可以操作文件。你实际上可以使用像 ffmpeg 这样的东西来,你知道,从视频文件中截取单独的屏幕截图,然后分析这些截图,嗯,并实际梳理并找到所有相关信息,以创建一个质量相当高、嗯,且非常个性化、高接触(hightouch)的提案,这对于像这样的小企业来说,在以前是难以想象的,对吧?如果你想想那些真正高端的室内设计师或者可能是非常高端的园林绿化师,也许他们能够为像一个 500 万美元以上的大客户做这样一个非常定制化的方案,你知道,如果他们是为一家大型酒店或其他机构工作的话。但是,为一个小后院的园林绿化提案做这么深度的高接触服务,实际制作出真实的模拟图像和一份真正的客户推销材料,这在以前是闻所未闻的。但这个人做到了,因为他们想出了如何以极其强大和自主的方式使用智能体,来完成原本人类无法完成的工作。

Original English

So here um the surveyor agent is going and looking at the original imagery uh or even video that the um submitter gave them. Um hyper agent uh in our case works with a fully um capable sandbox VM. So every agent in fact every thread with an agent is able to fully write code. It can manipulate files. You can actually use things like ffmpeg to like, you know, clip out individual screenshots from a video file and then analyze those um and actually go through and like find all the relevant information to create a fairly high quality um and hightouch uh proposal that would have been unfathomable for a small business like this before, right? If you think about like a really high-end interior designer or maybe like a very high-end landscaper, maybe they could do a very bespoke proposal like this for like a $5 million plus uh client, you know, if they're working for a big hotel or something. But it would have been unheard of to do something this hightouch and actually mock up real imagery and a real client pitch deck for, you know, just a small backyard landscaping proposal. But this person has done it because they've they figured out how to use agents in an extremely capable and autonomous way to do the job that otherwise would have been undoable by humans.

人类在智能体工作流中的新角色:疏通与决策

Speaker: 那么接下来发生的事情是,你知道,测量智能体去创建了这个提案和推销材料,现在这种介入必须要回到人类手中,比如,依然有一个步骤需要人类来为智能体扫除障碍(unblock)。Sarah Guo 几周前发表了一篇很棒的文章,关于,你知道,就描述了在这个 AI 领域随着事物的发展和模型变得越来越聪明,价值将聚集在哪里。我完全同意她的观点,那就是,你知道,能够部署有用智能体的终极障碍仍然存在着人类因素,你知道,尽管模型越来越聪明,但智能体被人类疏通的能力——那些仍然掌握着政策和决定,并且非常重要的是,可以把关,你知道,真正关键的、高影响力的行动——比如在这个例子中,实际向拥有一份具有约束力的报价的客户进行推销,嗯,或者比如同意接下这个工作,这就变成了人类然后要介入的事情,就像一个优秀的经理从他们的团队那里获得向上的、嗯,汇报或工作更新,然后再去为团队扫清障碍一样。

Original English

And so what happens next is you know the surveyor agent it it went and created this proposal and pitch and now the intervention has to go back to the human like there is a step where humans still need to unblock the agents right Sarah Guo put out a great post uh a few weeks ago about you know just describing like where the value is going to acrue in this AI landscape as things evolve and models get smarter and I firmly agree with her take which was you know the ultimate blocker to still be able to deploy useful agents there's still a human element to that you know As much as models are getting smarter, the ability for agents to get unblocked by humans that still own policies and decisions and really importantly just kind of gate, you know, really critical um high impact actions like in this case actually making a pitch to a customer that has a binding quote um or perhaps like agreeing to do the work that becomes something that the human then intervenes at just like a great manager who's getting upwards uh reports or you know upwards work updates from their team to then unblock.

Speaker: 所以在这里,我们有 Sage,他回到了,呃,这个企业真正的人类老板那里并说,嘿,这里有一个线索,我已经对这个项目进行了测量,这里有一份漂亮的推销演示文稿和一份方案,并且基于我已经对这个特定项目进行测量的所有因素,我得出了一个现实的报价,把它交给你,由你现在来批准将其发送给客户,看看他们是否接受。对吧?你可以看到,比如,即使是方案本身,也是这种非常漂亮的交互式网页。这对于这样的企业来说,能够向每一个客户发送这种方案,即使只是针对一个小任务、呃,或者是像 10,000 美元的园林绿化提案这样的工作,再次强调,这在以前是不可想象的,对吧?嗯,你知道,更好的是,当你与这些智能体互动时,你知道,就像 Gary 谈到的 Gbrain,以及创造这种能够从每次互动中学习的非常有粘性的上下文层(sticky context layer)。我完全同意。我认为这在某种程度上是一个普遍原则,也是未来有用智能体的一种表现形式,即智能体不应该是静态的。它们不应该只是,你知道,要么是静态的 LLM 调用,或者是甚至只带有一个提示词(prompt)和一些静态技能的静态智能体。它们真的需要通过每一次互动来学习,对吧?你不能去等待你的自有模型的下一次,你知道,微调(fine-tuning)运行——也就是你收集了一堆数据或评估标准(rubrics)去重新训练模型的时候。它应该实时学习,对吧?它应该积累记忆和技能更新,并直接接受用户反馈,真正从现实世界的表现中学习,从而像人类一样不断调整并变得更好。

Original English

So here we have Sage who has come back to uh the real human uh owner of this business and saying hey here's a lead I already surveyed the the project here's a beautiful pitch deck and a proposal and have come up with a realistic quote based on all the factors I've already surveyed in this uh particular project to give to uh to you to now approve sending to the client and see if they accept it right and you can see like even the u the proposal itself is this really beautiful kind of interactive uh web page That again would have been unfathomable for a business like this to be able to send to every single client even for small tasks uh or jobs like a $10,000 landscaping proposal, right? Um you know, better yet, as you interact with these agents, you know, as as Gary talked about with Gbrain and and just kind of like creating this very um sticky context layer that learns with every interaction. I completely agree. I think that is a you know kind of a universal principle and kind of a form factor for useful agents going forward is agents should not be static. They should not just be you know either static LLM calls or static even agents that just have a prompt and you know some static skills. They really need to learn with every interaction right and you can't wait for the next you know kind of fine-tuning run of your own models where you have you know a bunch of collected data or rubrics to to retrain the model. It should learn in real time, right? It should accumulate the memories and skill updates and just take user feedback and actually learn from real world performance to constantly adjust and get better just like humans do.

Speaker: 所以在这个例子中你可以看到,你知道,在 Slack 里,或者你知道,在我们的例子中,无论是在 Slack、电子邮件还是其他调用方式里,你基本上可以像跟一个真人一样与智能体来回交流,并说,嘿,比如,你知道,记得下次你要把这个考虑到,你有点,你知道,这里应该有不同的格式设置,或者是思考这个提案的一种不同方式,或者你漏掉了、呃,在这个特定项目中的这个额外成本杠杆。所以,你知道,下次记住这一点。我认为至关重要的是,在下一代智能体中,智能体需要记住这些东西,对吧?嗯,它们应该感觉是流动的、并且是,嗯,不断演进的,就像人类一样,只有这样它们才能充分实现最新前沿模型能力赋予它们的潜力,在自主性方面去实际付诸行动。嗯,最后,我认为在所有这一切的最后,随着所有这些智能体各自独立地具备能力甚至彼此交互,发生的结果是:我认为人类的工作越来越多地变成去为智能体扫清障碍,从而使它们能够高效地开展工作。

Original English

And so in this case you can see you know within Slack or you know in our case within uh either Slack or email or other um invocation methods you can basically go back and forth with the agent just like a real human and say hey like you know remember to take this into account next time you kind of you know here's something differently uh formatted or a different way to think about the proposal or you missed kind of um this additional cost lever in this particular project. So, you know, remember that for next time. And I think it's critical that in this next generation of agents, agents need to remember these things, right? Um they should feel fluid and and um evolving just like humans are in order for them to fully realize the potential that the latest frontier model capabilities enable them to to actually go and do in terms of autonomy. Um, and then finally, I think what ends up happening at the end of all of this is with all of these agents that are each individually capable and even interacting with each other, I think the job of the human becomes increasingly about unblocking agents to be able to do work effectively.

从单线程开发到智能体舰队编排

Speaker: 所以,你知道,我把这与开发工作流进行了比较,开发工作流是从,你知道,最初在任何,嗯,你知道,AI 辅助编码出现之前演变而来的。显然,我们那时只是坐在电脑前。作为人类,我们是相当单线程的,对吧?你会盯着一个代码文件看,你会思考一个问题,你会,呃,敲出一些代码,也许休息一下,然后再回到代码上;然后,随着第一代 GitHub Copilot 和类似工具的出现,我们有了增强编码,我认为那实际上只是编码上的一种代码补全形式。你输入一些代码,它可以自动补全额外的几行;接着,我们有了像 Cursor Composer 1.0 这样的东西,那是最初的,呃,体验,不是那个 Composer 模型,嗯,而是聊天体验,你可以和这个智能体交谈,它会进行更复杂的更改,这带来了更大的杠杆效能,对吧?但现在我们正迈向这样一个阶段,你知道,我认识的最优秀的、使用智能体的前沿开发者们——你们中的许多人就坐在这个房间里——实际上只是在监督一支智能体舰队,对吧?比如事实上,如果晚上睡觉前你没有让你的智能体去执行有用的隔夜工作,就会感觉你吃了个大亏,感觉好像你的团队没有在工作,因为你没有为它们扫清障碍。

Original English

And so, you know, I I compare this to the development workflows that uh have evolved from, you know, originally before any um, you know, kind of AI enabled coding, obviously, we just sat in front of a computer. We as humans are pretty singlethreaded, right? You'd stare at one code file, you'd think about a problem, you'd kind of bang out some code, maybe take a break, come back to it to then we had um augmented coding with the first generation of GitHub copilot and similar tools where I mean I think of that as really just a completions form factor on coding. You type some code, it can autocomplete a few extra lines to then we had things like cursor composer 1.0 0 like the original uh experience not the the composer model um but the chat experience where you could talk to this agent it would do more complex changes and that was even more leverage right but we're now moving to this point where you know the best frontier agentic developers I know and many of you are sitting in this room are are really just kind of overseeing a fleet of agents right like in fact going to sleep without setting off your agents to perform useful work overnight feels like you're you're taking this huge loss like your your team is not working because you're not unblocking

Speaker: 因此,你如何监督这些智能体的用户体验(UX),无论是在开发工作流中还是在像这个园林绿化例子这样的通用目的工作中,也需要发生转变。所以,我们对此的思考方式是,你需要看到这种像编排控制平面(orchestration control plane)一样的东西,在这里你可以一目了然地看到你所有的智能体都在处理什么,它们在哪些地方受阻了,谁正在把工作移交给谁;然后你的工作几乎变成了,你知道,向外缩小视角来俯瞰整个类似《模拟城市》(Sim City)的图景,在宏观层面上对所有发生的事情在你不同的智能体之间进行编排,而不是必须得去微操,把视角放大聚焦到每一个单独的项目或每一项单独的任务中。嗯,所以我认为这……

Original English

And so the UX of how you oversee these agents, whether it's in development workflows or general purpose work like this landscaping example needs to shift as well. And so the way we've thought about that is you need to see this kind of orchestration control plane where you can see at a glance what all of your agents are working on, what they're blocked on, who's handing off work to each other, and your job becomes almost like, you know, zooming out to see the entire Sim City landscape and orchestrating across everything going on at a macro level between your different agents rather than going and having to play micro and zoom into each individual project or each individual task. Um and so I think this

迈向智能体编排的新时代

Speaker A: 这是一个非常令人激动的时刻,因为我们已经不仅从自动补全和对现有的人类工作进行轻微增强中转移出来,而是真正实现了跨越式的飞跃——人类将成为智能体的编排者,而这些智能体也能自我编排。随着这些智能体变得越来越强大,找到与它们互动的正确方式,不仅是创造巨大成就的机遇,在我看来,更是各行各业的每家公司赖以生存的筹码。这将成为你在行业中领先并蓬勃发展的方式。如果不这样做,正如黄仁勋(Jensen)所说,取代你工作的将不是 AI,而是使用 AI 的人抢走你的工作或你的业务。因此,我真的非常激动,不仅是因为这个机遇,也是因为在座的各位能够参与到这场变革中,参与到向 AI 下一种形态迈出的这一大胆飞跃。说了这么多,我们非常希望大家能尝试一下 Hyper Agent。访问 hyperaggent.com/ie,你可以获得 1000 美元的推理额度。这些额度可以用于我们所有的不同模型服务中,包括最新的 Opus 4.8、Fable 5,你也可以在 GLM 5.2 或 GBD 5.5 等模型上使用它。我们非常期待看到你在 Hyper Agent 或其他平台上用智能体做出的成果。感谢大家的到来。你最近怎么样?

Original English

Speaker A: is a really really exciting moment as we have shifted from you know not just completions and kind of the slight augmentation of existing human uh work but really a radical now leap into humans as orchestrators of agents that also can orchestrate themselves. And finding the right way for us to interact with those agents as they become more and more capable is going to become not just the opportunity to boil the ocean, but in my opinion, table stakes for survival for every company in every industry. This is going to be the way to get ahead and thrive in your industry. And if you don't, you know, as Jensen put it, you know, it's not going to be AI taking your job. It's going to be, you know, somebody using AI who takes your job or takes your business. So really really excited both for the opportunity and for everyone in this room to be part of this um transformation and this really kind of bold leap into the next form factor of AI. Um all that being said, we'd love to see you try Hyper Agent. So hyperaggent.com/ie and you get $1,000 of inference credit. Um you know that can be spent on all of our different model offerings including uh you know the latest uh Opus 4.8, Fable 5, you can use it on GLM 5.2 uh on GBD 5.5 etc. and we'd love to see what you do with uh with agents on Hyper Agent or elsewhere. Thank you all for being here. How are you again?

Speaker B: 感谢你的参与。

Original English

Speaker B: Thank you for joining us.

Speaker C: 谢谢。在我们活动的尾声,几个月前,我看到 Hyper Agent 正在发起这场初创公司竞赛。在 AIE,我们一直想推出某种类似于竞技场的比赛,为人们提供一个角逐的平台。但我们之前真的没有任何机制去评判它,或者为它提供资金之类的支持。然后 Hyper Agent 出现了,他们带来了这个“500 Founders”(500 创始人)项目。

Original English

Speaker C: Thank you. Um so for the close of our events uh a few months ago I saw that hyper agent was launching this startup competition and at AIE we've always wanted to feature some kind of battlefield for like a competition and a contest for people uh and we didn't really have any sort of mechanism by which to judge or to fund it or anything like that and hyper agent came along and they had this uh 500 founders y

Speaker B: 对的,那么这个项目的灵感是什么?

Original English

Speaker B: right what was the inspiration for that

Speaker A: 你知道,我认为,当我思考这种多层次的颠覆叠加时,正如我所说,我们正在经历一种转变。以前,人类做着相同的工作,只是被 AI 稍微增强了一点。而现在,人类在监督由智能体组成的舰队去完成以前不可能完成的工作,比如像那位园林设计师,以一种根本不同的方式经营他的业务。再到现在,一个由原生智能体公司组成的全新经济体正在建立。所以,“Founding 500”的核心在于尽可能押注这个经济体中最前沿的部分,对吧?也就是那些最具智能体杠杆效应的公司。它们不仅在颠覆自己的领域,还在为这个新时代的公司该如何运营制定一套全新的剧本。而且我认为,我们已经从这些以这种方式运营的公司身上看到了巨大的杠杆效应,就像 Gary 所说的那样,这是极其深远的。我的意思是,这远远超过了作为一家早期由软件驱动或互联网驱动的公司的影响。所以我真的非常非常坚信,这部分经济子集将会迎来爆炸式的增长,其增速将远远高于任何其他板块。并且,它将横跨所有行业,而不仅仅是科技或软件行业。

Original English

Speaker A: you know I I think um as I think about the multiple layers of disruption stacking, you know, we we're going as I said from, you know, just humans doing the same work they were augmented a little bit by AI to now humans overseeing fleets of agents doing work that wasn't before possible like this landscaper running his business in a fundamentally different way to now this entirely new economy of agent native companies being built. And so, you know, really the founding 500 is about betting on the most frontier part of that economy as possible, right? It's the most agentically leveraged companies that are not only out disrupting their space but kind of defining an entirely new playbook for how companies in this new era will be run. And I think the leverage we're already seeing from some of these uh these companies that are running in this way like Gary said like it's profound. I mean it's far more than even being like an early software leveraged or internet leveraged company. So really really you know kind of big believer that this subset of the economy is going to have explosive growth far higher than any other sector and it's going to cut across not just you know the tech industry software industry but every industry.

初创竞赛与项目展示

Speaker C: 是的。在我在伦敦的主题演讲中,我经常谈到工作智能体是如何吞噬其他工作,并对非技术团队变得触手可及的。事实上,我的团队现在就在后面,因为 AIE 运行在 Airtable 上。他们非常兴奋能见到你们,也非常期待能与 Hyper Agent 合作。从 500 名创始人中,你们预先筛选出了 20 家公司。他们今天在创业竞技场上进行了角逐。最终有三家脱颖而出成为入围者。让我们来介绍一下他们。我不知道我是否应该点击这个。不过首先,我来介绍一下我们亲爱的评委们。我想我们有 Theo 和 Joshua。谢谢,他们来自 Hen。我想我们稍后会请他们在那里进行评判。但我们也会介绍一下比赛的形式和背景。所以基本上,我们要做的就是,Hyper Agent 团队非常慷慨地提供了 10 万美元的奖金。本质上这是一个比赛,旨在选出最令人印象深刻的获胜者。所以,冠军将获得 5 万美元,亚军 3 万美元,季军 2 万美元。我们准备了一段宣传视频,非常感谢这是由 Hyperframes 和 Hunen 制作的。我们会有有趣的推介环节,以及与评委的问答环节,然后我们也会继续进行下一个推介。也就是三个快速推介背靠背进行。基本上,我们决定我们需要某种评判机制,来衡量什么才是令人印象深刻的,这些技术的使用方式,以及它为什么重要。这与其说是一个投资路演,不如说,它不像 VC 路演,更像是在探讨:你为什么要关心?为什么这是一个有趣的技术应用?以及为什么说你是一个值得人们认识的有趣创始人?因此,我非常激动地宣布获胜团队。我想首先,我们要请出,来吧,Spiro,上台来。好的。谢谢你们。我想接下来就交给你们来展示了。

Original English

Speaker C: Yeah. Uh in my London keynote I always often talk about how uh work agents are eating the rest of work and being accessible to nontechnical teams. In fact my team is right there back now because AIE runs on air table uh and they're so excited to meet you and so excited to to work with hyper agents. Um so from 500 founders we you guys pre-seelelected 20 they competed today in the start of battlefields. Uh three have emerged as finalists. Let's introduce them. Um so I don't know if I'm supposed to to click this. Uh but uh first of all I'm just going to go to introduce uh our dear judges. I think uh we're going to have Theo and Joshua. Um thank you from Hen. And uh we I think we'll we'll we'll sort of uh have them uh judging there. Uh but also we're going to uh go into the format and the background as well. Uh so basically what we're going to do uh we um the Hyper Asian team has kindly put up $100,000 in prizes and basically what this is is a competition for whoever is going to be the most impressive uh winners. So uh 50k to the winners um 30k to runner up 20k to second runner up. Uh we're gonna have a hype video that is uh thankfully um created by Hyperframes and Hunen. Uh we have fun pitches and Q&A with the judges and then we'll also get to the next uh uh pitch as well. So just three quick pitches back toback. Um and basically we decided that we really wanted some kind of judging in terms of like what is impressive and usage of this technology and like why it matters. Um it's not an investment pitch so much as like you know it's not like a VC pitch so much as like well why should you care why is this an interesting use of technology and and why um you know why are you like an interesting founder that people should know um and so I'm very excited to announce like the winning teams. Uh so I think first of all we're going to introduce um come on Spiro come on the stage. All right. Thank you. Um I think you guys are gonna take over and present. Yeah.

Spiro Anakos: 好的。

Original English

Spiro Anakos: All right.

Speaker C: 你们有翻页笔吗?或者你们准备好开场了吗?好的。

Original English

Speaker C: Do you have your clickers or you got you got your you're gonna kick it off? Okay.

Spiro Anakos: 大家好,我是 Spiro Anakos,Kamad 的创始人之一。这位是 Dom。

Original English

Spiro Anakos: Hey everyone, I'm Spirro Anakos, one of the founders of Kamad. This is Dom

Nick Nakos: 我是 Nick Nakos。

Original English

Nick Nakos: and I'm Nick Nakos.

Spiro Anakos: 我们是 Kamad 的创始人。Kamad 已经重塑了实体大宗商品的贸易。从你吃的大米,到你戴的金链,到你智能手机里的芯片,到你给车加的油,再到贯穿我们数据中心的每一根铜管。实体大宗商品为世界提供动力。但是全球贸易体系已经崩溃了。每年有数万亿美元在流转。然而,由于欺诈、流程碎片化,以及协调全球贸易的极端复杂性,每年都有数千亿美元损失。Kamad 接管了这种复杂性,并将其转化为一个由智能体构成的执行层。我们专有的 CFC 状态机管理着专门的智能体,它们可以完全自主地验证、执行和结算交易。然后,人们会问,为什么这会这么复杂?如果你想一想就会发现,这涉及到世界上的每个国家、每一种资源、它们所有不同的货币,以及所有不同的银行法律。所以,这不可避免地会变成一场灾难,对吧?但是有了 AI,这终于可以被解决了。

Original English

Spiro Anakos: We are the founders of Kamad. Kamad has identified physical commodity trade. From the rice you eat to the gold chain you're wearing to the chip in your smartphone to the gas you're putting in your car to every single copper pipe running through our data centers. Physical commodities power the world. But the global trade system is broken. Trillions move annually. Yet hundreds of billions are lost annually due to fraud, fragmentation, and the sheer complexity of coordinating global trade. Kamad has taken that complexity and transformed it into an agentic execution layer. Our proprietary CFC state machine governs specialized agents that verify, execute, and settle trades entirely. And then, you know, people ask, why is this so complicated? And if you think about it, it's like, well, it's every country in the world, every resource, all of their different currencies, and all of their different banking laws. So, of course, it's inevitably going to be a disaster, right? But with AI, it's finally solvable.

Nick Nakos: 所以,我们非常高兴能邀请你们来到这里。我们的第一位客户也在现场的观众中,他来自阿联酋迪拜的一家贸易公司,他每年在贸易中流通的资金量高达数亿美元。他已经证明了这款产品。因此,我们很激动今天能站在这里,感谢给我们这个机会。

Original English

Nick Nakos: So, we're very excited to have you here. We also have our our first customer in the crowd out of the UAE U trading company based in Dubai, and he moves several hundreds of millions of volume in trade every single year, and uh he's proved the product. So, we're excited to be here today and thank you for the opportunity.

Dom: 每年有价值 17 万亿美元的实体大宗商品在全球流转。世界依靠它们运转,然而这一切仍在依靠纸张、电子邮件和握手协议来运作。没有人知道谁是真的,或者文件是否属实。因此损失了数十亿美元。Kamad 可以直接接入你现有的技术栈。Gmail、Slack、Stripe、Salesforce、WhatsApp 等等。当一笔交易开始时。智能体会审查双方。在几秒钟内完成 KYC(了解你的客户)、身份验证和制裁筛查。取证技术瞬间捕捉文件欺诈。接着开启一个私密的交易室。航运和物流合作伙伴通过实时报价链接进来。然后,一个智能体集群在并行的状态下构建整个交易。合规、法务、货币、银行、航运、结算,过去需要数周的工作,现在只需几分钟。这个集群服从唯一的协议,即 CFC,这也是我们已申请专利的状态机。智能体只有在清除了所有的否决条件后才能推进流程。在证据确凿之前,它是不会行动的。满足结算条件后,它从不触碰资金,而是发出信号,由你的银行来转移资金。原本需要 30 天的结算周期,现在几分钟内就能完成。这就是大宗商品交易。

Original English

Dom: 17 trillion dollars in physical commodities move every year. The world runs on them, yet it still moves on paper, email, and a handshake. No one knows who is real or if the documents are the result. Billions lost. Kamad plugs into the stack you already run. Gmail, Slack, Stripe, Salesforce, WhatsApp, and more. A deal begins. Agents screen both sides. KYC identity sanctions in seconds. Forensics catch document fraud instantly. A private deal room spins up. Shipping and logistics partners link in with live pricing. Then an agent swarm structures the deal in parallel. Compliance, legal, currency, banking, shipping, settlement, weeks of work in minutes. The swarm obeys one protocol, the CFC. Our patent filed state machine. Agents advance only by clearing every deny condition. It cannot move until the evidence is real. Settlement eligible. It never touches the funds. It emits the signal. Your bank moves the money. 30 days closed in minutes. Commodities.

Speaker C: 在监控屏幕上看不到视频。

Original English

Speaker C: Can't see the video on the confidence monitor.

Spiro Anakos: 现在该怎么做?

Original English

Spiro Anakos: What happens now?

Speaker C: 我不知道。

Original English

Speaker C: I don't know.

Spiro Anakos: 镜头是在拍我吗?

Original English

Spiro Anakos: Is it on me?

Speaker C: 是的。我刚意识到我们没有给你配麦克风。所以,你在那头还戴着麦克风吗?哦,你们都戴着麦克风。好的,太棒了。那么,我想我们接下来有一个简短的问答环节。

Original English

Speaker C: Yes. Um I I I realized that we didn't get you a mic. So, uh, are you're still miked up at the end of Oh, you're all miked up. Okay, great. Um, so I think we have a little bit of a Q&A session.

Spiro Anakos: 是的。

Original English

Spiro Anakos: Yeah.

Speaker C: 所以,我们应该提问还是怎样?

Original English

Speaker C: So, are we supposed to ask questions or what?

Spiro Anakos: 没错。

Original English

Spiro Anakos: Exactly.

Speaker C: 我的团队答应了这事,我可没答应。

Original English

Speaker C: My team agreed to this, not me.

Spiro Anakos: 等等,你们为什么要放这个?

Original English

Spiro Anakos: Wait, why are you guys playing this?

Speaker C: Joshua,你想先开始吗?

Original English

Speaker C: Joshua, you want to go?

Joshua: 好的。我对第二个项目 common.io 非常感兴趣。当然了,我们确实需要一个共享空间,让智能体们能够在一起协作。我只是有点好奇,你们有没有想过把这种格式做成 HTML 而不是……

Original English

Joshua: Uh, sure. I'm very interested in the um the second project, common.io. And you know certainly we need a share space for agent to collaborate together. Just curious have you thought about having that format to be HTML versus...

评委提问与大宗商品交易平台答辩

Speaker A: 类似 Markdown 文件?你们是怎么考虑这个问题的?

Original English

Speaker A: like markdown file? How you've been thinking about it?

Speaker B: 这是什么情况?

Original English

Speaker B: What is this?

Speaker C: 抱歉。

Original English

Speaker C: Sorry.

Speaker D: 没关系,没问题。

Original English

Speaker D: All right. Sure.

Speaker E: 是的,先别……

Original English

Speaker E: Yeah. Don't

Speaker F: 对,我刚才只是想说,你知道的,第二个项目 common.io 是关于……

Original English

Speaker F: Yeah. I was just saying that you know um um the second project common.io is about

Speaker G: 那是下一个项目了,是的。

Original English

Speaker G: That's the next one. Yeah.

Speaker F: 哦,那是下一个项目。

Original English

Speaker F: Oh, the next one.

Speaker G: 我们现在正在向他们提问关于第一个项目的事情。

Original English

Speaker G: We're asking them about the first one.

Speaker F: 呃,那么……

Original English

Speaker F: Uh so

Judge: 那我来问一个问题。你们这款产品属于那种开发难度极高的类型,因为它不仅需要大量特定领域的专业知识,还需要构建所有这些细分模块,以此来收集产品所需的各项信息。对于现代初创公司,我最大的担忧之一在于,他们正在构建的东西,最终可能会沦为一个足够智能的 AI 模型的某项功能。你们是否预见到这样一个未来:随着模型变得越来越强大,并且借助 Anthropic 和 OpenAI 赋予它们的工具,如果它们变得足够聪明并在正确的方向上进行构建,它们是否足以与你们竞争?

Original English

Judge: I'll ask a question. This is one of those types of products that's really hard to build because it requires so much domain specific knowledge as well as like the building out of all of these pieces to collect the information needed for the product. A lot of the concern I have with modern startups is the thing that they're building is eventually going to be a feature of a sufficiently intelligent model. Do you see a future where the models get capable enough and the tools they're given by anthropic and OpenAI could compete with you guys if they were to get smart enough and build in the right directions? ticket.

Commodities Founder 1: 在大宗商品交易中,所需要提取的数据是非常具体的。我们在编排层(orchestration layer)所做的工作,就是为每一个智能代理(agent)设定好权限与门槛(gates)。因此,它们是从真实的数据库中提取数据,并且自身拥有这些上下文信息。而你提到的那些交易,无论是黄金、石油、谷物还是其他任何商品,都有某些只有行业专家才真正了解的操作程序,并且只有他们才拥有引入这些数据所需的各类文件和信息。这就是为什么我们与 Connor 还有其他几个合作伙伴建立了合作关系,以此来整合所有这些数据。此外,我也拥有该领域的丰富经验,能够分辨出什么是真实的、什么是虚假的,并以此为基础来训练我们的智能代理。所以,回到你的问题,那种情况会发生吗?它们肯定可以在开放的互联网上学习到更多知识,但这是一个包含大量敏感信息的私有行业,它们必须基于某些特定的内部数据进行学习和训练,才能取得进步并成为该领域的专家。

Original English

Commodities Founder 1: So in commodities trade, it's very specific as to what data is being pulled. And what we've done with our orchestration layer is set the gates for each agent. So they're pulling from real databases and have that context themselves. Whereas you're talking about gold, oil, grain, whatever it may be, there's certain procedures that only industry experts really know and have the documents and the information to pull that in. Which is why we've partnered with Connor and a few other partners to bring in all that data. uh and I have domain experience as well to know what's real, what's not and to train the agents based upon that. So to your question, could that happen? They can learn more for sure on the open internet, but this is a private industry with a lot of sensitive information and they have to learn upon and train upon certain data in order to progress and become experts in that domain.

Judge: 所以听起来你们的服务在某种意义上是相当定制化的,因为每家公司的情况都略有不同。你需要为每一家公司配置好一切。那么,要想让这家公司成为一家估值十亿美元的企业,或者,比如说,达到一亿美元的年度经常性收入(ARR),你们认为需要多少客户才能实现这个目标?

Original English

Judge: So it sounds like relatively bespoke in the sense that like every company's a bit different. You need to get everything for them. How many customers would it take for this to become a billion dollar company or to hit like I don't know 100 mil ar like how many customers you think it would take to get there?

Commodities Founder 1: 大概只需要 100 个客户,总共 100 个客户就够了。

Original English

Commodities Founder 1: It's about 100 uh 100 customers total

Judge: 这数字真是惊人。

Original English

Judge: huge.

Commodities Founder 1: 同样重要的一点是,在这个行业里,每一单交易的金额规模都是极其庞大的。这跟 Stripe 处理小额微交易完全不是一个概念。这里的一笔交易可能就是 2 亿美元甚至更高。而且像炼油厂这样的合作伙伴,他们中很多每月的交易额都在 90 亿到 120 亿美元之间。因此,如果我们能在这样的资金体量下介入,并按照单笔交易收取费用,那将彻底改变我们能够赚取利润的整个格局,也会彻底改变我们公司的潜在估值。

Original English

Commodities Founder 1: It's also important to mention as well like the tickets in this industry are massive on on deal sizes. It's not like a Stripe for example processing microtransactions. One transaction might be you know $200 million or north of that. and one partner such as a refinery, many of them transact 9 to12 billion dollars a month. So getting in at that level, charging per deal on that changes the whole landscape of what we're able to earn and what we're able to be valued at.

Commodities Founder 2: 我想再补充一点。我们的解决方案是端到端的。如你所知,这个解决方案服务于整个供应链,但切入这个行业的最终突破口(wedge)在于交易员和银行。多年前,在欺诈行为变得如此猖獗之后,各家银行纷纷退出了实物大宗商品交易领域。因此,通过我们为大宗商品及交易本身提供的审计追踪和不可篡改性,这也将释放出极为庞大的流动性。所以,这项业务有很多值得期待的地方。

Original English

Commodities Founder 2: I just want to add one more piece as well. Our solution is endtoend. So you know the solution serves the entire supply chain, but the wedge ultimately into the industry is the traders and the banks. Banks pulled out of physical commodities trading years ago after the fraud became so rampant. So this unlocks a tremendous amount of liquidity as well by the audit trails and the immutability that we give uh into the commodity and the trade itself. So there's a lot to love about this business.

Judge: 我觉得这在很大程度上满足了所有标准,这就好比是典型的、最理想的垂直智能代理(vertical agent)商业模式。如果你看看 Y Combinator 关于垂直代理的需求方向(RFS),这简直就是在合适的领域里最完美的应用案例。因为那里有大量的资金可以争取,同时你也必须处理大量特定领域的专业知识,我认为这就形成了一个很好的、具备防御性的护城河。不过我的视角更多是倾向于,你知道,你该如何应对不断变化的竞争格局?随着模型的不断进化,你需要如何优化你们的产品以保持竞争力?对吧,这不仅仅关乎它今天是不是一门相关且出色的生意,更多的是在接下来的两三年里,甚至明年发生的变化会如何影响你。你知道,这并不一定意味着 Anthropic 或 OpenAI 会抢了你们的饭碗,但你们需要做些什么来保持,你知道,始终站在技术前沿并提供更好的产品。

Original English

Judge: I think like in many ways it's like checks all the boxes on like the classic like here is the ideal vertical agent business to build. Like if you look at the the YC kind of like RFS for vertical agents like this is like the perfect application in the right the right kind of sector. like there's lots of money up for grabs and there's a lot of like domain specific stuff that you have to do that creates I think a good defensible moat. Um I think my perspective is more like you know what do you do to keep up with the changing competitive landscape and just you know as models progress how do you need to evolve your product to stay competitive right so it's less about is this a relevant and great business today and more about what changes over the next two three years or even like next year that impacts you in a way that you know it doesn't necessarily mean like enthropic opening eye eats your lunch but what do you need to do to stay you know kind of at the frontier and be a better product.

Commodities Founder 1: 是的,我认为唯一真正具备防御性的,就是成为你所在品类中最具创造力的那一个,不管你正在打造的是什么产品,对吧?说真的,除了成为最具创造力的人之外,没有任何东西是绝对不可攻破的。现在任何人都有自己的产品,但我们实际上很推崇“Anthropic 会吃掉你的午餐”这种模式。我们喜欢那种导致物种灭绝式发布的模式(ship extinction model),就是我们一旦发布产品,第二天就会让 10 家公司面临淘汰。所以,是的,核心在于极其激进的发布节奏以及保持产品的绝对优势,但最终我们将通过分销渠道和市场推广来赢得这场竞争。

Original English

Commodities Founder 1: Yeah, I think like the only thing that's actually defensible is being the most creative in your category regardless of which product you're creating, right? Like truly nothing's defensible other than being the most creative, right? Anybody these days has a product, but we actually do like the anthropic eat your lunch model. Like we like the ship extinction model where we ship and then 10 companies go extinct the next day. So um yeah, it's aggressive shipping and product superiority, but then ultimately we're going to win on distribution.

Judge: 如果可以分享的话,在你们即将推出的产品路线图上,最让你感到兴奋的是什么?

Original English

Judge: What are you most excited about on your upcoming road map if you can share?

Commodities Founder 1: 成为银行。

Original English

Commodities Founder 1: becoming the bank.

Judge: 是的,这确实令人兴奋。非常感谢你们。太棒了。真的很出色。把掌声送给他们,接下来,我们要邀请下一个团队上场了。呃,如果你们想回到这边的话。

Original English

Judge: Yeah, it's exciting. Thank you so much. Awesome. Awesome. Give it up and come on. Uh we're going to invite the next team on now. Uh if you want to come back here.

Host: 谢谢。非常感谢。是的,太棒了。谢谢。接下来我们要有请 Max Windbomb 以及 Common I.O. 的团队。上来吧,Max。哦,你要播放视频是吧。好的,上来吧。好的,播放视频。就是那个,那是正确的……好的,欢迎加入。大家快上来吧。哦,谢谢。

Original English

Host: Thank you. Appreciate it. Yeah, that's awesome. Thanks. Um and up next we have Max Windbomb and the team from Common I.O. Come on, Max. Oh, you're going to play the video. All right. Come on. All right. Play the video. That was That's the right All right, welcome in. Come on, guys. Oh, thank you.

Common I.O. 团队登场与介绍

Host: 嗯,如果可以的话,我们这里遇到了一点技术故障,导致评委们无法看到视频。好的,所以在你正式开始整个提案之前,我们可能需要你稍微描述一下,就当是做一个简单的引导。好的,你就把我们当成是一个非多模态的模型吧。我们只接收纯文本信息。

Original English

Host: Um, if you could, uh, so we have a little bit of a technical snafu where the judges couldn't see the video. Okay. Uh we have to describe a little bit just to just to prompt things off before you if you sort of before you get into the whole pitch. Yeah. Okay. Pretend like we're a non multimodal model. We just we only accept text.

Common I.O. Team: 所以,是的,text.io 是一款……它是一个多人协作的 Markdown 编辑器。它致力于打造现代工作环境中最优秀的文档平台,也就是供人类和智能代理共同使用的环境。你在视频里本来会看到一个大家非常熟悉的文档编辑器界面,但是在这个界面里,人类和各种智能代理会不断弹出来。他们互相发送通知。每个人都可以把任何一个人或代理召唤进来。嗯,基本上这就是它的核心理念。

Original English

Common I.O. Team: So uh yeah text.io is a uh it's a multiplayer markdown editor. It's uh focused on making the best documents um for the modern work which is people in agents. But what you would have seen on the video is a familiar document editor interface, but people and agents are popping in. They're notifying each other. Everyone can call anyone in. Um, and that's that's basically the gist.

Judge: 太棒了。好的。

Original English

Judge: Awesome. All right.

Common.io Co-founder: 所以,想象一下,现在是周五下午,你正准备下班回家,正在收拾你的笔记本电脑,突然之间,你听到办公室里到处都是提示音:服务器宕机了,发生了一次大规模的系统故障。于是,你放下背包,走进会议室,拿出你的笔记本电脑。你脑海中只有一个问题:到底发生了什么,Claude?然后你把这个问题输入进去。你等了几分钟,系统正在汇总信息或者做其他的处理,然后它给了你一些想法。你做的第一件事,就是把这些想法告诉会议室里的所有人,好让他们把这些信息输入到他们各自的 Claude 模型中。接着,另一个人有了不同的想法。他们告诉你之后,你又把这些新想法输入到你的 Claude 中。然后你突然意识到,我们怎么会变成这样?工作怎么会变成了这种沟通模式?肯定有一种更好的办法。而且确实有。这种办法就是让正在执行工作的智能代理和工程师们,以及对这个问题负责的人们,能够实时共享上下文。但具体怎么实现呢?想想你们平时都在使用的生产力软件和协作软件。你能不能直接把你用来编写代码的智能代理丢进这些软件里,让它们互相协作?不,你做不到,因为那些软件是为人类打造的,而不是为智能代理设计的。这就是为什么我们要开发 Comment。它是一个多人协作的 Markdown 编辑器。我们极度专注于让它成为现代工作中最出色的文档平台,同时服务于人类和智能代理。它是完全开放的,且支持插件扩展。这对企业来说至关重要。这对开发者来说也是至关重要的。而且,所有这些都被包裹在一个极其熟悉、让人感到亲切的用户体验之中。我在生产力软件领域开始了我的职业生涯,曾为 Microsoft Word 引入过协同工作的功能。Max 和我曾经在 Texio 共事过,那也是首款 AI 原生的写作软件。在 ChatGPT 出现之前的六年里,我们就已经在向《财富》世界 500 强企业进行公司范围内的 AI 写作软件部署了。我们在这个领域深耕了很长很长的时间。

Original English

Common.io Co-founder: So, um, it's Friday afternoon, ready to go home, packing up your laptop, and all of a sudden you're hearing dings around the office, servers down, big outage. So, you put down your bag, you head into the conference room, and you're pulling out your laptop. There's one question on your mind. What is going on, Claude? And you're entering that in. You're waiting a few minutes. Give me some ideas. It's culminating or whatever. And then it gives you some ideas. And the first thing you do is you tell everyone else in the room the ideas so they can put it into their clots. And then someone else has a different idea. And then they let you know, and then you put it into your clot. And you real, how did we get here? How is this the job? There has to be a better way. And there is. It's agents and the pe who are doing the work and people who own the problem sharing context in real time. But how? Think of the productivity software that you all use, the collaboration software. Can you just like drop your coding agents into them and have them working with each other? No, you can't because they were built for people, not for agents. And that's why we're making comment. It is the multiplayer markdown editor. It is we're laser focused on making it the best document platform for modern work for people and for agents. It's open. It's pluggable. That's key for enterprises. It's key for it's key for developers. And it's wrapped in a user experience that is super familiar. I started my career in productivity bringing collaboration features into Microsoft Word. Max and I worked together um at Texio, the first AI native writing software. Six years before chat, we were deploying companywide rollouts of AI writing software to the Fortune 500. We've been doing this a long time.

Common.io Co-founder: Comment 的市场正随着智能代理的普及而不断扩张。随着越来越多的人使用智能代理,也会有更多的人需要一种能够让智能代理共享上下文的能力,而这种能力必须承载于一个专为智能代理打造、但又是为了人类在最高风险和极高压力环境下协同工作而设计的平台中。

Original English

Common.io Co-founder: The market for comment is expanding with agents. As more and more people use agents, more people are going to need the ability for agents to share context in a product that is built for agents but designed for people working together even in the highest stake circumstances.

Judge: 太棒了。我非常喜欢这项业务的形态,因为它就像是我们刚才从那几位做大宗商品(commodities)的小伙子那里听到的模式的倒影,对吧?他们是走了一条极其深度的垂直领域路线。

Original English

Judge: Awesome. I love the shape of this business because it's like the inverse of what we just heard from the um the commod guys, right? Like they went super vertically deep.

产品形态的思考与代理写作软件的挑战

Speaker A: 他们试图从一个非常深的垂直领域中榨取大量价值,而你们的产品则走的是超级广泛的路线,并且按理说可能更浅一些,但这是有意为之的,对吧?我这么说完全不是在贬低你们,我觉得思考当前市场格局的一种方式是,代理经济(agent economy)将会变得非常庞大,会有海量的资金在代理之间流动,实际上就是代币的消耗,对吧?如果你们的产品能接触到哪怕只有 1% 到 2% 的代币消耗量,那它也会变得非常有价值,即使每次调用、每次使用或者某种指标的利润提取或收入相对较低,你们依然可以成为一个非常便宜、非常快速、应用非常广泛的产品,并且仍然能获得极大的规模效应。所以我觉得这是一种非常非常有趣的产品形态。那么,从产品设计或产品功能的角度来看,你们发现的最令人惊讶的经验是什么?能够让这款产品在供代理使用时,比人类设计的同类产品更好?对吧,比如为什么不直接让代理在 Google 文档上协作,对吧?或者甚至尝试使用 GitHub 这样的代码仓库来处理不仅仅是代码,也包括 Markdown 的内容呢?

Original English

Speaker A: They're trying to extract a lot of value out of a very deep um you know domain and you guys are going super broad and arguably maybe a little thinner but by intention right and I don't say that as a as a detriment at all like I think you know one way to think about the landscape here is like the agent economy is going to be so large there's just so much money flowing through agents like literally the token spend right and if you can be a product that like even a one to you know 2% uh amount of those tokens touches like you can be a very valuable even if like the per uh per call or per usage um you know kind of uh metric uh you know kind of margin extraction or revenue is is relatively low you can be a very cheap very fast very um broad product and still gain a lot of scale right so I think it's a really really interesting shape of product um what have you found is like the most surprising learnings from like a product design or product affordances standpoint to make this product better for agents than like the human designed equivalents right like why not just have agents collaborate on a Google doc right or even try to use like GitHub as the uh you know kind of repo not just for code but for markdown.

Speaker B: 今天所有的写作软件对于代理来说使用体验都非常糟糕。像 Google 文档这样的软件格式极其糟糕。它有一个 MCP(模型上下文协议),你可以向那里发送一份文档,你也可以让你的代理从那里读取文档,但你无法让它们进行实时协作。而这对于所有需要完成的实际工作来说都是至关重要的。随着越来越多的代理加入劳动力队伍,它们需要与你的文档待在一起。并且正如我们与用户交谈时所了解到的那样,那些喜欢在文档中工作、日常工作严重依赖文档的人,也需要引入他们的代理,因为他们也开始使用代理了。这已经不再仅仅是开发者的专属了。

Original English

Speaker B: All writing software today is horrible for agents to use. Software like Google Docs has a really terrible format. It has an MCP. You can send a document there. You can have your agent read a document from there but you can't have them collaborate in real time. And that's critical for all kinds of real work that needs to happen. And as more agents join the workforce, they need to be in there with your documents. And as we've talked to our users, people who love working in documents, who rely on documents for their day-to-day, need to bring their agents because they're starting to use agents, too. It's not just developers anymore.

Speaker A: 好的,我觉得……

Original English

Speaker A: Okay. I think

Speaker C: 我来说一下吧。

Original English

Speaker C: I'll mention it.

Speaker B: 哦,最后的总结陈词。我本来只想说,我要提到的另一件事是,当我们在这次会议上与人们交谈时,令我感到非常惊讶的是,有多少人表示过:“我自己开发了一个可视化的 Markdown 编辑器,能让其他人看到,这样大家就能一起查看”,然后你会觉得:“啊,但我放弃了,我干脆直接渲染一个 HTML 文档了”,因为这是一个很难解决的问题,你需要集中精力投入才能把它做好,而这正是我们正在做的事情。

Original English

Speaker B: Oh, final closing words. I was just going to say the other thing I'll mention is that I've been super surprised about um talking to people at this conference when we've talked to them is how many people have said I've rolled my own like visual markdown editor that other people can look at so people can look and then you're like ah but I gave up I just rendered an HTML document instead like it's a hard problem and you need dedicated like focus on making it great and that's kind of what we're doing.

Foundry:将内容创作者转变为创始人

Max: 太棒了。好吧。干得漂亮,伙计们。非常感谢你们。接下来上场的是最后一支参赛队伍。我们将播放 Foundry 的视频。

Original English

Max: Awesome. All right. Nice job guys. Thank you so much. Uh next for and last team to compete. We're going to play the video by Foundry.

Foundry Video: 内容创作者没有时间去建立一家企业,所以他们必须雇佣一家中介机构。这种方式昂贵、缓慢,而且风险很高。Built by Foundry 取代了中介机构,让创作者能够以更快、更便宜的方式推出自己的产品。我们的 AI 代理会端到端地构建整个业务。创作者只需负责发布,而我们将持续增加收入。只要输入任何一位创作者的用户句柄(handle),我们的代理就会成为他们最铁杆的粉丝——观看他们的视频,阅读他们的评论,并找到他们的受众实际上愿意付费购买的“止痛药”(痛点解决方案)。我们已经为各种规模的创作者提供过这项服务,有超过五十万人使用了我们的代理所构建的产品。开发软件的成本正在趋近于零。剩下的核心要素就是信任和品味,而没有人比创作者拥有更多的信任和品味。我们正在构建一个代理产品团队,它将把每一位创作者转变成创始人。

Original English

Foundry Video: Content creators don't have time to build a business, so they have to hire an agency. It's expensive, slow, and risky. Built by Foundry replaces the agency and lets creators launch their own products faster and cheaper. Our AI agents build the business end to end. The creator launches it. We keep growing the revenue. Type in any creator's handle and our agents become their biggest fans, watching their videos, reading their comments, and finding a painkiller their audience would actually pay for. We've done this for creators of every size and over half a million people use products our agents built. The cost of building software is approaching zero. What's left is trust and taste and no one has more of it than creators. We're building the agentic product team that is going to turn every creator into a founder.

Max: 哇哦!Foundry 团队,请上台。

Original English

Max: Woo! Foundry team, come on up.

Weston Belgettes: 谢谢。谢谢你,Max。

Original English

Weston Belgettes: Thank you. Thanks, Max.

Weston Belgettes: 大家好,Aie。我是 Weston Belgettes。我是 Built by Foundry 的创始人兼 CEO。我们将把每一位内容创作者转变为属于他们自己的公司的创始人。今天我有一个好消息要宣布。我认为我们已经找到了一项能够盈利的业务的公式。你只需要把你独特的专业知识转化为产品,并且用你自带的分发渠道来解决冷启动问题。Theo,你在 T3 Chat 上就做过这样的尝试。

Original English

Weston Belgettes: Hello, Aie. I am Weston Belgettes. I'm the founder and CEO of Built by Foundry. We're going to turn every content creator into a founder of their very own company. And I have great news today. I think we found the formula for a winning business. You take your unique expertise and turn it into a product. And you solve cold start with your built-in distribution. Theo, you you did this with T3 Chat.

Theo: 是的,那很糟糕。那根本行不通。对于我们、我们的用户来说它可能很棒,但是 Gary 在 GStack 和 Gbrain 上也做了尝试。我们发现的是,如果我们帮助创作者退后一步,解决他们的用户所面临的一个真正的痛点——所以我们不只是为此再做一个 Patreon、将这些内容设为付费墙,我们是在帮助他们解决其背后真正的痛点。而这正是我们的代理所擅长的:在用户的问题中找到痛点。我们的公司已经在许多不同的细分领域、为不同规模和类型的创作者构建了产品。实际上,我们正在构建的是一种将权力还给内容创作者的机制。他们陷入了困境——我在 TikTok 工作了两年,我见到了数百名创作者,他们只能无奈地接品牌代言、发联盟营销链接,或者在他们的个人主页链接里售卖 20 美元的 PDF 文件,他们距离建立自己的企业只有一步之遥。他们只是需要有人来帮助他们。而这正是我们的代理要做的事情。在 Built by Foundry,我们部署了一批代理团队来为内容创作者打造能够产生经常性收入的企业,这样他们就可以专注于发布内容、不断地发布内容,而我们可以不断地帮助他们增长收入。对于这一点,我有很多想法。

Original English

Theo: Yeah. It sucked. It doesn't work. It works great for our our our users, but uh Gary did it with with G uh GStack and Gbrain. And uh what what we find is that if we help uh creators to take a step back and solve a real pain point that their users have, so we're not just making another Patreon for that, paywalling this content. We're helping them solve a real pain point behind that. And that's what our agents are experts in is finding the pain point within the users's problems. And our companies have uh built over uh many different niches and uh different sizes of sha and shapes of creators. And really what we're building is we want to give the power back to the content creator. They're stuck in these I worked at Tik Tok for two years and and I met with hundreds of creators who are stuck just doing brand deals or affiliate links or selling $20 PDFs in their link in bio and they are one step away from building their own business. They just need someone to help them. And that's what our agents do. At Built by Foundry, we deploy teams of agents to build recurring revenue businesses for content creators so that they can focus on posting and posting and posting and we can keep growing their revenue. I have a lot of thoughts on this one.

Weston Belgettes: 好的。

Original English

Weston Belgettes: Okay.

Theo 评委的反驳与深度分析

Theo: 那么,作为一个创作者,即使是面对我合作过的品牌,我发现的问题在于,存在一个饱和点(saturation point),当到达这个点时,再向我的受众去谈论某一款特定产品就不再有用了。比如拿我来说,就拿 Vercel 举个例子吧,很快就到了一个阶段:我受众里的每个人要么已经是一名 Vercel 用户,而且不再希望我继续向他们展示这款产品;要么就是他们永远都不会成为 Vercel 用户,并且不想再听到我谈论它了。非常容易就会达到这个饱和点,而最好的解决办法就是,拥有一系列更广泛的产品可以展示给你的受众。我为自己做过的最正确的事情之一,对我个人、我的职业生涯、我的业务以及我的收入产生了最积极影响的那个决定性时刻,就是将我的赞助商从 4 个增加到了 80 个。因为现在有种类更加丰富的东西可以带给我的受众。即使这其中 80% 的内容对他们来说并不相关,但剩下的那 20% 相关的部分,现在就转化成了我原本无法获得的转化率。我能带给受众的产品多样性,绝对是我发现的最大价值所在,并且我合作过的每一位创作者,最终也都得出了类似的结论。

话虽如此,你们确实触及了一些非常关键的点。创作者确实比大多数创始人更有品味。创作者确实比大多数公司拥有更好的分发渠道。创作者也确实比大多数企业更了解他们的受众。而且你们构建了一个可以识别出这一点的机器机制,但它并没有提供针对“另一个问题”的解决方案,那个问题就是——你必须给你的受众带来多样性。

我能看到这个模式可以发挥巨大作用的场景,是作为一种工具供企业使用:当企业发现某位创作者的受众群体,与他们试图寻找以帮助企业为像 Theo 这样的用户构建更好产品的目标客户原型相重合时。我已经把这个道理告诉了无数想要与我合作的企业:去看我的直播,去看我的视频,找到我面临的问题,然后再向我展示你们的产品是如何解决这些问题的。这比你作为一名 YouTuber 去尝试创办自己的公司,是一种要好得多的推介方式。作为一名(据我所知)创办公司数量排名第二多的 YouTuber,这根本行不通。它只会把你消耗殆尽,也会让你的受众感到疲劳。你能做的最好的事情,就是与这些公司合作,教会他们你的受众想要什么,并让他们带着资金来找你,好把他们的产品展示给那部分受众或那些受众成员。

Original English

Theo: So, the problem I find as a creator, even with the brands I work with, is that there is a saturation point where it stops being useful to talk about a given product with my audience. Like with me, for Versell, for example, there was quickly a point where everybody in my audience either was a Verscell user and doesn't want me to show it to them anymore, or they won't be a Versel user, and they don't want me to talk about it anymore. It is very easy to hit the saturation point and the best solution is to have a much broader set of things to show your audience. One of the best things I ever did for myself, the single moment that has had the most positive impact on myself, my career, my businesses, and my income has been going from having four sponsors to 80 because now there's a wider variety of things to bring my audience. And even if 80% of them aren't relevant to them, the 20% that is is now conversion I'm getting that I wouldn't have otherwise. the variety of what I bring to my audience is where I find the most value by far and every creator I've worked with ultimately has come to a similar conclusion. That said, you're touching on really important pieces here. Creators do have more taste than most founders. Creators do have better distribution than most companies. Creators know their audience better than most businesses. And you've built a machine that identifies that, but it it doesn't provide the solution to that other problem, which is you have to bring variety to your audience. Where I could see this working really well is as a tool that businesses use when they identify a creator is overlapping with the archetype of customer they're looking for to help that business build better things for Theo. And I've told this to so many businesses that want to work with me. Watch my streams, watch my videos, find the problems I have, and then show me how your product solves them. That's a way better pitch than going to try and start your own company as a YouTuber. As the YouTuber who has started the second most companies, as far as I'm aware, it doesn't work. It just burns you out and it burns your audience out too. The best thing you can do is work with companies to teach them what your audience wants and for them to come to you with money to show it to those audience or those audience members.

Weston Belgettes: 我明白,并且我想说,你作为一名科技区创作者所在的这个细分领域,是一个极具挑战性的领域,因为有太多的选项可供选择了。而我们发现的情况是,对于我们所有的创作者来说,我只举几个例子。有一位创作者名叫 Brandon,他主要做的是“自给自足式农庄(homesteading)”方面的内容,而他总是会忘记他的冻干机上的最佳设置,这在咱们听起来可能有些荒唐,但对于他的 200 万粉丝来说,这是一个实实在在的痛点。而现在,他们会向他支付每月或每年的订阅形式的经常性收入,来解决那个特定的痛点。所以,我认为不幸的是,Theo,你选了一个糟糕的细分赛道。很抱歉,但是如果你想转型去做农庄或园艺类内容的话,也许我们就能在一起合作了。但我主要想表达的是,无论用户的受众属于哪个垂直细分领域,只要不是科技区,也许除了喜剧演员之外,我们都能够做到……

Original English

Weston Belgettes: I understand and I will say your niche as a tech creator is one that is very challenging because there's a lot of options. Um and what we have found is that for all of our creators, I'll just few examples. Um, one one creator's name is Brandon and he does homesteading content and he always loses track of uh the best settings on his freeze dryer, which sounds ridiculous to us, but to his 2 million followers, that's a real pain point. And now they pay him a recurring revenue monthly or monthly or yearly subscription to solve that specific pain point. So, I think unfortunately, Theo, you've picked a bad niche. I'm sorry, but um if you want to pivot to homesteading or gardening, maybe uh we could we could work together. But I I'd say primarily whatever the users niche is apart from tech and maybe com comedians, we've been able to

Startup Battlefield Finale & Awards

Startup Pitcher: 退一步讲。我们的智能体(agents)能找到他们愿意为此付费的痛点,并且在我们整个公司的历史上,没有任何一位创作者离开过我们的公司,因为他们总是能保持盈利状态。

Original English

Startup Pitcher: take a step back. Our agents find a painoint that they're willing to pay for and zero creators across our entire company history have ever left our our company because they always are profitable.

Judge: 你的话非常有说服力。好的,非常棒。那么大家把掌声送给——

Original English

Judge: >> You're very convincing. Okay, excellent. Then give it up for

Host: 是的,谢谢大家。嗯,接下来我们将进入一小段评审商议时间。评委将到后台讨论几分钟。在此期间,抱歉,好的。在此期间,我们准备了一小段视频,回顾一下这次世界博览会(World's Fair)的精彩瞬间。所以,评委们,慢慢来。这是我们有史以来规模最大的AI工程师活动。我们一直在不断加入新技术供大家参考。热度。热度。嘿,嘿,嘿。非常感谢。好的。我想我们现在邀请所有人回到舞台上。把掌声送给战场(Battlefield)的环节。嗯,恭喜大家。是的。好的。那么,我们要颁发一些奖项。我想,后台的工作人员正在进行一些核对。获得季军(second runner up)的是Kimode。Theo,我想由你来颁奖。Theo。是的。好的,我们也会为团队拍一些照片。

Original English

Host: >> Yeah. Thank you. Um, so we're going to have a little bit of a deliberation moment. Uh, the judge is going to go backstage for a couple minutes. In the meantime, sorry. Okay. In the meantime, we've, uh, prepared this little video recapping the World's Fair. So, take your time, judges. This is our biggest AI engineer. event ever. We're constantly adding new techniques for you to draw. Heat. Heat. Hey, hey, hey. Thank you so much. All right. Um I think we're inviting everyone back on stage. Give it up for the start of Battlefield. Um, congrats guys. Yes. Okay. So, uh, we have some prizes that we're going to hand out. I think, um, the folks back there are doing some checks. Um, in second runner up place, we have Kimode. Uh, Theo, I think you're presenting. Uh, Theo. Yes. Okay, we'll take some pictures as well for the team.

Host: 好的。接下来,我想这同时揭晓了名次。获得亚军(runner up)的是Foundry。

Original English

Host: >> All right. Um, next, and I guess this also reveals the order. Uh, we have Foundry as runner up.

Judge: 演讲很棒,Max。非常棒的演讲。

Original English

Judge: >> Great pitch, Max. Great pitch.

Host: 恭喜。

Original English

Host: >> Congrats.

Host: 太棒了。好的,现在大奖没有悬念了。欢迎 comment.io。你们的视频也非常棒。

Original English

Host: >> Awesome. All right, no secrets now for the grand prize. comment I io welcome. Great video as well.

Judge: 恭喜。实至名归。

Original English

Judge: >> Congratulations. Well deserved.

Host: 你们也想拍张合影吗?我想我们大家一起拍——

Original English

Host: >> Do you want to also do the group photo? I think we're doing

Someone: 我们所有人。

Original English

Someone: >> all of us.

Host: 嗯,是的。好的。好的。

Original English

Host: >> Um, yes. Okay. All right.

Someone: 就让他们拍吧。

Original English

Someone: >> Let it be just them.

Someone: 他们是创始人。我们只是来陪衬大家的。

Original English

Someone: >> They're the founders. We're just here to com everyone.

Someone: 好吧。那我们溜进去。

Original English

Someone: >> Fine. We'll sneak in.

Someone: 抓住他们。快进来。

Original English

Someone: >> Get them. Come on in.

Someone: 好的。

Original English

Someone: >> Yes.

Host: 好的。非常感谢大家使用 hyper agent 进行开发,并且——

Original English

Host: >> All right. Thank you so much guys on building with hyper agent and

Winning Team Member: 是的,当然。嗯,有没有什么最后的话你想评论一下的?我想你们刚才已经简要提了一下。是的。

Original English

Winning Team Member: >> yeah of course um any last words that you wanted to comment on. I think you guys uh briefed a little bit. Yeah.

Another Winning Team Member: 是的,我只是想感谢……很抱歉关于——

Original English

Another Winning Team Member: >> Yeah. I just want to thank uh sorry about

Another Winning Team Member: 感谢在这里(AIE)的每一个人。在整个期间,我们得到了非常棒的反馈。感谢 hyper agent 邀请我们参加这次活动,也感谢 hyper fans 帮助我们制作了视频。

Original English

Another Winning Team Member: >> thank everyone here at AIE? We got amazing feedback um through the whole um time and uh hyper agent um for inviting us to this and hyper fans for helping us with the video

Another Winning Team Member: 总体来说,真的非常感谢大家。

Original English

Another Winning Team Member: >> and uh overall like so much.

Another Winning Team Member: 去看看我们的产品吧。它是免费的。

Original English

Another Winning Team Member: >> Check us out. It's free.

Another Winning Team Member: comment.io。

Original English

Another Winning Team Member: >> comment.io.

Host: 好的,就这些了。谢谢。女士们先生们,请再次用掌声欢迎我们的主持人,Replit的开发者关系工程师,Ralph Shabri 回到舞台。

Original English

Host: >> All right, that's it. Thank you. Ladies and gentlemen, please welcome back to the stage our MC developer relations engineer at Replet, Ralph Shabri.

Wrapping Up AI Engineer World's Fair 2026

Ralph Shabri: 太棒了。好的,让我们把掌声送给创业战场(Startup Battlefields)的决赛选手们。好的,他们总共赢得了价值10万美元的奖金,用来推动他们的项目发展。所以,这真是太棒了。嗯,好的。让我们同样把掌声送给所有的评委,Josh、Howie和Theo。好的,我的下一张幻灯片。翻页笔好像不起作用了。请切到下一张幻灯片。翻页笔坏了。好的。可以了。 随着这些环节的结束,2026年AI工程师博览会(AI Engineer World's Fair 2026)圆满落幕。我们一路走来,取得了长足的进步。今年的展会是我们有史以来最具雄心的一届。历时4天,7000名参会者,多场主题演讲,无数的工作坊、对谈和演示,共计40个分会场。我们希望大家觉得这些内容既有用又有帮助,并希望你们能带着灵感满载而归。但我们最希望看到的,超越其他一切的,是这个社区的人们走到一起,建立新的联系、新的友谊,交流思想,并共同建设。能够成为这其中的一份子,我们非常感激。所以,请大家把掌声送给你们自己。好的。 我们也对我们的赞助商表示极其深切的感激。所以,请大家把掌声送给我们的冠名赞助商,微软(Microsoft)。同时也要感谢我们的实验室及白金赞助商。请掌声不要停。感谢我们的黄金赞助商以及白银和青铜赞助商。谢谢大家。我们非常感激,因为没有这些赞助商,这次活动是不可能举办的。此外,我还要感谢一群默默无闻在后台工作的幕后人员,你们可能永远看不到他们,但如果没有他们,这次活动就不可能顺利进行。所以,感谢所有的工作人员和志愿者,让我们把掌声送给他们。 好的。那么,请切回上一张幻灯片。好的。马上就到了。好嘞。不对。好的。好的。虽然2026年世界博览会到了尾声,但我们还没有完全结束。我们希望能在这个十月于纽约见到大家中的许多人。而且,明年我们还会在这里相聚,举办下一届的AI工程师世界博览会。但在我们庆祝美国生日之前,我们还有最后一件事要向大家展示。热度。热度。把热度炒起来。热度。热度。

Original English

Ralph Shabri: Hot. Okay, so let's give it up to our startup Battlefields finalist. All right, they collectively won 100K worth of prize to boost out their the their project. So, this is great. Um, all right. Let's give it up also to all our judges, Josh, Howie, and Theo. All right, my next slide. Clicker not working. Next slide, please. Clicker not working. Okay. All right. And with that, AI Engineer Welfare 2026 is a wrap. And we've come a long way. This year's edition is our most ambitious one. 4 days, 7,000 attendees, keynote sessions, countless of workshops, conversations, demos, 40 tracks. And we hope that you find this content useful and helpful, and that you're going to go back home inspired. But what we love to see more than anything else is the community coming together, creating new connections, new friendships, exchanging ideas, and building together. And we're extremely grateful to be part of this. So, let's give it up for you guys. All right. We're also very extremely grateful to our sponsors. So, please let's give it up to our presenting sponsor, Microsoft. and also to our lab and platinum sponsors. Please, let's keep it going. Our gold sponsors and our silver and bronze sponsors. Thank you. We're extremely grateful because without these sponsors, this event wouldn't be possible. And also I would like to to thank an invisible group of people who are working backstage who you you never see but without whom this event wouldn't be possible. Uh so thank you all crew members, volunteers, let's give it up for them. All right. So uh previous slide please. Okay. Almost there. There you go. Nope. All right. Okay. So, this is the end of Worlds Fair 2026, but we're not done yet. We hope to see many of you in New York in October. And also, we're going to be back together here next year for another edition of AI Engineers World Fair. But before we celebrate America's birthday, we have one more thing to show you. Heat. Heat. Heat up Heat. Heat.

📌 文中提及的人物和组织

关键字: harness-engineering software-factory code-verification model-iteration