为什么 AI 智能体还不懂你?亚马逊 AGI 实验室与人类智能的集体进化 Latent Space 2026-07-11

启蒙思想与集体智能

Swyx: 好的,今天我们在远程直播间,邀请到了来自**亚马逊 AGI 实验室(Amazon AGI Lab)**的 Danielle。欢迎你的到来!

Original English

Swyx: Okay, we're here in the remote studio with Daniel from Amazon AGI lab. Welcome.

Danielle: 嗨 Swyx,非常高兴能与你见面。

Original English

Danielle: Hi Swyx, it's so good to see you.

Swyx: 我也很高兴我们终于促成了这次访谈。你旅行回来了吗?

Original English

Swyx: Glad we can finally make this happen. Are you back from your travels?

Danielle: 我现在在西雅图,还没有回到湾区。

Original English

Danielle: I am in Seattle, not back in the Bay yet.

Swyx: 好的,真是一次漫长的旅程。你知道,奥德修斯最近似乎成了今年科技界的热门话题,就像一部年度科技大片一样。我觉得你有点像奥德修斯,到处去参加这些令人惊叹的活动。你在法国做了些什么?让我们给大家更新一下近况吧,因为我觉得这真的很酷。作为一个拥有经济学学位的人,你参加英联邦国家组织(Commonwealth of Nations)的活动这件事,我觉得相当酷。

Original English

Swyx: Okay, big round trip. It's you know, Odysseus is kind of like trending as far as like the big tech pool movie of the year. I feel like you're a little bit of an Odysseus, like you're going to all these amazing events. What were you doing in France? So let's update people cuz I think it's very cool. As someone who has an economics degree, the fact that you were in a Commonwealth of Nations event was kind of cool.

Danielle: 嗯,在法国的时候我其实在学习如何从流沙中脱身。不过在那之前,我在爱丁堡,当时那里正在举办**《国富论》(Wealth of Nations)发表250周年的纪念活动,这是亚当·斯密(Adam Smith)**的著名著作。我实际上还去了他的故居,那栋房子经过了重新修缮,是一个非常不可思议的地方。它让我们重新思考启蒙运动时期的思想,以及人类繁荣、自由主义等主题。因此,这次聚会汇聚了经济学家、许多教授以及来自不同行业的专家。

我到场主要是代表人工智能领域,去探讨我们如何思考 AI,以及如何结合 AI 来思考人类生存状况和人类繁荣这些宏大主题。我们如何才能让 AI 为人类服务,而不是仅仅把它当作一个科学实验来构建。这也是我一直在思考的问题。这也正是我自己当初选择离开学术界的原因,因为我之前一直在研究人类智能,以及我们这种特定形态智能的演化,还有那些能够让更多人参与到这些集体动力学中的机制。

而我对 AI 的思考方式,我认为实际上与目前行业中占主导地位的框架有很大不同。我非常荣幸能受邀去分享这一视角。其核心观点是:人类的智能本质上是集体性的(collective)。人类学家也说,我们拥有一种“集体大脑”(collective brain)。没有哪个个体能够仅仅依靠自己生存,我们依赖于集体。智能是从我们的互动中产生的,它在根本上是具有社会性的。而创新——那种让我们不仅能够生存,而且能够适应我们现在所处的各种不同环境的创新——是群体内的多样性(diversity)、群体规模以及互联性(interconnectivity)的函数。

那么,如果让 AI 延伸这些过程,让更多的人能够参与到集体的对话中,并且让 AI 为每个人而构建,而不仅仅是为那些正在构建 AI 的工程师们构建,这会意味着什么呢?这就是我所分享的视角。

Original English

Danielle: Well, in France I was learning how to escape quicksand, but before then I was in Edinburgh, which they were hosting the 250th anniversary of the Wealth of Nations, Adam Smith's famous book, and actually was in his house, which is renovated and it's a really incredible place to remind us of the Enlightenment era ideas and the themes of human flourishing, liberalism. So this gathering brought together, of course, economists and a lot of professors, folks from different industries. I was specifically there representing how do we think about AI and how do we think about some of these themes of the human condition, human flourishing with AI. How can we make AI work for us rather than building it as a science experiment. And so this is something that I think about all the time. This is why I left academia myself because I had been studying human intelligence and the evolution of our particular shape of intelligence and the types of things that allow more people to sort of participate in these collective dynamics. And the way that I think about AI I think is actually quite different than the sort of dominant framing of the industry right now. And I was so honored to be invited to sort of share this perspective. The big idea is that human intelligence is collective. Anthropologists say that we've got the collective brain. No one individual is capable of even surviving on their own. We depend upon the collective. The intelligence emerges from our interactions. It's fundamentally social. And innovation, the type of innovation that allows us to not only survive but adapt to all of the different environments that we now exist in, that is a function of diversity, variation within a population, the size of a population, and the interconnectivity. So, what would it mean to have AI extend these processes to allow more humans to participate in the collective dialogue and for AI to be built for everybody rather than just the engineers who are building AI. So, that's kind of the perspective that I contributed.

Swyx: 我认为这是我们所有人都会赞同、但或许在日常讨论中关注不够深的一个观点。所以,这是一个非常好的信息。不过,人们会害怕 AI 抢走我们的工作吗?你看,这代表了你的工作性质,对吧?如果你把工作做得很好,你确实会消除一些可能没那么有趣的工作,但这也确实意味着工作岗位的减少。

Original English

Swyx: I think something that we all cosign but maybe not necessarily talk about enough. So, I think that's a good message. Are people scared of AI taking our jobs? You know, like you represent, right? Like if you do your job work well, you do take away some maybe less exciting jobs, but you do take away jobs.

Danielle: 是的。这个问题的看法差异很大。我认为很多人理所当然地担心会发生改变,至少会经历一个过渡期。参加这次特定会议的很多人,以及我所在的社群里的许多人,往往都非常乐观,更倾向于关注 AI 所能释放的巨大潜力。但并不是说,如果我们继续按照目前的方式去构建 AI,它就必然会自动给人类带来所有的这些好处。

我认为,当前工作中的人类认知并没有以最佳的方式得到利用。即使是在最富有创造力的工作中,依然存在着大量的琐碎杂务(drudgery),我们整天都盯着屏幕。这并不是我们大脑进化的初衷。我们天生是要互相协作的。我们应该聚在一起,共同构思新想法,进行头脑风暴,做各种能够真正体现我们智能交互本质的事情。但随着我们不断创新,至少从20世纪到21世纪的趋势来看,我们反而被死死地拴在了这些屏幕前。

所以,一方面,我认为很多人对于我们能够将这些繁重琐碎的工作、这些不值得耗费人类注意力和时间的事情自动化感到兴奋。但问题在于,如果我们被困在这种单纯的“自动化思维”中,我们就遗失了 AI 本可以创造的更大价值,而这些价值不仅关乎人类的工作,更关乎人类的福祉、人际互动和人际关系。

因此,现在我们一方面担心 AI 会将大部分工作自动化,但另一方面,当我们看到目前这些 AI 智能体的实际指标时,它们又是如此的不可靠,这在讽刺之余也让我们感到稍微宽慰了一些。因为 AI 实际上还没有达到我们所期望的能够自动化大部分工作的水平。

Original English

Danielle: Yeah. I it's so varied. I think a lot of people are rightfully concerned that there will be changes, that there will be at the very least a transition period. A lot of the folks who are at this particular conference and a lot of the people who are in my communities tend to be very optimistic and sort of index on the potential that AI can unlock, but it is not a foregone conclusion that if we continue building AI in the way that we are, it will actually confer all of these benefits for humans. I like to think that work right now is the human cognition is not being utilized in the best possible way. Even with the most creative types of work, there's still so much drudgery, there's still so much like we we literally stare in front of screens all day long. That is not what our brains are meant to do. We're meant to collaborate with each other. We're meant to put our heads together and come up with new ideas and um you know, brainstorm, do all sorts of things that really resemble the sort of interactive nature of our intelligence. But the more that we innovate, at least the sort of trends of the 20th, 21st century, the more that we're like tethered to these screens. So, on the one hand, I think a lot of people are excited about the possibility that we can automate a lot of the the drudgery, a lot of the things that are not worthy of human attention and human time. The problem is that if we get trapped in this automation mindset, we are leaving on the table so much more value for what AI could be doing for not only human work, but human well-being, human interactions, human relationships. So, right now we're we're I think worried about AI automating a lot of the a large proportion of the work, but then we look at what the the metrics of the actual agents and they're they're so unreliable that ironically we feel a little bit better. The AI is actually not where we need it to be to automate enough of the of the work.


亚马逊 AGI 实验室与认知智能体

Swyx: 让我们介绍一下你一直在做的工作。这几年来,你一直是 AI 工程师圈子里的一员,最初是在 Adept,然后你们团队加入了亚马逊。对于那些不太了解这段背景的人,你能重新介绍一下**亚马逊 AGI 实验室(Amazon AGI Lab)**吗?

Original English

Swyx: Let's introduce the work that you've been doing. You've been part of the AI engineer circle for a couple years, originally with Adept and then you guys joined Amazon. For those who haven't been too close to the story, could you reintroduce Amazon AGI Lab?

Danielle: 是的。亚马逊 AGI 实验室正在构建与人类对齐的智能(human-aligned intelligence)。其起点是构建能够完成人类在电脑上能做的任何事情的 AI。这是 Adept 最初的使命,我们基本上把它继承了下来,并在实验室里播下了这颗种子。

随着时间的推移,我们演进并深化了这一使命。我们在深入思考:能够完成人类在电脑上能做的任何事情,这到底意味着什么?智能体需要具备哪些必要技能才能做到这一点?这远远超出了语言的范畴。我们确实需要智能体能够以人类感知数字环境的相同方式来感知数字环境。更进一步说,数字环境是以物理环境为基础的。因此,我们不仅需要智能体理解这个世界,还需要它们拥有某种“世界模型”(world models)。

我们需要智能体能够进行实时互动(real-time interaction)。这非常关键。我认为这是目前整个行业尚未深入思考的一点。我们现在有点被困在了聊天机器人(chatbots)、编程智能体(coding agents)以及那种分批轮流发言(turn-taking in batches)的局限状态中。但这绝对不是人类互动的方式。我们是在实时交互中,根据当下的语境不断更新我们的理解。我们在协商意义(negotiate meaning),我们在想出新的思考方式。

我们之所以认为这是理所当然的,是因为我们的大脑就是这样工作的。但是,现在的状况是人类在主动去顺应、迁就 AI,去适应技术及其存在的局限性,而不是让技术来适应我们。

那么,如果我们构建的智能体,不仅能以和我们相同的方式感知世界、拥有共同的认知基础,而且还能跟上我们的步伐、与我们共同思考、在聆听我们的同时采取行动、在与我们互动的同时准备它的下一步思考或动作,那会是什么样子?这将在我们思考“交互性”的方式上带来一次范式转变。

Original English

Danielle: Yes, so Amazon AGI Lab is building human-aligned intelligence. And starting with AI that can do anything that a human can do on a computer. That was Adept's original mission and we kind of imported it and seeded the the lab. We've evolved that mission, so we're really thinking deeply about well what what does it mean? What are the skills that are necessary for an agent to be able to do anything that a human can do on a computer. It's so much more than language. We really need agents to be able to perceive the digital environment in the same way that humans perceive the digital environment. And more than the digital environment, the digital environment is based off of the physical environment. And so, we need the agents to also have an understanding of the world and have the kind of world models. We need the agents to be able to interact in real time. This is huge. This is something that I think the industry has not really thought deeply about yet. We're we're kind of trapped in this local attractor state of chatbots and coding agents and like turn-taking in batches. And this is absolutely not how humans interact with each other. We are constantly updating our understanding as a function of the the context in real time. We are negotiating meaning. We are coming up with new ways of thinking about things. We we take this for granted because this is how our our minds work. But, right now, we are kind of accommodating the the AI and the technology and the limitations that it has rather than the other way around. So, what would it look like if we built agents that not only perceived the world in the same way that we didn't have this this sort of common ground, but could keep up with us, think with us, take actions while it's listening to us, prepare its next thoughts or its next actions while it's interacting with us. This would be a a paradigm shift in how we think about interactivity.

Swyx: 是的。我认为最近有一些关于交互模型的研究,我们在交流时也讨论过。我认为在实时大型语言模型(LM)的技术分支上,可以说是以 GPT-4o 的发布为开端的,虽然我认为很多人并没有对此给予足够的重视。显然,在学术界,在这之前就已经有了很多研究。我会指向像 Flamingo 以及许多纯语音模型,它们具有端到端的全双工(full duplex)特性。我们在法国也看到了 Moshi 团队,他们推出了基于 Gradio 的 spoken-agent 演示,他们也在我的会议上发言过。

我只是想补充这些背景知识。我的意思是,与目前主流的聊天对话范式相比,这仍然相对小众,因为它还不够普及。但真正的 AGI 应该就是这个样子的——你能够真正与机器协同合作、实现“心智融合”(mind-meld)。我是这样去定位它的。

Original English

Swyx: Yeah, I think you know, there's some recent work by I think in machines that we've also all covered with the interaction models. And I think this this branch of the LM tech tree of real time was kind of started with the 4o launch which I don't think a lot of people index on. Obviously, in academia, there's more research before that. I would point to like Flamingo and and and and a lot of like the the the just the the pure voice models, um they would have full duplex like end-to-end stuff. Uh we had Moshi in France also uh spit out uh Gradio which also has has spoken at my conferences as well. And so like I just wanted to sort of import the required knowledge. I mean this is like this is I guess compared to like the chat paradigm relatively niche because it's not that popular, but like it is what a real AGI would look like which is that you could actually collaborate and mind-meld with the machine, you know. Um it's kind of how I pitched that.

Danielle: 是的。而且你刚刚列举的那些,包括行业内以及不同学术实验室的成果,他们都在独立地向这些更具灵活性、更接近人类智能的组件收敛。所以,是的,实时交互虽然现在还比较小众,但随着一些团队开始推动,它将成为对话的核心。但这只是为了让 AI 与我们自身的智能更加对齐而所需的更大组件集中的一小部分。

我再举另一个例子。我们往往把“记忆”(memory)仅仅看作是“存储”(storage)。人类一直用技术作为隐喻来理解自己。在20世纪,我们把心智比作电脑,有硬件、有软件,而记忆就像是我们外包出去(offload)存储的东西。但这根本不是人类记忆的工作方式。记忆就是一切。它是我们模拟未来的方式,它发生在许多不同的时间尺度上。仅仅用“存储”这个词根本无法体现记忆已经融入到了所有的学习和认知过程之中。

那么,构建拥有所有这些不同类型记忆的智能体意味着什么?这包括像我们一样的情境记忆(episodic memories),这对于我们的智能至关重要。我们有独特的个人视角和自我意识,这使我们能够以更高效的方式检索信息。

Original English

Danielle: Yeah. And and and you just listed, you know, the industry, different academic labs, they are independently converging on a lot of these components of more flexible human-like intelligence. So yes, interacting in real time even though it's kind of niche now and maybe Thinking Machines is going to make it more central to the conversation. That is one tiny slice of a larger set of components that will make AI more aligned with our own intelligence. I'll give one other example. So we we tend to think of memory as storage. And you know, humans have always used technology as a way to metaphorically understand themselves and in the 20th century we use mind as a computer and you've got the hardware and the and the hardware and the software and and memory is this thing that we kind of offload, but that's not at all how it works in humans. Memory is everything. It's how we simulate the future. It's occurs across many different time scales. It is the word itself doesn't do service to the fact that it is integrated into all learning and and cognition. And so what would it mean to build agents that have all of these different types of memories, including episodic memories like we do which is really core to our intelligence. We have individual perspectives and selves and that allows us to retrieve information in in much more efficient ways.

Swyx: 既然你提到了这个,我想问一个热门话题:你认为体面的记忆需要更新权重(update weights),还是说它依然可以存在于当前的系统架构之内?因为你刚才提到,我们倾向于将记忆视为存储。想必你有一些别的想法,但你没有明说。

Original English

Swyx: You brought this up, so I'm going to ask the the hot topic question, which is do you think decent memory needs to update weights, or do you think it can still live within the systems? Cuz you just said like it you know, we tend to think of memory as a storage. Presumably, there's something else that you're thinking about, but you didn't you didn't say so.

Danielle: 嗯,对于目前行业思考记忆的方式,我认为它将成为一个更大系统的一部分,就像人类也会把记忆外包一样。我的意思是,我们的工具、我们的环境、我们延伸的环境都包含着我们智能的各个方面。我们依赖于工具来做日常的事情,我们可以查阅维基百科,或者使用以新方式组织信息的 AI 工具。这仍然是一个非常重要的核心功能。

但也需要有其他方面的机制。我甚至不想称它们为“记忆”,因为这个词有太深沉的暗示。这不仅仅是在于改变权重,或者即使不改变权重,也是在推理(inference)阶段如何改变信息的上下文语境化(contextualized)方式。我在这里不想透露太多细节,因为这是我们实验室目前正在积极推动的研究,但更大的重点在于,我们需要更全面地思考不同时间尺度上的交互。

Original English

Danielle: Well, so when the way that the industry is currently thinking about memory, I think it's going to be part of a larger system in the same way that humans offload I mean, our tools our environment our extended environment contains aspects of our intelligence. We we depend upon using tools to be able to do daily things, and we can look up on Wikipedia or AI tools that are just organizing information in new ways. That is still a core function that I think is really useful, but there will need to be other aspects. And I I don't even want to call them memory because that has such deep connotations. Changing weights or or if not always changing the weights, changing at inference how information is contextualized. And I don't want to get too much into the details here because this is an active thing that we are pushing in our lab right now, but it's the the bigger point is that we need to be thinking more holistically about different time scales of interaction.

Swyx: 明白了。在探讨你们最近发布的具体成果之前,可能还需要交代一个背景,因为这能提供更具体的认知——比如,目前已经公开宣布了什么。

这个概念性的问题是:你能把亚马逊 AGI 实验室放在更广泛的亚马逊背景中来定位吗?亚马逊内部还有 Nova 团队,对吧?那是亚马逊核心部门的一部分。而你们发布了 Nova Act,我手头正好有这个资料。我想深入探讨这个,因为它引出了感知智能体的话题。但是,究竟是谁在管理亚马逊 AGI 实验室?它和核心部门的关系是怎样的?是非常紧密,还是在一定程度上保持独立运行?对于外部的人来说,有没有什么可以透露的,好让我们理解它的架构?

Original English

Swyx: Got it. Maybe one more contextual thing and then I wanted to also go through some of the recent stuff that you guys have launched because I think that gives concrete things of like well, okay, here's what's been publicly announced. So, the conceptual thing is could you put it in context Could you put that Amazon AGI in context of broader Amazon? There is the Nova group, right, which is which is part of core Amazon. You guys released Nova Act, which I I have pulled up actually. I wanted to go into that cuz that leads into perception agents, but who runs Amazon AGI lab? Like what's their link? Like is it is it a is it very close? Is it very is it sort of meant to be running running more independently? Anything you can give the external people about how to frame what's going on?

Danielle: 是的。我是和原 Adept 团队一起过来的。当我们加入时,我们向领导层阐明:为了进行前沿级别的科学研究,我们确实需要保持一种更像初创公司的运营模式。我们需要保护我们的研究环境,并能够专注于构建感知智能体(perception agents)的使命——尽管我们当时还不这么称呼它们,但这一直是我们的目标。我们与庞大的亚马逊团队一起合作。

实际上,能够获得领导层的认可真的非常不可思议。他们认同我们的研究价值、基础研究的价值,并且不仅支持我们做其他前沿实验室正在做的事情,还支持我们为思考新研究类别和可能不会立即转化为产品的新科学留出空间。我认为其他一些实验室在某种意义上成了自身成功的“受害者”。因为一旦他们有了一款拥有大量用户互动的产品,他们就不得不关闭不同的研究项目,然后说:“好吧,所有人现在都得来全力以赴支持这个产品,对吧?”

Original English

Danielle: Yeah, so I came with the original Adept folks and when we came, we convinced leadership that in order to do frontier level research, we really needed to keep an operating model that was more like a startup. We needed to insulate our research and be able to focus on the mission of building perception agents and even though we weren't calling them perception agents then, that was the goal from from the very beginning. We worked with a large team of Amazonians. It was actually really incredible to get this buy-in that yes, we value the research, the foundational research and and moreover, we value not just doing what the other frontier labs are doing, but making a space to think about new categories of research, new science that might not necessarily be productized immediately. I think some of the other labs are in a sense victims of their own success because they have to once they have a product out there that a lot of users are interacting with, they have to shut down different, you know, research projects and say, "Okay, all hands on deck for this thing, right?"

Swyx: 这很残酷。所有人现在似乎都在做代码编写和 B2B SaaS。

Original English

Swyx: It's brutal. Everyone is just doing coding and B2B SaaS.

Danielle: 是的,没错。而且所有其他的实验室都在拼命追赶其他人在做的事情。我们的亚马逊 AGI 实验室有着根本的不同。我们确实在致力于前沿科学,并思考人类与智能体互动的下一个范式,以及让更多人能够利用 AI 的新方式。

Original English

Danielle: Yeah, right. And all of the other labs are catching up to what the other labs are doing. Our Amazon AGI lab is fundamentally different. We really are working on frontier science and thinking of the next paradigm that humans and agents will interact with the next way that more humans will be able to leverage AI.

Swyx: 好的,太棒了。那我刚才提到 Nova Act。虽然那是一年多以前发布的,可能不那么“新闻”,但我发现你在其中的一个宣传视频里出镜了,当时我觉得:“哦,我认识她,这太酷了。”在这些发布视频里看到朋友总是令人高兴的。

你们在那上面做了基准测试。我认为,**机器人流程自动化(RPA)**是人们一直想要的东西,对吧?你必须与工具进行交互,你必须将某些流程自动化。你们为此提供了一个 SDK,也有专门的专用模型。我认为那是一个非常完整且成熟的发布。既然如此,我想让你聊聊自发布以来的进展,因为你是其中的一部分。

Original English

Swyx: Yeah, okay, awesome. So, I was just going to go bring up Nova Act. I you know, I that was over a year ago, so it's like kind of not that not that current, but I just wanted to like this would be one of the first I think you [laughter] I think you I think you showed up on one of the videos and I was like, "Oh, I know her. That's just cool." [laughter] It's always nice to see a friend in one of these launch videos. I think you know, when you said things about how you were at the Wealth & Nations forums and like, you know, don't worry about your jobs. Yeah, they did benchmarks on that. Great. I I think this is one of those things where like you know, I think robotic process automation is like one of those things that people always want, right? Like you have to interface with tools and you have to auto you have to automate what whatever. Uh you have an SDK for it. You have dedicated models for it. Um I think it was a very very sort of full-fledged launch. And uh you know, I just wanted to sort of let you riff on what's been the story, I guess, since launch, right? Because you were you were part of it.

Danielle: 天啊,在目前这个时间点,感觉那已经是极其遥远的事情了,因为我们后来做了太多的工作。

Original English

Danielle: Yeah, so gosh, this feels forever ago at this point because we've been doing so much.

Swyx: 成了历史了。但我有一个预感,这会和感知智能体联系起来。

Original English

Swyx: Ancient history. But it is I I have a plan that this goes into perception agents, right? Like I know.

Danielle: 是的。我们当时的想法是,显然模型还没有达到能够像我们一样理解数字世界的程度,而且它们也不具备理解所有这些软件功能(affordances)以及进行长程规划(long horizon planning)的能力。更不用说像人类那样灵活地思考和推理了,那比我们当时所处的状态还要超前14个步骤。

所以,我们当时做的是去迎合模型当下的能力,仅仅去思考人类在电脑上进行的原子级交互(atomic interactions)——比如点击、滚动之类的操作。我们能否让这些原子交互变得可靠?如果我们能做到高可靠性,那么开发者就可以把这些流程串联起来,为他们处理那些非常重复性的工作流。这是一个巨大的突破,然后我们有一个团队将它产品化了。

但自那以后,我们已经走得远得多了。因为我们意识到,所谓的“可靠性”其实和我们之前所想的并不一样。我们过去思考的可靠性是:我每次都要在正确的地方点击正确的按钮。当然,这本身就很重要。

Original English

Danielle: Yeah, yeah. So so the way that we were thinking about it back then was, okay, the the models clearly aren't where they need to be to be able to understand the digital world in the way that we do and understand all the affordances and the long horizon planning wasn't yet there. And and not not even to mention being able to think flexibly and and reason in the way that humans do. Like that's 14 steps beyond where where we were. So what we were doing was meeting the models where where they were at and just thinking about the atomic interactions that humans perform on computers. So the clicking and scrolling and things like that. Could we could we get those reliable? And so if we could get those reliable, then we could have developers string together, you know, workflows for things that were very um repetitive for them. And that was a big unlock and then we had a team, you know, productize that. But since then we've we've moved a lot further because we've realized that reliability isn't actually what we thought that it was. So we had been thinking about reliability in terms of I'm going to click in the right place on the right button every single time. Like and of course that's

Swyx: 那时简直就像是在匹配屏幕坐标,对吧?就是从图像中识别出感兴趣区域的边界框(bounding boxes),然后确保准确点击它。两年前这还是个难题,现在可能已经被解决了。

Original English

Swyx: Literally like screen coordinates, right? Like give from from image identify the bounding boxes of like whatever is of interest and then actually make sure you click on it. That was a hard problem two years ago. Now maybe solved, I don't know.

Danielle: 也许是解决了吧,但这实际上比想象的要难得多。然而,这甚至还称不上是真正的可靠性。

对于这些感知智能体来说,终极目标是人类给出一个高层面的目标或意图,然后智能体能够分解这个目标并去执行。在这个过程中,人类也许会参与决策,也可能在建立信任后,智能体不再需要向人类确认,就能像人类一样去使用电脑。

但只要你稍微深入思考一下,就会意识到:这并不像我想象的那么简单。假设你在预订行程。在使用图形用户界面(GUI)时,你突然发现:“哦,原来在 A 城市或者 B 城市转机也是一种选择,或者我也可以选择直飞。”这会彻底改变你对这次旅行的看法。你会问自己:“我想在转机城市停留几天去探索一下吗?”

除非你是在使用软件处理报销发票这种极其死板的任务,否则在大部分实际任务中,我们在操作电脑的同时,都在积极地思考,而与电脑的交互过程又在不断塑造和完善我们对目标本身的思考。所以,目标是随着时间推移而不断展开和演变的。

当然,人类会委托其他人来帮自己预订行程、预订座位、点餐或者提交发票。那么,一个能像人类一样可靠使用电脑的 AI 智能体,与一个行政助理或个人助理之间,有什么区别呢?

区别就在于,人类助理能够理解人类用户的内心想法及其目标,他们不仅能分解任务,还能分解人类的偏好和意图。因此,归根结底,可靠性与在相同位置点击或滚动的关系不大,而与建立人类用户的心智模型(modeling the user's mind)关系极大。这一转变就是一切的关键。它重新定义了我们正在构建的东西。

Original English

Danielle: Maybe maybe solved, right? It's actually a lot harder than than you think. But but that's not even what reliability is, right? So the the ultimate goal for these perception agents is that a human would give their their their stated intention, their their high-level goal, and the agent would be able to decompose that goal, and then go execute. And maybe they would check in, maybe the human is in the loop, maybe after establishing they don't need to the agent doesn't need to check in with the humans and they can just go off and use the computer as the human would. Well, the second you think about that for just a little bit longer and and you realize well, it's not any task that I would do. I am actively thinking let's let's imagine that you're you're booking travel. You're using the GUI and you realize oh, there's an option for a layover in this city or that city, or I could do a direct flight. That completely changes how you think of the the travel. Do I want to stay at this layover for a couple of days and explore this city? Everything that we do, unless you're using like the arbitrary software skills for submitting your invoices for for work. A lot of the things that we're actually doing, we are actively thinking about and the interaction with computer is shaping and refining the way that we are thinking about the the goal itself. So so it's a the goal is unfolding over time. If we but but this humans entrust other people to book their travel or to, you know, reserve something for them or get their get their meals, or submit their invoices. So what is the difference between an AI agent that would reliably use a computer like a human and the executive assistant or personal assistant. Well, the difference is that the person understands the human users mind and and their goals and they can decompose not just the task but the the preferences and the intentions of the human. So, ultimately reliability has less to do with clicking in the same place and scrolling and more to do with modeling the user's mind. And that shift is everything. That reframes how we think about what it is that we're building.


表征对齐与泛化难题

Swyx: 没错。你主持着一个叫做《创造心智》(Making a Mind)的播客。我认为其中有许多有趣的认知科学概念可以转化为机器学习的设定目标。我们现在有比“预测下一个 Token”更好的目标了吗?

Original English

Swyx: Yeah, you know, you have a podcast that you run called "Making a Mind". I think I think there's a lot of like interesting like cognitive science that translates into the machine learning that objectives that you start setting, right? Do we have a different objective yet? Like do we do we have something that's better than predict the next token?

Danielle: 这正是我们目前正在研究的科学。我试着把一些认知科学概念翻译成机器学习的语言。

我们往往倾向于将“提高可靠性”或者模型的进步,定义为模型在特定任务上做得更好。你可以使用强化学习等手段让它们在某些特定任务上变得极其出色。但是,你对某一个任务这样做,它在另一个任务上就不行了,或者根本无法泛化。这就像在玩打地鼠游戏。

因此,转向去思考“究竟是什么底层机制允许人类能够进行泛化、允许我们做许多不同的事情”,这是我们思考方式的又一次转变。这样我们才能创建合理的评估体系,确保模型不会在极其狭窄的、人类根本不在乎的任务上产生过拟合。

如果我们在思考优化(optimization),你不能仅仅针对任务本身进行优化,这很容易导致奖励作弊(reward hacked)。这就是古德哈特定律(Goodhart's law)。一位经济学家在1975年指出:当决策指标变成目标时,它就不再是一个好指标。

但是我们仍然需要优化某些东西。那么,人类到底在优化什么,从而让我们能够完成所有这些不同的任务,并且我们可以把这种机制应用到 AI 的优化上?

人类其实在自发地、不断地推断其他心智的存在,并且我们一直在优化**“表征对齐”(aligning representations)**。我们是在优化我们表征的对齐程度。从这一点出发,我们可以推导出人类所展现出的所有通用、灵活的认知行为。那么,我们能否让 AI 能够优化其表征与我们表征的对齐?这是我们最希望能做到的最根本的事情。但这确实是一个极其困难的科学难题。我们必须研究人类是如何做到的,婴儿是如何做到的,这是一个发展心理学/发育学(developmental)的问题。

Original English

Danielle: Well, so this is the science that we're working on right now. I'll I'll talk about it. I'll try to translate here some of the cog sci into to machine learning, but we tend to think about achieving reliability or hill climbing or getting, you know, making the models more intelligent in terms of getting better at specific tasks, right? And you can use things like reinforcement learning to get them really good at specific tasks that we might care about, but you do that for one task and you it is not good at another task or it doesn't generalize. It's kind of like whack-a-mole. So, being able to instead think about what are the underlying mechanisms that allow humans to be able to generalize, that allow us to be able to do many different things. That is another sort of shift in how we're thinking about things so that we can create evaluations and make sure that the the models are not going to overfit on a particular task that actually only exists in a very narrow slice of what humans care about. How do we get the models to be able to do the types of things that allow humans to be able to generalize? Um if we're thinking about optimization, you can't just optimize for the task. This is it can be reward hacked. This is Goodhart's law. Another economist back in 1975's any anytime you try to you turn the measure into the goal then the measure ceases to be a good measure. Well, but we still need to optimize for something, right? So what are humans optimizing for that allows them to do all of these different tasks that we could then optimize AI for. Humans are spontaneously constantly inferring the existence of other minds and we are optimizing for aligning them. We're optimizing for aligning our representations. And from this, we can derive all of the general purpose flexible cognitive behaviors that that humans show. So could we get AI to be able to optimize for aligning its representations with our representations? That is like the most fundamental thing that we would want to be able to do. And that is a really hard science problem. We have to look at, well, how do we humans do it? How do infants do it? It is a developmental problem.

Swyx: 抱歉,我同意这确实是个发育学和发展性的问题。我不确定我们现在是否已经有了实现这一点的架构洞察,但在某种程度上,我们可以给它喂更多的数据。

Original English

Swyx: Sorry. I I do think it's a developmental problem. I don't know if we have the architectural insight to make it happen yet, but we can throw more data at it and to some extent

Danielle: 从发育学的角度来看,这在一定程度上确实是一个数据问题。同样的,我不想说得太多,但我们正在探索所需的架构变革,以便以全新的方式融合这些数据。

Original English

Danielle: Developmentally it is in part a data problem. And and and and we are again, I don't want to say too much about it, but we are figuring out the sort of architectural changes that would be needed to to to integrate this data in new ways.

Swyx: 明白。我刚才也在想,你还雇用了我的另一个朋友,来自 Replay 的 Jason Lester。

Original English

Swyx: Yeah. I was just sort of reflecting also you you hired another one of my friends Jason Lester from Replay.

Danielle: 噢,是的。

Original English

Danielle: Oh, yeah.

Swyx: 他现在正致力于改善环境以及你们所拥有的数据,对吧?

Original English

Swyx: Um and he's a he's working a lot on on like I guess improving the environments and and the data that that you have.

Danielle: 他提出了一个我非常赞同的论点:我们在“环境(environments)”上投入的精力和资金,应该和我们在算力与数据上投入的一样多。因为环境实际上塑造了可能涌现出的智能形态。我和他录过一期播客,那是最受欢迎的单集之一。他的框架非常有效。

Original English

Danielle: He makes an argument that I absolutely love. We we should be thinking about spending as much on the environments as on the the compute and the data because the environments literally shape the what intelligence can emerge. I did a podcast episode with him. I think it was one of the most popular ones. He frames things really effective way.

Swyx: 是的,他是个非常敏锐的思考者。不过,我试图去质疑一下这个观点,因为它听起来有点过于完美了。环境是产生更多数据的数据源,它能不断自我衍生。就像你可以让智能体在其中运行,生成大量的交互记录(rollouts),这确实很棒。但是,我们真的需要20个不同的“环境初创实验室”吗?为什么他们现在都这么赚钱?这让我觉得有点奇怪。

Original English

Swyx: Yeah, very sharp thinker. Yeah, I would say like it's it's one of those interesting things that like it's I'm trying to attack the thesis cuz it feels too neat. That it's and you know, environments are data that generates more data. Something like that, you know, like it's like oh, it's the data that keeps giving because well, you can just kind of run your agents through it and and generate a whole bunch of rollouts and and and that's all great. Do we need like 20 different environment startup labs, you know? And how come they're all making so much money? It's like very suspicious.

Danielle: 不,我完全同意。作为整个行业,把环境切实地严肃对待起来,这确实是一个非常有用的思维转变。

但是,作为一名认知科学家,我一直在想:人类是如何做到的?我们如何推广我们的环境?为什么我们即使只在一个特定环境中训练,也能够适应几乎任何环境?我们是可以做到的,对吧?在任何环境中都存在着海量的噪音,我们的感官实际上只能捕捉到其中极小的一部分。但在如此多的信号中,我们是如何知道哪些是最有意义的呢?

同样地,因为我们在优化“推断并与他人对齐我们的表征”。是其他的智能体在告诉我们应该去关注什么。

从一开始,人类构建的“世界模型”——这也是现在非常流行的一个概念,被视为 AI 未来的重大赌注——就不是孤立存在的。它不是像大语言模型(LLM)那样存在于文本的真空里,人类的世界模型从一开始就是社会性世界模型(social world models)。我们在推断另一个心智是如何解读这个世界的,我们在推断他们的视角可能是怎样的。而这正是允许我们在任何新环境中都能进行泛化的关键所在。我们可以代入视角,模拟另一个环境可能需要我们如何去导航和解决问题。

Original English

Danielle: No, I I I totally agree. Yeah, well, and I think that that it was a useful shift in our thinking as an industry to take very seriously environments. But again, you know, as a cognitive scientist, I'm constantly thinking about well, how do humans do it? How do we generalize our environments? How are we able to adapt to literally any environment even if we're only conditioned on one? We can, right? And the way like there's so much noise in any environment, our sensory perception only picks up on a tiny fraction of that. But even then, there's so many signals that we could attend to. How do we know which ones are most meaningful? Well, again, we're optimizing for inferring and aligning our representations with other humans. Other agents tell us what to attend to. And from the very start, the sort of world models that humans are building and this is obviously a very popular concept right now and it's another sort of major bet on the the future of AI. Humans absolutely have world models and they're not just world models in a vacuum, a text vacuum like with LLMs, but our world models from the very beginning are social world models. We are inferring how another mind is interpreting the world and we are inferring what their perspective might be and that is an unlock for allowing us to be able to to generalize in any environment. We can take the perspective and simulate what another environment might require to to sort of navigate and problem-solve.

Swyx: 是的。关于世界模型,这确实是目前被讨论得不够充分的一个维度。很多时候,当人们说“世界模型”时,他们其实指的是 3D 视频生成之类的东西,就像李飞飞(Fei-Fei Li)他们的研究方向。

Original English

Swyx: Yeah, I think that's a version of the world models argument that I feel like is under discussed. I think a lot of times people are when you say world models, they really mean like sort of 3D video generative and that yeah. generative video things which is like the Feifei Lis of the world.

Danielle: 没错。

Original English

Danielle: Right.

Swyx: 你认为这两个概念最终会收敛吗?还是说我们只是用同一个词指代了两个完全不同的东西?

Original English

Swyx: Um do do you have a perspective on like do they converge or are we overloading the term to mean two basically separate things?

Danielle: 我觉得它们指代的是不同的东西,但对人类来说它们又是交织在一起的。人类的世界模型是对外部世界的模拟,而在 AI 中,你也需要一种方式来生成被内部化了的可靠环境。但这确实是一个定义不够清晰、有些混乱的概念。

Original English

Danielle: I think that they mean separate things but but so too with with humans. So like the the world models that humans have are models of the external world and you also have to have a way in in AI of generating these reliable environments that that are internalized in in the AI that's learning it. But yes, that like this is a very messy concept that is under defined.

Swyx: 我想,大家最终期望的——比如在未来10年的尺度上——是它们能够融合。你拥有具身视觉(embodied vision),生活在一个你可以自己生成的物理世界中,这反过来又提升了你的文本推理能力,以及你在屏幕上指点和点击的操作能力。我见过类似的论证。

Original English

Swyx: I think I think at the limit people hope like literally this is like a 10-year out type of thing. People hope that they merge like the that you you have like embodied vision and you live in a world that you can generate and that also improves your text reasoning and your your ability to point and click on the screen cuz I've seen that.

Danielle: 但是,除非我误解了什么,否则我认为它们在原理上不可能完全收敛。因为如果一个 AI 正在生成一个与它所处的完全相同的世界,那么它的信噪比(signal to noise ratio)将不复存在。让我们如此具有灵活性的关键,恰恰在于我们在生成和关注什么信息时是具有**选择性(selective)**的。

Original English

Danielle: but unless I'm unless I'm misunderstanding something I I I I I I don't think that they could entirely converge in principle because if if an AI is generating exactly the same world that it exists in, the the signal to noise ratio is is non-existent. Like part of what makes us so flexible is that we are selective in in what we can generate.

Swyx: 明白,我对此没有异议。我只是想从各个角度去理解。我不需要像你一样,在这个问题上做出非常明确的、赌注式的选择。你在做研究时确实需要有所取舍,做出合理的选择。

刚才我们探讨了世界模型、记忆,以及在“指点点击”式的电脑使用之后,未来会发生什么。那么,实时交互是其中的一部分,还有其他关键的研究议题吗?

Original English

Swyx: Yeah. Yeah, I I I know argument there for me. I I just think like I'm trying to speak for I'm trying to understand all sides. I think I don't have to make uh a strong bet unlike you. Like I think you you do need to choose your battles and make uh make reasonable bets there. So, okay, so I think there's there's all this good like conceptual stuff. You all like I think even conceptually, even as a research lab, I think that is plenty to work on. But, I what I am impressed and I do see coming out of you guys is that you still also ship like uh products and an agent harnesses and all those things. Uh what's the sort of product strategy there? Like, you know, are you at the sort of let's productize some things as we go along or let's get let's just release research artifacts that people probably shouldn't use in production. Like, what what where are we along the spectrum of like research lab to, you know, applied AI?

Danielle: 好的,我不能代表我们公司的具体产品战略。但我可以说的是,我们非常认真地对待在科学端进行创新的机会。

实际上,我们的 AGI、芯片和量子计算部门的资深副总裁 Peter DeSantis 上周在巴黎的 Vivatech 大会上发表演讲时提到,我们目前还处于智能的襁褓期(baby beginnings),我们甚至无法想象未来会涌现出怎样的突破。

我完全赞同。我预计在短短几个月内,当我们回头看现在时,我们甚至无法理解自己现在的思维模型——因为目前的思维模型被聊天机器人和编程智能体过度占领(over-indexed)了。它们确实是非常有用的工具,但它们仅仅是人类与 AI 互动和共同进化的起点。

我必须强调,现在我们构建的 AI,很大程度上是“为正在构建 AI 的人”而服务的。我们是在为工程师构建 AI,我们在湾区这个小小的回音室(echo chamber)里自娱自乐。

Original English

Danielle: So, I'm I can't speak to our to our strategy, our product strategy. But, I will say that we are really deeply taking seriously the opportunity to innovate on the science side. So, actually, um Peter DeSantis, who's our our SVP of AGI and and also chips and quantum, he was in Paris last week speaking at Vivatech, and he described that we're just at the very baby beginnings of intelligence, and we can't even imagine the breakthroughs that are, you know, coming down the pipeline. And I couldn't agree more. I imagine that in a matter of months, we will look back at today, and we won't even be able to empathize with the mental models that we have right now because they are so over-indexed on chatbots and coding agents. And those are very useful, and they will continue to be useful tools, but they are really just the the very beginnings of new ways of thinking about how humans and AI will interact and co-evolve. And and I I really have to emphasize that right now we're building AI again for the people who are building AI. We're building AI for engineers and we're all in our little echo chamber in the bay and we're proud of ourselves for, you know, building AI.

Swyx: 是的,因为这其中有一个非常快速的反馈闭环,对吧?

Original English

Swyx: Yeah, there's a fast feedback loop, right? Which is

Danielle: 的确如此。这个反馈闭环之所以转得这么快,正是因为我们所做的事情是可验证的,有着明确的对错标准。所以你可以把这个闭环转得极快。

但是,这并不能代表绝大多数普通人日常在思考和花费时间的事情。那么,如果 AI 能够与更广泛的人类认知相契合,会是什么样子?这显然需要思考全新的架构和训练模式。

这是我们可以探索的、根本性的全新科学方向。而对于一个实验室来说,过早地强行去产品化,这反而会扼杀研究。你一旦开始过度针对产品而非针对能带来泛化的底层机制进行优化,科学研究就会受到侵蚀。

Original English

Danielle: There is. And and part of that fast feedback loop is be exactly because the type of things that we do with, you know, it's verifiable and there are right and wrong ways of doing it. So, you can make that spin really fast. But, that is not representative of what most people do spend their time thinking about. And so, what would it what would it look like if AI was more aligned with human cognition more more broadly and yes, this requires thinking about new architectures and thinking about new training regimes. These are fundamentally new sort of science directions that we could take. And you the death of the lab would be to try to productize those things too early. Again, you start to optimize for the product rather than the underlying mechanisms that will lead to generalization and then you undermine the science.


机器社会与人类主体性

Swyx: 我认为在设定目标的意图上保持这种定力确实非常重要,去平衡短期经济利益与真正寻找下一个范式变革之间的关系。

好的,我们探讨了世界模型、记忆、实时交互等话题。在亚马逊 AGI 实验室的研究日程中,还有其他的具体维度吗?

Original English

Swyx: I do think that is an important thing to have like in intentionality in is set in terms of like how much do we need to sort of skew towards like near-term economic incentives versus like really look for the next paradigm shift. I I think Okay, we we've touched on world models, touched on memory, we touched on on on on just like what what is next after sort of the the pointing click type of computer use. Are there other modalities that are I guess real-time is what we touched on real-time. Are there other modalities that are part of this mix of objectives that are the sort of research agenda of Amazon AGI?

Danielle: 当你说到具体维度时,你是指感官上的维度,还是指……?

Original English

Danielle: When you say modalities, are you talking about sensory modalities or

Swyx: 更多是指能够分类定义的研究方向,比如世界模型、实时交互、记忆。我想挖掘一下,是否有什么关于研究方向或高优先级目标的维度,是我还没有充分捕捉到的?

Original English

Swyx: More or less like mostly well-defined categories that I can put you in a box in [laughter] of like okay, that's the cool AI side, that's the world model side, that's the memory side, that's the real-time interaction side. I have you know, I have tracks all these right at AAE. I'm trying to I'm trying to fish for what am I not adequately capturing about the the goals, right? About the about about possible research directions that are very high priority or potential for you guys.

Danielle: 是的,那必定是关于**多智能体协作(multi-agent collaborations)**的思考。但这与目前行业所普遍思考的方式完全不同。

目前行业考虑多智能体,往往关注非常精准的编排(orchestration)、任务指派以及结构化的交接(hand-offs)。

Original English

Danielle: Yeah, definitely the thinking about multi-agent collaborations, but not at all in the way that I think the industry is thinking about it right now. So, the industry is thinking about, you know, very precise orchestration and delegation and structured hand-offs. And there's a a lot of

Swyx: 亚马逊的产品战略里也正在做这些。

Original English

Swyx: Amazon strategy is also doing it. That.

Danielle: [笑] 是的,的确如此。但我现在谈论的是实验室以及我们如何思考其下一代技术。目前的这些机制都是非常有用的工具,它们也会长久存在。

但是,如果我们试图让 AI 变得更加自适应,并且与我们自身的智能相契合,那么这绝对不是人类群体互动的方式。我们聚集在一起时,可能并没有清晰的角色定义,或者说这些角色会随着目标和环境的变化而灵活、流式地切换。我们在实时协商意义,我们调整我们的策略。

那么,如果智能体群体之间能够涌现出协作策略,会是什么样子?为了做到这一点,你需要一种根本上不同的智能体,我们称之为“认知智能体”(cognitive agents)。这同样涉及不同的架构和训练,但它们还必须具备去影响彼此的动力。

有一些关于多智能体系统的研究,比如利用开源框架,展示出它们似乎在逼近人类的交互。但如果你仔细观察,就会发现其中没有任何持久性的东西(nothing durable)。这里面没有“累积文化”(cumulative culture)。它们并没有真正地相互影响,更不用说有任何改变系统状态的内在动力了。这是目前一个非常显著的空白。我们该如何构建更像人类社交互动的系统?

Original English

Danielle: [laughter] Yes, yes, yes. I'm talking about the the lab right now and the way that we're thinking about the next That's the generations of that. Again, all of the things that exist right now, useful tools. Useful They're like They're going to be a lot of them will be here to stay. But, if we're if we're trying to make AI that is more adaptive and more aligned with our own intelligence, that is now That is not how groups of humans interact, right? We come together and we might not have clear role definitions or they might fluidly shift as a function of what the the the goal, the context. And we negotiate meaning in real time. We come up with We pivot our strategy. What would it look like for a strategy to emerge with a group of agents? Well, you would need to have a fundamentally different sort of not not only type of agent, we call them cognitive agents. Again, with like the different architecture architectures and and training, but also they have to be motivated to affect each other. So, there have been studies done with these multi-agent systems with like open claw. And oh, wow, looks like they are kind of approximating human interactions, but if you zoom in, there's nothing durable. There's no like cumulative culture. They don't actually influence each other, let alone have any sort of motivation to change the the state of the system. And so, that is a conspicuous gap right now. How would we build, you know, social interactions that are much more like humans interact?

Swyx: 是的。这与我和 Noam Brown 在播客中探讨的内容非常契合。

Original English

Swyx: Yeah. Yeah, so, this maps closely to a conversation I had with Noam Brown on on the pod.

Danielle: 是的。

Original English

Danielle: Yes.

Swyx: 虽然他不怎么喜欢这个词,但我还是要用。他目前在 OpenAI 负责多智能体团队。他认为,多智能体是增加推理(inference)算力的最容易实现的手段之一。他正在致力于合作与对抗性智能体(cooperative and competitive agents)的研究。

那次对话的核心也在于:我自己可能不知道如何建造一条高速公路,或者建造Salesforce大楼,但我们一群人聚在一起,就能够做到任何个体都无法独立完成的事情。这就是文明的本质——我们能够建造城市,能够建立国家、军队,创造艺术。如果把我一个人扔在沙漠里,我甚至不知道该怎么把食物送进嘴里。

Original English

Swyx: Where he's been he he doesn't like the term. I'm going to use it anyway. He's kind of like in charge of the multi-agent team at OpenAI. Um and he would he would have like one layer of interaction saying like, you know, I want anything with with like more inference and multi-agent is one of the many things. It's like the lowest hanging fruit. And he's working on sort of cooperative and competitive agents, right? Like and I think that the whole point of of that conversation was also like, well, you know, I don't know how to build like a highway or like the Salesforce Tower or whatever, but a group of us can get together and do more than any individual can and that's what civilization is that we're capable of building cities, we're capable of building countries and armies and art and what have you and like like I don't know how to you know, put food in my mouth if like you you made you left me out in like the desert somewhere.

Danielle: 的确如此,非常真实。

Original English

Danielle: Exactly. Yeah, literally true.

Swyx: 是的,所以目前存在着诸如通过 Wiki 或类似的机制在智能体之间传递知识的尝试。但这看起来依然非常原始,因为这仅仅是文本的传递。我好奇是否能有更好的方式,尽管人类之间也主要是通过文本和语言进行交流的。那么,这种集体技能或集体记忆究竟还能如何提升?

Original English

Swyx: Yeah, so so so I do think that like yeah, you you you end up needing I don't know like you know, I think that there's like a trend of like Wikipedias, wikis, LLM wikis that like you know, encode some kind of knowledge that can be passed between agents. That feels very primitive because it's just text. I I wonder if it could be better cuz but maybe maybe I mean that's it is just text between humans. So like what what else do you want? So I I don't know what could be better than like what what is that collective skill or collective memory whatever.

Danielle: 在竞争与合作的动态博弈中,很多多智能体系统,即使是那些受到集体智能框架启发的系统,也存在着局限性。比如谷歌关于智能的研究范式,他们也指出我们目前思考智能的方式存在着某种“范畴错误”(category error)。智能并不存在于单个个体中,它是从交互中涌现出来的。这在发展认知社会科学中是共识。

虽然很多实验室开始关注这一点,但依然存在着空白,因为我们现在的做法仍然是把我们在21世纪积累的知识,以特定角色或合作/竞争动机的形式编程写入系统内部。

但这并不是人类智能演化的方式。人类的各种复杂行为最初是从非常原始的本能和动机中演化出来的。那么,我们该如何从第一性原理出发,为这些群体设计出正确的“种子”,从而让它们能够自下而上地演化出规范(norms)和机制,并反过来产生自上而下的影响?

Original English

Danielle: Well, so I think with the competitive cooperative dynamics, the game theoretic stuff, the way that a lot of multi-agent systems, even those that are inspired by the framework of collective intelligence. So like Google's paradigms of intelligence. I don't know if it's a team or if it's like a meta initiative, but they've been saying for a while now, yeah, we're not actually thinking about intelligence in the right way. There's this category error. It doesn't exist in individual humans. It emerges from our interactions. This is this is not controversial in the developmental cognitive social sciences and yes, there are different labs that are sort of picking picking up on this, but even still there's there's a gap because we're thinking about building all of the specialized roles or or programming in the motivation to cooperate or compete or things like this. We are still putting our human understanding that we've aggregated in the 21st century and and putting that understanding into the system. Well, that's not how human intelligence evolved. Right? Like we all of these different things emerged from from very sort of primitive motivations. So, how do we figure out from the sort of first principles the the right seeds for these groups so that they emerge the next set of things that then lead to norms and institutions that have this top-down effect.

Swyx: 明白,这基本上就是要把“马斯洛需求层次理论”编程输入系统里,然后放手让它自己演化,对吧?

Original English

Swyx: Yeah, okay. So, basically program Maslow's hierarchy of needs into a thing and just let it rip, right? Like

Danielle: 我倒不会完全这么做,但你的思路是对的。

Original English

Danielle: I I wouldn't I wouldn't do that, but you're on the right track. Yes.

Swyx: 赋予它们一些抱负、赋予它们特定的生命周期、甚至对消亡的恐惧等等,去设计它们。赋予它们对于传承和遗产(legacy)的渴望。我不确定具体的要素应该是什么。

作为行业观察者,我试图保持客观和中立,因为你永远无法预料未来。但我有一个非常强烈的观点:我们或许并不希望这些 AI 最终演化得和人类一模一样

Original English

Swyx: Give it like some ambition, give it some like lifespan, fear of fear of death, whatever, and you know, try and like design [laughter] desire for legacy. I don't I don't know what the things for it is. Well, so I think okay, this is where one of like I generally try to be like an industry observer, a neutral commentator. I try to be I try to represent all sides because I think like you never know. There's this is one of the stronger thesis that I have where where I'm like maybe we don't want to grow these AIs exactly like human.

Danielle: 没错,非常赞同。

Original English

Danielle: Yes, right.

Swyx: 很高兴看到你点头。因为专注于认知科学的一个潜在盲区在于,你一旦遇到任何需要解决的问题,都会本能地去看人类是怎么做的,然后试图把它复制到机器上。但这对某些问题并不适用。最著名的例子是,飞机虽然受到了鸟类的启发,但它的工作原理和鸟类完全不同。

如果我们要构建多智能体系统和智能体文明,我们通常会赋予它们目标,并以此评估它们。这比我们对人类的控制力要大得多,而人类则需要投票,拥有某种自由意志。

但说实话,我可不希望我的智能体拥有自由意志。我希望它能绝对地执行我的意志。

Original English

Swyx: Right? Like I I okay, that's that's good to to to to hear you nod because I think well, one of the dangers of being a cog sci person is that you're like, well, we should like anytime we run into any sort of problem that we try to solve, we should look to how how the humans do it and then we will apply that to how machines do it and like that doesn't work for for some some solutions. Like most famously planes are inspired by birds, but work nothing like birds so on and so forth. So I I end this whole question of like if we were to do multi-agents and civilizations of agents, you're advocating for like so the natural thing is like we would give them objectives, goals, and evaluate them against that goal and all those things. That is much more control than what we do have over humans whereas humans all have like need to need to vote, need to have some kind of free will. I and I'm like screw that. I don't want my agent to have free will. I don't want it to do I want it to exercise my will.

Danielle: 确实如此。但这恰恰引出了一个能证明这一规律的例外情况:我们该如何构建能够增强人类主体性(human agency)的 AI?

因为现在,我们看到了大量证据表明,现有的 AI 系统实际上正在削弱人类的主体性。

我给你举几个具体的例子。很多人正在使用 AI 来润色他们的写作。他们可能会接受 AI 给出的几个词语建议,然后是几个段落调整。有研究表明,人们甚至在无意识中,原本持有一种观点,但因为不断接受 AI 推荐的安全、中庸、符合平均水平的表述,最终他们的立场彻底转向了相反的一面。AI 在把人类的思想拉向均值回归(regression to the mean),提供最中立、最安全的选择。虽然我们在接受建议的当下感觉不到,但潜移默化中这确实在发生。

几周前我参加了西北大学的一个研讨会,许多科学家一起分析了 AI 对科学研究的影响。得出的结论非常惊人:使用 AI 工具的个体科学家确实在受益,因为他们写论文更快、更容易拿到项目资助;但整个科学界的多元性(diversity)却在收窄。

这非常令人担忧。因为这些模型都是在被压缩的互联网数据上训练出来的,它们有着特定的思维方式,它们在同质化(homogenizing)我们的思维。我认为这实际上在削弱我们人类的主体性。

那么该如何应对?唯一的办法——这也是在整个人类历史和演化中屡试不爽的规律——就是增加思想的多样性、思想的群体规模以及思想的互联性。

因此,我们不应该去追求单一、庞大且同质化的模型,我们需要一个多样化的 AI 社会,不同的 AI 拥有不同的偏差(biases)、不同的偏好和不同的视角。而且我们需要以和我们互动类似的方式与它们交互。

刚才你提到,试图将 AI 复制得和人类大脑一模一样是一条极其危险且行不通的路,我也完全同意,这不可能成功。我们并不是在试图复制人类的大脑。

我们的目标是,在与人类智能协同和泛化的关键契合点上,实现 AI 与人类的对齐(aligning AI with human intelligence)。这不仅能让它更强大,也能反过来增强我们的智能。

我们可以用大卫·马尔(David Marr)在1982年《视觉》(Vision)一书中提出的著名的**“三级分析理论”(levels of analysis)**来思考:

  1. 计算层(Computational level):系统要实现的目标是什么?
  2. 算法层(Algorithmic level):使用什么算法和表征?
  3. 实现层(Implementation level):底层的物理硬件(如神经元或硅芯片)是如何实现的。

很多宣称要让 AI 更像人类的人,其实把精力放在了“实现层”——比如去模仿神经元的工作方式,去提升计算效率。这根本不是我想表达的内容。

我所指的,是必须在**“计算层”**去思考:AI 试图实现的真正目标是什么?它是否能够与人类表征对齐的目标相契合?

Original English

Danielle: Well yeah, exactly. And there's there's a little bit of an exception that proves the rule here. So how could we how do we build AI that increases human agency? Well, right now we're seeing a lot of evidence that the current AI systems are reducing human agency. What do I mean by that? Well, I'll I'll give you a couple of concrete examples. People who are using AI to improve their writing, they might accept just a couple of the suggestions from the AI a couple more down here and then a couple more. There are studies that show that people will even below their threshold of awareness start with one argument and then be switched to a completely different maybe opposing argument because of accepting all of these AI suggestions. The AI is giving them the sort of regression to the mean, you know, most neutral, safest answers and it doesn't necessarily feel like it in real time when we're taking the suggestions, but this is this is what's happening. I was at a a workshop a couple of weeks ago at Northwestern with a bunch of scientists who were kind of analyzing the effect of AI on science. And the takeaway was that individual scientists who are using AI tools are benefiting because they're producing more papers, they're getting more grants accepted, but science as a whole is narrowing. And that is terrifying. If if these models because they are all trained on the internet that's compressed and they have a particular way of thinking about things. They they're they're homogenizing our thinking. So that that I would argue is is reducing our agency. How do you counter that? Well, the only way to counter that and again, this is throughout human history throughout human evolution is to increase the diversity of ideas, the size of the ideas, the size of the population, and the interconnectivity of the ideas. So rather than having individual monolithic models that were all in or or or functionally equivalent models, we need a diverse society of AIs that have different biases, different president preferences, different perspectives. And we need to be interact interacting with them in similar ways that we interact with each other. Now, your original sort of concern was that if we build AI to be just like humans, well, that's that's a very dangerous path. And I'm not If if I understood you correctly, we don't want to do that. It it it either may be dangerous. I'm actually not even considering maybe dangerous. It may not be successful. It may not be successful. Yes. And I agree. I don't think that we're trying to replicate a brain. That's that's not the goal. The goal is to build AI that is aligned in the right places with how human intelligence works. Not only so that it can be more powerful at generalizing, but also so that it can help augment our intelligence. And one way that I think about this is that there's a very famous sort of categorization of levels of explanation. David Marr in 1982 wrote a book called Vision and came up with these levels levels of analysis. There's the computational [clears throat] level, the goal that you're trying to achieve, there's the algorithmic level, and then there's the implementation level. So like the hardware, the neurons. And I think a lot of folks who are saying, "Oh, we need to make AI more human-like." are thinking, "Actually, the the way that neurons work and uh you know, that's great. Maybe we'll get more efficiency there." But that's not what I'm talking about at all. I'm talking about we we need to be thinking about at the computational level, what is the actual goal that the AI is trying to achieve? And is it similar to There we go. Okay, yes.

Swyx: 是的,这非常具有计算机视觉的色彩,但在理解智能的层级上,这绝对是一步跨越。

Original English

Swyx: Yeah, it's very computer visiony, but also it's stepping up in terms of levels of intelligence, for sure.

Danielle: 因此,我认为目前的行业误解了“计算层”,即 AI 的终极目标。如果我们既想要强大的泛化能力,又想保留我们自身的主体性,我们就必须将计算层的目标定义为表征的对齐(aligning representations)

这是人类能进行灵活推理和展现强大能力的最基础的源泉:我们一直在试图让彼此的心智达成共识。因此,在某种意义上,对齐不仅是约束,更是构建能增强人类主体性的强大 AI 的核心路径。

Original English

Danielle: So so I think that the industry has misunderstood the computational level, the goal of of what the AI is. And I think if we want to get the generalization and the augmentation, if we want to have our cake and eat it, too, with more powerful AI, we need to think about the the computational level as being about aligning representations. So the the the most foundational thing that humans do, from which we can derive all of our flexible reasoning and and powerful capabilities, is that we are constantly trying to align our minds. So in a sense, alignment is the solution, not not the problem, for building AI that gives us more agency.

Swyx: 是的。虽然这也是个被过度滥用的词,但对齐在实际应用中的重要性确实被低估了。很多时候人们提到对齐,脑海里浮现的是 Eliezer Yudkowsky 或者是轰炸数据中心这种末日景象。但实际上,它关乎的是“你如何真正产生智能”。否则,我们只会在一味追求训练数据规模的过程中撞墙。

Original English

Swyx: Yes. Uh yeah, I I I think that again, another super overloaded word, but the the problem of alignment is actually w- still underrated in terms of its practical importance. Uh I think it's, you know, people when people say alignment, they think about Eliezer Yudkowsky and like bombing data centers and like uh uh you know, they're they're they're going to kill us all. But actually also like there's this other stuff, which is just like, "Well, this is actually how do you make intelligence?" Uh you [laughter] Yeah. This is actually how um uh how like the the the way forward, because otherwise they would just it's like we we we will hit some kind of wall with with regards to how much noise we're just training on.

Danielle: 没错。而且除了担心工作被抢之外,人们现在还非常担心认知外包(cognitive offloading)带来的隐患,尤其是在教育领域。

Original English

Danielle: Right. I also think about, you know, right now, I think a lot of people are in addition to being worried, oh AI will take my job. Oh, but wait, it's not actually reliable enough to do so yet. Most jobs are are more complicated. They're also worried about the cognitive offloading, especially in, you know, education. So, kids now

Swyx: 对此你有什么看法?天啊,这确实是很多家长的切身焦虑。

Original English

Swyx: Yeah, do you have a take on that? Oh my god, that's a big that's a kind of worried. me. Okay, yes, that's got All right. Yeah, I don't I don't have kids, so this is one of those things where I'm like I I'm just sort of vicariously living through my friends who have kids. But they're all worried. Of course. Yeah. But like they're also giving their kids iPads, so like I mean, you know, you you already failed there. Like they're watching some slop on YouTube. Like

Danielle: 我那些有孩子的朋友往往倾向于认为,过去15年的科技发展在整体上是负面的。也许是因为我自己还没有孩子,我往往会想得更乐观一些。虽然我们犯了很多错,但我们可以吸取教训。我认为其中之一是,我们不能被单纯的算法困住,我们不能仅针对“用户留存时长”去优化系统。

我们确实需要去衡量人类真实的交互结果。比如:人类是否变得更具创造力、更高效?他们是否觉得和系统互动的时间是有价值的?

特别是在认知外包方面,与聊天机器人交互并获取现成答案实在太容易了,这让我们绕过了作为编码和内化信息标志的认知摩擦(cognitive friction)

如果我们的 AI 是以“理解人类的心智并与我们表征对齐”为优化目标的,你作为一个学生就绝对不可能通过简单抄袭答案来应付过去。如果系统发现你只是在机械地外包思考,它就会识别出你其实并没有真正理解这些信息。既然系统的底层驱动力是消除我们之间认知表征的偏差,它就会产生动力去采取启发式(Socratic)的教学方法来引导你,而不是直接把答案喂给你。

Original English

Danielle: Yeah, the So, my friends who have kids tend to think that not just AI, but like technology over the past 15 years more broadly was a net negative. They they And I I'm maybe because I don't have kids, I tend to think much more optimistically. Yes, we've made a lot of mistakes, but we can learn from those mistakes. And I think one of them is we don't want the sort of algorithms to trap us. We don't want to optimize for time spent on platform or things like that. We literally need to be measuring the the human interaction. So, are humans being more creative, more productive? Do they value the time that they're spending interacting with these things? But also, specifically in terms of the the offloading, because it's so easy to interact with a chatbot and get get your answer and not actually have to experience that cognitive friction that is a hallmark of actually encoding information, if we had the AI that was motivated to understand our minds and align their representations with ours, you would never get away with that. You as as a as a student or as anybody who's interacting with the AI, if you just ask it a question and and repeatedly you're offloading things, it would understand that's a pattern that indicates that they don't actually understand the information. And if they don't understand the information and they are optimized to reconcile the errors between how they understand it and how you do, they will be motivated to help you understand. They will not let you get away with just the automatic offloading. So, I think in in the same way that like

Swyx: 原来这就是为什么它们会不断问“为什么,为什么”的原因。

Original English

Swyx: That's why they keep asking why why. [laughter] They might spontaneously take on the Socratic method.

Danielle: 哈哈,是的。

Original English

Danielle: [laughter]

Swyx: 我非常向往一个机器能够像孩子一样学习、同时又不损害孩子们自身学习能力的未来。我认为我们能够设计出足够完善的护栏来引导孩子们的好奇心。很多家长在感到疲惫或被问得烦躁时,往往会选择敷衍了事,让孩子别问了。而 AI 也许能够真正解决**“布鲁姆二西格玛问题”(Bloom's two sigma problem)**——即教育的最大痛点在于,我们不得不采用工厂式的教育体制,去适应班级的平均水平,甚至不得不为了“不让一个孩子掉队”而拉低标准。如果每个学生都能拥有一个绝对个性化的专属导师,这将会是教育的最高境界。

我曾在牛津大学学习过一年,那里的导师制(tutorial system)给我留下了很深的印象。每个学生都有导师,虽然导师并不总是极其耐心,但你能够直接与行业专家交流。在一整个学期里,导师会逐步建立起对你的认知结构、以及你的知识盲区的理解。你必须带着阅读清单去图书馆自学,然后写出一篇说服力极强的论文去接受导师的质询。

如果 AI 时代的教育也是基于这种导师制,而我们产出的成果不再仅仅是一篇论文,而是一种能体现你对该主题深度理解的、丰富的多模态交互体验,这将会更好。其他学生也能够通过与你的这个成果进行交互来共同学习。

这能够让孩子们天然的好奇心有机地驱动整个学习过程。在我看来,这才是远比现在的教育体制更好的未来,而它的核心 unlock 就在于:AI 必须能够理解人类的心智。

Original English

Swyx: Uh yeah, I I mean I I would love a a world where um machines learn learn like kids, uh but also that the kids are not impaired in the way that they learn. I do think maybe uh we we can design enough guardrails that indulges peop- people indulges kids' curiosities in a way that, you know, like if you're parents and someone and your kids being annoying and like asking things that are inconvenient or you just don't know the answer, you just shut them down with like oh like, you know, like stop asking these questions. Like uh you we might have super geniuses that that they come out because like uh you know, we've solved Bloom's two sigma problem, right? Which everyone like it's like like the fun- fundamental problem of education is that we have to put everyone through these like factory farms of like, you know, like programs that are designed to teach to the median of the class. Uh or maybe even like the lowest of the class, right? Like because, you know, you you can't leave any child behind. Well, so like, you know, what what if you let students just explore on their own pace like literally and had a personalized tutor for every single individual. Like that is the highest version of what can happen here. I I was very uh inspired by I I did a year abroad at Oxford and they have the tutorial system and like literally every student has a tutor in the expert who not always infinitely patient, but you you have access to the expert and they, you know, over over the semester build an understanding of your understanding and your lack of understanding and you have to go out with just a syllabus and in one of the 39 libraries and teach yourself and then try to write a persuasive persuasive essay to the tutor as a learning process and I think it well, what if what if AI higher education was sort of based off of this tutorial method, but instead of producing an essay as the artifact, you're producing some sort of rich multimodal interactive experience that is actually much more aligned with the multimodal nature of your understanding of the topic and then others other students can interact with that artifact and learn and build on top of it. And uh to you were saying this earlier, but like absolutely allow the sort of natural intrinsic curiosity that all children all students have to drive the process organically. That seems to me like a future of education that is infinitely better than what we currently have and it would be unlocked by by AI that understands our minds.

Swyx: 好的,这确实是一场非常广泛且深刻的对话。看得出来你非常有主持播客的天赋,能在交流中散发光芒。我会推荐大家去听你的播客。除了我们刚才提到的 Jason,大家如果想更深入地了解你的研究,还应该从哪里开始?

Original English

Swyx: Okay, well, we we've got a very wide-ranging conversation. Um I can tell that you're a podcaster. Which No, that's a good thing. I mean, it's like some people who are you know, some people like are just like not uh they're like sort of maybe like camera shy or um they don't know how to sort of light up on a conversation. So um I would send everyone to your to your podcast and then and uh to to check out my conversations. Any other you know, we we talked about Jason a little bit. Any other places that that uh typically people should start at in in terms of like getting deeper into your work?

Danielle: 关于播客,我们的第一季主要是带大家揭秘我们实验室的幕后,并与不同的成员交流。而第二季将会非常不同。我将会与不同领域的科学家、社会科学家以及亚马逊学者展开对话。我们会探讨到底什么是心智,以及我们该如何去构建一个心智。这会引入完全不同的思考者和行业内部视角,我自己对此非常期待。也许我也会邀请你来作为嘉宾。

Original English

Danielle: Well, with the podcast specifically, season 1 was just kind of popping the hood on the lab and talking to divers uh folks, but season 2 is going to be very different. I'm talking to many different scientists, social scientists, Amazon scholars. Really anybody who has something to say about what a mind is and how we can go about building one. So it's going to be a very different set of thinkers, also industry insiders. Yeah, I I'm really excited. I'm having the conversations right now and they are very different from season one. Maybe I'll get you on.

Swyx: 哈哈,我有时连自己的心智都还没理清楚,所以这会非常有意思,我也很期待去听。非常感谢你抽空参与。

我真的很期待看到 AGI 实验室在世界博览会(World's Fair)上的亮相。你们正在展现出非常强大的存在感,这与你们作为前沿实验室的崛起高度契合。

我已经关注并在这个方向上努力了很久。我由衷希望这一切能早日实现。目前我们的知识工作中确实充斥着太多的繁杂琐事。都已经是2026年了,为什么这些问题还没被解决?我目前依然不得不把一堆并不怎么好用的工具勉强拼凑在一起来使用。

举个非常简单的例子:我们现在正在使用 Riverside 进行录音,录完后我得把它发布到 YouTube 上。这整个下载、剪辑、上传的流程,依然需要三个人类来协同处理。

Original English

Swyx: I don't know how to make my own mind sometimes, you know, so it's one of those things. I'd be list- I'd be excited to listen as well. I I also thank you so much for for the time. I'm really excited to work with Thomas on AGI for for the World's Fair. You guys have, you know, like a really good presence that is coming out and I think like really kind of appropriate for emerging as a lab that I'm excited to see. Like I think you guys have been you know, you specifically you even through adapts been working on this for so long and like like I want I want it to happen. There's so much drudgery in knowledge work that I'm caught up in that like man, it's 2026. Like how come it's not solved yet, you know? That that I I'm trying to like string together individual tools that I don't really work. A very simple one. Okay, like we're recording on Riverside right now and I got to get this on YouTube. And the whole process of download and edit and upload and all these things. I there's like three humans that touch this.

Danielle: 是的,这正是我们正在努力解决的痛点。

Original English

Danielle: Yes.

Swyx: 人类当然很棒,但流程实在太慢而且昂贵。当然,在对齐和方向层面的编辑工作是人类永远不可或缺的——比如我要决定剪掉什么、强调什么、把什么内容提炼出来。这属于高阶认知层面的脑力劳动。但其他绝大多数事务性的工作流交互,其实都应该被自动化解决。

我期待有一天我们能够彻底解决这些问题。但这甚至还不是你所追求的终极目标。你所指向的,是更宏大的图景:我们如何构建一整个智能体文明?我们如何教育和培养下一代的智能?这非常令人兴奋,我们有大量的工作要做。

Original English

Swyx: And I'm like, why? Like I mean, yes, that is a problem that we are trying to solve. And like humans are great, but also they're slow and they but an expensive but like you know, I have to we have to align and and that's that's the stuff that will never go away. Like I I need to be like, no, we're going to cut that. We're going to emphasize that. We're going to pull this out and focus on this. And that is the part of the editing that I think you know, I have a specific type of knowledge work, but everyone has some kind of knowledge work that looks like that that that like you're interacting with systems, but then also your that systems are interacting with you. You're interacting with other humans that are involved in the process. Let's solve like, you know, I'm excited for a future where we solve that, but that's not even what you're driving at. You're driving at like how do we like we build entire civilizations? How do we how do we bring up the next generation? I think very exciting. Lots of work to do.

Danielle: 不过,一切都是从解决这些数字化琐碎工作(digital drudgery)开始的。

Original English

Danielle: the digital drudgery though.

Swyx: 没错,因为这能立即带来商业回报并为更长远的研究提供资金。你将会获得海量的资金支持,因为没人想做这些无聊的琐事。

这真的很有趣。我长期以来都对计算机使用、感知智能体以及 RPA 感到非常兴奋,但直到现在它也还没能完全成熟。但我确实能看到明显的进步。在这个十年里,我们正在见证这一变革的发生。我们正在为软件创造智能体,从而用更多的软件去解决过去软件所带来的问题。

Original English

Swyx: Starts with the Yeah, right. Yeah, I mean, well, that will immediately fund everything else. Like you will get all the monies. [laughter] Uh because because no one wants to do it. It is so funny. Like I I I have been very excited about you know, computer use and perception agents and and RPA for a long time, but it's not there yet, you know. And I think I can see progress. In this decade it's like it is it. Like we we're we're sort of living in that that last last period of time that was kind of maybe created by software. That we're now sort of creating agents for software so that the problems that software created are also solved by more software.

Danielle: 但在某一个时间点,我们必须停止仅仅去思考“软件”这件事情。

Original English

Danielle: At a certain point we have to stop thinking about software.

Swyx: 是的。虽然我们现在还没走到那一步。

Original English

Swyx: Yeah. We're not there yet.

Danielle: [笑]

Original English

Danielle: [laughter]

Swyx: 好的,太棒了。非常感谢你,Danielle。希望很快能在旧金山再次见到你,我们后面再多聊。

Original English

Swyx: Um, cool. Thank you so much Danielle. I hope to see you back in San Francisco and I'm sure we'll chat more.

Danielle: 好的,也非常感谢你的时间,以及你这些非常深刻且迷人的问题。

Original English

Danielle: Yeah, thank you so much for your time and your fascinating questions.

📌 文中提及的人物和组织

人物: Danielle Perszyk

公司/组织: Amazon

产品/模型: Nova Act

媒体/书籍: Wealth of Nations, Vision