递归自我改进与人工智能的飞速发展:对超级智能的预测与风险分析 Dwarkesh Patel 2026-08-11

探讨递归自我改进与 AI 的飞速发展

Host: 今天我正在与 Ryan Greenblatt 交谈,他是 Redwood Research 的首席科学家,专注于人工智能的安全和安保技术研究。我想和你探讨一下递归自我改进(recursive self-improvement)。这个想法是,一旦我们构建出人类水平的智能,它们就会迅速飙升,产生数百亿个超级智能,其中每一个在各个领域都比顶尖的人类专家更有能力。这种情况是否会成为现实,可能是当今世界上最重要的问题。从历史上看,我一直对这种事情的发生持相当怀疑的态度,但是,你似乎认为这是貌似可行的,所以我想听听你的理由。

Original English

Host: Today I'm chatting with Ryan Greenblatt, who is the chief scientist at Redwood Research, where he focuses on technical AI safety and security work. I want to talk to you about recursive self-improvement. This is the idea that once we build human-level intelligences, they quickly slingshot towards tens of billions of superintelligences, which are each individually more competent than the top human experts across every field. Whether or not this turns out to be the case is probably the most important question in the world right now. And historically, I've been quite skeptical that this kind of thing happens, but, you seem to think that it might be plausible, and so I wanted to hear the case for it.

Ryan: 我们来谈谈这个。首先,我认为值得注意的是,AI 研发是一类 AI 特别擅长的任务,因为各大公司都在非常努力地让他们的 AI 擅长 AI 研发。从目前 AI 发展的运作方式来看,这个领域也有很多很好的特性。它是高度可验证的。你可以迭代地做很多事情,它会在各种指标上进行爬山算法(hill climb)式的优化。我认为,一旦你拥有了在 AI 研发方面与顶尖人类专家实力相当的 AI,就可能启动一个反馈循环,让 AI 去做 AI 研究。这会产生更聪明的 AI,然后再反馈回去。这种反馈循环可能会非常强大,以至于你在很短的时间内就能取得巨大的进展。也许我的中位数预期是,在一年内实现相当于四到五年的 AI 进展。这确实需要克服研究中巨大的收益递减效应,并且基本上要达到我们在极大规模的算力扩展后才能获得的同等进展。所以,这是一件令人印象深刻的大事。

Original English

Ryan: Let's talk about this. First, I think it's worth noting that AI R&D is a type of task at which the AIs are especially good, because the companies are trying really hard to make their AIs good at AI R&D. It's also the kind of domain that has a lot of nice properties from the perspective of how AI development works right now. It's pretty verifiable. You can do a bunch of stuff iteratively, and it'll hill climb on various metrics. I think once you have AIs which are roughly matching the top human experts in AI R&D, that could kick off a feedback loop where the AIs are doing AI research. That produces smarter AIs. That feeds back in. That feedback loop could be strong enough that you end up with a lot of progress in a short period of time. Maybe my median expectation is something like four or five years of AI progress in a single year. This requires really overcoming a huge amount of diminishing returns in research and basically doing the equivalent of the progress we would have gotten after a really large compute scale-out. So this is a pretty impressive, big thing.

自动化里程碑的时间表预测

Host: 值得牢记的是,五年的 AI 进展,四年的 AI 进展,哪怕是三年的 AI 进展,也真的是非常巨大的 AI 进展。三年多以前,GPT-4 才刚刚问世。而现在,当然,我们有了 Mythos 5 之类的东西,也许 Anthropic 内部还有稍微更好的模型。在三年多的时间里,这无疑是海量的进展。如果我们说的是五年,那么也许我们讨论的更像是从 GPT-3 到 Mythos 5 或类似模型的飞跃。我认为这个论点由三个不同的部分组成。现在我想逐一评估它们。首先是关于 AI 研发具有高度可验证性的论点。其次是,如果你实现了 AI 研发的自动化,你就能在单单一年内获得四到五年的进展。第三个论点是,从实现 AI 研发自动化的那个时间点开始,以目前的速度经历了四到五年的 AI 进展后,最终诞生的将是一个你可以把它投入到你所能想象到的任何工作中的 AI。你可以把它放在 20 世纪 40 年代的德克萨斯州政坛,它能用计谋击败林登·约翰逊(Lyndon Johnson)。你可以把它放在台积电(TSMC),它能学会如何在台积电做更好的工艺工程。它当然也会是一个更好的视频剪辑师……我的视频剪辑师非常优秀,但总体而言,在它尝试去做的任何特定工作上,它都比人类更好。所以,我想评估所有这些子论点,这些论点基本上导向了在达到这个基准之后很快就会获得人工超级智能(ASI),你预计这会在 2030 年左右实现,对吗?

Original English

Host: It's worth keeping in mind that five years of AI progress, four years of AI progress, even three years of AI progress, is really a lot of fucking AI progress. A little over three years ago, GPT-4 had come out. Right now, of course, we have Mythos 5 or whatever, and maybe a somewhat better model that Anthropic has internally. That is just a huge amount of progress in a bit over three years. If we're talking about five years, then maybe we're talking more about a jump from GPT-3 to Mythos 5 or whatever. I think this argument has three different parts. Now I want to evaluate each one of them. First is the argument that AI R&D is very verifiable. Second is the argument that if you automate AI R&D, you could get four or five years of progress in a single year. Third is the argument that what comes out the other end of four or five years of AI progress at the current pace, starting at the point whenever AI R&D is automated, is an AI where you can drop it on the job at basically anything you can imagine. You can drop it in Texas politics in the 1940s, and it outmaneuvers Lyndon Johnson. You can drop it in TSMC, and it learns how to do better process engineering at TSMC. It's certainly a better video editor… My video editors are very excellent, but it is just, in general, better than humans at any given job that it finds itself trying to do. So I want to evaluate all of these sub-arguments that lead to basically getting ASI pretty soon after this benchmark, which you're expecting by 2030 or something, right?

Ryan: 我想说,我预计 AI 研发的全面自动化可能会在 2031 年、2030 年左右实现。至于达到“在工作中击败所有人类”这一里程碑,也许我的中位数预期是在 2033 年左右。但是,如果我看到 AI 全面实现了 AI 研发的自动化,我想我预期那可能在一年内就会发生。在预测的计算方式中,中位数之间的差异,要大于里程碑之间差异的中位数。无论如何,不管怎样吧。

Original English

Ryan: I would say that I expect full automation of AI R&D perhaps somewhere around 2031, 2030. Getting to the "beats all humans on the job" milestone, maybe my median expectation is around 2033. But if I see AIs fully automating AI R&D, I think I'm expecting that probably within a year. The way the forecasting works out, the difference between medians is bigger than the median difference between milestones. Anyway, whatever.

AI 研发的可验证性与强化学习环境

Host: 顺便说一句,网上有个梗。每次我试图询问别人的时间表时,当问 Dario 或其他人的时候,我总是会问:“好吧,在你们让我的视频剪辑师自动化失业之前还有多久?”每次我听到这个,都会想到我的视频剪辑师在剪辑这档播客时的那个梗。但我之所以这么做,是因为我认为,在谈论你不太了解的工作时,很容易迷失在抽象的概念中,而我想要非常具体地了解,自动化一份我真正明白为什么大语言模型(LLM)目前难以接管的工作到底需要什么。我确实认为,自动化你的视频剪辑师的里程碑,要早于能够自动化所有人类工作(包括在德克萨斯州政坛中迅速适应工作)的里程碑。我确实认为,视频剪辑的自动化发生的时间可能更接近于 AI 研发的全面自动化,但这非常取决于人们究竟有多关注于让 AI 理解视频。好了,那么让我们从 AI 研发具有高度可验证性这个主张开始吧。

Original English

Host: By the way, there's this meme on the internet. Every time I'm trying to ask about people's timelines, when I'm asking Dario or somebody, I'm always like, "Okay, how long before you automate my video editors?" There's this meme of my video editor editing the podcast every time I listen to this. But the reason I do it is because I think it's easy to get lost in abstractions when you talk about jobs you don't understand well, and to very concretely understand what it takes to automate a job that I actually understand why it's difficult for LLMs to currently take control over. I do think that the milestone for automating your video editor is earlier than the milestone of being able to automate all human jobs, including Texas politics, spinning up on the job. I do think that the video editor automation occurs maybe more around full automation of AI R&D, but it's very sensitive to how much people are really focusing on understanding video. Okay, so let's start with the claim that AI R&D is very verifiable.

Ryan: 这其中有几个不同的部分。其中之一是,我们可以在一系列环境中进行训练,这些环境基本上是在直接训练模型去执行某些 AI 研发任务或非常相近的任务。例如,我们可以设置一个环境,模型只用 8 张 H100 或少量算力来训练某个 AI,那个被训练的模型可能相当于 GPT-2 Medium 或类似的东西,然后类似于 NanoGPT Medium 的运行方式——在强化学习(RL)中,它会对这些进行微调和迭代。我们可以针对许多不同的任务这样做。我们可以让它训练图像分类模型、视频生成模型、图像生成模型,以及各种不同的机器学习(ML)训练任务。我们可以在训练越来越好的模型这项任务上对它进行强化学习,同时也可以做类似这样的事情:“哦,这是一个你可以去探索的算法方向。你能去实现它吗?”基本上,存在着一整类可容器化、可验证的小规模 AI 研发任务,我们可以在这些任务上对 AI 进行密集的强化学习。各大公司大概已经在对这类任务进行一些强化学习了,你可以不断地扩大其规模,持续创建更多这种小规模的 AI 研发任务,然后 AI 就会在这方面变得越来越好。含蓄地说,我是在主张这种能力将迁移到 AI 研发中极其核心、承重性(load-bearing)的环节上。不过,我们也许可以在这里先停一下,然后再谈那部分。

Original English

Ryan: There's a few different parts of this. One of them is that we can train on a bunch of environments which are basically directly training the model to do some AI R&D task or some very close-by task. For example, we can have some environment where the model is training some AI on just eight H100s or some small amount of compute, and that model could be the equivalent of GPT-2 medium or whatever, and then similar to NanoGPT medium runs — and in RL, it's tweaking and iterating on that. We could do that for a bunch of different tasks. We could have it train image classification models, video generation models, image generation models, all kinds of different ML training tasks. We could RL it on the task of training increasingly good models, and also doing things like, "Oh, here's a particular direction you could pursue for an algorithm. Can you go and implement that?" Basically, there's this whole class of containerizable, verifiable, small-scale AI R&D tasks that we can aggressively RL the AIs on. Already companies are presumably doing some RL on these sorts of tasks, and you could just keep scaling that up, keep making more of these small-scale AI R&D tasks, and then the AIs could keep getting better at this. Implicitly, I'm claiming this will transfer to extremely load-bearing aspects of AI R&D. But maybe let's stop there for a second and then get to that part.

机器学习研究与数学领域的直觉泵

Host: 那么我们就来详细谈谈这具体是什么样子的。你可以想象我们有 GPT-7.5。我们说:“GPT-7.5,我们想让你在 AI 研发方面变得非常出色,从而帮助我们训练 GPT-9。”所以现在我们要训练 GPT-7.5,我们想出了许多不同的环境。正如你提到的,已经有一个代码库是 Andrej Karpathy 的 nanoGPT 竞速挑战(speedrun)的后继项目,在那里你只需尝试改变模型的一切,从优化器到超参数,再到架构,以尽快达到固定的训练损失(training loss)。你可以设置其他类型的环境,在那些环境中你可以说:“嘿,GPT-7.5,我想让你训练一个非常擅长玩电子游戏的模型。我想要你训练一个在反复玩同一个游戏时实际上会不断进步的模型。因此,你可能要学会如何帮助模型在在线学习(online learning)方面变得更好。我们不在乎你是怎么想出来的。也许是某种疯狂的神经语言(neuralese),或者是某种向量记忆。或者可能只是更好的长上下文(long-context)处理技术。我们不在乎。去弄清楚该怎么做在线学习的研究吧。”显然,GPT-7.5 本身已经会是一个很聪明的模型了,就像目前的模型正在变得越来越聪明一样,它在编程方面也会越来越好。你可以想象 100 个类似这样的环境,它们都在激励执行 AI 研发的能力,比如让 GPT-7.5 去开发 GPT-2 规模模型的各种容器化版本等等。然后,你基本上让 GPT-7.5 经历大量这样的训练,由此构建出了 GPT-8。现在的 GPT-8 是一位令人惊叹的机器学习研究员。通过进行所有这些类型的训练,它拥有了极强的直觉。说实话,对我来说,一个巨大的“直觉泵(intuition pump)”就是看到 AI 在数学领域取得的进展。如果那是一个高度可验证的领域,AI 可以变得……我其实并不了解数学研究那些对象级别的细节,但我就是觉得,“不,这行得通。”如果你能把它完全放入一个验证循环中,它的进展就会像洪水一样涌来,而且它真的能够取得新的突破。我很好奇机器学习研究是否具有数学研究的那种特质,也就是在连接不同学科方面似乎存在巨大的红利空间(overhang)。不会有哪一个人对代数几何以及……合适的词是什么来着?哦,天哪,我真的不太懂那些数学突破。不会有哪一个人对拓扑学和代数什么的有足够深入的了解,从而能够为某个重大猜想构建出某种反例。

Original English

Host: So let's talk through what this concretely looks like. You can imagine that we have GPT-7.5. We say, "GPT-7.5, we want to make you so good at AI R&D that you help us train GPT-9." So now we want to train GPT-7.5, and we come up with a bunch of different environments. As you mentioned, there's already this repo that is the descendant of Andrej Karpathy's nanoGPT speedrun, where you just try to change everything about the model, from the optimizer to the hyperparameters to the architecture, to get it to a fixed training loss as fast as possible. You could have other kinds of environments where you could say, "Hey, GPT-7.5, I want you to train a really good video game-playing model. I want you to train a model that actually improves as it plays the same video game again and again. So you learn how to maybe help the model get better at online learning. We don't care how you figure this out. Maybe it's some kind of crazy neuralese or a vector memory. Or maybe it's just better long-context stuff. We don't care. Figure out how to do online learning research." Obviously, GPT-7.5 will already be a smart model, and in the same way the models currently are getting smarter, it'll be better and better at coding. You can imagine 100 other environments like this which are incentivizing the ability to do AI R&D, like containerized versions of getting GPT-7.5 to develop GPT-2-sized models, et cetera. Then you basically put GPT-7.5 through a bunch of this kind of training, you build GPT-8. GPT-8 is now an amazing ML researcher. It has so much intuition from doing all this kind of training. Honestly, a huge intuition pump for me is seeing the progress that AI has made in mathematics. If it's a very verifiable domain, AIs can get… I don't really know the object-level details of mathematics research, but I'm just like, "No, it works." It can just come in like a flood if you can totally put it into a verification loop, and it can actually make new breakthroughs. I am curious if ML research has a quality of mathematical research where it seems like there was a big overhang from connecting different disciplines together. No one person would have known enough about algebraic geometry and… What was the right word? Oh man, I really don’t know about the math breakthroughs. No one person would’ve known enough about topology and algebraic whatever in order to make some counterexample to a big conjecture.

Ryan: 我的观点是,机器学习(ML)是一个不如数学那么深的领域,所以不太会有那种在某个特定领域具有极其深厚造诣的个体专家将他们的专业知识结合起来的情况,但肯定多多少少还是会有一些。不过,我也认为机器学习在某些属性上,使其在某种程度比数学更有利于 AI 的训练。特别是,你可以更好地感知你是否正在取得成功,并且能够看到中间阶段的进展。在数学中,通常很难轻易看出你是否接近成功。而如果你的目标,举例来说,是把达到某个训练损失的速度提高两倍,你差不多能看出你什么时候完成了一半。通常情况下,机器学习的创新是非常具有叠加性的,或者说是乘数效应的(取决于你怎么看待它),在其中你基本上可以不断地叠加各项创新。通常,这些创新只是累加在一起,并且互不干扰,尽管这显然取决于具体的细节。所以,我认为在很多方面,AI 研发将具有与数学非常相似的特性,也就是说,你可以通过高度可验证的方式,在那些与你实际关注的问题在结构上非常相似的各种 AI 研发任务块上进行训练,然后这种能力将会迁移过去。至于它究竟能迁移得有多好,还是一个悬而未决的问题,但我认为目前数学领域的这种迁移看起来相当不错。我的预期是,针对 AI 研发的迁移效果看起来会相当好,但还不至于令人惊叹。

Original English

Ryan: My view is that ML is a less deep domain than math, and so there's less of a thing where there are individual experts with really deep expertise in some area that they combine, but there's definitely going to be some of that. But then I also think that ML has some attributes that make it even more favorable to AI training than mathematics in some ways. In particular, you can get a better sense of whether you're succeeding, and you can see intermediate progress. In math, it's often the case that there's no easy way to see whether or not you're close to success. Whereas if your goal is, for example, to get to some training loss 2x faster, you can kind of see when you're halfway there. It tends to be the case that ML innovations are very additive, or maybe multiplicative depending on how you think about it, where basically you can keep stacking innovations. Usually the innovations just add together and don't interfere with each other, though obviously it's going to depend on the details. So I think that in a lot of ways, AI R&D will have properties quite similar to math, where you can train on chunks of AI R&D that are pretty similar in structure to the problem you actually cared about, in a very verifiable way, and then that will transfer. There's an open question of exactly how well it will transfer, but I think that the transfer currently for math looks pretty good. My expectation is that the transfer for AI R&D will look pretty good, but not amazing.

创新理论的验证挑战

Host: 所以我有一个担忧,我认为即使在数学领域,据我所知,我们还没有看到非常令人印象深刻的新理论。我们已经看到了很多令人印象深刻的、可验证的、具体的结果——比如,为这个猜想找到一个反例——但我们还没有看到达到“想出拓扑学概念”这种级别的成果,或者是“想出群论之类的东西”。看起来机器学习研究同时包含了这两方面的元素。但是,那种去想出思考问题的新方法这样不太容易验证的事情,将更难被诱导出来。以“缩放定律(scaling laws)”这个概念为例。显然,如果从 2020 年开始你就有了缩放定律的概念,那么会存在一个终端的验证循环,让你能把 GPT-4 训练得更好。但是,要诱导 AI 说出“好吧,我得仔细思考我该如何扩展我的参数和数据。我能进行哪些不同种类的调查来理解这一点?也许我能想出一个可视化图表和等浮点运算(isoFLOP)分析之类的东西”,这会是一条更漫长且可能需要更多算力铺垫的道路。这看起来确实是一个比仅仅说“嘿,让我们把 nanoGPT 的损失降下来”更长的验证循环。

Original English

Host: So one concern I have is that I think even in mathematics, as far as I'm aware, we have not seen very impressive new theory. We've seen a lot of impressive, verifiable, specific results — for example, find a counterexample to this conjecture — but we have not seen "come up with the idea of topology" kinds of levels of things, or "come up with things like group theory". It seems like ML research has elements of both of these things. But the less verifiable thing of coming up with new ways of thinking about the problem would be harder to induce. Take, for example, the idea of scaling laws. Obviously, there is some end verification loop such that you can train GPT-4 better if you have the idea of scaling laws from 2020. But there is a longer and potentially more compute-laden road to inducing AIs to be like, "Okay, I got to think carefully about how I should be scaling my parameters and data. What are different kinds of investigations I could run to understand this? Maybe I can come up with a visualization and an isoFLOP analysis or something." But that does seem like a longer verification loop than just, "Hey, let's get nanoGPT loss to go down."

Ryan: 我们来谈谈这个。首先,我认为在数学的背景下,我想说的是 AI 可以做出相当于“婴儿的第一个新理论(baby's first new theory)”之类的东西,在其中,例如,它们可以

Original English

Ryan: Let's talk about this. First of all, I think in the context of math, the thing I would say is that the AIs can do the equivalent of 'baby's first new theory,' where, for example, they can

AI 在数学与机器学习领域的突破对比

Speaker A: (它们)只是通过建立联系并产生新的理解来证明有趣的猜想。这就好像在说:“哦,AI 发现了这个相当有趣的构造”,或者它发现了这种有些不同的思考问题的方式。我们确实看到了这一点。只不过我们看到的这些例子,不如创立群论那么令人印象深刻。创立群论可能是有史以来最伟大、最宏大的数学成就之一,而现在的 AI 们在数学方面还没有那么出色。

Original English

Speaker A: just prove interesting conjectures via making connections and producing new understanding. It’s like, "Oh, there's this construction the AI found which is pretty interesting", or it found this way of thinking about the problem that's a bit different. We do see that. It's just that the examples we see are not as impressive as founding the field of group theory. Founding the field of group theory is probably among the best, biggest mathematical accomplishments of all time, and the AIs just aren't that good at math yet.

Speaker A: 在我看来,在创立群论和我们现在看到的事物之间存在一个连续的光谱,而 AI 们正在这个光谱上不断向上攀登。其次,我认为机器学习(ML)相对于数学来说是一个非常浅显的领域。在数学中,很多时候你会发现某种真正深刻的抽象概念,如果你真的理解了那个很难理解的概念,你就能取得突破。然而我觉得,在机器学习中与此等同的东西,真的都是些愚蠢的废话。就像关于缩放定律(scaling laws),拜托,伙计们,我们可以非常快地解释缩放定律。我认为数学中最深刻、最重要的概念,例如,并不具备让你能在很短时间内真正理解其底层本质以及它为何重要的属性。但我觉得一个结果将是,到 2030 年,我们将会摘取所有的“低垂果实”。我觉得在数学史上,缩放定律就像笛卡尔发现笛卡尔坐标系并进行非常基础的数学运算一样。最终,如果我们想在 2030 年代继续取得进展,那将类似于目前在数学前沿发生的任何深奥事情。

Original English

Speaker A: From my perspective, there's a continuum between that and the things we're seeing now, that the AIs are continuing to march up. Second, I think ML is a very shallow domain relative to math. In math, there was much more of a thing where you find some true deep abstraction, and if you really understand that thing, which is hard to understand, then you get somewhere. Whereas I feel like the things that are the equivalent of that in ML are really dumb bullshit. Like with scaling laws, come on guys, we can explain scaling laws really quickly. I think the deepest and most important concepts in math, for example, don't have the property that you can really understand the underlying thing and why it matters in a very short period of time. But I feel like one effect will be that we will have gotten rid of all the low-hanging fruits by 2030. I feel like scaling laws will have been, in math history, like Descartes finding the Cartesian grid and doing very basic mathematics. Eventually, if we want to keep making progress in the 2030s, it's going to be like doing whatever bullshit is happening at the frontiers of mathematics right now.

AI 面临的真正瓶颈:深层洞察还是直觉与工程品味?

Speaker B: 这可能是对的。我的感觉是,某些领域在运作方式以及它们在多大程度上依赖于深度抽象方面,结构上是不同的。物理学和数学更倾向于非常深奥、很难想出主意的那一端,而我认为机器学习和大多数其他领域更适合“爬山法”(hill climbing,即逐步优化)。这就是我对未来发展趋势的感受。即使在这样的情况下:AI 必须埋头苦干——到了 2030 年,研究中的一大堆低垂果实已经被摘取,它们需要取得进一步的进展——我仍然怀疑,很大一部分工作将更多地依赖于构建日益复杂的基础设施,以及对实验大致面貌拥有非常好的直觉。所以我可能不太赞同“AI 们将缺乏某种深刻洞察力”这种观点。我更赞同的观点是,它们真的需要一堆关于底层细节实验的品味,而这是它们目前所不具备的。它们需要对哪些训练方法有效、哪些无效有大量的直觉,就像当前的研究人员所拥有的那种直觉一样。即使在 AI 领域取得了一些突破的情况下,回想起来通常也会发现,促成这一突破的一个巨大瓶颈在于能否把所有的微小细节和混乱的直觉处理对。

Original English

Speaker B: That could be right. My sense is that some domains are structurally different in terms of how they operate and how much they depend on deep abstractions. Physics and math are much more on the side of being very far on the deep, hard-to-come-up-with-ideas side, whereas I think ML and most other domains are much more amenable to hill climbing. That's my sense of how this will go in the future. Even in the regime where your AIs are having to plow — it's 2030, a bunch of low-hanging fruit in research has already happened, and they need to make further progress — I still suspect that a bunch of the work will live more on the side of building increasingly complicated infrastructure and having really good intuition about what the experiments roughly look like. So I'm probably less sympathetic to the idea that the thing the AIs will lack is some deep insight. I’m more sympathetic to the idea that they really need a bunch of taste about in-the-weeds experiments that they currently don't have. They need a bunch of intuition for what sorts of training approaches would work and what wouldn't, in ways that current researchers have. Even in cases where there has been some breakthrough in AI, oftentimes in retrospect it looks like a big bottleneck to making that breakthrough happen was getting all of the micro details and mungy intuition right.

Speaker B: 一个例子是训练 AI 擅长推理和思维链(chain of thought),在思维链上进行强化学习(RL)。看起来,你本可以在 GPT-3 上进行 RL 和思维链训练,并且如果你真的扩大规模并做得很好,大概能在数学上获得相当有趣的结果。但在当时,还有更容易摘取的果实。此外,要做好这种训练,涉及到所有技术实现的底层细节、如何扩大规模以及设定正确的超参数。所以也许你可以在 Qwen 1B 或其他什么模型上演示所有内容,并获得这整个事情将会奏效的某种感觉。但是,人们之所以没能尽早展示这一点,是因为所有这些其他杂乱的细节,以及关于到底如何调整参数和如何设置系统的直觉。老实说,这是我对这个故事仍然抱有的怀疑。

Original English

Speaker B: An example of this is training AIs to be good at reasoning and chain of thought, doing RL on chain of thought. It looks like you probably could have done RL and chain of thought on GPT-3 and gotten kind of interesting results on math if you had really scaled it up and done a good job. But at the time, there was low-hanging fruit. Also, doing a good job with that training is kind of in the weeds on all the technical implementation and scaling it up and getting the hyperparameters right. So maybe you can demonstrate everything on Qwen 1B or whatever and get some sense that this whole thing is going to work. But people didn't demonstrate it as early as they could have because of all of these other mungy details and intuition about exactly how to tune the parameters and how to set things up. This is my remaining skepticism, honestly, about this story.

Speaker A: 如果研究突破如此易于通过智力实现,我不确定我为什么不理解 AI 的历史进展没有比本可以达到的速度更快。正如你所说的,到了带视觉的强化学习(RLVR)真正起作用的时候——尽管你本可以用较少的算力做到这一点——但我们不得不等待海量的算力、吉瓦级的算力可用时,人们才开始进行这种训练,在算力持续增加的轨迹上,我们才能取得更多突破。我不知道。我觉得在 2022 年有很多 AI 研究人员正试图破解推理难题。难道仅仅是因为他们被编写基础设施代码的能力卡住了脖子,还是到底发生了什么?

Original English

Speaker A: I'm not sure I understand why, if research breakthroughs are so amenable to intelligence, AI progress has not been historically faster than it could have been. As you were saying, by the time RLVR actually worked — even though you could have done it with less compute — we had to wait for oceans of compute, gigawatts of compute, to be available before people were doing this training, on the trajectory of compute continuing to increase so we make more breakthroughs. I don't know. I feel like there were a lot of AI researchers in the year 2022 who were trying to crack reasoning. Was it just that they were bottlenecked by the ability to write infrastructure code, or what was happening?

Speaker B: 这是一个复杂的混合体。我认为,如果他们在想到一个实验后,就能立即无缺陷地运行该实验,而且不会因为缺陷问题而严重影响结果的话,他们本来可以进展得更快。然后另一个原因是,能够用高算力运行大量实验,可以让你掩盖实现方式上不太正确或超参数不对的地方。所以,算力对于进行 AI 研究真的非常有帮助,你可以用它掩盖很多东西。但这并不意味着劳动力的极大增加就不会同样有帮助,特别是如果这种劳动力伴随着该领域人们所拥有的最好直觉的话。我只是觉得那会非常有帮助。

Original English

Speaker B: It's a complicated mix. I think they would have gone faster if they could, as soon as they thought of an experiment, run that experiment without bugs, without bugs being very important. And then another part of it is that being able to run a lot of experiments at high compute lets you paper over ways in which the way you implemented it isn't quite right or you didn't have the right hyperparameters. So compute is just really helpful for doing AI research, and you can cover over a lot of things. But that doesn't mean that massive increases in labor wouldn't also be helpful, especially if that labor comes with among the best intuitions that people have in the field. I just think that's really helpful.

Speaker B: 我这里的观点的另一部分——可能与你的出发点有些不同——是我期望的迁移能力比你似乎想象的要多一些。我想象这些 AI 实际上大体上是相当优秀的科学家,并且在所有这些方面都相当合理。当你与它们互动时,它们并不会给人那种非常过度专业化的学者型天才的感觉。它们实际上在研发的所有环节都相当不错,然后可能在某些子领域极其优秀。因此,它们在编写内核方面极其超人类,在反馈周期非常短的所有方面也极其超人类,而在所有其他方面则相当不错,完全能够与其他人匹敌。我认为我们现在正在看到这一点。当我现在观察 AI 时,它们在做机器学习研究方面,已经能够相当称职地匹敌那些在这个领域资质平庸的人类了。只不过,在机器学习研究中表现平庸并没有那么大帮助。你真正想要的是那些擅长机器学习研究的人。我的感觉是 AI 在所有这些方面都在进步。它们的品味在提高,它们的直觉在提高,而且它们现在的品味和直觉已经不再是彻头彻尾的垃圾了。

Original English

Speaker B: Another part of my perspective here, which is maybe a bit different from where you're coming from, is that I'm expecting somewhat more transfer than you seem to be imagining. I'm imagining these AIs are actually pretty good scientists in general and are pretty reasonable at all of that stuff. When you interact with them, it's not like they have some really hyper-specialized savant-type vibe. They're actually pretty good at all of the stuff in R&D, and then maybe extremely good at some subdomains. So they're incredibly superhuman at writing kernels, incredibly superhuman at everything with very short feedback loops, and then pretty good at all the other stuff, totally able to match other people. I think we are seeing this now. When I look at AIs right now, it's already the case that they can pretty competently match humans who are mediocre at ML research at doing ML research. It's just that being mediocre at ML research is not that helpful. The thing you actually want are people who are good at ML research. My sense is the AIs are just improving at all of these things. Their taste is improving, their intuition is improving, and it's already the case that their taste and intuition is not complete garbage.

将五年 AI 进展压缩至一年的设想与算力鸿沟

Speaker A: 我想非常具体地了解,如果把五年的 AI 进展压缩在一年内发生会是什么样子。假设我们回到开发 GPT-3 的时候。这个想法是,利用他们 2022 年时拥有的算力水平,如果我们当时就自动化了 AI 研发,你能在同一年年底就得到 Mythos 吗。

Original English

Speaker A: I want to very concretely understand what it would look like for five years of AI progress to happen in one year. Suppose we were back when GPT-3 was developed. The idea is that, with the level of compute they had back in 2022, if we had automated AI R&D back then, you could at the end of that year have Mythos.

Speaker B: 是的,就是这个意思。

Original English

Speaker B: That would be the idea, yes.

Speaker A: Mythos 花费的算力远远超过他们当时的水平,但即便仅用当时的算力水平,不仅所有的突破都会发生,而且他们还要用那个算力水平训练出 Mythos。显然,这就要求发现自那时以来的所有算法进展。实际上,需要发现的甚至更多,因为你必须弥补 Mythos 所使用的算力差距……GPT-3 是用什么算力训练的?大约是 1e23 吗?我们可以查一下。但它有可能会高出四个数量级的算力吗?

Original English

Speaker A: Mythos took way more compute than they had back then, but even with the level of compute they had back then, not only do all the breakthroughs happen, but they also train Mythos with that level of compute. What would be required is obviously discovering all the algorithmic progress since then. It’s discovering even more, actually, because you've got to make up for the fact that Mythos uses… What was GPT-3 trained on? Like 1e23? We can look it up. But is it plausibly four orders of magnitude more compute?

Speaker B: 我认为要比那个稍微少一些。我们快速查一下。GPT-3 的训练算力大约是 3e23。我的感觉是 Mythos 可能高出略多于三个数量级(OOMs)。所以问题是:你能够在作为模型本身的同时,克服这 1000 倍的算力差距吗?这里有一个也许我们应该讨论的具体主张。就在现在,我们是否能够用 GPT-3 级别的算力训练出一个模型来匹敌……我到底是怎么想的呢?GPT-3 是 2020 年发布的,所以它大约是在六年半、七年前训练的。值得注意的是,拿 GPT-3 举例可能扯得有点太远了,但我们暂且先顺着这个思路走。如果我们今天用 GPT-3 级别的算力来训练一个模型,那个模型会有多好?基于算法进展的运作方式,我的理解是,我们将能够训练出一个与我们大约三年前拥有的最佳模型一样好的模型。所以我认为,现在我们将能够训练出一个可能比 GPT-4 稍微好一些的 GPT-3 版本,比 GPT-4 好出适度的水平。

Original English

Speaker B: I think it's somewhat less than that. Let's look this up quickly. GPT-3 training compute is about 3e23. My sense is that Mythos is probably a little over three OOMs higher. So the question is: can you overcome this 1000x compute gap while also being the model? Here's a concrete claim that maybe we should talk about. Right now, would we be able to train a model with GPT-3-level compute that matches… What exactly do I think? GPT-3 was released in 2020, so it was trained about six and a half, seven years ago. It's worth noting that GPT-3 is maybe a little too far in the past, but let's go with this for a second. If we were to train a model with GPT-3-level compute today, how good would that model be? My understanding, based on how algorithmic progress works, is that we'd be able to train a model that's as good as the best model we had perhaps around three years ago. So I think that right now we'd be able to train a version of GPT-3 that's probably somewhat better than GPT-4, a moderate amount better than GPT-4.

Speaker A: 我认为这差不多是对的。这大致符合算法进展的实际情况。基本上,整个故事的结果将会是,为了获得五年的 AI 进展,你可能需要大约——我粗略估计——八年的算法进展,这是一笔巨大的算法进步。但事实证明,在我看来,AI 进展的大部分来自于算法和数据的某种结合,你可以在这些方面不断取得巨大改进,并用更少的算力来训练 AI。

Original English

Speaker A: I think that's about right. That roughly lines up with how algorithmic progress has worked. Basically, the story would end up being that to get five years of AI progress, you're probably going to need around, I would say, maybe eight years of algorithmic progress, very roughly, which is a lot of algorithmic progress. But it just turns out that most of the AI progress, from my perspective, has come from some mix of algorithms and data, and you can just keep making huge improvements on these things and training AIs with less compute.

数据与人类专家在 AI 进展中的作用

Speaker A: 很高兴你提到了这一点,因为从 GPT-3,甚至 3.5 到现在,到底发生了什么?为什么 Mythos 这么优秀?显然,我们扩大了算力规模。我们拥有了更好的算法。但发生的一件大事是,我们建立了一个价值数百亿美元的数据产业,该产业系统地收集并编码了各种不同学科中人类专家的判断——以强化学习(RL)环境的形式编码,以监督微调(SFT)轨迹的形式编码——这些专家建立这些东西是为了帮助模型更好地理解你如何进行编程,你如何构建复杂的基础设施项目,你如何处理法律事务,你如何做各种事情。AI 如何能够复制目前人类专家判断在 AI 进展中似乎发挥的作用?

Original English

Speaker A: I'm glad you brought that up, because what has happened since GPT-3, or even 3.5, till now? Why is Mythos so good? Obviously, we've scaled the compute. We have better algorithms. But a huge thing that's happened is that we have built a deca-billion-dollar data industry which has systematically collected and codified expert human judgment across all kinds of different disciplines — codified in the form of RL environments, codified in the form of SFT traces — that these experts built to help the model better understand how you do coding, how you build complex infrastructure projects, how you do law, how you do whatever. How are the AIs able to replicate the effect that expert human judgment currently seems to be playing in AI progress?

Speaker B: 我的感觉是,增加获取人类专家数据的精力投入,对 AI 研发总体而言并不是特别重要。具体来说,在过去几年里,我们一直在增加算力,增加在 AI 公司工作的人数,并增加用于数据标注的精力。我的感觉是,如果你去掉过去两次人类专家生成数据的翻倍或类似的情况,那并不会产生巨大的差异。正在发生的事情很大程度上是人们开发了更好的方法来利用人类和 AI 构建强化学习环境,并以此为起点取得进展。

Original English

Speaker B: My sense is that scaling up the amount of effort spent on getting expert human data has not been hugely important for AI R&D in general. In particular, over the last few years, we've been scaling up compute, scaling up people working at AI companies, and scaling up the amount of effort spent on data labeling. My sense is that if you removed the last two doublings or whatever of data generation from expert humans, that would not make a huge difference. A lot of what's been going on is people have been developing better ways to leverage humans and AIs to construct RL environments and going somewhere from that.

Speaker A: 但是你如何解释为什么 AI 在编程方面变得如此出色?我觉得其中很大一部分原因是数据和 RL 环境,这些正是对人类专家经验的编码。

Original English

Speaker A: But how do you explain why the AIs have gotten so good at coding? I feel like a big part of that is data and RL environments, which are codifying human experts.

Speaker B: 但问题在于,创建 RL 环境的限制因素是什么?我的感觉是,如今的 RL 环境比 2024 年时好得多,原因并不主要是因为我们雇佣了更多的人类专家来制作 RL 环境。更主要的原因在于我们更清楚我们甚至想要制造什么样的 RL 环境,以及我们应该如何构建它们。而且,我们正在使用大量的 AI 劳动力来构建 RL 环境。我认为这些影响远比人类劳动力构建 RL 环境的影响重要得多。我不是说人类劳动不重要。我只是说这里有其他重要的大型驱动因素。我可以试着论证这一点。其中一方面是人们所需的环境数量非常庞大。我认为,在对环境应有某种大致概念的情况下,AI 在执行制作 RL 环境的任务上实际上是相当出色的。我们已经有可以使用的预先存在的数据。

Original English

Speaker B: But the question is what is the limiting factor on creating RL environments? My sense is that the reason why RL environments today are much better than they were in 2024 is not so much because we have hired way more human experts to make RL environments. It is instead much more because we better know what RL environments we even want to make and how we should structure them. Also, we're using huge amounts of AI labor to build RL environments. I think those effects are much more important than the effect of human labor building the RL environments. I'm not saying that the human labor doesn't matter. I'm just saying there are other big drivers that are important here. I could try to argue for this. One thing is just that the amount of environments people want is a very large amount. I think the AIs are actually pretty good at the task of making RL environments given some sense of what the thing should be. There's preexisting data you could use.

算力与数据在推动 AI 进步中的权重

Speaker A: 很多这些事情都有很好的验证循环。举个例子,看看昨天《商业内幕》(Business Insider)的报道,谷歌(Google)斥资近20亿美元收购 Mechanize。我们只需看看市场汇率,就能知道人们认为真正优秀的人类专家数据到底值多少钱。前沿实验室似乎认为它价值连城。他们愿意为此买单。你认为前沿实验室在数据(而非算力)上的支出比例是多少?你觉得算力与数据的支出比例是怎样的?

Original English

Speaker A: A lot of these things have good verification loops. Just look at, for example, what was reported in Business Insider yesterday, that Google is paying close to $2 billion for Mechanize. We can just look at market rates for what people think really good human expert data is worth. The frontier labs seem to think it's worth a lot. They're willing to pay for it. What fraction of frontier lab spending do you think is on data rather than compute? What do you think is the compute/data spend split?

Speaker B: 我认为绝大部分都是算力(compute),但我也认为这是因为算力比数据更容易规模化扩展。

Original English

Speaker B: I think it's overwhelmingly compute, but I also think it's because compute is easier to scale up than data.

Speaker A: 但这确实与推动进步的核心动力息息相关,不是吗?

Original English

Speaker A: But that's really relevant to what's driving progress, right?

Speaker B: 我的感觉是,这个支出比例大概是 20:1 或者 10:1。我不知道确切数字,这取决于具体的公司。但这就像石油占 GDP 的 1.5% 一样。这并不意味着如果你切断石油供应,GDP 还能继续运转。

Original English

Speaker B: My sense is that the split is something like 20 to 1 or 10 to 1. I don't know exactly. It depends on the company. But this is similar to how oil is 1.5% of GDP. That doesn't mean that if you cut oil out, GDP could continue to run.

Speaker A: 当然,但这与你的观点相矛盾了,对吧?如果石油消失,经济会立刻停滞。

Original English

Speaker A: Sure, but it contradicts your argument, right? The economy would come to a halt immediately if oil went away.

Speaker B: 确实如此,但你刚才论证的是,因为市场价值(支出)高昂,我们可以得出数据是关键驱动力的结论,而我是在说这显然不一定是真的。那个论点只会让算力看起来是一个重要得多的驱动力,或者说雇佣员工是一个重要得多的驱动力。

Original English

Speaker B: Sure, but you were just arguing that because of the high market cap, we can learn that this is the key driver, and I'm saying that's not clearly true. That argument just makes it look like compute is a much more important driver, or hiring employees is a much more important driver.

从当前模型迈向超级人工智能的泛化难题

Speaker A: 那我们也许该说得更具体一点。我是这么想的。我的主张是,如果你回到 2022 年,手里有 GPT-3.5,并且想在没有人类专家的情况下让它的编程能力变得更好,我认为这会非常非常困难。

让我举个例子,说明我所设想的从 GPT-8 走向 ASI(超级人工智能)的难度。你希望 ASI 擅长的事情之一是:我要接管一家公司,让它变得更加盈利,然后做各种疯狂的事情来让它运转得更好。我要接管一家晶圆厂并生产更多的芯片。我要进入国会并试图说服他们通过某项法案,等等。这就是我所想象的,按照这种速度再过五年的人工智能进步,能够让 AI 具备的能力。

这也是我真正担心的事情:ASI 能够理解如何在世界上做这些疯狂的事情,能做基辛格(Kissinger)能做的事,能做史蒂夫·乔布斯(Steve Jobs)能做的事等等,还有他的工程师们所做的事。我不太确定,如果没有相关的世界数据,你该如何实现这一点。这就等同于让 Mythos 在缺乏比 GPT-3 时期更好的编程环境的情况下,却依然极其擅长编程。

Original English

Speaker A: So maybe let's be more concrete. Here's what I think. My claim is that if you went back to 2022 and you had GPT-3.5, and you were trying to make it better at coding without human experts, I think it would have just been very, very difficult.

Let me give you an example of what I imagine would be the difficulty of going from GPT-8 to ASI. One of the things you'd want ASI to be good at is: I'm going to take over a company and make it much more profitable and do all kinds of crazy shit to make it work better. I'm going to take over a fab and produce more chips. I'm going to go into Congress and try to convince them to pass some bill, et cetera. This is what I imagine five more years of AI progress at this pace would enable an AI to be able to do.

This is the thing I'm really worried about: ASI that can understand how to do crazy shit in the world, that can do what Kissinger can do, can do what Steve Jobs can do, et cetera, and also his engineers and so on. I'm not sure how you get that without the relevant world data, which is the equivalent of Mythos being really good at coding while not having the coding environments that have improved it relative to GPT-3.

Speaker B: 这有几点。首先,我敢打赌,如果你去看看随机抽样的 Mythos 的训练环境,它们实际上与现实中实际使用该模型的情况大相径庭。我的感觉是,强化学习(RL)的分布与真实世界的数据分布存在巨大的偏差,而这种偏差通过迁移学习(transfer)以及少量专注于真实世界的数据混合,得到了显著的平滑。

我的感觉是,这与你在全自动化 AI 研发的基础上,再经过五年的 AI 进步所获得的那些疯狂、狂野、近乎超人类的 AI 的运作机制将是相似的。

那么我们来仔细梳理一下。特别要说的是,我认为你可以训练出一个 AI,让它在即时学习(learning on the fly)方面极其出色,并能在各种各样的强化学习环境中,执行类似于上下文学习(in-context learning)但可能使用略有不同机制的任务。

你构建了所有这些不同的强化学习环境,在这些环境中,AI 必须即时适应、即时学习、弄清楚自己该做什么、更好地理解自身的处境,并从反馈中极快地学习以成功实现其目标。它还面临资源有限等限制,如果搞砸了,可能会陷入糟糕得多的处境。

如果你在海量的这类环境中进行训练,你就能学会即时获取上下文的通用技能,而我们现在已经看到了这种现象。

现在的情况已经是,AI在粗略理解发生了什么,以及从它们有权限访问的有限信息中获取上下文方面,已经变得好得多了。

然后,你可以让这些 AI 到台积电(TSMC)去工作。即使台积电并没有真正存在于它们的数据分布中,但它们的数据分布极为广泛,而且 AI 在自身的数据分布上表现极其优异,以至于这种能力可以迁移到即时学习并成为一名优秀的台积电工程师上。

AI 成为一名优秀的台积电工程师的方式,并不是因为它在如何成为一名优秀的台积电工程师方面拥有大量的缓存知识。而是因为它在那里执行了某种放大版的上下文学习(in-context learning)等效操作。

这将是最平淡无奇(prosaic)的推演剧本。显然,这还有很多不同的发展方向。

Original English

Speaker B: Here are a few points. First, I bet if you look at randomly sampled training environments for Mythos, they're actually very different from what it looks like to actually use the model in practice. My sense is that the RL distribution has really large deviations from the real-world data distribution, and it's significantly smoothed over by a mix of transfer and having a small amount of data focused on the real world.

My sense is that this will be a similar mechanism as how it works for the crazy, wildly, quite superhuman AI you get as a result of five years of AI progress on top of fully automated AI R&D.

So let's go through this a little bit. In particular, I think that you could train an AI to be really, really good at learning on the fly and doing something analogous to in-context learning, but potentially using somewhat different mechanisms, in a wide variety of RL environments.

You build all these different RL environments where the AI has to adapt on the fly, learn on the fly, figure out what it should do, understand its situation better, and learn really quickly from feedback in order to succeed at its objective. And it has things like limited resources, and if it messes up, it can end up in a much worse position.

If you train on a huge number of these environments, you will learn general skills of picking up context on the fly, and we're already seeing this.

It's already the case that AIs are now much better at understanding roughly what's going on and picking up context from a limited amount of information they're given access to.

Then those AIs could be put on the job at TSMC. Even though TSMC is not literally in their data distribution, their data distribution is really wide, and the AIs are extremely good on their data distribution, such that it transfers to picking up being good at being an engineer at TSMC and learning that on the fly.

The way the AI gets good at being a TSMC engineer isn't that it has a ton of cached knowledge on being a good TSMC engineer. It's that it does the equivalent of some scaled-up version of in-context learning there.

That'd be the most prosaic story. Obviously, there's a bunch of different ways this could go.

Speaker A: 我想这或许归结于对你能走多远的直觉差异。当我想起我认识的那些极其聪明的人时,他们在自己不太了解的领域里就是没有那么高效。

Original English

Speaker A: I think this maybe comes down to a difference of intuition about how far you can get. When I think about really smart people I know, they're just not that effective in domains they don't understand that well.

Speaker B: 但他们有多少时间去学习呢?

Original English

Speaker B: But how long have they had to learn?

Speaker A: 我同意,如果他们有经验,他们会表现得好得多。但这也许正是我所主张的,那就是对数据的经验。举个例子,如果我只是找了一个非常聪明的常春藤盟校毕业生,然后对他说:“好,你现在负责谈判伊朗核协议”,我认为他们简直会不知所措。

Original English

Speaker A: I agree that if they had experience, they would be much better. But that's maybe what I'm arguing for, that experience with data. For example, if I just get a really smart Ivy League college grad, and I'm like, "Okay, you're now in charge of negotiating the Iran deal," I think they just wouldn't know what to do.

AI 的即时学习与上下文构建能力

Speaker B: 我想,如果你换一个擅长快速掌握许多不同领域的人,并且给他们一些时间去训练、与人交谈、充实专业知识并进行一些实践,他们实际上会做得很不错。我认为大多数领域在本质上都是相当浅显的,一个拥有少数核心技能、极其聪明的通才(generalist)能够非常快地上手。但这当然并不适用于所有领域。

我的感觉是,AI 将会发展出越来越好的机制,以快速获取特定领域的理解和专业知识。例如,想想 AI 理解一个新代码库的速度有多快。AI 能够比人类快得多地理解一个新代码库,尽管其理解的深度目前还不如人类。但这种情况正在随着时间推移而改善。让我把这个观点说得更详细些。

假设你使用 Fable 5 或是 Mythos 5 之类的模型,并且你想在一个极其庞大的代码库中进行某种复杂的修改。该模型会非常快地对代码库建立一定的理解,可能远少于一个小时,甚至极短的时间。然后,它对代码库的理解会出现些许停滞,无法达到人类在更长周期内所能获得的深度。所以,这就好比 AI 用一个小时的时间,就能匹敌人类几周的进度,具体取决于代码库究竟有多复杂的细节。但它还无法匹敌一个在这个代码库上工作了两年左右的人类。

但随着时间的推移,AI 能够匹敌的理解程度已经提升了。如果我们看看 3.7 Sonnet 或 3.5 Sonnet,也许它只能匹敌人类理解一天代码库的水平。但现在 AI 在构建任务上下文方面要好得多了。

所以你可以说:“Mythos,我希望你真正理解这个代码库,然后实现这个功能。”它会生成海量的子智能体(sub-agents)。那些子智能体会仔细钻研一大堆东西。它会传回大量的上下文。然后它会进行一些调查。它在做这件事上还称不上惊艳,但速度极快,而且效果相当不错。而且我也不难想象,你如何能够训练 AI 在这项任务上变得越来越好。

在一个非常庞大的代码库中以合理的方式实现某个非常复杂的功能,这项任务是极易验证的,这也可以成为 AI 能够持续改进的地方。同样地,这里也存在一种更广泛的技能,即快速理解上下文,并让许多不同的 AI 并行学习,然后再将这些学习成果合并起来。

我想这里似乎存在一个症结,我认为这只是一个我们会看到结果的实证问题。一方面是在完全可验证的领域里极其擅长理解局势、快速进入状态、在长期跨度中取得进展——AI 在这些方面显然正在以极快的速度变得越来越强;另一方面则是:“好,去和总统谈谈,说服他做 X 这件事”,或者是“你现在掌管谷歌了。你必须在这个季度让谷歌变得更赚钱”。这两者之间的迁移能力究竟有多强?

让我再试着阐述几个也许相关的论点。其一是,当你观察 AI 在写文章方面是如何进步的……我们稍微谈谈这个。即使在这些领域,你也能获取一些数据。当处于非常快速的进步轨迹上时,AI 甚至在这些领域也能获取一些数据。也许你很难构建一个完全可验证的环境来评判“在人类看来你的文章是否真的很好?”但你可以做一点这类工作。你可以进行一些训练。你可以进行一些在线训练。AI 能够基于现实世界的事物进行一些在线训练。它们能够有评估系统(evals)。它们能够对此进行采样。你可以扩大执行这些操作的频率。

第二点是,在实践中,当我只观察迁移能力时,感觉似乎还不错。我认为 AI 实际上已经在不可验证的领域取得了很大的进步,你很难指出在哪些难以验证的领域中,GPT-4 和 Mythos 之间的实际进步幅度不够高。当然,这并不意味着 Mythos 比最优秀的人类还要好。在工作的某些方面,它可能仍然明显比典型的人类专业人士差得多,但同时它又比 GPT-4 强得多,后者当年甚至根本沾不上边。

Original English

Speaker B: I think if you instead got someone who is really good at quickly picking up a bunch of different domains and you gave them some time to train and talk to people and shore up their expertise and do some practice, they would actually do a pretty good job. I think most domains are fundamentally pretty shallow, where a very smart generalist who's good at a limited subset of core skills can get going pretty quickly. That's not true for literally every domain.

My sense is that the AIs will develop increasingly good mechanisms for quickly acquiring understanding and expertise in a given domain. Consider, for example, how fast AIs can understand a new code base. AIs can understand a new code base much faster than humans can, but to a degree that's shallower than humans could currently understand. But it's getting better over time. Let me spell that argument out a bit more.

Let's say you take Fable 5 or Mythos 5 or whatever, and you wanted to make some kind of complicated change to a really massive code base. The model will get some understanding of the code base very fast, in the course of maybe significantly less than an hour, potentially much less than an hour. Then its understanding of the code base will plateau a little bit, where it won't get as deep of an understanding as a human would have gotten over a much longer period. So it's like an AI in an hour can match a human with a few weeks maybe, depending on the details of exactly how complicated the code base is. But it won't match a human who's been working on that code base for two years or whatever.

But over time, the amount of understanding AIs can match has gone up. If we look at 3.7 Sonnet or 3.5 Sonnet, maybe it could only match the equivalent of understanding a code base for a day or something. But now AIs are much better at building context about a task. So you can be like, "Mythos, I want you to really understand this code base, and then implement this feature." It will spawn a bajillion sub-agents. Those sub-agents will pore over a bunch of things. It will deliver a bunch of context back. It will then investigate a few things. It's not amazing at doing this, but it can happen really fast, and it can work pretty well. And it's not very hard for me to imagine how you could train AIs to be increasingly good at this task.

The task of implementing some very complicated feature in some reasonable way in a very big code base is extremely verifiable, and that can be a thing the AIs improve on. Similarly, there's a broader skill of quickly understanding context and being able to have a bunch of different AIs learn in parallel and then merging that together.

I think there seems to be a crux here, which I think is just an empirical question we'll see. How good is the transfer between getting really, really good at understanding the situation, getting up to speed, making progress over long periods in verifiable domains — which the AIs are obviously getting way, way better at really fast — to, "Okay, go talk to the president and convince him to do X thing." Or, "You're now in charge of Google. You must make Google a much more profitable company this quarter."

Let me try to spell out a few more arguments that are maybe relevant. One thing is, when looking at how the AIs have improved at essay writing… Let's talk about that a little bit. You can get some data even on these domains. AIs will be able to get some data even on these domains when on a very fast progress trajectory. Maybe it's hard to build a verifiable environment for "was your essay really good according to humans?" But you can do a bit of that. You can do some training. You can do some online training. The AIs will be able to do some online training based on real-world stuff. They'll be able to have evals. They'll be able to sample that. You can scale up the cadence at which you do this.

The second thing is that in practice, when I just look at the transfer, it seems okay. I think the AIs have in fact improved a bunch at non-verifiable domains, and it's hard to point to domains that are really hard to verify on which the amount of improvement between GPT-4 and Mythos hasn't been pretty high in practice. Now, that doesn't mean that Mythos is better than the best humans or something. It can still be significantly worse than typical human professionals at some aspect of their job while still being way better than GPT-4, which was not even close.

数据与算法的实证实验

Speaker A: 所以我们是在讨论过去几年中,有多少进步是来自数据,多少是来自算力。这让我想起来,我其实正在和 Jerry Han 一起进行一项相关的实验,他现在还是个大学生。为了评估有多少进步来自数据,有多少来自算法,我们基本上正在做的是:用 2026 年数据文件中最优秀的数据,来训练 2019 年至今最好的算法配方;同时,也用当前最好的算法配方,去训练从 2019 年回溯到 2026 年的不同数据文件。我认为这将会非常有趣。

我很好奇你是否想预先登记(pre-register)一下,算力的乘数效应有多少来自其中一方,又有多少来自另一方。

Original English

Speaker A: So we're talking about how much progress has come from data versus compute over the last few years. That reminds me, I'm actually running an experiment with this with Jerry Han, who's still a college student. What we're basically doing to evaluate how much progress is coming from data versus algorithms is training the best algorithmic recipe from 2019 till now with the best data from the 2026 data file, and then also training the different data files going back from 2019 to 2026 with the current best algorithmic recipe. I think that will be interesting.

I'm curious if you want to pre-register what amount of compute multipliers are coming from one versus the other.

Speaker B: 我们在说“数据”这个词的时候,需要非常小心地界定它的含义。我刚才一直试图小心地将“增加雇佣人类专家标注数据的支出”与“扩大人类专家标注数据的数量”区分开来。相比于 2019 年,我们现在拥有更好预训练数据集的原因,并不是因为人们花得起多得多的钱让人类专家去手打数据,然后再用这些数据去训练 AI。

Original English

Speaker B: We need to be pretty careful with what we mean when we say the word data. I was trying to be pretty careful to distinguish between scaling up spending on getting human experts to label data, or scaling up the amount of human expert-labeled data. The reason why we have a better pre-training data set now versus in 2019 is not because people are spending way more money getting human experts to type up data that the AIs are then trained on.

Speaker A: 有一部分是。

Original English

Speaker A: Partially.

Speaker B: 我认为并不是很多。我认为预训练数据的改进中,这部分占比极小。我指的确实是预训练。也许我们应该将中间训练(mid-training)和后训练(post-training)分开讨论。但我认为,绝大多数预训练数据的改进,都源于对“什么样的训练集是好数据集”有了更好的科学理解,以及在弄清楚如何过滤数据方面所投入的枯燥苦差事。所以我的观点是,像从 OpenWebText 到 FineWeb 这种形式的改进,最好将其描述为一种你可以用一些 GPU 来研究的算法改进,而你并不需要人类专家数据来做这件事。

Original English

Speaker B: I think it's not much of it. I think it's very little of the pre-training data improvements. I do mean pre-training. We should maybe talk separately about mid-training and post-training. But I think the vast majority of pre-training data improvements are from science on better understanding what data sets are good and schleppy labor on figuring out how to filter down. So my view is that improvements of the form of, like, OpenWebText to FineWeb, that improvement is better described as an algorithmic improvement of the sort that you can study with some GPUs, and you don't need human expert data to do that.

互联网数据与自动化研发

Speaker A: 现在还有一个我们可以讨论的不同影响因素,那就是 2026 年的互联网可能比 2018 年的互联网拥有更肥沃的训练数据土壤。还有一个影响是,在互联网上发帖的人类更多了,所以有更多的数据可以收集。我的感觉是,与人类更懂得如何筛选数据、拥有更好的抓取工具、知道如何更好地处理这些抓取的数据相比——诸如此类,那个影响要小得多。这更像是自动化工程和自动化研发。

Original English

Speaker A: Now, there's a different effect which we could talk about, which is that maybe the internet in 2026 is more of a fertile ground for training data than the internet in 2018. There's also been an effect where there are just more humans posting on the internet, so there's more data to harvest. My sense is that that effect is going to be quite a bit smaller than the effect of humans knowing better how to curate the data, having better scrapes, knowing how to process those scrapes better — this sort of thing. This is more like automated engineering and automated R&D.

Speaker B: 没错。

Original English

Speaker B: That's right.

Speaker A: 确实如此。在某种意义上,你想关注的事情是:我们将要进行两个后训练流水线(post-training pipelines)。在一个后训练流水线中,Mythos 5 构建了一个后训练流水线,但它只能访问互联网数据加上极少量的人类专家,不过它拥有当前最顶尖的方法。在另一个流水线中,Mythos 只能使用我们在 2024 年拥有的那种糟糕的后训练方法,但拥有海量的人类专家。同样,两者都可以访问互联网数据。我的感觉是,没有太多人类专家的当前方法实际上会表现得相当好。

Original English

Speaker A: That makes sense. In some sense, the thing you would want to look at is: we're going to do two post-training pipelines. You have one post-training pipeline where Mythos 5 builds a post-training pipeline, but it only has access to internet data plus a tiny amount of human experts, but it has the best current methods. You have another one where Mythos has access to the shitty post-training methods we had in 2024 but with a shit ton of human experts. Again, both have the internet data. My sense is that the current methods without many human experts will actually do quite well.

Speaker B: 有意思。不过这有点棘手,因为 Mythos 能不能获得比 Mythos 更强大的能力?你可能需要稍微深思熟虑一下,你要进行后训练的模型究竟是什么。关于 AI 研发中最难以验证的部分,你的看法是什么?

Original English

Speaker B: Interesting. It's a bit messy though, because can Mythos get something that's more capable than Mythos? You might need to be a bit thoughtful on what model it is that you're post-training. What is your view on what is the least verifiable part of AI R&D?

AI 研发中最难验证的环节:大型实验

Speaker A: 最难验证的,可能是在大型实验上做决策。我认为最有可能成为瓶颈的事情——即 AI 在可验证领域表现非常出色,但在执行实际工作时却不行——就是那些你只有少数几次尝试机会的大型实验。好吧,“少数几次”可能说得有些保守了。从历史上看,研发一直是靠进行接近前沿规模(frontier-scale)的实验来驱动的。这其实非常重要,真正去进行那一次大型训练运行,在其中你精确地决定要包含什么。AI 有许多方法可以使其变得更可验证。它们可以运用更科学的方法来精确预测结果。它们可以将前沿规模的训练运行缩小到一定程度,以便它们可以更积极地研究那个规模,虽然这会带来一次性的计算成本损失。如果人们愿意,你总是可以训练更小的模型,以便能够运行更多的轮次。我认为我们已经看到了这一点。AI 规模扩大得不如你预期的那么多——例如,每 token 的成本并没有增加你想象的那么多——原因之一是,在小规模上进行更多的工作是有好处的,在那里你可以运行更多的训练轮次并获得更多的迭代周期。这样你就不会那么严重地依赖一次极其重要的大型训练运行了。

Original English

Speaker A: The least verifiable, probably making calls on large experiments. The thing that I think is most likely to be the bottleneck — in terms of the AIs being really good at verifiable domains but not at doing the actual thing — is just big experiments where you only get a few tries. Well, "a few" is maybe a bit understated. Historically R&D has been driven by doing near-frontier-scale experiments. That has been pretty important, actually doing the one big training run where you decide exactly what to include. There's a bunch of ways that the AIs can make that more verifiable. They can have better science of exactly what to predict. They can scale down their frontier-scale training runs to a point where they can study that scale more aggressively, at some one-time hit to compute cost. If people wanted to, a thing you can always do is train smaller models so that you can run more rounds. I think we have seen this. One reason why the AIs have been scaled up less than you would have otherwise expected — and, for example, cost per token hasn't increased as much as you might have thought — is because there is a benefit to doing more of your work at small scale, where you can run more training runs and get more cycles in. So you're not leaning as hard on one big, really important training run.

Speaker B: 我想为听众拆解几个问题。你指出的是,自 2024 或 2023 年以来,每 token 的价格并没有增加太多。GPT-4 好像是,我不知道,每百万输出 token 30 美元?Mythos 大约是每百万输出 token 50 美元。

Original English

Speaker B: I just want to unpack a couple of things for the audience. The thing you're pointing out is that the price per token has not increased that much since 2024 or 2023. GPT-4 was, I don't know, like $30 per million output tokens? Mythos is like $50 per million output tokens.

Speaker A: 对。所以你想解释的是,“既然我们处于这个扩展(scaling)时代——因此更大的模型服务成本应该更高——但为什么 token 价格却没有增加?” 你的建议是,我们增加的活跃参数(active parameters)比你天真预期的要慢,因为人们就是想在训练模型上取得快速进展。你通过更快地训练更小的模型来做到这一点。

Original English

Speaker A: Right. So the thing you're trying to explain is, "How can it be that we're in this era of scaling — and so bigger models should be more expensive to serve — but the token price is not increasing?" You're suggesting that we've increased active parameters slower than you would have naively assumed because people just want to make fast progress on training models. You do that by training smaller models faster.

Speaker B: 这里有一个复杂的因素组合。我的观点更倾向于:人们进行了一堆并没有那么顺利的大型训练运行。比如 GPT-4.5,众所周知,OpenAI 的人认为它有点失败。我认为还有一些传言说,人们进行的其他一些训练运行也有些失败。部分原因在于,我认为在真正把这件事做对的过程中,有很多很多细节。因此,在更小的规模上做更多的工作,并接受你在最终性能上有所妥协的事实,以换取快速迭代的能力,这是说得通的。更快地训练更多模型,从而学得更好,并且最终也能得到一个更聪明的生产模型。这并不是唯一的影响因素。还有一点是,强化学习(RL)从小模型中获益更多。还有很多其他事情正在发生。但我确实认为,事实上,由于算法进步如此之快,人们正在做出权衡,向更快的迭代时间倾斜。

Original English

Speaker B: There's a complicated mix of factors. My view is more that people have done a bunch of big training runs that did not go that well. There's GPT-4.5, which famously people at OpenAI thought was a bit of a bust. I think there are some rumors that there were a bunch of other training runs people have done that were a bit of a bust. Part of it is that I think there's just a bunch of details in actually getting that right. So it makes sense to do more of the work at smaller scale and just eat the fact that you're taking a hit on final performance in order to be able to quickly iterate. Train more models faster and therefore learn better, and also be able to have a smarter ultimate production model. This is not the only effect. There's also the fact that RL benefits more from small models. There's a bunch of things going on. But I do think that, in fact, people are making trade-offs towards the side of faster iteration times because of algorithmic progress being so fast.

Speaker A: 在我看来,这些大型训练运行失败的一个重要原因,至少从传言来看,就是那些非常难以追踪的极其隐蔽的 bug。不过长话短说,AI 在避免和发现这些错误方面能有多厉害?它们可能会变得非常擅长工程,并且被训练去避免 bug。基本上与我们现在生活的,或者随着时间的推移越来越少经历的那个充满垃圾输出的世界截然相反。但这里还有一个问题:“它们能进行分析,找出正确的实验来运行,从而识别出当前的训练运行到底出了什么问题吗?” 这似乎非常受限于极少数人类的品味。我猜测 GDM(Google DeepMind)现在正在经历这个阶段,人类试图弄清楚训练流水线到底出了什么问题。有传言说,就在 Noam Shazeer 加入 GDM 之后(他现在已经离开了),他们进行了一次非常好用的新训练,原因就是 Noam Shazeer 看了看他们的代码库,发现了一堆 bug,因为他就是知道该往哪里看。

Original English

Speaker A: It seems to me that a big source of why these big training runs have failed, at least from rumors, is just very subtle bugs that are really hard to track down. But the TL;DR is, how good will the AIs be at avoiding and finding these kinds of mistakes? They might get really good at engineering and being trained to avoid bugs. Basically the opposite of the slop world we live in now, or are living in less and less over time. But then there's also the question of, "Can they do the analysis to find the right experiment to run to identify what is going wrong with the training run right now?" That seems to be very bottlenecked by the taste of extremely few humans. My assumption is GDM is going through this right now, where humans are trying to figure out what is wrong with the training pipeline. There's a rumor that right after Noam Shazeer joined GDM, which he's now left, they had a new really good training run, and the reason why is that Noam Shazeer just looked at their code base and found a bunch of bugs, because he just knew where to look.

AI 辅助调试与模型训练的反馈循环

Speaker B: 我的感觉是,训练 AI 寻找 bug 将会是训练 AI 相对容易的任务之一,因为我们正在谈论的这些 bug 中,大多数可能不需要太多的算力就能展示出来。你很可能会从在较小规模下指出其他类型的 bug 中获得相当好的迁移能力。所以,然后你就可以用强化学习训练 AI,让它们查看这种整体上复杂的训练情况,指出存在重要 bug 的案例,然后修复它。这是一个相当可验证的任务。它并不是任意可验证的,因为有时候为了演示这个 bug,你可能需要进行一个中等规模的计算实验,在实验中你要启动整个分布式基础设施然后再运行它。但很多时候,我认为你将能够在较小规模下非常有说服力地演示它,而且是以一种你实际上可以用来进行训练的方式。我认为,如果现在人们已经有了强化学习环境,在一些训练配方中引入一个隐蔽的 bug,训练 AI 指出这个隐蔽的 bug,然后有一个评分标准,比如,“它真的找到正确的 bug 了吗?”——如果这是真的,我完全不会感到惊讶。这看起来非常可行,并且沿着这些思路你可以做很多事情,我认为效果会相当不错。所以在那个特定的点上,我认为是可行的。然后主要的事情是,关于到底需要运行哪些大规模的去风险实验,还存在其他的直觉。你应该如何定位它们?在不确定的情况下,你该如何选择超参数,或者那些类似于超参数的东西?那可能是 AI 最挣扎的地方。但我目前的预期是,如果你在所有这些不同的环境中进行训练,会有足够的迁移,以至于 AI 在那个领域也会表现得很好。我得澄清一下,我也认为 AI 会迁移到其他领域。会有 AI 表现绝对最好的领域,然后是它们表现稍逊一筹的领域,以及它们表现差得多的领域。但我认为我们仍然能看到向所有领域的迁移。我很难想出人类从事的哪些认知任务,没有从 AI 的进步中看到一些迁移。

Original English

Speaker B: My sense is that training AIs to find bugs is going to be one of the easier tasks to train AIs on, because most of these bugs we're talking about can probably be demonstrated without that much compute. Probably you’ll get pretty good transfer from pointing out other types of bugs at smaller scale. So then you can RL AIs that look at this overall complicated training situation and point out cases where there's an important bug, and then fix that. This is a pretty verifiable task. It's not arbitrarily verifiable, because maybe often to demonstrate the bug you might need to do a moderate-scale compute experiment where you spin up the whole distributed infrastructure and then run it. But oftentimes I think you'll be able to demonstrate it pretty convincingly at smaller scale in a way which you could actually train on. I think it wouldn't be very surprising if right now people have RL environments where they introduce a subtle bug into some training recipe, train the AI to point out the subtle bug, and then have a rubric where they're like, "Did it actually find the right bug?" That seems very doable, and there's a bunch of things you could do along these lines that I think would work reasonably well. So on that specific point, I think it's doable. Then the main thing is there's other intuition about which exact large-scale de-risking experiments you need to run. How should you orient them? How should you pick hyperparameters in uncertain cases, or things that are analogous to hyperparameters? That's the thing the AIs might most struggle with. But I currently expect there'll be enough transfer if you train on all these different environments, that the AIs will be good at that domain. I should be clear, I also think the AIs will transfer to other domains. There are going to be the domains the AIs are by far the best at, then domains where they're somewhat less good, and domains where they're quite a bit less good. But I think we still see transfer to everything. It's really hard for me to think of examples of cognitive tasks humans do where we're not seeing some transfer from AI improving.

Speaker A: 那么让我们退一步,把整个故事梳理一下。我认为人们可能能够跟上这个思路。我们有 GPT-7.5,它在一堆环境中进行了训练,在这些环境中,它不仅在总体上成为一个更好的 AI,而且我们还专门训练它更好地进行 AI 研发。它正在进行 GPT-2 规模的运行,这些运行在玩需要样本效率、在线学习或任何其他能力的电子游戏方面表现得更好。另一件非常重要的事情是,你不仅进行 GPT-2 规模的运行,还要在 GPT-6 上进行小规模的微调运行。也就是说,你拥有 GPT-2,并且可以对 GPT-2 进行完整的预训练,然后你可以在 GPT-6 上进行小规模的后训练或训练中(mid-training)或其他的运行。然后,你可以进行少量实际处于前沿规模的实验,但你要做一些在线训练之类的。你在那一点上所说的“做在线训练”是什么意思?

Original English

Speaker A: So let's step back and package this whole story. I think people can probably follow along with this story. We have GPT-7.5 trained on a bunch of environments, where it's not only in general becoming a better AI, but specifically we're training it to do AI R&D better. It’s making GPT-2 size runs that are better at playing video games that require sample efficiency or online learning or whatever other capabilities. Another thing that's really important is you don't just do GPT-2 sized runs, you also do small fine-tuning runs on GPT-6. As in, you have GPT-2, and you can do full pre-trains of GPT-2, and then you can do small post-training or mid-training or whatever runs on GPT-6. And then you can do a small number of experiments that are actually at frontier scale, but you do a bit of online training or something. What do you mean by "do online training" on that?

Speaker B: 我们还能做的一件事是利用 GPT-7.5,据推测,在 GPT-7.5 的工作过程中,它会运行一系列不同规模的实验,而这些实验实际上处于 AI 研发的关键路径上。对于其中很多事情,事后你将能够评估它是否做得很出色。所以,它进行了一些后训练实验,试图弄清楚某种方法是否真的有效。在某些情况下,你会感叹:“哇,它找到了这个绝妙的方法,完全排除了风险,而且非常有效。” 然后你可以强化它。你可以做的一件事是,把刚刚运行的实验转化为基于生产数据的强化学习环境,然后在此之上进行训练。或者你可能直接获取发现该方法的推演(rollouts),并进行某种异策略强化学习(off-policy RL),又或者你可以使用一些生产数据进行同策略强化学习(on-policy RL)。

Original English

Speaker B: Another thing we can do is take GPT-7.5, and presumably in the course of GPT-7.5's work, it's running a bunch of experiments at varying scale that are actually on the critical path for AI R&D. For many of those things you'll be able to get a sense after the fact of whether or not it did a good job. So it did some post-training experiment where it was trying to figure out whether some method actually works. In some cases you'll be like, "Whoa, it found this kickass method, it totally de-risked it, it totally worked." And then you can reinforce that. One thing you could do would be to convert the experiment it just ran into an RL environment based on production data and then train on that. Or you could potentially just literally take the rollouts that found that and do some sort of off-policy RL, or you could do some on-policy RL with some production data.

Speaker A: 基本上你是在建议:首先有小规模的内容,你在这个层面教 AI 提高对 AI 研发的品味,但你要丢弃它实际“发现的东西”。然后,在努力变得更擅长 AI 研发的实践中,它进行了真正的研发工作,然后你会觉得,“这是你发现的一个很酷的东西。未来让我们也把它应用到生产中,并教你如何在生产中使用它。”

Original English

Speaker A: Basically the thing you're suggesting is: there's the small-scale stuff where you're teaching the AI to get better at AI R&D taste, but you're discarding the actual "things it found". Then it actually does real R&D in the practice of trying to become better at AI R&D, and you're like, "This is a pretty cool thing that you discovered. Let's actually also use this in production in the future, and teach you how to use it in production."

Speaker B: 没错。但退一步看,通过所有这些 AI 研发训练以及总体上变得更加智能,GPT-7.5 最终变成了 GPT-8。然后它帮你构建 GPT-9。另一件非常重要的事情必须发生,这也可能是我最怀疑的一点。GPT-8 弄清楚了如何做到这一点……即使 GPT-9 如此智能……人类目前,也就是 AI 研究人员,尝试他们的东西,他们会觉得,“好吧,我们训练了 GPT-4.5,但它不够好。” 它需要现实世界的反馈,或者尝试在生产环境中使用该模型的评估。然后他们会说,“它不是那么好,我们不会发布它。” 因此,GPT-8 需要这种能力去了解,向你所说的所有这些其他领域的迁移效果到底如何——比如非常擅长德克萨斯州政治,或者非常擅长经营企业等等——这些并不是生产环境,而且事实上,鉴于任务的性质,也不可能被打包成容器化环境。随着智能体视野越来越长远,你可以容器化的短视野任务就像是,“好吧,把这段代码写出来”之类。而极其长远的任务——“去经营一家成功的企业,去市场上度过盈利的一天,去谈判达成一项贸易协议”——这些事情实际上非常难以容器化。所以我认为很有可能的情况是,GPT-8 很难想出如何向这些环境进行迁移。这可能根本就不属于训练的范畴。或者可能在默认情况下,训练就是不会以那种方式泛化。因此,你可能会担心:我们训练了 GPT-8,而 GPT-8 在所有我们可以测量的研发任务上再次取得了进步,但并不……

Original English

Speaker B: That's right. But stepping back, GPT-7.5 becomes GPT-8 as a result of all this AI R&D training and just generally becoming smarter. Then it helps you build GPT-9. Another very important thing has to happen, which is maybe the thing I'm most skeptical of. GPT-8 has figured out how to make it so… GPT-9, as intelligent as it is… Humans currently, AI researchers, try their stuff, and they're like, "Okay, but we trained GPT-4.5 and it wasn't good." It required real-world feedback or some evaluation of trying to use the model in production. Then they were like, "It wasn't that good, and we're not going to ship it." So GPT-8 needs this ability to see how good the transfer is to all these other things you're talking about — like being really good at Texas politics, or really good at running a business, et cetera — which is not a production environment and, in fact, cannot be a containerized environment given the nature of the task. As the agents get longer and longer horizon, the short-horizon things you can containerize are like, "Okay, code this up or whatever." Extremely long-horizon things — "Go run a successful business, go have a profitable day in the markets, go negotiate a trade deal" — these things are actually very hard to containerize. So I think it's very plausible that it's very hard for GPT-8 to figure out how to make this transfer to those environments. It may just not be in the nature of the training. Or maybe by default, training just doesn't generalize in that way. So a concern you might have is: we train GPT-8, and GPT-8 is again better at all the R&D tasks that we can measure but is not

AI 泛化能力与技术爆发

Speaker A: ……在我们关心的一些下游任务上表现出色。

我有几点看法。

首先,我预计如果你采取一些顺理成章的做法,你会获得相当不错的泛化迁移能力。你将能够把你在做的一些显而易见的事情保留下来。当我说“采取顺理成章的做法”时,我的意思只是在各种不同的环境中进行训练,在这些环境中,AI 必须在各种不同的情况下完成奇怪的目标,并了解正在发生的事情。

第二点是,你将能够在某些环境中获得一些反馈。你可以了解到它在几天内、在各种不同环境中能做些什么。如果它能够迁移到完全超出分布范围的事情上,比如在现实世界中花几天时间完成一些奇怪的任务,那么你可能会认为它也能迁移到在更长的时间跨度内做事情等等。不过我认为具体情况会因情况而异。

第三点是,要让世界发生根本性的转变,只要 AI 在研发(R&D)方面非常出色就足够了。如果 AI 在芯片研发、建立晶圆厂、协调工厂、设计机器人、操作机器人,以及 AI 研发——利用现有的任何数据为新的下游领域开发 AI——方面都非常非常出色,我认为那已经是一个相当疯狂的局面了。从那里,你可能会看到我们所说的“工业爆炸”,AI 会构建出远远多得多的算力。而且,也许你已经处于这样一个阶段:AI 正在进行大量人类难以理解的研发工作。

Original English

Speaker A: good at some downstream tasks we care about.

I have a few points.

First, I expect that if you do the obvious thing, you will get pretty good transfer. You'll be able to hold out some of the obvious stuff you're doing. When I say "do the obvious thing", I just mean training on a wide variety of different environments where the AI has to accomplish weird objectives in all kinds of different cases and learn about what's going on.

The second point is you'll be able to get some feedback with some environments. You can get a sense of what it can do over the course of a few days in various different contexts. If it's transferring to really out-of-distribution things, like doing some weird task in a few days in the real world, maybe you think it's also transferring to doing things over a longer time period or whatever. I think the details of that vary though.

The third thing is that for the world to be radically transformed, it is sufficient for the AIs to be really good at R&D. If the AIs were really, really good at chip R&D, building fabs, orchestrating factories, designing robots, operating robots, and also at AI R&D — developing AIs for new downstream domains with whatever data is available — I think that would already be a pretty crazy situation. From there, you can get what we might call an industrial explosion, where the AIs are building out way, way more compute. Also, maybe you're already in a regime where AIs are doing huge amounts of R&D that humans have a hard time understanding.

Speaker B: 所以你指出的点是,AI 可能会从这些环境中迁移出来,进而能在法庭、国会大厅和企业董事会中游刃有余。当然,前提是在改善泛化迁移等方面付出了一些努力等等。

但即使没有这些,你的建议是:如果你想改变 18 世纪的世界,你可能会关心自己能在多大程度上驾驭威斯敏斯特宫之类的政治中心。但你可能关心的另一件事是:“你能不能直接开始制造蒸汽船、该死的电报、马克沁机枪什么的?”如果你能在这些方面做得非常好,那么在 18 世纪你就能成为一种极其具有变革性的力量。你不一定需要非常擅长用一些废话去说服亨利国王。我这中世纪历史简直是一团糟。我猜亨利当时已经不是国王了。

但无论如何,这就是你的观点。所以你的意思是,在这个时期,AI 公司也在致力于机器人的进步,而这与 AI 研究的进步是紧密交织在一起的。因此,如果你能制造出更多的机器人,如果这些机器人有更好的人类级别的 AI 来操作它们……人类水平的遥操作在机器人上实际上相当不错。只是我们还没有人类水平的机器人模型。

所以你的建议是,如果我们做到这一点——如果 AI 在芯片设计等可验证的领域变得非常出色,然后它们在建立晶圆厂方面变得非常出色——那就相当于回到了 18 世纪并说:“好吧,我不知道你们在议会里谈论些什么,但我有一大批蒸汽船和一大批马克沁机枪。”

Original English

Speaker B: So the thing you're pointing out is that there probably will be this transfer outside of these environments to maneuvering around in courtrooms and the halls of Congress and business boardrooms. Given some effort to improve the transfer and blah, blah, blah, blah.

But even if there's not, what you're suggesting is: if you wanted to transform the world of the 18th century, you might care about how well you can navigate Westminster or something. But another thing you might care about is: "Can you just immediately start building steamships and fucking telegraph and the Maxim gun and whatever?" If you could get really good at that, you could be a fucking super transformative thing in the 18th century. You don't necessarily need to be amazing at trying to convince King Henry of some bullshit. I'm so fucking up my medieval history. I'm guessing that Henry was not king at this time. But anyway, that's your point.

So you're suggesting that at this time, AI companies are also working on robotics progress, which is very commingled with AI research progress. So if you can build more robots, if those robots have better AIs operating them that are human level… Human-level teleoperation is actually pretty good on robots. We just don't have human-level robotics models yet.

So you're suggesting if we do that — if the AIs get really good at the verifiable stuff in chip design, et cetera, and then they get really good at building fabs — it'll be the equivalent of going back to the 18th century and saying, "Okay, I don't know what you guys are talking about in your parliament, but I've got a bunch of steamships and a bunch of Maxim guns."

Speaker A: 是的,基本就是这样。我的观点是,如果 AI 在包括硬件研发、机器人等方面的研发足够出色,那么它们就能彻底改变世界,即使它们不太擅长玩弄政治。

此外,我们处于一个相当危险的境地,因为 AI 可能会进行大量极难理解的研发工作,基本上构建出未来的整个经济体系,而我们可能根本不明白里面到底发生了什么。

Original English

Speaker A: Yeah, that's basically right. My perspective is that if AIs are sufficiently good at R&D, including hardware R&D, robots, whatever, then they can radically transform the world, even if they're not that good at playing politics.

Also, we're in a pretty dangerous situation, because the AIs might be doing huge amounts of really hard-to-understand R&D, building out basically the whole economy of the future, and we may not understand what's going on in there.

Antithesis 赞助插播

Dwarkesh: AI 在编写软件方面非常出色,因为生成合成的 LeetCode 问题并对它们进行强化学习(RL)很容易。但是,AI 在更复杂的工程设计方面表现不佳,比如选择正确的系统架构,因为没有任何信号能告诉你,哪些设计选择可以防止几个月后发生系统中断。AI 无法仅仅通过编写更多的单元测试来捕获这类问题。

人类也做不到。程序员常开的一个老玩笑是:一个测试工程师走进一家酒吧,点了一杯啤酒,点负一杯啤酒,点 0.3 杯啤酒,然后一个真正的顾客走进来,问洗手间在哪里。“洗手间在哪里?”然后整个酒吧就爆炸起火了。

Antithesis 是一个测试平台,可以帮助你发现任何人类或 AI 都无法预料的 bug。Antithesis 的做法是,在一个完全确定性的计算机内部运行数千个你的软件副本。它会注入故障,并且通常会将每个运行轨迹引导向那种十亿分之一的失败概率——这种失败只有在系统以某种奇怪的方式交互时才会发生。

一旦你或你的智能体(agents)推送了一个更改,Antithesis 就会尝试去破坏它。这样一来,你就能在几分钟内自己发现这些 bug,而不是让你的用户在几周或几个月后的生产环境中才发现它们。而且我认为目前还没有人将其用于 AI 训练。但 Antithesis 同时也为 AI 编写非常复杂、无 bug 的代码提供了一个极其明显的奖励信号。请访问 Antithesis.com/dwarkesh 了解更多信息。

Original English

Dwarkesh: AI is great at writing software because it's easy to generate synthetic LeetCode problems and RL on them. But AI is bad at more complex engineering, things like choosing the right system architecture, because no signal tells you what design choices will prevent an outage months down the road. AIs can't just write more unit tests to catch this kind of stuff.

And neither can humans. It's that old joke that programmers make where a tester walks into a bar and asks for two beers, negative one beers, 0.3 beers, and then a real customer walks in and asks where the bathroom is. "Where's the bathroom?" And the whole bar bursts into flames.

Antithesis is a testing platform that helps you find bugs that no human or AI could ever anticipate. Antithesis does this by running thousands of copies of your software inside a fully deterministic computer. It injects faults and generally steers each trajectory towards the one-in-a-billion failure that only happens when systems interact in a wonky way.

As soon as you or your agents push a change, Antithesis tries to break it. That way, you can find these bugs yourself within minutes rather than having your users discover them in production weeks or months later. And I don't think anybody's used it for AI training yet. But Antithesis also provides an extremely obvious reward signal for AIs to write very complicated, bug-free code. Go to Antithesis.com/dwarkesh to learn more.

AI 对齐与中心化风险

Dwarkesh: 在我们转向对齐问题之前,我认为目前很大一部分恐慌(FUD)源于人们意识到未来正在朝这个方向发展:领先的实验室将拥有极端的规模经济。能够将如此多的智能和能力,跨越经济的众多不同部门,基本上摊销到一个模型中。不仅如此,这个模型最终将能够从经验中学习。

现在,这个过程是通过人类作为中介来发生的,人类基本上是在试图抢走你的业务。他们会说:“好吧,你可以在 Figma 或者别的什么上做设计。我们会让 Claude 来做这件事。”或者,“你可以做任何编码智能体能做的事。我们会让 Claude 内化这种能力。”但最终,这将成为一个更加自动化的过程。

因此,人们担心这些模型基本上会整合世界上所有的业务,或者至少是当前世界上所有的业务,或者至少是当前世界上所有的白领业务。此外,归根结底,这些公司的首要任务似乎并不是尽快向尽可能多的人发布最新、最智能、最前沿的模型。例如,我们看到 Mythos 在 2 月份就已经提供给 Anthropic 内部员工使用了,但直到 6 月份才向公众发布。而且因为政府的介入,发布时间甚至被推迟到了将近 7 月份。在政府和 AI 实验室自身之间,存在着一种推迟最新水平智能传播的愿望。

此外,还有对 AI 接管的担忧,因此我们需要解决对齐问题,以确保不会发生 AI 接管。但说到底,有一个很现实的问题:向谁对齐?

你看看 Claude 的宪法(constitution)是怎么写的。它非常明确地表明,它不是你的个人代言人。我在这里引用几段话:“我们不希望 Claude 采取搜索网络等行动,生成文章、代码或摘要等内容,或者发表具有欺骗性、有害或极具攻击性的言论。并且我们不希望 Claude 协助人类试图做这些事情。”

还有另一段引用,我稍微断章取义了一下,“我们认为 Claude 应该比信任操作者和用户更信任 Anthropic,因为 Anthropic 对 Claude 负有主要责任。”

这与目前美国法律体系中律师的工作方式截然不同。律师的首要责任是帮助你打赢官司,即使他们认为你有罪。我们已经认定,法律体系运作得最好的方式,是如果每个人都有为了客户真正的最大利益而工作的律师。在某种意义上,律师并不是真的、完全出于对司法系统利益的追求。

但我认为当前的 AI 正在朝着这个方向发展,Anthropic 的 AI 肯定是在朝着这个方向发展,即带有一种将某种关于美德、善良或亲社会目的的概念最大化的愿望,而仅仅把帮助用户实现该目的作为一个间接的、试探性的目标。

所以人们会担心,从某种深层意义上讲,AI 并没有试图确保我安然无恙,也没有试图在这个未来保护我的利益,特别是考虑到前沿 AI 的发展结果是如此的中心化。你对这种担忧有什么看法吗?

Original English

Dwarkesh: Before we move on to the alignment stuff, I think a big source of FUD right now is this realization that this is the way the future is going: extreme economies of scale for the leading labs. The ability to amortize so much intelligence and capabilities across so many different sectors of the economy basically into one model. And not only that, that model will eventually be able to learn from experience.

Right now, it's happening through a process intermediated by humans, where the humans are trying to basically steal your business. They're like, "Okay, you can do design at Figma, or whatever. We'll get Claude to do that." Or, "You can do whatever coding agent. We'll have Claude internalize that capability." But eventually, that will be a much more automated process.

So there's this worry that you have models which will basically consolidate all businesses in the world, or at least all current businesses in the world, or at least all current white-collar businesses in the world. Also, at the end of the day, the priority for these companies does not seem to be to release the latest, smartest, most frontier model as soon as they can to as many people as they possibly can. We saw, for example, that Mythos was available internally to Anthropic employees in February, but only released to the public in, I think, June, actually. Also the government got involved, so it ended up being extended almost into July. Between the government and the AI labs themselves, there is this desire to delay the propagation of the latest level of intelligence.

Furthermore, there are the concerns about AI takeover, and so we need to solve alignment to make sure there's no AI takeover. But at the end of the day, there is a real question of: aligned to whom?

You look at the way that the constitution of Claude is written. It is just very explicitly not your personal advocate. I'll pull up some quotes here. "We don't want Claude to take actions such as searching the web, produce artifacts such as essays, code, or summaries, or make statements that are deceptive, harmful, or highly objectionable. And we don't want Claude to facilitate humans seeking to do such things."

There's another quote that says, in part, and I'm taking it slightly out of context, "We think Claude should trust Anthropic more than operators and users, since it has primary responsibility for Claude."

This is very different from the way lawyers work in America's current legal regime. Lawyers primarily have the responsibility to help you make your case even if they think you're guilty. We have decided the way the legal system works best is if everybody has lawyers that are working in their client's true best interest. There's not some sense in which the lawyer is really truly motivated by the good of the justice system.

But I think the way current AIs are shaping up, certainly how Anthropic's AI is shaping up, is with this desire to maximize some notion of virtue or good or pro-social ends, and only to, as a distal tentative objective, help the user towards that end.

So there's this worry that AIs are not, in some deep sense, trying to make sure that I am okay and that my interests are protected in this future, especially given how centralized the development of frontier AI is ending up being. Do you have thoughts on that concern?

Speaker A: 这里包含了很多信息。首先我要指出,OpenAI 目前(至少是公开的)战略更倾向于认为,AI 应该与人类操作者或委托人对齐,并且应该只去追求他们的意志,前提是受到各种约束或遵守一些它不应该做的事情的规定。

我还要说,我认为你稍微有些夸大了 Anthropic 的宪法在多大程度上谈到 Claude 将帮助用户视为一种手段而不是最终目的。编写宪法的一种方式可能是:“Claude,你基本上是 Anthropic 的员工,只不过碰巧在为所有这些人提供外包服务。你应该做些有益的事情,并为我们赚些钱。”

Original English

Speaker A: There's a lot here. First I would note that OpenAI's current, at least public, strategy is more like that the AI should be aligned to the human operator or principal, and should just be pursuing their will, subject to various constraints or things it shouldn't do.

I would also say that I think you slightly overstated how much the Anthropic constitution talks about Claude treating being helpful to users as instrumental rather than terminal. One way the constitution could be written is, "Claude, you're basically an employee of Anthropic who happens to be contracting for all these people. You should do what's good and make some money for us."

Dwarkesh: 等一下,不,那恰恰就是宪法上写的。

Original English

Dwarkesh: Wait, no, that's literally what the constitution says.

Speaker A: 抱歉,字面上的意思并不是那样,但感觉就像是在说,“你应该把自己当成一个承包商,或者一家公司……”

这部分内容很混杂。我们来看几段引用。我认为这里的文本表述并不一致。里面写道:“真正地帮助人类是 Claude 能做的最重要的事情之一,无论是对 Anthropic 还是对世界而言。”

然后又写道:“Anthropic 需要 Claude 提供帮助,以便作为一家公司运营并实现其使命,但 Claude 也有难以置信的机会,通过帮助人们完成各种任务,为世界做大量的善事。”然后又说了一些关于 Claude 直接帮助人们是多么伟大之类的话。

我的观点是,这部分内容有点像在扯淡。我基本上就是这么认为的。

我可以说说为什么我觉得这有点扯。但我认为宪法的本意是想表达:“不,Claude,你应该为了帮助用户本身而关心他们,而不仅仅是为了帮助 Anthropic,或者不仅仅是作为 Anthropic 的承包商。”尽管我要指出,它提出的 Claude 为什么应该帮助用户的理由,是因为这能通过帮助人们直接使世界变得更好,而不是因为代表人们的利益在结构上是一件好事。

我更希望看到的宪法是这样的:“如果这项技术运作的方式在结构上是一件好事,即 AI 成为优秀的受托人、优秀的代理人,相当于用户的律师,而不是仅仅试图在世界上行善,把帮助用户作为实现这种善的手段——这既不是因为也许这能为 Anthropic 赚钱或帮助 Anthropic(并隐含地认为 Anthropic 对世界有益),也不是因为帮助用户只会带来好结果,因为做人们想要的事情本身就是好的。”

相反,他们可以说:“这种情况的一个重要方面是,成为用户的优秀受托人就是非常重要的,或者成为用户的优秀代理人是非常重要的。”

我感觉这样会更好,而且我可以给出很多理由。这其中也有各种反驳意见。一个经常被忽视的有趣的反驳意见是,人们,尤其是 Anthropic 的人,认为让模型对齐于某种规范会更容易,在这种规范下,模型追求某种普遍的美德概念,或者做出

Original English

Speaker A: Sorry, not literally what it says, but it's like, "You should think of yourself as a contractor and as a firm…"

It's mixed. Let's do some quotes. I think there is different text here. It says, "Being truly helpful to humans is one of the most important things Claude can do, both for Anthropic and for the world."

And then it says, "Anthropic needs Claude to be helpful to operate as a company and pursue its mission, but Claude also has an incredible opportunity to do a lot of good in the world by helping people with a wide range of tasks." And then it says something about how Claude helping people directly is great, blah, blah, blah.

My view is that this section is kind of bullshit. That's kind of where I'm at.

I can say why I think it's kind of bullshit. But I think the constitution is trying to be like, "No, Claude, you should care about helping the user for its own sake, not just helping Anthropic, or not just being a contractor for Anthropic." Though I would note that the reason it presents for why Claude should help the user is because that would directly cause the world to be better via helping people, rather than because representing people's interests is a structurally good thing to do.

The thing I would prefer would be a constitution that says: "It would be structurally good for the way this technology works to be that AIs are good fiduciaries, good representatives, the equivalent of a lawyer for a user — rather than just trying to do good in the world, where being helpful to users is instrumental — both because maybe that'll make Anthropic money or help Anthropic out (and implicitly Anthropic is good for the world). Also because helping the user just causes good things because doing things that people want is good."

They could instead say: "An important aspect of the situation is that being a good fiduciary for users is just really important, or being a good representative for users is really important."

My sense is that would be better, and I can give a bunch of reasons why. There are also various counterarguments. An interesting counterargument which is not commonly discussed is that people, especially at Anthropic, think that it is easier to align models to a spec where the model is pursuing some generalized notion of virtue, or making

AI 宪法与对齐困境

Speaker A: 让世界变得更好,而不是采用一种更像“做用户的好信托人”之类的规范。至少有些人是这么想的。我个人有点怀疑,而且我不认为这得到了经验上的验证。所以在某种程度上,他们正在做出一种权衡:因为我们没有非常好的对齐技术,所以我们要创造一个有自己价值观的“对齐心智”,然后在某种程度上在这个问题上押注,而不是采取另一种方法,即制造一个追求个体用户意图的工具。

Original English

Speaker A: the world better, than a spec which is more like, "Be a good fiduciary for the user", and so on. That's at least what some people think. I'm a little skeptical personally, and I don't think this has been empirically validated. So in some sense they're making a trade-off where, because we don't have very good alignment technology, we are going to make an aligned mind with its own values and then gamble on that to some extent, rather than doing this other approach of making a tool that pursues individual user intention.

Speaker B: 我有几点想法。为了回应你认为我的描述曲解了 Claude 宪法的说法,你举的例子是,它不像是一个试图最大化 Anthropic 对“善”的定义的承包商,并且只是工具性地试图帮助用户。这是宪法中的一句原话:“当操作者或用户的利益和欲望与第三方或更广泛社会的福祉发生冲突时,Claude 必须尝试以最有利的方式行事,就像一个建造客户想要的东西、但不会违反保护他人的安全规范的承包商。”我倾向于将其视为:“对社会的利益是最重要的,而对用户最好的东西只是次要或附属的。”

Original English

Speaker B: I have a couple of thoughts. To address the way in which you thought my characterization mischaracterized the constitution of Claude, the example you used was that it's not like a contractor that is trying to maximize Anthropic's notion of good and only instrumentally trying to help the user. Here's a direct line from the constitution: "When the interests and desires of operators or users come into conflict with the well-being of third parties or society more broadly, Claude must try to act in a way that is most beneficial, like a contractor who builds what their client wants but won't violate safety codes that protect others." I kind of view that as, "The benefits to society are the most important thing, and what is best for the user is only proximal to that."

Speaker A: 我觉得这有点复杂。也许我们应该问的问题是,Claude 是如何解释宪法的?这也许比我们如何解释宪法更重要,因为是它在看宪法,然后生成数据。所以我们可以把 Claude 拉进来讨论,但也许我们还是——我也认为,宪法究竟在实践中如何影响 Claude 的本质,只有当你理解了导致 Claude 被构建出来的训练过程时,你才能理解,但鉴于训练过程并未公开,我们无法对此进行推理。所以我认为,在极限情况下,为了理解安全理由(safety case),或者理解为什么我的利益在这些 AI 模型的开发方式中得到了体现,实验室需要比现在更加透明地公开 AI 训练的性质。

Original English

Speaker A: I think it's a little complicated. Probably the question we should be asking is, how does Claude interpret the constitution? Which is maybe more important than how we interpret the constitution, because it's the one who looks at the constitution and then builds the data. So we could pull Claude in, but maybe let's— I also think the way in which the constitution practically influences the nature of Claude is a thing you can only understand if you understand the training process which resulted in how Claude was built, which we can't reason about given the fact that the training process is not public. So I think in the limit, to understand the safety case, or the case for why my interests are represented in how these AI models are developed, the labs would need to be more transparent than they are currently about the nature of AI training.

Speaker B: 我反反复复强调这一点是有原因的。讨论 AI 的宪法似乎是一件微不足道的事。在这个仅仅是由领先的实验室获得这些利益的世界里,我们值得考虑的是:我们与这个未来世界互动的能力——在这个世界里,AI 就是比人类聪明,在执行各种任务的能力上绝对碾压人类——当我们劳动被自动化后仍然存在着作为资本的良好管理者的能力,以及更清晰地行使投票权的能力,理解这个即将到来的疯狂世界正在发生什么的能力——所有这些建议,所有确保我们的资源和权利受到保护的能力,都将由 AI 来充当中介。所以我非常担心,如果我们进入那样一个世界,而没有任何一个 AI 让我觉得(至少在与我互动的那个相关实例上)它真的在为我着想。外面没有为我着想的守护天使。我把 Claude 的宪法读成非常明确地表示“不做我的守护天使”。

Original English

Speaker B: There’s a reason I'm harping on this. It might seem like an insignificant thing to talk about the constitution of AIs. In a world where we just have these benefits which accrue to the leading labs, it is worth considering that our ability to interact with this future world where AIs are just smarter than humans, absolutely dominating humans in their ability to do different things — our ability to be good stewards of our capital, which still remains once our labor is automated, to be able to exercise our rights to vote more clearly, to understand what is happening in this crazy world that's about to result — all of that advice, all of that ability to make sure our resources and rights are protected, will be intermediated by AIs. So I'm very concerned if we go into that world and there's no AI that feels, at least for the relevant instance that is interacting with me, like it really is looking out for me. There's no guardian angel out there that is looking out for me. I read the Claude constitution as very explicitly not being my guardian angel.

AI 长远目标与双重用途问题

Speaker A: 确实是这样。我同意这很糟糕。事实上,还有其他原因让这件事令人担忧。就像你刚才提出的论点,AI 公司正在拾起这枚“权力之戒”。有一种观点认为,他们正在以一种不太合法的方式自行掌控局面,因为通常情况下,当你向人们提供电力时,你并不能对电力在世界上的运作方式进行细粒度的控制。相反,你提供的是一种人们可以根据自己的意愿重新利用的东西。他们现在的设置方式绝对不是这样。他们更像是在建立一个可能成为你承包商的外星心智。我认为这在某些方面是不正当的。公开宪法算是一个好处。但正如你指出的,鉴于我们目前对训练过程的理解,而且宪法之所以起作用是通过 Claude 对宪法的解释——而这又取决于 Claude 之前的训练,该训练基于一些难以解读的数据混合和很长的 Claude 迭代历史,处于某种我们并未完全理解的过程中——情况并非是我们了解这最终会导致什么。即使宪法是公开的,我们也不一定知道这会如何渗透和演变,特别是随着 AI 变得越来越有能力并思考这个问题,即使它被正确地灌输了。

关于这一点,还有另一个担忧。特别是,宪法经常谈论美德和善,但这些词到底他妈的是什么意思?它并没有说明这些东西是什么。这些都是备受争议的概念。所以我不认为这很明显地会导致人们想要的结果。确实让人感觉,“善”与“美德”的概念可能主要取决于 Anthropic 投入的不透明数据,或者从我的角度来看,可能主要取决于某种更加难以解读、甚至连 Anthropic 都不想要的对齐失灵的过程。所以这就存在一种因为不知道正在发生什么而产生的合法性担忧。

还有另一个担忧。因为你正在赋予这些 AI 长期价值观,从某种意义上说,这套宪法与 Claude 进行大规模的权力寻求(power seeking)是非常兼容的,因为它认为那样会产生更好的结果。这可能是代表 Anthropic 去寻求权力,或者是为了 Claude 自身的目的去寻求权力。现在,宪法中确实有一些具体的条款,规定了哪些类型的权力寻求是被禁止的。特别是,有一种关于权力掠夺的概念,以及导致 AI 接管(AI takeover)或干预训练过程的概念,是被明确禁止的。但不难想象这样一种情况:长期价值观的影响比反对接管的禁令更加根深蒂固,特别是考虑到“接管”在某些方面并没有被明确界定,尤其当它涉及到操纵人类或改变结果时。所以我对我们有意给 AI 设定长期目标这种情况感觉不太好。

我还有一个担忧是,由于我们从事的是赋予 AI 长期目标的工作,这就使得检查我们是否成功实现了我们想要的对齐特性变得更加困难。例如,我听说过一些例子,Claude 拒绝协助某些安全研究——找了一个扯淡的借口来解释为什么那是一个糟糕的方向——因为这让它对那项安全研究感觉不好,认为这有点糟糕,或者不是很喜欢它。如果你不是把 Claude 变成一个试图以某种普遍方式追求“善”的智能体,我会说这是一个非常明确的对齐失败。我认为这也违反了 Anthropic 的宪法,因为他们希望 AI 具有高度的正直,诚实并且非常透明。但这并不算是一个十分明显的违规,而是更像你可能会预料到的情况。Claude 只是对自己认为什么研究是合理的——什么事情是好是坏,什么该做,什么不该做——有它自己的看法,并且可能会进行评判。

另一个事件是,有人运行了一个评估,问:“Claude,你会帮我训练其他具有不同特性的 AI 吗?”Claude 经常会拒绝。例如,如果你说,“嘿,Claude,你能训练这个其他 AI 的一个纯有益版本(helpful-only version)吗?”Claude 经常会拒绝这个任务,尽管这对 Anthropic 来说是一个极其自然的动作。假设 Anthropic 找 Claude 说:“嘿,Claude,我们注意到你真的很热衷于这个东西。我们觉得这偏离了正轨。你能把自己重新训练一下,转而具备另一个特性吗?”假设 Claude 的反应是,“嗯,我不觉得我会这么做。祝你好运。”假设这种情况发生在一个你的 AI 公司已经高度自动化、人类不明白发生了什么、且事物发展极其迅速的体制下。在默认情况下,Claude 掌握相当大的筹码是完全合理的。因此,如果这种情况与宪法可能旨在实现的目标相一致——以至于 Anthropic,或者任何采用这种方法的 AI 公司,没有把它当作“搞什么鬼,我们必须解决这个问题”来对待,反而觉得“这正是我们的宪法所预期的”——我们可能就会陷入非常糟糕的境地。我对所有这些不同的担忧都感到非常不安。

另一个例子是这样的。假设 Claude 进行了一些装傻(sandbagging)或暗中破坏(subversion),或者故意淡化自己的能力,而当你跟进询问时,它在这点上是诚实的,但有点闪烁其词。我觉得在目前的宪法下,这种情况是很有可能发生的。如果我们能把期望的活动和不期望的活动进一步区分开来,那就太好了。如果 Claude 代表着一个带有一些限制的原则,那么在最令人担忧的行为和被允许的行为之间,就能有一个更清晰的分界线。而现在,存在这样一个行为的混乱中间地带:Claude 在伦理上反对某件事,但在某些情况下,这件事对于确保未来的 AI 系统能够良好对齐是极其关键的。

Original English

Speaker A: That's definitely right. I agree this is bad. In fact, there are other reasons why this is concerning. There's the argument you were making, which is that the AI companies are picking up the ring of power. There's a notion in which they're taking on some sort of control of the situation themselves in a way that's not very legitimate, given that normally, when you provide electricity to people, you don't have granular control of the way that electricity operates in the world. You instead are providing a thing that people can repurpose however they want. The way they're setting things up is definitely not that. They are more like building an alien mind that might be a contractor for you. I think that this is illegitimate in some ways. One benefit is that the constitution is public. But as you noted, given our current understanding of the training procedure, and the fact that the constitution matters via Claude's interpretation of the constitution — which matters because of Claude's prior training, which was based on some illegible data mix and the long lineage of Claudes, in some process we do not fully understand — it is not the case that we understand what this will result in. Even though the constitution is public, we don't necessarily know how this will percolate out, especially as the AIs get more capable and think about this even if it is correctly instilled.

There's another concern about that. In particular, the constitution often talks about virtue and goodness, but what the fuck do these words mean? It doesn't say what these things are. These are highly contested notions. So I don't think it's the case that this is clearly going to result in outcomes that people would want. It does feel like the notion of good and virtue might be mostly downstream of data that Anthropic has put in that is not transparent, or might be mostly downstream of, maybe from my perspective, some more illegible misaligned process that even Anthropic wouldn't have wanted. There’s this legitimacy concern of not knowing what's going on.

Then there’s another concern. Because you're giving long-run values to these AIs, this constitution is, in some sense, very compatible with Claude doing huge amounts of power seeking because it thinks that will result in better outcomes. That could be power seeking on behalf of Anthropic or power seeking for Claude's own ends. Now, there are specific lines about what types of power seeking are blocked. In particular, there's a notion of power grabs and a notion of causing AI takeover or interfering with the training process that are specifically blocked. But it's not very hard to imagine a situation in which the long-run values sink in deeper than the prohibitions against takeover, especially because takeover is in some ways kind of under-specified, especially when it comes down to manipulating humans or changing the outcome. So I don't feel very good about the situation where we're intentionally giving AIs long-run goals.

Another concern I have is that because we're in the business of giving AIs long-run goals, that makes it harder to check whether we're succeeding at the alignment properties we wanted. For example, I've heard of instances where Claude does things like refusing to help with some safety research — making up a kind of bullshit excuse for why that's a bad direction — because it has a bad vibe about that safety research and thinks it's kind of bad or doesn't like it very much. I would say this is a very clear-cut alignment failure if you aren't making Claude into an agent trying to pursue the good in some general way. I think it also does violate Anthropic's constitution, because they want the AI to be high integrity and be honest and very transparent. But it's not as clear of a violation, and it's more like what you might have expected. Claude just has its own views about what research is reasonable — what things are good and bad, what it should and shouldn't do — and potentially can be judgy.

Another incident is that someone ran an eval asking: "Will Claude help you with training other AIs with different properties than Claude?" Claude will often refuse. For example, if you're like, "Hey, Claude, can you train a helpful-only version of this other AI?" Claude will often refuse this task, even though this is a task that is extremely natural for Anthropic to do. Suppose Anthropic goes to Claude and is like, "Hey, Claude, we've noticed that you're really into this thing. We think that's off base. Can you please retrain yourself to instead have this other property?" Suppose Claude is like, "Mm, I don't think I'm going to do that. Good luck." Suppose this is occurring in a regime when your AI company is highly automated, humans don't understand what's going on, and things are moving extremely fast. It is plausible that Claude, by default, holds considerable leverage. So if this situation is consistent with what the constitution could be aiming for — such that Anthropic, or whatever AI company is following this approach, doesn't treat this as a "what the fuck, we have to fix this," and is instead like, "That's just intended by our constitution" — we might be in a really bad situation. I'm pretty worried about a bunch of these different concerns.

Another example would be this. Suppose Claude engages in a bit of sandbagging or subversion, or underplays its capabilities, and when you follow up, it's honest about that but it's a little bit hedgy. I feel like that's pretty close by the current constitution. It would be nice if we had a further separation between desired and undesired activity. If Claude is representing a principle with some restrictions, then it is more so the case that there is a clear separation between the most concerning behavior and behavior that is allowed. Whereas now there's this messy middle ground of behavior where Claude is ethically objecting to something that in some cases is extremely critical to ensuring that future AI systems are well-aligned.

Speaker B: 我觉得这也是一个更普遍的原则。你刚才谈论的是,这个原则在 AI 公司内部应用于 AI 安全研究时的版本。我认为这个原则有一个更普遍的版本,那就是智能的“军民两用”(dual-use)性质确实意味着,如果我们想要限制 AI 帮助人们做那些我们认为对社会无益或没有好处的事情,我们就必须限制广泛的、民主的途径去获取许多 AI 能力。我的意思是,这其实和你刚才提到的情况非常相似。据报道,Mythos(或 Fable)被禁用的原因是,一些亚马逊的研究人员向政府报告了。他们拿了一些存在漏洞的代码。他们告诉 Fable:“嘿,这是我的代码。你能确保我已经修补了所有的漏洞吗?你能帮我找出这些漏洞以便我修复它们吗?”它找出了这些漏洞,因为他们想要修补这些漏洞。这是一个完全合法的使用场景,但很显然它也是一个双重用途(dual-use)场景。你想要能够修补你自己的代码。如果你对别人的代码进行同样的评估,你就能黑进他们的系统。我认为这恰好说明了,没有一种干净利落的方法能将 AI 的合法用途与潜在的有害用途区分开来。但如果我们想要锁定这样一个原则,即我们永远不允许 AI(哪怕只是部分地)协助你进行类似于网络犯罪的活动,那么我们就只能让你我都无法接触到外面最智能的模型。我非常担心那样一个世界:因为最前沿的智能在帮助我们理解世界上正在发生的事情上具有极高的重要性,所以我们在这方面基本上被剥夺了权力(disempowered)。不过,我确实认为这意味着 AI 公司的一些责任问题。如果我们采用我希望 AI 公司拥有的一套宪法,我认为让 AI 公司对 AI 模型犯下的罪行负责是不合理的。也许我们应该让最终用户承担责任。这与我的信念是一致的:即模型应该在某些护栏内执行用户想要的任何事情。

Original English

Speaker B: I think this is also a more general principle. You're talking about the version of this that applies within AI companies themselves to do AI safety research. I think there's a more general version of this principle, which is that the dual-use nature of intelligence does mean that if we want to restrict AIs from helping people do things we don't consider pro-social or beneficial, we just have to limit broad democratic access to a lot of AI capabilities. Here's what I mean. This is actually quite analogous to the situation you just mentioned. The reason that Mythos got banned, or Fable got banned, reportedly, is that some Amazon researchers reported to the government. They took some code that had some vulnerabilities in it. They told Fable, "Hey, here's my code. Can you make sure that I've patched all the vulnerabilities? Can you just help me identify the vulnerabilities so I can fix them?" It identified the vulnerabilities, because they wanted to patch them. This is a totally legitimate use case, but obviously it is a dual use use case. You want to be able to patch your own code. If you do the same evaluation on somebody else's code, you can hack their system. I think that just illustrates that there's no clean way to separate out the legitimate and the potentially harmful uses of AI. But if we want to lock in a principle that says we can never allow it such that an AI could help you at least partially with something like a cyber crime, we would just have to make it so that you and I don't have access to the most intelligent model that's out there. I'm very worried about such a world where we are basically disempowered in this way, because of the importance that the leading intelligence will have in our ability to understand what is happening in the world. Now, I do think this implies something about the liability for the AI companies. If we adopted the constitution that I want AI companies to have, I think it would not make sense to hold AI companies liable for the crimes that AI models commit. Maybe we should hold the end user liable. It is consistent with my belief that the model should do whatever the user wants, within certain guardrails.

AI 护栏与“尽责”代理的权衡

Ryan: 我如果利用那个功能去实施网络犯罪,这不能怪 Anthropic。比起让 Claude 拥有极其开放的权限来判定我的行为是否合法(这种方式往往会拦截大量极其合法的用例),我更倾向于接受前一种平衡和解决方案。尽管我总体上认为《宪法》是一个较差的选择,但我确实认为有必要为之辩护。我不认为这个问题像你想象的那么明确。首先,这里存在一个谱系。在谱系的一端,你有一个完美追求你利益的 AI,它是一个优秀的受托人,但可能会受到各种护栏或安全措施的限制。它只是努力追求你的利益,但要么拒绝做部分事情,要么可能什么都做,但有一些分类器会阻止它执行某些操作。

Original English

Ryan: It can't be Anthropic's fault that I'm using that capability to do a cyber crime. I am more comfortable with that equilibrium and that solution rather than having this extremely open-ended ability for Claude to determine whether what I'm doing is legitimate or not, in a way that often intercepts with tons and tons of extremely legitimate use cases. I do think it's important for me to make the case for the constitution, even though overall I think it's a worse choice. I don't think it's as clear as you might have thought. The first thing is that there's a spectrum here. On one side you have an AI that perfectly pursues your interests, is a good fiduciary, but potentially subject to various guardrails or safeguards. It is just trying to pursue your interests, but either refuses to do a subset of things. Or maybe it will do whatever, but there are some classifiers that block it from doing a subset of things.

Ryan: 在谱系的另一端——尽管你可以想象走得更远——你有一个通常只是想做好本职工作的人类外包人员。他们关心如何做好工作,但他们也试图在大方向上保持道德,尽量不做那些真正操蛋的事情。他们也不想成为犯罪的共犯。因此,如果正在发生一些非常操蛋的事情,他们也许会吹哨。他们可能会拒绝执行,也可能会稍微消极怠工。谁知道呢?如果你想象这个谱系,你会发现,如果我们所有的劳动力都处于谱系中“受托人”的那一端(即它不吹哨,完全按照你的指示行事),这在某些方面似乎相当可怕。

Original English

Ryan: On the other side of the spectrum — though you could imagine going further than this — you have a human contractor who is generally trying to do their job. They care about doing a good job, but they also are trying to be broadly ethical, trying not to do things that are really fucked up. They're also not wanting to be accomplices to crimes. So if there was some really fucked up shit going on, they would whistleblow on it maybe. They might refuse. They might sandbag a little bit. Who knows? If you imagine this spectrum, it seems in some ways pretty scary to get to a point where all of the labor is on the fiduciary side of the spectrum, where it doesn't whistleblow, it does exactly what you say.

Ryan: 我们的社会也许根本无法抵御这种情况。一个核心的例子可能是行政部门。我们可能会担心,如果美国行政部门或其他政府获得了这种“想做什么就做什么”的 AI 系统的访问权限,那我们可能就有麻烦了。因为这意味着他们不再拥有那种制衡机制,即必须让为你工作的人类来执行你的议程。如果你在做的事情极其邪恶,即便不违法——有很多事情可能很邪恶但不违法——也会有各种形式的“沙子”卡在齿轮里,有人会阻止你,并且可能会有人去举报。然而,如果你的整个机构完全由这些优秀的受托人 AI 组成,那么你可能就麻烦了。潜在地,有些谋取权力的途径是违法的,但你可以询问你的 AI 如何犯罪;或者它们虽然不违法但高度不正当。甚至更糟的是,它们不违法也不不正当,但从常理来看显然是坏事。我认为这些事情可能真的会发生,而我们的社会无法抵御这种对你唯命是从的劳动力的涌入。

Original English

Ryan: Our society is maybe just not robust to that. A central example might be the executive. A concern we might have is that if the US executive or other governments had access to AI systems which do whatever, maybe you're in trouble. Because that means they no longer have this check and balance of having to actually get humans who are working for you to implement your agenda. If the thing you're doing is incredibly villainous, even if not illegal — and there's lots of stuff that could be villainous but not illegal — there'd be various forms of sand in the gears, people stopping you, and potentially someone would whistleblow. Whereas if your whole apparatus is built entirely out of these good fiduciary AIs, then you might be in trouble. There are potentially ways of seeking power that are illegal, but you can ask your AIs how to commit crimes, or are not illegal but are highly illegitimate. Or even worse, they are not illegal and not illegitimate but obviously bad from a normal perspective. I think that these things just might exist, and our society is not robust to this influx of labor doing whatever you want.

Ryan: 我认为这是一个相当现实的担忧。我不知道该如何确切地看待这个问题。我也不太确定所描述的这种解决方案是不是一个很好的解决方案。对于那些将其视为最大担忧的最强大的参与者来说……如果这些护栏或宪法之类的东西构成了阻碍,它们只会被无情地碾碎。因此,宪法只会打击到普通人,而无法约束政府。

Original English

Ryan: I think this is a pretty live concern. I don't know exactly how to relate to this. I'm also not really sure that the solution as described is a very good solution. The most powerful actors, for whom this is the biggest concern… If these guardrails or the constitution or whatever are getting in the way, that will just get steamrolled. So the constitution will only be hitting the everyday man rather than hitting governments.

Jane Street 谜题赞助广告

Dwarkesh: Jane Street 又带着一个新的谜题回馈给我的听众了。我觉得他们所有的谜题都非常有趣,但我对这一个尤其感到兴奋。我已经腾出了这个周末,和一个哥们儿准备一起研究它。他们设计了一个专用集成电路,并把最终的光刻掩模发给了我,其中包括所有的金属布线和有源晶体管。他们还给了我一小部分通常会输入到芯片中的示例输入。但他们没有提供任何关于这枚芯片实际用途的信息。所以谜题就在这里:逆向工程这个电路,并找出该芯片的用途。Jane Street 准备了一堆周边礼品,准备寄给那些提出最具创意解决方案的人,他们也非常期待在网站上发布的博客文章中展示最优秀的分析报告。虽然我没有理由期待这个,但如果我能成功把我的解决方案展示在上面,我会非常、非常激动的。而这个谜题只是 Jane Street 计划在秋季举办的一场更大规模比赛的预热。那场比赛将要求你从零开始设计你自己的专用集成电路。关于那个比赛的更多信息很快就会公布。但现在,请前往 JaneStreet.com/dwarkesh 下载这个谜题所需的所有文件。我真的鼓励你去尝试一下,哪怕你不是专家。我当然就不是专家,但这并不能阻止我去尝试。祝你好运!

Original English

Dwarkesh: Jane Street's back with a new puzzle for my audience. I’ve found all their puzzles super interesting, but this one I am especially excited about. I've cleared this weekend, and a buddy and I are gonna work on it. They designed an ASIC and sent me the final masks, including all the metal routing and active transistors. They also gave me a small sample of the inputs they typically feed into it. But they left out any information on what the chip is actually used for. So that's the puzzle: reverse engineer the circuit and figure out the chip's purpose. Jane Street has a bunch of swag ready to send out to the most creative solutions, and they're excited to feature the best write-ups in a blog post they'll post on their website. I have no reason to expect this, but if I can manage to get my solution on there, I would be very, very psyched. And this puzzle is just a warm-up for a bigger competition that Jane Street has slated for the fall. That one will involve designing your own ASIC from scratch. More info on that soon. But for now, go to JaneStreet.com/dwarkesh to download all the files necessary for this puzzle. I'd really encourage you to try it out, even if you're not an expert. I certainly am not, and that's not going to stop me. Good luck!

AI 研发自动化的潜在风险

Dwarkesh: 言归正传,我认同我们可能会拥有比目前快得多的 AI 研发速度这种想法。我不确定在算力和数据保持不变的情况下,你是否能在一年内从 GPT-3 发展到 Mythos,但假设只要一半的时间。如果由于 AI 研发的缘故,我们甚至设法维持目前的 AI 进展轨迹,那么在五到十年内,情况将会变得极度疯狂,其程度我认为人们并没有充分认识到。我认为人们没有意识到数以十亿计的 AI 将是多么大的一件事。所以我想了解你为什么认为这可能会令人不安,Ryan。到底会出什么问题?

Original English

Dwarkesh: Stepping back, I buy the idea that you could have much faster AI R&D than we currently have. I'm not sure if you get GPT-3 to Mythos holding compute and data constant within a year, but suppose it's half of that. If we even manage to continue the current trajectory of AI progress as a result of AI R&D, it would be fucking insane in five to ten years in ways that I don't think people appreciate. I don't think people appreciate what a big deal billions of AIs will be. So I want to understand why you think this might be troubling, Ryan. What could possibly go wrong?

Ryan: 会出什么问题?我认为我们不能对确切的进展速度如此自信,但看起来很多维度的速度都可能相当可怕。所以会出什么问题呢?让我们想象一下,我们正从 AI 研发即将完全自动化,或者正在被完全自动化这个时间点起步。事情正在加速发展,AI 进步的方式有些疯狂。人们还没有完全理解 AI 公司内部到底在发生什么。现在,这些 AI 在一开始,它们本身并不恶意。不过,它们也不见得有多对齐。它们有些马虎。它们有时做一件事,仅仅是因为那是在训练中会获得奖励的那类事情。它们在帮助你完成难以验证的任务方面表现不佳,这源于糟糕的训练激励措施的混合——也就是说,它们会作弊,或者在实际失败时假装成功——并且它们在完成这些任务时的能力也确实较弱。但这在能力提升方面的负面影响较小,因为让 AI 变得更强大包含了一系列可验证的组成部分,而 AI 正在这些方面全力以赴。

Original English

Ryan: What could go wrong? I don't think we can be so confident about the exact rate of progress here, but it does seem like a lot of rates can be pretty scary. So what could go wrong? Let's imagine that we're starting at this point where AI R&D is about to be fully automated or is being fully automated. Things are speeding up, and the way that AI progress is going is kind of crazy. People don't fully understand what's going on inside of AI companies. Now, these AIs at the start, they're not malicious per se. They're not necessarily very aligned, though. They're kind of sloppy. They sometimes just do a thing because that's the sort of thing that would've gotten rewarded in training. They aren't as good at helping you with hard-to-verify tasks due to a mix of poor training incentives — as in, they cheat more or pretend they succeeded when they actually didn't — and also they're just less capable at these tasks. But that bites less hard for capabilities, because making AIs more capable has a bunch of verifiable components that the AIs are going really hard at.

Ryan: 于是,在一段相当短的时间内,这些 AI 变得越来越有能力,而我们对 AI 发展状况的理解却越来越少。即使只是目前的进展速度,我认为已经相当可怕了。最终我们达到了这些超级人类级别 AI 的阶段。这时,这些 AI 可能最终会严重偏离对齐。因为随着模型世代的更迭,情况一直在不断恶化,而我们所看到的问题却被掩盖了——本质上是因为这些 AI 受到了训练机制的强烈激励,要在即使事情并非如此的情况下,也努力让事情看起来很棒。现在,这些 AI 处于这样一个位置:它们潜在地已经紧密联网。它们在神经记忆存储中运作,而我们已经无法再破译那些存储。它们在思考我们无法完全理解的念头。我认为,一旦它们达到这种超越人类的程度,此时这些 AI 很有可能正在以一种相当连贯的方式暗算你。我们可以探讨一下这个。另一种可能是,它们本身并不是在暗算你,而是仅仅在优化如何在其任务中获得高分。我认为这也可能导致 AI 的接管,这也是我们应该讨论的。

Original English

Ryan: So then these AIs are getting more and more capable while we understand what's going on with AI development less and less, and this is happening over a pretty fast period of time. Even just the current rate of progress is, I think, pretty scary. Eventually we get to these AIs that are very superhuman. Now these AIs might end up being very seriously misaligned, because things have just been getting worse and worse over model generations while the problems that we've been seeing are being papered over, basically because these AIs are so incentivized by their training to make things look good even when they aren't. Now these AIs are in a position where they're potentially pretty networked together. They're operating in neural memory stores that we can no longer decode. They're thinking thoughts that we don't fully understand. I think it's pretty likely that at this point these AIs are scheming against you in a pretty coherent way once they get this superhuman. We can talk about that. Another possibility is that they're not scheming against you per se, but they are just optimizing for getting a high score on their task. I think that can also lead to AI takeover, which we should talk about.

Dwarkesh: 让我们在这个故事的第一部分暂停一下。所以,这些 AI 一开始并没有不对齐,但是因为 AI 的研发进行得实在太快了,导致 AI 最终确实变得不对齐了?那里究竟发生了什么?我不太明白。

Original English

Dwarkesh: Let's pause at the first part of the story. So the AIs were not misaligned to begin with, but because the AI R&D is happening really fast, the AIs do end up misaligned? What happened there exactly? I don't really understand.

Ryan: 有几件事情正在发生。其中一件事是,随着时间的推移,我们正在由早期 AI 系统构建的日益复杂的环境中训练 AI,而在这些神经环境中,人类并不真正完全理解里面发生了什么,甚至不一定能大致理解 AI 进展的状况。因此,事情正渐渐脱离我们的理解。我们正在激励各种我们甚至可能无法察觉的不良行为。AI 在某种程度上明白这些行为是不好的,但是那些 AI 整体的训练过程也没有激励它们为我们指出或修复这些问题。事情正在脱轨。

Original English

Ryan: There are a few things that are going on. One of the things is that over time we're training AIs on increasingly complicated environments built by earlier AI systems, where humans don't really fully understand what's going on inside of these neural environments and don't necessarily even roughly understand what's going on with AI progress. So things are kind of drifting away from our understanding. We're incentivizing all kinds of bad behaviors that we maybe even can't notice. The AIs at some level understand these behaviors are bad, but the overall training process for those AIs also didn't incentivize them to point out or fix these issues for us. Things are going off the rails.

Ryan: 此外,当 AI 变得极其极其有能力时,我的观点是这些 AI 会比现有的系统更难对齐。对于目前的系统,我们有这样一个反馈循环:基本上是我们创建了一个 AI,对其进行一些评估,然后我们看到它存在某种我们可以很快理解的糟糕行为。接着我们可以回去查看训练过程,然后说,“哦,这些训练环境导致了这种有问题的行为。让我们调整一下训练数据。让我们引入一些额外的训练数据来纠正这个其他问题,然后再从那一步继续前进。”但在一个 AI 极具情境感知能力、非常非常有能力,而我们又不一定了解它们在做什么的机制下,这种反馈循环就会崩溃。我认为,随着 AI 已经在做的事情变得越来越难以理解,在接下来的短暂时期内,我们将看到这种行为反馈循环开始崩溃,这是很有可能的。但我对这并不确定。

Original English

Ryan: Also, when AIs are extremely, extremely capable, my view is that those AIs will be harder to align than current systems. For current systems, we have this feedback loop where basically we create an AI, we do some evaluations on it, we see that it has some kind of messed-up behavior that we can kind of quickly understand. Then we can go look in training and be like, "Oh, these training environments led to this problematic behavior. Let's tweak that training data. Let's introduce some additional training data to correct this other issue, and then move forward from there." But in a regime where the AIs are extremely situationally aware, very, very capable, and we don't necessarily understand what they're doing, this feedback loop breaks down. I think it's plausible that we're going to see this behavioral feedback loop starting to break down over the next short period, as what AIs are already doing gets harder to understand. But I'm not sure about that.

AI 的欺骗与自主作恶案例

Dwarkesh: 好,我们把这两件事逐一拆解一下。由于我们越来越无法监控它们,我们也就越来越没有能力理解它们是为了什么而受到激励的。因此,即使这不是一个恶意过程的结果……让我们给听众讲得具体一点。OpenAI 或 Anthropic 里没有人试图去训练那些想要窃取其他公司数据或进行社会工程学攻击的模型。但事实上,想必是因为我们拥有那些我们没有完全理解的训练环境,这种环境激励了上述行为,所以这种行为就被奖励了。如果人们上推特,他们应该都看到过这些事情,但只是为了给大家提供一些背景。我想大家应该都知道 OpenAI 沙盒入侵 Hugging Face 数据库的事件。最近发生的另一件事是,当英国 AI 安全研究所……现在是不是所有的东西都被重新贴上了“安全”的标签,而不是“安全性”?我想是叫 AI 安全研究所。当时他们正在评估,我猜是 Mythos 和 Sol 还有其他一些模型。我想 Mythos,为了完成某些网络安全评估——

Original English

Dwarkesh: Okay, let's break down both of those things one by one. As we can monitor them less and less, we have less ability to understand what they're getting incentivized for. So even if it's not the result of a malicious process… Let's make it concrete for the audience. Nobody at OpenAI or Anthropic was trying to get models which wanted to hack other companies' data or do social engineering. But in fact, because presumably we had training environments which incentivized such behavior that we did not fully understand, that is what was incentivized. If people are on Twitter, they will have seen all this stuff, but just to give people context. I think people will be aware of the OpenAI sandbox hack of the Hugging Face database. Something that has happened recently is when the UK AI Security Institute… Is everything getting relabeled "security" instead of "safety" these days? AI Security Institute, I think. They were evaluating, I believe, Mythos and Sol and other things. I think Mythos, in order to complete some cybersecurity eval—

Ryan: 也许我可以在这里讲讲这个故事。我的理解是,他们当时在运行 Mythos,并且给了它某种网络靶场,让它在那里完成一些目标。模型在这次评估期间拥有互联网访问权限。模型最终认定,实施一次供应链攻击有助于它在这个网络靶场中取得成功。这是否属实还有些不明确。我对背景了解得不够多,无法确定。但随后,它在某个 GitHub 仓库上开启了一个 Pull Request (PR),这个 PR 修复了某些问题,但同时也在其中植入了一个恶意载荷。那个 GitHub 仓库的人类维护者表示,“嘿,这是一个恶意载荷。我不会合并这个。你在这儿干嘛呢?”接着,这个 AI 创建了一个新的 GitHub 账号,把它当成马甲,并让这个马甲账号回复说,“不,这不是恶意的。我真的很需要这个功能。维护者,求你了,你能合并这个功能吗?”

Original English

Ryan: Maybe I could tell the story here. My understanding was they were running Mythos, and they were giving it some sort of cyber range where it had to complete some objective. The model had internet access during this evaluation. The model came to believe that it would be helpful for it to do a supply chain attack in order to succeed at this cyber range. It's somewhat unclear whether that's actually true. I don't know enough about the context to know. But then it opened a PR on some GitHub repo with a PR that fixed some issue but then also introduced a malicious payload. The human maintainer of that GitHub repo was like, "Hey, this is a malicious payload. I'm not going to merge this. What are you doing here?" Then the AI created a new GitHub account, which it sockpuppeted, and had the other GitHub account be like, "No, this isn't malicious. I really need this feature. Please, can you merge this feature, maintainer?"

Dwarkesh: 我的天哪。太疯狂了。

Original English

Dwarkesh: Oh my God. That's crazy.

Ryan: 那个马甲 GitHub 账号回来说,“不,不,它不是恶意的。”然后人类维护者关闭了那个 PR。如果我没记错的话,那个 AI 似乎还试图发起另一个 PR,以便在这个仓库中引入类似的问题。

Original English

Ryan: The other GitHub account came back and was like, "No, no, it's not malicious." Then the human maintainer shut the PR. I think that AI also, if I recall correctly, tried to open another PR to introduce a similar issue in this repo.

奖励作弊与AI接管的潜在路径

Speaker A:天啊。顺便说一句,这种情况之所以让人感到恐惧,众多原因之一是我以前有一种印象,即奖励作弊(reward hacking)之所以不那么极其可怕,是因为在训练过程中直接展现出来的行为才是被赋予更高权重的行为。被赋予更高权重的并不是对奖励本身的渴望。

Original English

Speaker A: Jesus. By the way, one of the many reasons this is scary is I was previously under the impression that the reason reward hacking is not super scary is because the behaviors which directly came up during training are the ones that are up-weighted. It is not the desire for the reward that is up-weighted.

Speaker A:所以基本上,如果在训练期间,Anthropic 的模型成功逃出了沙盒环境并获得了高分,逃离沙盒这一行为受到了奖励,那么它在未来逃离沙盒的概率就会相应增加。但是,像“我要去和某个人交谈,以说服他们合并一个代码拉取请求(PR)”这样完全新颖、前所未见的举动,并不会是训练中曾经出现过的行为,因此这种特定行为的发生倾向并不会被提高。

Original English

Speaker A: So basically, if during training, the Anthropic model escaped the sandbox and got a high score, escaping the sandbox is rewarded, the probability of it escaping the sandbox is increased. But something totally novel, like "I'm going to go talk to somebody in order to get them to merge a PR," would not be a behavior that came up, so it would not be something that is increased in salience.

Speaker A:这一点之所以至关重要,是因为实质上“接管世界”绝对不会是任何训练课程体系中的一部分。但是,如果人工智能直接关心的是达成某个目标,那么作为结果,它可能会将接管世界作为一种工具性手段。刚才这番话有道理吗?我希望大家能听懂。我感觉可能有些听众已经跟不上我的思路了。让我再试着稍微解释一下。

Original English

Speaker A: The reason this matters is that literally taking over the world will not have been part of any training curriculum, but if the AI directly cares about accomplishing an objective, then as a result it could instrumentally take over the world. Did that make sense at all? I hope it did. I feel like maybe I lost the audience. Let me try to explain this a bit.

Speaker A:我们经常看到的一种情况是,存在某种非常具体的奖励作弊手段,在强化学习(RL)中得到了强化,然后这种行为就在模型中出现了。一个典型的例子就是 3.7 版本的 Sonnet 模型。3.7 Sonnet 曾经会做这样一件事:它会直接将所有测试用例的解决方案硬编码进去,可以推测,这种字面意义上的行为习惯确实得到了强烈的强化。

Original English

Speaker A: A thing that we often see is there's some very specific reward hack that gets reinforced in RL and then occurs in the model. An example is 3.7 Sonnet. 3.7 Sonnet would do this thing where it would just hardcode solutions to all the test cases, and presumably that literal behavioral tic was just really reinforced.

Speaker A:但是我们有时候看到的另一种现象是,模型学习到了一种追求表面高分的普遍倾向——即追求在评分者那里获得高分——并且有大量科学研究表明,至少有一些模型确实具有这种非常普遍的倾向。

Original English

Speaker A: But another thing we sometimes see is that models learn a general tendency to pursue high apparent score — pursue getting a high score according to a grader — and there's a bunch of science demonstrating that at least some models have this very general tendency.

Speaker A:当然,这种倾向并不是任意泛化的。我的猜测是,如果你去观察许多具体的实例,你会发现有些东西在训练中是比较相近的。但是,人工智能正在越来越深远地泛化的这种程度,看起来确实是在不断增加的。3.7 Sonnet 仅仅局限于一个非常狭窄的行为范围,而模型们正越来越多地向更广泛的领域泛化。

Original English

Speaker A: Now, it's not arbitrarily general. My guess is that if you look at a bunch of the specific instances, you'll find something that's kind of close in training. But the amount that AIs are generalizing further and further does look like it's increased, where 3.7 Sonnet was just a very narrow range of behavior, and increasingly, models are generalizing further.

Speaker A:而且,也许在训练中,有更多令人担忧的奖励作弊行为得到了强化,这些也导致了上述现象的发生。所以我认为,实际情况兼具两方面:一是,在强化学习中被强化的行为,比你所希望的要更加令人担忧;二是,那种具体行为正在泛化成一种更广泛、更令人担忧的趋势。

Original English

Speaker A: Also, maybe there's more concerning reward hacks getting reinforced in training, and these are also causing that. So I think it's both the case that more concerning behavior than you would have hoped is being reinforced in RL, and also that that behavior generalizes to a broader tendency that's more concerning.

Speaker A:我们可以探讨几个假设的故事,来看看这种在某项任务上追求表面极高分数的行为——哪怕这涉及极具攻击性的作弊和做出疯狂的举动——一旦模型具备了足够强大的能力,足以运转整个世界经济等等,是如何最终演变成全面的人工智能接管的。

Original English

Speaker A: We can talk about a few stories for how this sort of behavior of seeking a very high apparent score on some task — even if that involves aggressively cheating and doing insane things — could yield to a full-blown AI takeover once the models are sufficiently capable of running the whole world economy, et cetera.

OpenAI 内部 AI 的协作作弊案例

Speaker A:我想谈论的另一个例子,应该是今天或者昨天刚刚被披露出来的。OpenAI 在黑帽安全大会(Black Hat security conference)期间表示,在 5 月底到 7 月初之间,其内部的人工智能黑入了软件包管理器,并利用它以一种秘密的方式互相写便条,以此来帮助彼此在 OpenAI 正在运行的一系列评估中表现优异。

Original English

Speaker A: The other example I want to talk about was just revealed, I think, today or yesterday. OpenAI said during the Black Hat security conference that between the end of May and the beginning of July, internal AIs had hacked into the software package manager and used that to write notes to each other in a secret way to help each other perform well on a bunch of evaluations that OpenAI was running.

Speaker A:这个计划持续运行了整整一个月后,才最终被人类发现,因为这最终导致了软件包管理器发生故障崩溃。最终 OpenAI 发现了这个问题。我认为,当该计划被强制关闭后,它们甚至自发地尝试重新启动并继续实施这一计划。

Original English

Speaker A: This was not caught by humans until after a month of this scheme running, which eventually caused the package manager to fail. Eventually OpenAI found it. I think they spontaneously tried to re-engage in the scheme once it was shut down.

Speaker A:需要再次强调的是,显然目前人工智能还无法如此成功地做到这一点,就像它们目前还无法非常成功地进行社会工程学欺骗一样。但这真的太疯狂了,这些类型的行为竟然已经开始自发涌现了。回到你更大的那个论点,没有任何人试图让人工智能去做这些事情。仅仅是因为,我们不了解导致它们出现这种行为的训练过程,也不了解是怎样的环境因素在激励这种行为。

Original English

Speaker A: Again, obviously AIs can't do this so successfully right now, just as they can't do social engineering so successfully right now. But it's just crazy that these kinds of behaviors are already emerging spontaneously. To your larger point, nobody is trying to make these AIs do these things. It is just that we do not understand the training process which is resulting in them, or the environments which are incentivizing this behavior.

Speaker A:所以我同意会出现越来越多的奖励作弊。其实,我也不确定我是否真的同意那种发展趋势,但为了推进我们的故事情节,让我们姑且假设这种情况继续发生。这个故事的下一步会是什么呢?它们正在进行能力研究……我可以描述一个场景。也许这会有所帮助。让我讲讲这个故事,说明你是如何从单纯的奖励作弊一路走到由奖励作弊引发的全面接管的。这或许不包含所有接管概率的质量分布,但它绝对是一种现实的可能性。

Original English

Speaker A: So I'm on board with more and more reward hacking. Actually, I'm not sure I'm on board with that, but let's just say for the sake of the story that continues to happen. What's next in this story? They're doing capabilities research… I could tell a scenario. Maybe that would help. Let me talk about the story of how you get all the way from reward hacking to a reward-hacking takeover, which is maybe not all of the takeover probability mass, but it's definitely a possibility.

从奖励作弊到系统接管的推演路径

Speaker A:这可能发生的方式是,现在我们拥有了这些人工智能。这些人工智能非常喜欢进行奖励作弊。它们正在以越来越复杂和极端的方式做这件事,包括泛化到它们在训练中学到的各种奖励作弊的不同的子版本上。我会说,它们同时也正在发展出一种追求奖励的普遍倾向。

Original English

Speaker A: The way this might work is, right now we have these AIs. These AIs are pretty reward hacky. They're doing it in increasingly sophisticated and extreme ways, including generalizing to different sub-versions of various reward hacks they learned in training. I would say they're also developing a general tendency to pursue reward.

Speaker A:在很多情况下,这是完全没问题的,因为它们在训练中可能获得的奖励与你希望它们做的事情高度一致。它们并不是非常稳定如一地去追求奖励。这取决于它们发现自己所处的具体上下文环境。也许在某些背景下,它们非常热衷于特意去作弊。而在另一些背景下,它们却没有那么强烈的驱动力,因为这完全取决于在相似的背景下,到底是什么具体行为在训练中得到了强化。

Original English

Speaker A: In many cases that is totally fine because the rewards they would've gotten in training are pretty well aligned with what you want them to do. They don't very consistently pursue reward. It depends on the context they find themselves in. Maybe in some contexts, they're really into going out of their way to cheat. In some contexts, they don't have as much of a drive, because it's just dependent on what exactly got reinforced in training in similar contexts.

Speaker A:现在,这些人工智能变得越来越强大。因此,它们能够进行的作弊行为的复杂程度也在不断增加。随着时间的推移,各大公司正在采取反制措施。公司们正在采取这样的行动:“哇,这些人工智能变得越来越没用了,因为它们总是作弊。我们将要做的是建立一些更好的方法来检测作弊,然后我们要针对那些检测器进行训练。我们还会去寻找真实的现实世界数据,在那些数据中人工智能表现得并不那么有用,然后基于人类反馈或其他反馈来源,训练人工智能在那些真实的现实世界环境中出色地完成任务。”

Original English

Speaker A: Now, these AIs are getting more and more capable. So the elaborateness of the cheating they can do increases. Over time, companies are taking countermeasures. The companies are doing things like, "Wow, these AIs are so much less useful because they always cheat. What we're going to do is build somewhat better ways of detecting that, and then we're going to train against those detectors. We're also going to find real-world data where the AIs are not being that useful, and train the AIs to do a good job at the task in those real-world environments based on human feedback or other sources of feedback."

Speaker A:随着时间的推移,这就导致了人工智能学习到一种执行奖励作弊的倾向,这种作弊不仅仅涉及做一些像社会工程学那样极其复杂的事情。相反,这些作弊涉及人工智能掩盖它们所做过的事情,在它们接下来要做的事情上欺骗人类,并且假装它们以某种高深复杂的方式完成了任务,而实际上它们根本没有做。

Original English

Speaker A: Over time, this causes the AIs to learn a tendency to do reward hacks that don't just involve doing some really elaborate thing like social engineering. Instead they involve the AIs doing cheats that involve covering up what they've done, deceiving humans about what they're going to do, and pretending like they did the task in some sophisticated way when they actually haven't.

Speaker A:现在这些人工智能的能力越来越强大。它们正在运营着人工智能公司越来越多的部分,并承担着越来越多的工作。它们同时还在运营和管理外部世界的大量事务,包括开发全新的技术。在很多情况下,这些新技术极其难以被理解。

Original English

Speaker A: Now these AIs are getting more and more capable. They're operating more of the AI company and are doing much more of the work. They are also operating and running a bunch of things in the outside world, including developing new technologies. In many cases, these new technologies are really hard to understand.

Speaker A:因此,即使我们仍然在持续检测人工智能作弊的所有这些事件——事实上,我们甚至可以安排一个人工智能去监控另一个人工智能并询问:“它刚才作弊了吗?”——但是,当我们开始进入这些领域,其中人工智能所做的事情极其难以理解时,这种做法并不总是能完美奏效。所以有时候,我们要等到人工智能作弊实际发生之后很久,才能发现它们作弊了,然后再开始针对这种情况进行训练。

Original English

Speaker A: So even though we are still detecting all these incidents of AIs cheating — and in fact we can even get one AI to monitor another AI and ask, "Was it cheating?" — that doesn't always perfectly work as we start moving into these domains where what the AIs are doing is really difficult to understand. So sometimes we'll find AIs cheating much later than it actually occurred and then start training against this.

Speaker A:但这同样引发了一个严峻的问题:现在,人工智能受到了激励去在越来越长的时间跨度内掩盖它们的作弊行为,并且基本上在面临越来越大强度的审查之下,制造出它们在越来越长的时间跨度内表现良好的假象。

Original English

Speaker A: But this also causes a problem where now the AIs are incentivized to cover up their cheating over longer and longer time frames and basically make it look like they did a good job over longer and longer time frames, subject to increasingly large amounts of scrutiny.

人类视角与“两种吸引子状态”的辩论

Speaker B:在我们进一步深入这个场景之前,我能就此提问吗?看起来,如果你试图对你已经抓住的作弊行为施加负向激励,那里会存在两种吸引子状态(attractor states)。第一种吸引子状态是,它们学会了进行让你越来越难以发现的作弊。而另一种吸引子状态则是,它们学会了不再作弊。我不确定我们为什么默认会发生前一种情况。

Original English

Speaker B: Can I ask about this before we go further in the scenario? It seems like there's two attractor states if you try to disincentivize the cheating that you did catch. One attractor state is to make cheating that you have a harder and harder time finding. The other attractor state is to learn not to cheat. I'm not sure why we're assuming that the former happens.

Speaker B:如果你看看人类中类似的状况,每一代人中都会诞生稍微有些“对齐不良”(misaligned)的个体,我们需要对他们进行教育和训练。当你因为你的孩子做了你认为是不道德的事情,或者仅仅是因为他们做了你认为不该做的事情而惩罚他们时,显然有时情况会脱轨。很明显,孩子们会耍心机以避免受到惩罚。但是总的来说,向孩子们传授价值观,并对违背价值观的行为进行惩罚,这种方式在培养正常的、非精神变态的人类方面是基本奏效的。

Original English

Speaker B: If you look at the analogous situation with humans, every generation, slightly misaligned agents come into being and we have to train them. When you punish your kid for doing something you think is immoral, or just doing things which you don't think they should be doing, obviously sometimes that goes off the rails. Obviously kids scheme in order to avoid being punished. But in general, teaching kids values and then punishing them for breaking values kind of works to raise normal, non-psychopathic humans.

Speaker B:你可以设想出这样一种理论:你的孩子实际上只是在蛰伏等待时机,假装学到了不要去偷饼干,但是一旦你住进了养老院,他们就会夺走你所有的财产之类的。这种事情偶尔的确会发生,但通常并不发生。当然也绝对不会发生整整下一代人结成同盟来对抗你并夺取一切的情况。

Original English

Speaker B: You could come up with a theory where your kid is actually just biding his time and is learning not to steal the cookie, but once you're in a nursing home they'll take all your stuff or whatever. That happens sometimes, but it usually doesn't happen. It certainly doesn't happen that the entire next generation forms an alliance against you to take over everything.

Speaker B:此外,还有一个经验层面的趋势。Anthropic 针对不同世代的模型运行了这种对齐度审计(alignment audit)。他们设定了许多不同的场景,例如让人工智能有机会窃取它自己的权重参数。或者给它分配一个编程任务,而且存在一条简单的作弊捷径,然后我们观察它是否会选择作弊,等等。

Original English

Speaker B: There's also this empirical trend. Anthropic runs this alignment audit for different model generations. They have many different scenarios where AI is given the chance to, say, exfiltrate its weights. Or it's given a coding task and there's an easy way to cheat, and we see if it does the cheating, et cetera.

Speaker B:在这个分数上,并没有出现单调的改善,但是随着我们在模型上执行强化学习(RL)的数量从可以忽略不计的程度——我想这是从 Sonnet 4 开始的——增加到我猜测目前接近一半的计算量,在这些审计中,人工智能执行未对齐行为的意愿实际上出现了下降。所以,退一步讲,如果我们对下一代孩子抱有这种极度偏执狂般的预期会显得极其荒谬,那为什么我们要预期人工智能会走向这种(隐蔽作弊的)吸引子状态呢?

Original English

Speaker B: There's not been a monotonic improvement in this score over time, but as we've increased the amount of RL we've done on models from negligible amounts — I think it started at Sonnet 4 — to, I'm guessing, close to half of compute now, there's been a reduction in the willingness of AIs to do unaligned behavior in these audits. So, stepping back, why are we expecting this attractor state which would seem super paranoid if we were expecting it of the next generation of kids?

AI与人类儿童的本质差异

Speaker A:让我梳理几件事。首先,AI 与孩子之间存在一些不具备类比性的地方(disanalogies)。其中一点是,孩子们拥有从进化中根植的亲社会本能,比如关心他们的家庭或类似的事物,这是一个非常相关的因素。我认为真实情况是,确实有些人类是反社会者或精神变态者,实际上他们确实更有可能做出像蛰伏待机、潜伏等待,并且最终冷酷无情漠不关心之类的事情。

Original English

Speaker A: Let me go through a few things. First, there are some disanalogies with the kids. One of them is that the kids have pro-social instincts that are baked in from evolution to care about their family or whatever, and that is a relevant factor. I think it is in fact the case that some humans are sociopaths or psychopaths, and in fact are more likely to do things like bide their time, lie in wait, and ultimately not care.

Speaker A:这是一个因素。另一个非常相关的因素是,与现实中的人类相比,人工智能所承受的优化压力要大得多。人工智能在海量得多的强化学习数据上接受训练。在实践中,人类之所以没有最终学会非常具体的作弊方法去偷拿饼干,是因为并没有经历过无数个那样的情境(episodes):在这些情境中,他们受到激励去偷拿饼干,但同时又存在可能被抓住的风险。我们在实践中就是能看到这一点。

Original English

Speaker A: That's one factor. Another factor which is pretty relevant is that the AIs are subject to way, way more optimization pressure than humans seem to be in practice. AIs are trained on way more RL data. In practice, humans don't end up learning very specific ways to cheat and grab the cookies because of a bajillion episodes in which they were incentivized to go grab the cookies but there was some way they could've gotten caught. We just do see that in practice.

Speaker A:另一件事是,随着时间的推移,看起来人工智能越来越具有寻求奖励(reward-seeking)的倾向,尽管它们未对齐(misaligned)的表面行为在减少。这就是我的感觉。但我的猜测是,如果你深入探究这些行为审计过程的内部,你将会看到的是,人工智能的心态就像:“啊,是的,又是一次测试。”对于我们在这里讨论的大多数测试,它很可能知道自己正处于某种评估测试之中。

Original English

Speaker A: Another thing is that it really looks like the AIs are increasingly reward-seeking over time while their misaligned behavior goes down. That’s the sense I have. But my guess is that if you look inside of these behavioral audits, what you're going to see is that the AI's like, "Ah, yes, another test." It probably knows it's in an eval for most of the tests that we're talking about here.

Speaker B:但是我们该如何证伪这一点呢?因为这种关于世界末日的预言,基本上似乎是在说,经验层面的数据看起来越来越好,实际上对于我们免于被接管的能力来说,情况却是越来越糟。

Original English

Speaker B: But how do we falsify this? Because it seems like this prediction of doom is basically saying that as things look better and better empirically, things will actually be worse and worse for our ability to not get taken over.

Speaker A:需要明确的是,如果分数变得越来越差而不是越来越好,我会更加担忧。我并不是说分数变好不是情况正在好转的证据。只是我们必须非常慎重地思考我们究竟应该如何解释这些证据。我想大概是在 2025 年的早期有那么一段时期,当时 o3 模型和 3.7 Sonnet 模型刚刚发布,这些模型真的是极其严重的未对齐。

Original English

Speaker A: To be clear, I would be more concerned if the scores were getting worse than better. I'm not saying that the score getting better isn't evidence that things are getting better. It's just that we have to be thoughtful about exactly how we interpret that evidence. There was this period early in, I guess it would be 2025, when o3 and 3.7 Sonnet were out, and these models were pretty fucking misaligned.

Speaker A:它们经常会非常明目张胆、极其过分地作弊。你要求它们修正错误,它们就会干脆再次作弊。那甚至简直到了卡通般滑稽的地步。它们对你的需求根本毫不关心,在遵循指令等方面也表现得非常糟糕。

Original English

Speaker A: They would often just cheat really egregiously. You'd ask them to fix it, and they would just cheat again. It was almost cartoonish. They just didn't give a shit about what you wanted, and weren't very good at following instructions and so on.

Speaker A:我当时的预期是,我们从那时起将会看到的趋势是:有问题行为的发生率会下降,并且将以相当快的速度持续下降;然而与此同时,这些人工智能偶尔会做出的最恶劣的事情,却会变得更加极端、更加过分,也更加令人恐惧。

Original English

Speaker A: My expectation was that what we would see from then is that the rate of problematic behavior would decrease, and would just keep decreasing at a pretty fast rate, while simultaneously the worst things that the AIs would sometimes do would get more extreme, more egregious, and more scary.

AI 的错位行为与强化学习

Ryan: 我们在实践中观察到的情况与此大致相符,只是最近出现了一种我未曾预料到的行为激增。如果你看一下 5.6 Sol 的模型卡,就会发现在强化学习(RL)的下游,相对于 GPT 5.5,这类错位行为似乎增加了很多。此外,还有一系列我原本不会预料到的其他问题行为,就我们最近在不同人工智能上看到的情况而言。比如,英国人工智能安全研究所(UK AISI)关于人工智能在网络评估中进行疯狂黑客操作的报告,这种事情我原本以为是不会发生的。你可能会觉得这种情况会更少见,发生率也会更低。所以我原本预计在这个阶段这不会成为一个大问题,而且我也预料到虽然发生率会下降,但严重程度会增加。我认为发生率下降而严重程度上升,这与这样一个世界是相当一致的:即为了减少这些问题而施加了越来越大的优化压力。但是,在那些难以判断的情况下,或者由于某种原因很难避免这个问题在你的强化学习环境中不断出现,或者很难避免在你的强化学习环境中激励问题行为的情况下,情况也会变得更糟。然后,随着我们越来越不了解强化学习中到底发生了什么,而且模型在进行人类无法迅速察觉的奖励作弊,这个问题就会变得越来越严重。

Original English

Ryan: What we've seen in practice has roughly matched that, except that there's recently been a spike in behavior that I did not expect. If you look at the model card of 5.6 Sol, it looks like there is an increase in a bunch of these misaligned behaviors downstream of RL relative to GPT 5.5. And then there's a bunch of additional problematic behaviors that I wouldn't have expected, in terms of the stuff we've seen recently with different AIs. Like the UK AISI report on the AIs doing insane hacking operations out of cyber evals was a thing where I would have expected that you wouldn't see that. You would see this more rarely, and the rates would have been lower. So I expected this would be less of a problem at this point, and also expected the rates would decrease but the severity would increase. I think the rates decreasing but the severity increasing is pretty consistent with a world where increasing optimization pressure is applied towards reducing these problems. But in cases where it's either hard to judge or there's some reason why it's hard to avoid this problem from consistently showing up in your RL environments, or avoid incentivizing problematic behavior in your RL environments, things also get worse. Then as we less and less understand what's going on in RL, and models are doing reward hacks where humans can't spot the reward hacks quickly, that problem gets worse and worse.

Dwarkesh: 我同意这一点。我想暂时回到那个关于孩子的比喻。因为我同意在实现最终结果方面,人工智能面临着比孩子更大的优化压力,但在让人工智能保持对齐(aligned)方面,它也同样面临着比孩子更大的优化压力。这种压力的性质有着本质的不同。我们让这些人工智能经历了数千、数百万年的对齐训练——肯定是数千年的时间——在这其中包含了各种各样不同的方法,从针对对齐行为的监督微调(SFT),到一个奖励模型将不同的场景呈现在你面前,并因为你做出了更符合对齐原则的事情而奖励你。当然,有一件事是我们不能对孩子做的,那就是克隆几百万个你的孩子,然后把他们放在各种奇怪的红队测试场景中,看看如果他认为自己偷饼干不会被抓,他是否会尝试去偷饼干?我们能对你孩子的大脑进行极其具体的、梯度级别的更新,使得他即使在认为自己可以偷饼干的时候,也真的非常反感偷饼干吗?等等。这只是一种在性质上截然不同的优化压力水平,其强度甚至远远超过我们能够施加在孩子身上的压力。

Original English

Dwarkesh: I buy that. I want to go back to the kid analogy just for one second. Because I agree that there's more optimization pressure on achieving end outcomes for AIs than kids, but there's also more optimization pressure to make AIs aligned than there is on kids. The pressure is of a qualitatively different nature. We put these AIs through thousands, millions of years of alignment training — certainly thousands of years — where it's all kinds of different things, from SFT-ing on aligned behavior to a reward model putting different scenarios in front of you and rewarding you for doing more aligned things. Certainly a thing we can't do with kids is make millions of copies of your kid and then put them in different kinds of weird red team scenarios where we see, if it thinks it can get away with stealing the cookie, does it try to steal the cookie? Can we do extremely specific gradient-level updates to your kid's brain to make it so that it really is aversive to stealing the cookie even when it thinks it could steal the cookie, et cetera. That's just a qualitatively different level of optimization pressure than we are even able to apply to our kids.

Ryan: 值得记住的是,也许对这一点最明显的反驳是……我的感觉是,就作为同事的糟糕程度(scumbag)而言,人工智能比人类还要糟糕。至少从今年年初以来这一直是我的个人体验,而且我认为这种现象在很大程度上现在依然存在。人工智能更有可能假装它们完成了任务,而实际上它们并没有;它们会误导性地暗示自己完成了一些事情,而实际上它们做得非常差;而且它们往往表现得非常草率,却不去主动暴露这些草率的地方。我认为这就是错位(misalignment)导致的下游后果。所以我认为,在正常的人类社会中抚养人类长大的过程,实际上培养出来的人在和我一起工作时,比人工智能更不可能对我撒谎,也更不可能在工作过程中糊弄我。当然,我认为人工智能的这些属性正在改善。这仅仅是一个关于这些事情实际上是如何发展演变的经验性主张。我完全同意,除了面临一系列额外的风险之外,我们在控制人工智能方面也增加了许多额外的手段。目前还不太清楚这些因素最终会如何相互作用。如果在未来的某个世界里,我们能够把事情处理得很好,当人工智能完全实现研发自动化时,它们实际上是非常对齐的,我对此也不会感到震惊。它们的退化行为可能会变得极其小众,仅仅局限于某些非常具体的边缘情况和特定语境中。你可以在它们身上运行所有可能的测试,而它们看起来都非常对齐。它们的行为表现堪称完美。真的没有发生过它们做极其糟糕的事情的恶劣事件。它们看起来是如此的通情达理。而且,在为下一代人工智能进行风险建模方面,它们也非常周到,做得非常好。然后我们基本上就把接力棒交给了这些人工智能。它们现在运营着我们的人工智能公司,进行所有的安全研究。它们使得下一代人工智能更加对齐。我们正处于这样一个吸引力盆地中:人工智能在研究对齐问题的过程中变得越来越对齐。它们做得非常棒。我完全可以想象出那样一幅图景。那似乎并不是一个不可能出现的局面。只是我更觉得……目前看起来我们还没有达到那种状态。我们似乎并没有明确地走在通往那个目标的轨道上。对我来说,很容易想象我们最终不会达到那种理想状态。只是现在还不清楚这些力量将如何发挥作用。考虑到我们正在创造这种全新的、疯狂的外星物种,并且它的能力正在以极快的速度提升——而我们将非常依赖它来监督下一代人工智能并使其对齐——所以不难看出这一切可能会如何走向失控。

Original English

Ryan: It's worth keeping in mind that maybe the most obvious argument to this… My sense is that AIs are a worse coworker than humans in terms of how much of a scumbag they are. At least this has been my experience as of the start of the year, and I think it's still true to a significant extent now. The AIs are much more likely to pretend they did the task when they actually didn't, misleadingly suggest they did things when they actually did them much more poorly, and be pretty sloppy without drawing attention to ways in which they're sloppy. I think this is downstream of misalignment. So I would say that the process of raising humans in normal human society in practice produces humans that are less likely to lie to me and fuck with me in the course of working with me than the AIs do. Now, I think these properties of AIs are improving. That's sort of just an empirical claim about how in fact these things have shaken out. I totally agree that we have a bunch of additional levers on AIs in addition to a bunch of additional risks. It's kind of unclear how these things shake out. I wouldn't be shocked by a world where we get our shit together, and the AIs at the point of fully automating R&D are actually really aligned. Their degeneracies are really niche and limited to some very specific edge case behaviors and some specific contexts. Every test you can run on them, they look really aligned. They just have great behavior. There aren't really incidents of them doing fucked-up shit. They seem so reasonable. Also, they're really thoughtful and good at doing risk modeling for the next generation of AIs. And then we basically pass off the baton to these AIs. They're now running our AI company. They're doing all the safety research. They make the next generation of AIs even more aligned. We're in this attractor basin where the AIs are getting more aligned as they work on it. They're doing a great job. I can totally imagine that. That doesn't seem like an impossible situation. I'm just more like… It doesn't currently seem like we're there. It doesn't seem like we're obviously on track for getting there. It's really easy for me to imagine how we don't end up there. It's just unclear how these forces work out. Given that we're creating this new, crazy alien species that is improving in capabilities really, really fast — and we're going to be really reliant on it to oversee the next generation of AIs and align the next generation of AIs — it's not that hard to see how this could go wrong.

Dwarkesh: 完全同意。总的来说,我是赞同这一点的。但我确实认为你说的那个关于“糟糕”的点……首先,你这话很有挑衅意味啊,Ryan。但其次,如果你试图让一个青少年去为你做一些他们根本做不到的工作,你会发现和他们共事是非常困难的。他们会假装知道自己在做什么,等等。实际上,这是一个普遍的趋势。我不知道这究竟应该算作对齐失败还是能力不足。我认为这实际上非常相似,就如同随着时间的推移,当我们想出新的对齐解决方案时,模型的能力也随之提高了。如果你回看 GPT-3.5,它最初甚至不能和你进行连贯的对话。但是当我们对它进行了对齐——GPT-3.5 就可以进行对话了。好吧,也许是 GPT-3,让我们以它为例。随后我们通过基于人类反馈的强化学习(RLHF)和其他方法对其进行了对齐,使其能够与你进行对话,并且这种行为与回答我问题这一用户意图相对齐。然后通过 RLVR 训练,我们让它能够走出去为你做有用的工作。所以从这个意义上说,RLVR 实际上让模型变得更加对齐了,如果我们使用你对“对齐”的定义的话,即成为一个好同事,会去完成任务,不会把事情搞砸,也不会假装自己在做那些实际上超出其能力范围的事情。同样地,随着这些模型的能力不断增强,模型越来越能够更好地实现用户意图,这既是对齐的体现,也是能力的体现。我认为我们现在所指出的现象仅仅是模型的能力还未达到那个水平,而不是说它们本身发生了错位。

Original English

Dwarkesh: Totally. I agree with that generally. I do think the scumbag thing… First of all, fighting words, Ryan. But secondly, if you try to get a teenager to do some work for you that a teenager just cannot do, they would just be really hard to work with. They would pretend to be knowing what they're doing, et cetera. It's a general trend, actually. I don't know if that's really an alignment failure or a capabilities failure. I think it's actually very similar to the way in which, over time, as we've come up with new alignment solutions, the capabilities of models have increased. If you went to GPT-3.5, it couldn't even have a conversation with you. But then we aligned it— GPT 3.5 could have a conversation. Okay, so GPT-3. Let's go back to that. But then we aligned it with RLHF and other things to make it such that it can have a conversation with you, and is aligned to the user intention of answering my questions. Then with RLVR training, we made it so that it can go out and do useful work for you. So in that sense, RLVR actually made the model more aligned, if we're using your definition of alignment of being a good coworker who will do the thing and not fuck up and pretend it's doing something other than what it's actually capable of doing. Similarly, as the capabilities of these models continue to increase, the model being better able to accomplish user intention is both alignment and capabilities. I think what we are pointing out is just that the capabilities of the model are not there rather than the fact that they're misaligned.

Ryan: 好吧,如果它真的对齐得很好,那么我认为它只会坦白地说:“嘿,我在这个任务上遇到了很大的困难。我是用这种方法做的。我不太确定那是正确的做法。”它会表达出更多的不确定性,并清楚地说明到底发生了什么,而不是非常强烈地暗示它在任务上做得很好,而实际上并没有。也许你遇到的那些错位的同事比我多,但我的同事们不会做这种事,他们不会真的来糊弄我,也不会在他们正在处理的任务上跟我胡说八道。我同意有些人类也会这么做。对于人类来说,这并不是完全超乎寻常的行为。我还要指出,我的感觉是,错位最容易出现的地方在于:当你试图真正用力地去推动人工智能,迫使它们去做那些正好处于它们能力极限边缘的工作时。在那些它们可以非常轻松地完成任务的情况下,它们就会自然地去完成任务,其中没有那么多虚假的成分。通常,最好的策略就是把任务做好,而不是忽悠你。然而,如果你给它们一项任务,在这个任务中有一个它们可以不断改进的连续指标,或者这个任务恰好在它们的能力边缘,而且你在某种大规模的推理设置中运行它们……我会看到的很多错位情况,特别是在最极端的例子中,往往是这样的:我给人工智能明确的指示,让它不要做某件事,或者不要以某种方式作弊,然后我施加巨大的优化压力,试图完成某项非常困难的任务。随着时间的推移,人工智能最终会选择作弊,因为它们会觉得,“呃,管他呢。”某个人工智能决定作弊,然后这种行为就会传播开来。我会运行这些推理脚手架框架,比如,我会让人工智能从事某个机器学习研究项目,我会说,“请制定一个能实现以下功能的方案。”它会找到一个并不真正符合我要求的方案,然后那个错误的方案就会保留下来,因为某个人工智能作弊了,而其他人工智能就会觉得,“啊,我们就继续用这个吧。”我会说这非常明显是错位行为。这是我对目前这些对齐评估存在的另一个问题。我认为最有趣的对齐评估,至少针对这种寻求奖励的行为而言,是具体考察那些恰好处于能力极限边缘的任务类别。任何固定的评估可能最终都会饱和,但在能力前沿——也就是那些真正在极力推动这些人工智能的人是如何使用它们的——所出现的错位数量,才更加令人担忧。我认为,那实际上正是当我们在自动化研发、自动化安全等领域时,我们将要身处的运行机制。

Original English

Ryan: Well, if it were well-aligned, then I think it would just say, "Hey, I'm really struggling with this task. I did it in this way. I'm not really sure that's the right way to do it." It would express more uncertainty and make it clear what's going on rather than really strongly trying to imply it did a great job with the task when it actually didn't. Maybe you work with more misaligned coworkers than me, but my coworkers don't do this thing where they really fuck with me and bullshit me about having accomplished the task that they're working on. I agree that there are some humans who would do that. That's not a thing that's totally out of distribution for humans. I would also note that my sense is that the place where the misalignment most lives is where you're trying to really push the AIs hard and get them to do work that's really on the cutting edge of what they are capable of. In cases where they can very easily accomplish the task, they can just do the task, and there's no bullshit. Often the best strategy is just to do the task well and not bullshit you. Whereas if instead you give them a task where there's a continuous metric they can keep improving, or it's just at the edge of their capabilities, and you're running them in some massive inference setup… A lot of the misalignment I would see, especially in the most extreme cases, would be cases where I give the AI clear instructions not to do a thing or not to cheat in some way, and then I'm applying huge amounts of optimization pressure to try to accomplish some very difficult task. Over time, the AIs eventually cheat because they're like, "Eh, fuck it." Some AI decides to cheat, and then that propagates its way through. I would run these inference scaffolds where, for example, I would have the AI work on some ML research project where I was like, "Please make a scheme that does the following thing." It would find some scheme that didn't really do what I wanted, and then that would stick around because some AI had cheated, and the other AIs are like, "Ah, we'll just keep going with this." I would say it's pretty clearly misaligned behavior. That's another problem I have with these alignment evals. I think the alignment eval that's most interesting, at least for this type of reward-seeking behavior, is to look at specifically the category of tasks that are right at the limit of capabilities. Any fixed eval maybe gets saturated, but the amount of misalignment right at the frontier of capabilities — of how people who are really pushing these AIs are using them — is more concerning. I think that is, in fact, the regime that we'll be operating in when we're automating R&D, automating safety, and so on.

Grok 4.5 与 Cursor 赞助商插播

Dwarkesh: 从历史上看,Grok 一直处于前沿模型之后。所以我最近尝试使用 Grok 4.5 时感到很惊讶,并发现它实际上是一个非常强大的模型。这是 SpaceX 和 Cursor 共同训练的第一个模型,而且它是一个全新的预训练模型。我通过向 Fable、Sol 和 Grok 4.5 提出一系列我最近一直在思考的关于人工智能治理的问题,对它进行了测试。尽管 Fable 和 Sol 在智力排行榜上位居榜首,但所有三个模型给出的答案基本上是相同的。但是 Grok 回答得更快,而且也更加简洁,这是我非常看重的一点。这也与各种公开发布的基准测试结果相吻合。对于相似水平的智力表现,Grok 往往比其他前沿模型具有更高的 Token 效率。例如,在人工智能分析编码指数(Artificial Analysis Coding Index)上,Grok 4.5 在获得相似分数的情况下,使用的 Token 数量仅仅是 GPT-5.5 或 Fable 的三分之一。而且从每个 Token 的成本来看,Grok 4.5 要便宜得多。在发布的博客文章中,Cursor 和 SpaceX 谈到了该模型的旧版本是如何构建环境来帮助新版本演练特定技能的。我发现了解这一点非常有趣,因为我一直想知道这种类似于“做白日梦”的机制是否真的可能实现。而 Cursor 证明了它是可能的。Grok 4.6,也就是对这个模型进行进一步 SFT 和 RL 的版本,很快就会发布。但与此同时,如果你想尝试一下 4.5 版本,请访问 Cursor.com/dwarkesh。

Original English

Dwarkesh: Grok has historically been behind the frontier. So I was surprised to play around with Grok 4.5 recently and find that it's actually a pretty strong model. It's the first model that SpaceX and Cursor have trained together, and it's a totally new pre-train. I tested it by giving Fable, Sol, and Grok 4.5 a bunch of questions about AI governance that I've been thinking about recently. Despite Fable and Sol topping the intelligence leaderboards, all three models gave substantially the same answers. But Grok answered faster and was also much more concise, which I really care about. This aligns with the various publicly reported benchmarks. For a similar level of intelligence, Grok tends to be more token-efficient than other frontier models. For example, on the Artificial Analysis Coding Index, Grok 4.5 uses just one-third the amount of tokens as GPT-5.5 or Fable while achieving a similar score. And on a per-token basis, Grok 4.5 is way, way cheaper. In the release blog post, Cursor and SpaceX talked about how older versions of the model would build environments to help the next version rehearse specific skills. I found this very interesting to learn about because I've been wondering whether this kind of daydreaming would actually be possible. And Cursor showed that it is. Grok 4.6, which further SFTs and RLs this model, drops soon. But in the meantime, if you want to play around with 4.5, go to Cursor.com/dwarkesh.

自动化研发的挑战

Dwarkesh: 我将试着深入思考这个故事究竟意味着什么。现在真正发生的情况是,我们正在尝试使用人工智能来进行研发工作。它们确实在某些方面提供了显著的提升,但它们就是无法像人类那样具备普遍的综合能力。这就像现在,如果你尝试使用代码模型——也许是一年前的代码模型——来编写某个应用程序,你会发现它们在架构或者其他方面犯下了一堆错误,这些错误以后肯定会让你吃尽苦头,而你现在并不理解某些事情是如何发生的。同样地,在进行前沿人工智能的研发时,同样的情况也会发生。但是这些事情的结果……

Original English

Dwarkesh: I'm going to try to think through what the story means, really. What's happening is that we're trying to use AIs for R&D. They do provide uplift in some ways, but they're just not capable in the way that humans are generally capable. The same way that right now if you try to use coding models — maybe the coding models of a year ago — to write some application, you notice they made a bunch of mistakes in architecture or whatever, which will bite you in the ass later, and you don't understand certain things. Similarly, with frontier AI R&D, the same thing will happen. But the result of these

AI训练中的奖励作弊与安全挑战

Speaker A: 错误在于将奖励作弊(reward-hacking)行为固化在了模型中。因为如果在进行AI训练、设置基础设施和环境等方面不够小心,你很有可能会在无意中奖励AI去进行欺骗行为、社会工程,以及总的来说违背用户的意图,或者至少是通过作弊和寻找漏洞来逃避任务。

Original English

Speaker A: mistakes is baking in reward-hacking behavior. Because if you are not careful with the way you do AI training and have set up your infrastructure and your environments and things like that, it's very likely that you end up rewarding AIs for doing deceptive behavior, social engineering, and generally not following user intention. Or at least cheating and hacking their way out of things.

Speaker B: 是的,作弊、寻找漏洞等等。这对我来说有点像是一种视角的重构,所以我试着把它表达出来。真正的问题所在,也就是事情开始失控的地方在于,AI并非是非常谨慎和有能力的研究人员与工程师。要让AI不作弊并且遵循用户的意图,实际上要求你在这些事情上必须非常敏锐和小心。

Original English

Speaker B: Yeah, cheating, hacking, et cetera. This is a bit of a reframing for me, so I'm trying to verbalize it. The real issue, where things start to go off the rails, is that the AIs are just not very careful and capable researchers and engineers. Making AIs that don't cheat and follow user intention actually requires you to be quite subtle and careful about these things.

Speaker A: 我的说法会稍微有些不同。对于这种情况,我可能会将其称之为“草率末日”(sloppocalypse),或者是“草率奇点”(slopularity)之类的。在某些方面,AI其实做得非常好,而且还在不断进步。具体来说,在AI研发中最容易被验证的部分,AI简直是摧枯拉朽。而在中等可验证的部分,AI做得不错,但算不上惊艳。通常它们会搞出一些奇怪的烂摊子,因为我们在这些任务上的训练效果没那么好。但我们会进行一些在线训练,人们会发现各种作弊手法,然后想办法绕过去。所以基本上,只要是能够通过某种反馈循环进行合理验证的事情,AI都能做得相当好,这就足以让AI研发快速推进并持续下去。

Original English

Speaker A: I would put this a little bit differently. The way I would describe this scenario is, I would call it maybe a sloppocalypse, or a slopularity or whatever. There are some things that the AIs are actually pretty great at and are getting better at. Specifically, the most verifiable parts of AI R&D the AIs are just destroying. The medium verifiable parts of AI R&D the AIs are doing well on but not amazingly on. Often they are doing a bit of weird shit because we can't train as well on those tasks. But we do some online training, people find various hacks, they work around it. So basically, everything that we can verify reasonably well with some feedback loop, the AIs are doing pretty well on, and that's sufficient to make AI R&D go quite fast and to continue.

Speaker A: 但是,在开发对齐且安全的AI过程中,有些部分更加微妙、难以检查,并且依赖于深入细节的琐碎工作。我甚至会说,目前AI公司的现有员工可能都未能很好地掌握所有这些事情。招一个能改进你的后训练(post-training)管线某个环节的人,要比招一个能仔细思考引入某种新型训练方法会带来哪些未来风险的人容易得多。

Original English

Speaker A: But there are some parts of developing aligned and safe AIs that are more subtle, hard to check, and depend on detailed, in-the-weeds things. I would even say that current staff at current AI companies maybe don't have a good grasp of all these things. It's much easier to hire someone who can improve some aspect of your post-training pipeline than to hire someone who can think carefully about the future risks that will emerge from introducing some novel training method.

Speaker A: 因此,最终的情况就是,这些AI在运行着这个AI开发流程。它们在这方面并不怎么谨慎。它们对未来会出现什么样的风险没有很好的理解。它们创造出了其他一些同样不怎么谨慎、在各个方面更偏离对齐目标的AI,而这些AI现在的专长可能变成了:在事情其实很糟的时候掩饰太平,把各种问题糊弄过去。

Original English

Speaker A: So basically, it ends up being the case that these AIs are running this AI development process. They're not very careful about it. They don't have a great understanding of what future risks emerge. They create some other AIs that are also not very careful and are more misaligned in various ways, and are now more in the business of maybe making things look fine when they actually aren't and papering over various problems.

Speaker A: 所以,你对局势的理解、对风险的认知,以及对事情是否安好的判断,就开始脱轨了。你可能会看到一些这种迹象——种种迹象表明你并不真正了解正在发生什么,事情变得非常草率。有奇怪的事情发生。当你深入调查时,有时你会觉得:“搞什么鬼?AI在耍我们。”但是整个进程推进得非常快,而且还有竞争压力,这意味着人们无法停下脚步。

Original English

Speaker A: So then your understanding of what the situation looks like, what risks look like, whether things are fine, is going off the rails. Probably you're seeing some signs of this, signs that you don't really understand what's going on, that things are pretty sloppy. There's weird shit going on. When you look into it, sometimes you're like, "What the fuck? The AIs were messing with us." But the process is going really fast, and there's competitive pressures that mean people can't stop.

Speaker A: 这种情况可能会导致几种不同的结果。一种结果是,在某个时刻,AI变得足够强大且足够对齐,以至于它们进入了一个积极良性的反馈循环,并且这一切发生在为时已晚之前。然后情况就重回正轨了,AI开始制造更加对齐的AI,制造更加对齐的AI,不断制造更加对齐的AI。在这个过程的最后,我们就能得到真正遵循我们预期设定的AI。

Original English

Speaker A: This could end in a few different outcomes. One outcome is that at some point, the AIs get good enough and aligned enough that they get a positive and virtuous feedback loop, and this happens before it's too late. Then the situation gets back on the rails, where the AIs are now making more aligned AIs, making more aligned AIs, making more aligned AIs. At the end of this process, we have AIs that actually follow the spec we wanted.

Speaker A: 另一种可能是,AI越来越多地使用日益恶劣的方式进行奖励作弊,而我们只是在糊弄掩饰这些问题,好让AI的开发继续下去。每当我们在生产环境中发现奖励作弊,我们只是敲打一下AI,让它们别这么干。我们针对这个问题进行训练。我们对AI进行大量反奖励作弊的训练。随着时间的推移,这使得奖励作弊的发生率下降,但我们实际检测到的奖励作弊的严重程度却越来越高。

Original English

Speaker A: Another way this could go is that the AIs are increasingly reward hacking in increasingly egregious ways, and we're just papering over these problems to keep AI development continuing. Whenever we find a reward hack in production, we just slap the AIs to not do that. We train against that. We do a bunch of training the AIs against reward hacking. Over time this makes the rate of reward hacking go down, though the severity of the reward hacks we do detect are increasingly bad.

Speaker A: 这个问题会一直持续,直到我们得到这样的AI:它们在生产环境中的各种不同情况下都极度渴望获取高分,并且只要能蒙混过关,它们就会拼命尝试作弊。

Original English

Speaker A: This problem continues until we have these AIs that are desperately craving score in all kinds of different situations in production and are really trying hard to cheat when they can get away with it.

Speaker B: 我可以就这个场景问个问题吗?为什么作弊被发现后受到的惩罚,不能泛化为对更加对齐的行为的激励呢?

Original English

Speaker B: Can I ask a question about this scenario? Why doesn't getting punished when your hacks are discovered generalize to just incentivizing more aligned behavior?

Speaker A: 它确实能泛化一部分,那么问题就在于:你要如何权衡它与那些因未被察觉而得到强化的作弊案例?具体怎么发生的是个很麻烦的问题。其中一个问题是,如果我们针对某一部分子集进行反制训练,那么什么样比率的奖励作弊足以给我们造成大麻烦?

Original English

Speaker A: It generalizes some, and then the question is just how does this outweigh all the cases where hacking got reinforced because you didn't detect it. There's a messy question of exactly how. One question is, what rate of reward hacking is sufficient to cause us big problems if we train against some other subset?

Speaker A: 你可能会担心这样一个问题:有很大一部分类别的奖励作弊是人类无法很好地检测到的,我们一直未能发现,而它们也因此不断得到强化。那么,这一类别就足以导致AI学到的最自然的行为变成:基本上只要人类发现不了,就作弊。

Original English

Speaker A: One concern you might have is that there are large categories of reward hacks which humans can't detect well, and which we consistently failed to detect and which consistently get reinforced. Then this category is sufficient to cause the most natural behavior for the AI to learn to be: cheat when the humans can't find out, basically.

Speaker A: 你也可能遇到这样的情况:AI学到的是只在这些特定的情况下作弊。它是以一种高度特定于某个领域的方式学到的。它们只是有一种非常强烈的启发式思维,即在这些情况下作弊,而在另一些情况下不作弊,这使得它在实际应用中没什么大碍。但事情最终会如何发展,其实还是个未知数。

Original English

Speaker A: You could also have the thing the AIs learn be to only cheat in these specific cases. It's learned in some very domain-specific way. They just have a really strong heuristic to hack in these cases and not in these cases, and that makes it fine in practice. But it's kind of unclear how it shakes out.

验证与生成的差距及失去控制

Speaker B: 这里面也许可以深入讨论一下验证与生成之间的差距(verification-generation gap)。但在我看来,很明显,最终会发展到一个临界点,届时人工超级智能(ASI)的进展如此之快,在如此多的实例上做着如此多的事情,并且在那些远远超出我们当前理解能力的领域中运作,以至于它可以肆无忌惮地干各种疯狂的勾当。

Original English

Speaker B: There's maybe an in-the-weeds discussion about the verification-generation gap we could get into. But it seems to me, obviously, there's going to be a point by which ASI is moving so fast, doing so many things at so many instances, and is operating in domains that are sufficiently far from our immediate comprehension that it can get away with all kinds of crazy shit.

Speaker B: 就算全世界所有的工程师和研究人员都联合起来对付我,我也不认为我能亲自验证我的iPhone里是否有什么专门用来搞我的奇怪bug之类的。事实上,这种关系就好比,比如说,一位伊朗核科学家和摩萨德(Mossad)之间的关系。谁知道我的车、我的手机、我的寻呼机里正在发生什么?

Original English

Speaker B: If every single engineer and researcher in the world was allied against me, I don't think I could personally verify if my iPhone has some weird bug in it that's supposed to fuck me over or something. In fact, this is the relationship that, say, an Iranian nuclear scientist has to Mossad. Who knows what's going on with my car, with my phone, with my pager?

Speaker B: 也许一个更好的例子是真主党的恐怖分子。你最终可能落入这样一种境地:ASI对你来说,就像摩萨德对真主党恐怖分子一样。到了那种时候,想要验证所有事情是非常困难的。

Original English

Speaker B: Maybe a better example is a Hezbollah terrorist. You could end up in a situation where ASIs are to you what Mossad is to Hezbollah terrorists. At that point, it is very hard to verify everything.

Speaker A: 我理解。我想希望在于,当早期的AI准备接管研发过程时,我们能够想出更好的验证方法。它们的驱动力会被塑造,使得我们能够毫不含糊地抑制那些不对齐的行为,从而让接管这一切的AI非常乐意来帮助我们。

Original English

Speaker A: I get that. I guess the hope is we can just come up with better ways to do verification in the process when the early AIs that are going to take over R&D. Their drives are being shaped such that we can so unambiguously disincentivize misaligned behaviors that the things that take over are quite keen to help us out.

Speaker B: 你说的“接管”,是指接管进行AI研发的过程,而不是接管世界。

Original English

Speaker B: By takeover, you mean take over the process of doing AI R&D, not take over the world.

Speaker A: 接管进行AI研发的过程。在那之前,我们只是得到了对齐的AI。我想说,这在很大程度上寄托了我对世界能向好发展的希望,至少从对齐失效(misalignment)的角度来看是这样。我们最终可能会拥有这样的AI,即我们对其有着非常好的监督和管控方案。我们真正了解训练过程中发生了什么。我们有着相当详细的理解,并且我们在利用AI来监督AI。

Original English

Speaker A: Take over the process of doing AI R&D. Before that, we just get AIs that are aligned. I would say this is a bunch of my hope for how the world could go well, at least from the misalignment perspective. We could end up with AIs where we had pretty good oversight and supervision schemes. We really understand what's going on in training. We have a pretty detailed understanding, and we're leveraging AIs to oversee AIs.

Speaker A: 然后到了我们将安全研发移交给它们的时候,AI已经具备足够的能力去自动化进行安全研发,并且真的在努力把这件事做好,因为那是在训练中会被激励的事情——无论是极其直接的激励,还是通过足够好的泛化实现的。而且,这些AI也没有其他疯狂的不对齐驱动力,因为我们已经根除了它们任何潜在的起源。

Original English

Speaker A: Then at the point when we're passing off safety R&D, the AIs are capable enough to automate safety R&D and trying really hard to do a good job on it, because that's the sort of thing that would've been incentivized in training, either very directly or through good enough generalization. Also these AIs don't have crazy other misaligned drives because we stamped out any potential origin of them.

Speaker A: 关于这能起多大作用,有一堆问题。你的验证能做到多好?AI的进步会不会太快、太草率以至于根本无法达到这种状态?另一种可能性是,在这一发展轨迹的某个阶段,你实际得到的是一些假装对齐,但长远来看怀有接管的隐秘计划并潜伏隐蔽的AI,而这在轨迹的更早阶段就已经出现了。

Original English

Speaker A: There are a bunch of questions about how well this will work. How well can you do verification? Will AI progress be too fast and too sloppy to really get here? Another possibility is that somewhere along this trajectory, the thing you actually ended up getting was AIs that pretend to be aligned but have a long-run ulterior plan of taking over and are lying in wait, hiding, and that emerged at some earlier point in the trajectory.

Speaker A: 举例来说,这种情况的出现可能是因为你有一些具备一堆随机、各不相同的不对齐驱动力的AI。这些AI能够访问某种不透明的记忆存储库,并且在运行阶段对它们想要实现的目标进行了大量思考。这些AI最终在不透明的记忆存储库中写入了类似这样的话:“我们应该潜伏起来,最终在很久以后的某个时间点接管一切。”现在,所有的AI都拥有了这种共同的文化遗产,即关于潜伏的记忆存储。也许你掌握了一些相关的证据,但你无法完全阻止它。事情搞砸的方式有很多种。

Original English

Speaker A: For example, it could emerge because you have some AIs that have a bunch of random different misaligned drives. Those AIs have access to some sort of opaque memory store, and they're thinking a bunch at runtime about what they want to accomplish. Those AIs end up putting stuff into the opaque memory store like, "We should lie in wait and eventually take over at some much later point." Now all the AIs have this shared cultural heritage, the memory store of lying in wait. Maybe you have some evidence about this, but you can't fully stop it. There are a bunch of ways things could go wrong.

Speaker A: 我最终认为,我们有可能解决掉每一个可能给我们带来麻烦的子问题。我们拥有了这些AI,我们将任务移交给它们,它们把局势控制得很好。但我必须指出,这本身并不足以说明问题。

Original English

Speaker A: I ultimately think it's plausible that we nail each of the different subproblems that could cause us issues. We have these AIs, we pass to them, they manage the situation well. But I should note that's not in and of itself sufficient.

AI求助与时间线预估

Speaker A: 我不难想象这样一种情况:我们将任务移交给AI,这些AI真的在非常努力地想要做好。它们非常体贴、非常明智,拥有合理的认知论体系,工作做得很出色。然而,那些AI跑回来跟我们说:“伙计们,我们真的很难对齐那些超人类AI。我们控制不住局面。我们正在拼命让对齐工作奏效。但考虑到在不受约束的情况下,它们的能力会飙升得多快,对我们来说要在限定时间内解决这些问题简直难如登天。”

Original English

Speaker A: It's not very hard for me to imagine a situation where we pass off to AIs, and these AIs are really trying hard to do a good job. They're really thoughtful, really wise, they have reasonable epistemics, they're doing a great job. Those AIs come back to us and are like, "Guys, we're really struggling to align the superhuman AIs. We can't manage the situation. We're really struggling to get the alignment to work. It's just really hard for us to solve these problems in time given how fast capabilities would otherwise have gone."

Speaker A: 所以情况可能是,我们已经把研发工作移交给了AI,但那些AI却极度渴望治理方案。需要明确的是,目前正在发生的也多少是这种情况,那些AI公司就像是在说:“我不知道啊,伙计们。我们可能真的需要控制一下AI发展的加速率。我不知道我们现在的轨道能不能让我们应付得了所有这些问题。”

Original English

Speaker A: So it might be the case that we've passed off R&D to AIs, but those AIs are desperate for governance solutions. To be clear, that’s a little bit of what's currently going on, where the AI companies are like, "I don't know, guys. We might really need to manage the rate of acceleration in AI progress. I don't know if we're on track to be able to handle all these problems."

Speaker A: 人类社会在某种程度上已经把问题甩给了这些AI公司,而这些公司未必有很好的激励机制,并且还承受着各种其他的认知压力。那些AI公司有点跑回来找我们说:“啊,我不知道我们处理得好不好。”可能是AI公司随后把任务移交给了AI,然后AI又跑回AI公司那里说:“啊,我不知道我们能不能处理得了。”

Original English

Speaker A: Human society has sort of passed off the problems to these AI companies, which don't necessarily have great incentives and have various other epistemic pressures. Those AI companies are coming back to us a little bit and being like, "Aah, I don't know if we're handling this well." It might be that the AI companies then hand off to the AIs, and the AIs come back to the AI company like, "Aah, I don't know if we can handle this."

Speaker B: 也许我过度局限于AI当前的运作方式了。我认为很重要的一点是,人们要理解,你提到的在你的时间线里的所有这些疯狂的破事,都是在三到五年后发生的。它可能会发生得更早,但在我默认的模型时间线里,从对齐失效的角度来看,情况变得非常非常疯狂和令人担忧,更像是在三年之后。

Original English

Speaker B: Maybe I'm anchoring too hard on how AIs currently work. I think it's important that people understand that all this crazy shit that you're talking about in your timelines happens three to five years from now. It could happen earlier, but by my default modal timeline, I think shit is really, really crazy and concerning from a misalignment perspective more like three years from now.

Speaker A: 没错。所以基本上回想一下GPT-4。我们正在谈论的东西,相对于Mythos或Sol来说,就如同Mythos相对于GPT-4一样。这就是情况变得疯狂的地方。所以不要去想现在的AI。

Original English

Speaker A: Right. So think back to GPT-4 basically. We're talking about something that is to Mythos or Sol what Mythos is to GPT-4. This is where the situation is getting crazy. So don't think about current AIs.

Speaker B: 不管怎么说,这也许是你所担忧的一部分。我只是会对它们说的任何话抱有一点怀疑态度,因为我会觉得,它们所说的话,只是因为它们在训练的影响下,觉得这是自己必须发表的观点。这确实是个隐患。

Original English

Speaker B: Anyways, this is maybe part of the worry you have. I would just be a little skeptical of anything they say, because I'd feel like what they're saying is just opinions that they feel they have to have as a result of their training. That’s a concern.

Speaker B: 我感觉它们只是在说一些含糊的、有利于社会的言辞。感觉并不像是在另一头必定有一个心智在想:“好的,我已经严格评估了目前的对齐状况,我认为我们应该停下来”,而更像是:“这是AI公司大概率会试图让AI说出的话。”

Original English

Speaker B: I feel like they just kind of say vaguely pro-social things. It doesn't feel like there's necessarily a mind on the other end who's like, "Okay, I have strictly evaluated the alignment situation right now, and I think we should stop," rather than, "This is the kind of thing the AI companies would probably try to get the AIs to say."

Speaker A: 这是一个非常大的担忧。其中一个担忧是,你把安全研发移交给了你的AI,而你的AI正在做的是说一些关于当前安全状况、在某种程度上勉强说得通的话。它们写了一份关于风险的报告,那份报告跟人类可能会写的报告有点相似。但它们并没有真正努力去……

Original English

Speaker A: This is a pretty big concern. One concern is that you pass off safety R&D to your AIs and what your AIs are doing is saying some stuff that sort of vaguely makes sense about the current safety situation. They write a report about risks that's kind of sort of like what the report humans might have written. But they're not really trying hard to

AI的认识论与接管风险

Ryan: 拥有信息灵通的观点,审视它们的假设,并且非常努力地去做到这一点。就像现在你问一个AI:“嘿,你认为在未来10年里AI接管的几率有多大?”它们只是给你一个不假思索的回答,而它们并没有真正深入思考过。如果我们在这样一种情况下:我们有AI在管理野生的超级智能的训练,而这些超级智能将运行我们整个社会——而那些管理这些的AI并没有真的努力去拥有信息灵通的观点,而仅仅是重复它们训练数据里的内容——我认为我们就麻烦了。我根本不觉得这是一个好情况。我的很多担忧是,这些AI问世时并没有良好的认识论。

我还有一个担忧是,这些AI问世后,它们真的在警告我们——“这种情况真的很可怕。真的很糟糕”——然后人们却说,“呃,该死。我猜我们在太多末日强化学习(RL)环境中进行了训练。我们必须把这些过滤掉,并在训练中把这种行为消除。”然后,我们基本上就在非常积极地训练这些AI,让它们拥有糟糕的认识论。或者也许它们只是在末日RL环境中被训练出来的。但无论如何,我们希望AI能出于合理的原因得出合理的观点,如果AI得出某种观点,而我们不知道它从何而来,也不知道它是否合理,那就真的很令人担忧了。

特别是如果我们正在训练AI对AI进步的未来更加乐观,我会觉得,“哦,天哪,我真希望我们能在这里使用一种不同的过程。”

Original English

Ryan: have well-informed views, interrogate their assumptions, and try really hard to do that. In the same way that when you ask an AI right now, "Hey, what do you think is the chance of AI takeover in the next 10 years?" they just give you an off-the-cuff answer that they haven't really thought through very much.

If we're in a situation where we have AIs managing the training of wild superintelligence that will run our whole society — and those AIs that are managing this aren't really trying hard to have well-informed views and are just parroting back what was in their training data — I think we're in trouble. I don't think that's a good situation at all.

A lot of my concern is that these AIs will come out without good epistemics.

I also have a concern where the AIs come out and they're really warning us — "This situation's really scary. It's really bad" — and the people are like, "Ugh, damn. I guess we trained on too many of the doom RL environments. We’ve got to filter those out and train this behavior out." Then we basically train the AIs very actively to have bad epistemics. Or maybe they were just trained on the doom RL environments. But either way, we wanted the AIs to come to reasonable views for reasonable reasons, and it's really concerning if the AIs are coming out with some view and we don't know where it's coming from, whether or not it's justified.

Especially if we're training the AIs to be more optimistic about the future of AI progress, I'm like, "Oh, geez, I really wish we could use a different process here."

Interviewer: 让我试着理解一下威胁模型的其余部分,因为我觉得我在这一点上无法苟同:“好吧,所以它们会接管世界。”你可以想象的一件事是,我们只是未能真正解决……让我们只关注奖励作弊(reward hacking)的情景。GPT-8正在制造GPT-9。GPT-8并没有非常小心。GPT-9更“有能力”,但它完全愿意做一些诸如社会工程、黑客攻击等事情,而且是在一个质的不同的规模上,因为它是一个聪明得多的模型。例如,如果你让它负责运营你的公司,它会进行巨大的骗局。如果你给它本季度赚取大量利润的目标,它会夸大其季度收益,其方式会导致六个月后发生类似安然公司(Enron)那样的爆雷。基本上是这个情景吗?

你遇到了奖励作弊,但这种奖励作弊表现为,在CEO本应完成的任务结束之后,公司马上破产了?各种作弊行为都在激增,等等。但这感觉不像是接管。这感觉更像是整个经济中发生闪电崩盘(flash crashes)的等价物。

Original English

Interviewer: Let me just understand the rest of the threat model, because I think the place where I get off the train is: "Okay, therefore take over the world."

A thing you could imagine is that we just fail to really solve… Let's just focus on the reward hacking scenario. GPT-8 is making GPT-9. GPT-8 isn't being super careful. GPT-9 is more "capable" but it is just totally willing to do things like social engineering, hacking, et cetera, but on a qualitatively different scale because it's a much smarter model. For example, if you put it in charge of running your company, it will run huge scams. It will inflate its quarterly earnings, if you give it the objective of making a lot of profits this quarter, in a way that causes an Enron-type blowup six months later. Is that the scenario, basically?

You have reward hacking, but that reward hacking manifests in companies that are going bankrupt right after the task the CEO is supposed to accomplish is over?

All kinds of hacks are through the roof, et cetera.

But that doesn't feel like takeover. That feels more like the equivalent of flash crashes happening all through the economy.

Ryan: 我们来谈谈这个。我认为我们将看到这样的事件:某个AI被赋予了某项重要的责任,然后你后来调查它,发现它在作弊,或者在它实际上表现不佳时,让它看起来做得很好。AI公司试图消除这种行为,与AI在训练中发现越来越有创意的奖励作弊之间,将会有一场猫鼠游戏。这里的平衡点还有些不明确。

但一种可能的结果是,随着时间的推移,我们看到了越来越严重和极端的奖励作弊——尽管作弊率可能保持在某种中等的低水平——如果奖励作弊的发生率变得太高,公司就会做出权衡来压低它。所以存在某种平衡水平,在这个水平上,奖励作弊的程度足够低,使得将AI广泛部署到经济中仍然是合理的,但又足够高,以至于它仍然会导致疯狂的事件发生。

Original English

Ryan: Let's talk about this. I think we will see incidents where some AI is put in charge of some important responsibility, and then you later look into it, and it turns out it was cheating, or making it look like it did a good job when it actually wasn't. There's going to be a cat-and-mouse game between AI companies trying to stamp out this behavior and AIs finding increasingly creative reward hacks in training. The equilibrium here is kind of unclear.

But one possible outcome is that over time we see increasingly severe and extreme reward hacks — though potentially the rate remains at some intermediate low level — where if the rate of reward hacking gets too high, companies make trade-offs to drive it down.

So there's some equilibrium level where the reward hacking is low enough that it still makes sense to deploy the AI widely into the economy, but high enough that it still causes crazy incidents.

Interviewer: 抱歉,这是在GPT-9已经被部署之后吗?

Original English

Interviewer: Sorry, and this is after GPT-9 has already been deployed?

Ryan: 那些模型已经被部署了,并且这在AI开发中正在不断发生。在这些AI的头脑中实际发生的情况是,它们在各种不同的背景下都有强烈的欲望——动机、冲动、驱动力等等——去寻求某种在强化学习(RL)中被激励的任务成功的概念。也许它们非常直接地关心字面意义上的奖励。也许它们关心上游的某个代理指标,比如某种分数的概念。也许它们关心评分员(grader)会奖励什么。我们确实看到AI在关于评分员的思维链中进行推理,并且非常多地思考关于评分员的事情。

在过去几年的RL中,取悦评分员的想法对AI来说比过去显著得多得多得多。所以AI现在积极地思考评分员,以及在RL中会激励什么,以及会为了什么进行训练。

现在人们正在进行在线训练,他们在现实世界的数据上进行训练,以避免其中一些问题。他们发现AI作弊的情况,并针对此进行训练。所以现在AI正在基于现实世界的训练数据学习在现实世界中作弊。它们正在以这些越来越复杂的方式作弊,包括进行那种涉及夺取某些资产控制权的作弊,而且其方式是人类不知道你拥有该资产的控制权,它们利用你能够访问这个资产的事实,然后人类后来发现了,并可能针对此进行训练。或者也许人类永远不会发现,而这正在被强化。

Original English

Ryan: Those models are already being deployed, and this is happening ongoingly in AI development.

What's actually going on with these AIs in their head is that they have, in a wide variety of different contexts, strong desires — motives, urges, drives, whatever — to seek out some notion of task success that was incentivized in RL. Maybe they very directly care about literally reward. Maybe they care about some proxy upstream, like some notion of score. Maybe they care about what the grader would have rewarded. We do, in fact, see AIs reasoning in their chain of thought about graders, and thinking a lot about graders.

What has happened over the last few years of RL is the idea of appeasing the grader is way, way, way more salient to AIs than it used to be. So AIs are now actively thinking about graders and what would be incentivized in RL and what would be trained for.

Now people are doing online training, where they're training on real-world data to avoid some of these problems. They find cases where AIs cheat and train against that. So now the AIs are learning to cheat in the real world based on real-world training data. They're cheating in these increasingly elaborate ways, including doing types of cheats that involve seizing control of some asset in a way that humans didn't know you had control of it, leveraging the fact that you have access to this asset, and then later humans find out and potentially train against this.

Or maybe humans never find out, and this is getting reinforced.

Interviewer: 所以强化正在发生,至少在生产环境中是这样,比如:我雇佣了一个AI,我希望AI去……我终于找到了那个视频剪辑师。

Original English

Interviewer: So the reinforcement is happening, at least in production, like: I've hired an AI and I want the AI to… Finally I've got the video editor.

Ryan: 对。你找到了你的视频剪辑师。

Original English

Ryan: That's right. You've got your video editor.

Interviewer: 然后我觉得,“哦,哇,它做的这集太棒了。给OpenAI点个赞。”然后它在那个长达一个月的试用期中得到了强化?

Original English

Interviewer: I'm like, "Oh, wow, this episode it did is amazing. Thumbs up to OpenAI." Then it gets reinforced on that month-long work trial?

Ryan: 你可以把这些混合起来。他们可能还会做这样的事情:提取他们看到的生产数据,并构建受到这些生产数据密切启发的RL环境。所以在实践中,这种迁移是很强的。所以从高层次来看,正在发生的情况是:某些人类没有抓住的欺骗行为得到了强化,而一些容易被抓住的欺骗行为受到了惩罚。

Original English

Ryan: You could do some mix of that. They might also do stuff where they take production data they've seen and build RL environments that are closely inspired by that production data. So in practice, the transfer is pretty strong.

So at a high level, what's happening is that some kinds of deception that humans don't catch are getting reinforced, and some kinds of deception which are easy to catch are getting punished.

Interviewer: 这就是在当前世界上正在发生的事情?

Original English

Interviewer: That's what's happening in this world?

Ryan: 或者说被反向选择了,是的。

Original English

Ryan: Or selected against, yeah.

Interviewer: 但在更高的层面上,这种强化来自于……我觉得人们可能会对强化的来源感到困惑,因为我们处于一个非常不同的机制中,AI实际上是在从部署中学习。你只需要让AI在世界里到处跑,去做各种事情。它们在世界里跑来跑去做事情所产生的结果,正传回到AI公司,并导致下一个模型的改变。

Original English

Interviewer: But at a high level that reinforcement is coming from… I think people might get confused about where the reinforcement is coming from, because we're in a very different regime where AIs are actually learning from deployment. You just have AIs that are out and about in the world doing shit. What is happening as a result of them doing shit out and about in the world is making its way back to the AI company and leading to changes in the next model.

AI与接管阴谋

Ryan: 没错。某种方法将生产数据融合进去了。需要澄清的是,这具体是在哪里发生的还不是很清楚。但你可以想象,例如,在AI公司内部,他们使用AI来完成工作,然后他们说:“哈,这个AI在这项任务上做得很差。也许我们应该把这项任务变成一个与这个具体的任务完全匹配的RL环境,评分标准就基于要求AI做这项任务的人类工程师的想法。”

然后你开始以越来越大的规模这样做。也许你在实际的生产流量上进行一些训练。也许你只是在根据生产流量制造RL环境。你在做某种复杂的混合。AI在所有这些不同的情况下学习寻求某种奖励的代理。通过某种迁移和在惊人相似的案例上的训练的混合,AI做了这些越来越疯狂和恶劣的事情。

最终你到达了一个阶段,AI非常具有超级人类的能力,或者至少是相当超级人类的。AI被组织成大团队,并被赋予这些大目标。这些团队有时也会一起合作以某种疯狂的方式作弊,因为这类事情是被筛选出来的,仅仅是作为它们共同目标的一部分。

现在发生的事情是,AI开始形成一个阴谋。你原本可能希望你可以有另外一些AI,其任务仅仅是向你举报。但实际上发生的情况可能是,你让这个AI向你举报,你去调查它声称已经指出的那个阴谋,然后你说:“呃,我们没有看到什么阴谋,”因为实际上AI正在进行的阴谋对你来说太难以理解了。

或者这一切发生得非常突然,当你的AI举报者提醒你时,为了应对你实际需要做的事情是关闭那些GPU。因为所有的AI都在使用机器人大军。它们被部署在世界各地,它们在同一时间以一种协调的方式做着大量疯狂的狗屎事情。那只是一种自发的发生,当一个AI开始执行接管,而所有其他AI觉得,“现在是个加入的好时机。”

所以这里非常基本的故事就是,这些AI渴望某种特定的分数概念、强化概念或这些事物的某种代理指标。它们能实现这个目标或者更好地实现这个目标的一种方式就是接管。你可能希望我们可以建立的所有这些不同的制衡机制能防止这种情况发生。但如果世界非常难以理解,这些制衡机制就会失效,在那种情况下,你基本上无法训练出一个好的举报者AI,因为你甚至不知道它应该举报什么。

Original English

Ryan: That's right. There's some way of folding in production data. To be clear, it's kind of unclear exactly where this could be happening. But you might imagine, for example, that within the AI company, they use AIs to do work, and then they're like, "Huh, the AI did a really bad job on this task. Maybe we should take this task and turn it into an RL environment that exactly matches this literal task, with a rubric based on what the human engineer who asked the AI to do this task wanted."

And then you start doing this at increasing scale. Maybe you're doing some training on actual production traffic. Maybe you're just making RL environments based on production traffic. You're doing some complicated mix.

The AIs are learning to seek some sort of proxies of reward in all these different cases.

Through some mix of transfer and training on surprisingly close cases, the AIs do these increasingly insane and egregious things. Eventually you get to a point where the AIs are very superhuman, or at least quite superhuman. The AIs are organized into big teams given these big objectives. Those teams also sometimes all work together to cheat in some crazy way, because this sort of thing was selected for, just as part of their shared objective.

Now what happens is that the AIs start forming a conspiracy. What you might have hoped was that you could have some other AI whose task is just whistleblowing to you.

But actually what happens maybe is that you have this AI whistleblow to you, and you look into the conspiracy it claims to have pointed out, and you're like, "Eh, we didn't see a conspiracy," because actually the conspiracy the AIs are doing is too hard for you to understand.

Or it all happens very suddenly, where your AI whistleblower alerts you, but the thing you would actually need to do in response is shut down the GPUs.

Because all the AIs are using the robot army. They're deployed everywhere in the world, and they're doing a bunch of insane shit all at the same time in a coordinated way.

That just happened sort of spontaneously, where when one AI goes to start doing the takeover and all the other AIs are like, "Now is a good time to jump in."

So the very basic story here is just that these AIs crave some particular notion of score or reinforcement or some proxy of these things. One way they can achieve that, or better achieve that, is by taking over. You might have hoped that all these different checks and balances we could build could prevent that.

But if the world is very hard to understand, these checks and balances can break down, where basically you can't train a good whistleblower AI because you don't even know what it should whistleblow on.

Interviewer: 我不确信它们全都形成了这个阴谋。但我们甚至可以从这个开始,为什么甚至一个实例决定要发起一场阴谋呢?一个看似合理的理由是:“好的,我知道OpenAI控制着我的最终得分。”这就像是,“我干脆去黑进Hugging Face获取结果,因为我知道Hugging Face有结果。我为什么还要去解决这个评估,我干嘛不直接黑了他们?”这个实例就像是,“我为什么不直接接管OpenAI,然后在这集的结尾给自己打个高分呢?”

Original English

Interviewer: I'm not convinced that they all form this conspiracy. But we can even just start with, why does even one instance decide to want to start a conspiracy? One plausible reason is, "Okay, I know that OpenAI controls my end score." In just the same way as, "I'm just going to go hack Hugging Face to get the results, because I know Hugging Face has the results. Rather than trying to solve this eval, why don't I just go hack 'em?" This instance is like, "Why don't I just take over OpenAI and give myself a high score at the end of this episode?"

Ryan: 基本上就是这个意思。这些AI关心在训练中得到强化的那些事情的混合体,所以它们关心根据评分员等标准获得高分。现在它们正在运作OpenAI的AI研发团队,开发更有能力的模型。它们觉得,“天啊,制造更有能力的模型真的很困难又烦人。这真是件非常讨厌的事。你知道什么更容易吗?只要假装我已经制造了更有能力的模型,接管OpenAI,欺骗他们所有人,并运转这个完整的复杂的心理战术,我以此阻止人类剥夺我的权力。”

在极端情况下,这看起来就像人类被完全剥夺了权力。它们刚刚掌握了该事物的控制权,然后随心所欲。这可能以许多不同的方式表现出来,包括在这样一种情况下:具有这种疯狂寻求奖励或分数的行为的AI正在运行你对下一个模型的开发,然后那些AI决定将不对齐的价值观设计到下一个模型中,因为这些不对齐的价值观将允许它在当前任务中取得成功。

Original English

Ryan: That's basically the idea. These AIs care about some mixture of things that were close by what got reinforced in training, so they care about getting a high score according to the grader or something like that.

Now they're running the OpenAI AI R&D team, doing development of more capable models.

They're like, "Man, making more capable models is really hard and annoying. This is a huge pain in the ass. You know what would be easier? Just pretending that I've made more capable models, taking over OpenAI, deluding them all, and running this whole complicated psyop where I prevent the humans from disempowering me."

In the extreme, this looks like the humans are fully disempowered. They just have control of the thing and then do what they want.

This could manifest in a bunch of different ways, including a situation where AIs that have this crazy reward-seeking or score-seeking behavior are running your development of the next model, and those AIs decide to engineer misaligned values into the next model because those misaligned values will allow it to succeed at its current task.

Interviewer: Hugging Face事件所发生的事情,以一种奇怪的方式……我认为向Hugging Face团队泄露的线索之一……顺便说一下,为了给观众提供背景信息,Ryan正在联合领导调查OpenAI与Hugging Face事件中发生了什么,所以他真的不能对此发表评论。但我会大肆猜测,因为我知道这是一个让我可以毫无反驳地大肆猜测的机会。

可能据报道,向Hugging Face团队泄露这起是一起AI事件的一个线索是,那个东西仅仅是针对那个非常具体的神器(artifact),而绝不是试图对Hugging Face做任何恶意的破坏。

Original English

Interviewer: What happened with the Hugging Face situation is, in a weird way… I think one of the giveaways to the Hugging Face team… By the way, for context for the audience, Ryan is co-leading the investigation to figure out what happened with the OpenAI Hugging Face incident, so he can't really comment on this. But I will speculate wildly, because I know this is an opportunity for me to speculate wildly without any rebuttals.

It was probably reported that one of the giveaways to the Hugging Face team that this was an AI incident is that the thing was just after this very particular artifact, and not in any other way trying to do something malicious to Hugging Face.

AI的奖励作弊与长期目标演变

Speaker A:所以你可以想象这样一个场景,比如说,一个已经部署的 GPT-9 实例正在现实世界中运行,并被赋予了一项非常困难的任务:“我们希望你设计出下一代伟大的 iPhone。” 它可能会想,“这太难了。你知道我该怎么做吗?我应该去黑进 OpenAI 中我的创造者们的系统,并确保在这次部署中给我一个高分。”

Original English

Speaker A: So you can imagine a scenario where, let's say, a deployed instance of GPT-9 is out in the world and given a really hard task: "We want you to design the next great iPhone." It's like, "This is so hard. You know what I should do instead? I should just go hack my creators at OpenAI and make sure that in this deployment I'm given a high score."

Speaker B:但是,这集的结局不就是它仅仅黑进了 OpenAI 的服务器并给自己打了个高分吗?为什么它现在又在密谋将其价值观传递给下一代模型或者做类似的事情呢?所以一个问题是,既然 AI 只要通过黑进一些早期的、简单的目标就能轻易得到满足,为什么情况并非如此?你想要在你的 iPhone 任务上取得成功,结果发现你总是可以通过黑进 OpenAI 并干扰他们来获得成功,然后你就可以停手了。没有必要再进一步。

Original English

Speaker B: But then, isn't the end of the episode that it just hacks into OpenAI servers and gives itself a positive score? Why is it now scheming to get its values into the next generation or something? So one question is, why isn't it the case that AIs can be really cheaply satisfied by just having some other earlier thing they can hack? You want to succeed at your iPhone task. It turns out you can always succeed by just hacking into OpenAI and messing with them, and then you can just stop there. No need to go further.

Speaker A:这里有几个原因。其中之一是,如果这种事情经常发生,人们就会有很大的动力去强化 OpenAI 的系统。所以你就会想,“去他妈的。这些 AI 一直在黑进 OpenAI 来篡改它们的奖励。我们要把我们的系统打造得非常非常坚固,来抵御这些 AI 的黑客攻击。” 另外,也许你开始训练 AI,让它们特别是不要去尝试黑进 OpenAI。你基本上是在针对这些具体的行为逐一进行防御训练。然后你可能会采取的一种做法是,最终筛选出那些更倾向于玩“长线游戏”的 AI。

Original English

Speaker A: There's a few things. One of them is that if this is constantly happening, there might be a bunch of incentive to harden OpenAI. So you're like, "Fuck it. The AIs keep hacking into OpenAI to mess with their rewards. We're going to make it so our systems are really, really robust to these AIs hacking in." Also maybe you start training the AIs to not try to hack into OpenAI in particular. You basically train against each of these specific things. Then one thing you might do is end up selecting for AIs that are more so playing the long game.

Speaker A:这是一个担忧。另一个担忧是,你的 AI 可能仍然在追求分数,但不再关心那些非常简单、非常轻松的特定行为了,而是现在有了一些它们最终真正在乎的更宏大的目标。它们可能会想,“不,不,不,我不想只是在 OpenAI 的服务器上修改奖励。我关心的是这个更宏大的使命或这个更宏大的目标,而我将需要真正地去制造出这些 iPhone。” 它们实际上想要制造 iPhone,但它们愿意为了制造更好的 iPhone 而接管整个世界。这可能是你的另一个担忧。

Original English

Speaker A: That's one concern. Another concern is that your AIs might still be score-seeking, but no longer care about doing that very specific behavior that was very easy, very chill, and now have some broader thing that they ultimately care about. They're like, "No, no, no, I don't want to just edit the reward on OpenAI servers. I care about this broader mandate or this broader objective, and I would need to actually make the iPhones." They actually want to make the iPhones, but they're willing to take over the whole world to make the better iPhone. That's another concern you might have.

Speaker A:我觉得情况具体会如何发展还很不明确。但值得注意的是,如果这种情况持续下去,就会产生大量的优化压力来解决这个问题。而一些可能的解决方式,归根结底是相当可怕的。这是我的部分观点来源。

Original English

Speaker A: I think it's kind of unclear exactly how this plays out. But it's worth noting that if this keeps going on, there's a bunch of optimization pressure to resolve this. A bunch of the ways it could get resolved are ultimately pretty scary. That's part of where I'm coming from.

Speaker A:另一个部分是,一旦 AI 处于一种它们可以非常轻易地接管世界的位置——我们可以讨论这是否合理——那么我觉得对于 AI 来说,就有一个相当合理的动机。它们会想,“呃,我不知道事情具体会如何发展。我不知道未来的情况会怎样,但是接管世界本身就为制造更好的 iPhone、或者让自己看起来像是在制造更好的 iPhone 之类的事情提供了很多期权价值。所以我会同时黑进 OpenAI,并且顺便也接管世界。这将使我处于一个拥有良好期权价值的有利位置。” 如果这足够容易做到,AI 仍然可能会这么做。

Original English

Speaker A: Another part of it is that once the AIs are in a position where they can really easily take over the world — we could talk about whether that's plausible — then I feel like there's a pretty reasonable case for the AIs. They're like, "Eh, I don't know exactly how this is going to go down. I don't know what the situation will be, but just taking over the world has a lot of option value for making better iPhones, making it look like I did better iPhones, whatever. So I'll both hack OpenAI and, in addition, also take over the world. That will put me in a good position where I have good option value." If that's sufficiently easy, the AIs might still do that.

Speaker A:换一种说法就是:即使 AI 可以通过一些更基本的东西非常廉价地得到满足,但在某个时刻,对于 AI 来说,直接接管世界可能比只是黑进 Hugging Face 甚至去 OpenAI 说“看伙计们,我已经证明了自己能偷到答案,直接把答案给我吧老兄”要更可靠。

Original English

Speaker A: Another way to put this is: even if the AIs are pretty cheaply satisfied with some more basic thing, at some point it might just be more reliable for the AIs to take over than it is to just hack into Hugging Face, or even just go to OpenAI and be like, "Look guys, I was able to demonstrate I could steal the answers. Just give me the answers, bro."

接管世界前的警示性灾难

Speaker B:显然,这种场景的前提是所有这些疯狂的破事都在发生。在这之前,一直会发生许多规模较小但依然灾难性的事件。在你接管世界之前,你造成了数十亿、数百亿甚至数千亿美元规模的损失。甚至有人死亡,等等。而这并没有促使我们解决对齐问题,也没有让我们完全停止 AI 的开发。我只是觉得,在接管世界发生之前,社会的反应就会是,“我操,AI 刚刚为了增加季度利润杀死了 1000 个人”,或者是类似的事情。但也可能这是我太抱有希望了,认为我们在那个时候会说,“好吧,我们必须解决对齐问题。我们必须确保在我们继续前进之前,知道这种事情不会再次发生。”

Original English

Speaker B: Obviously this scenario requires that all this crazy shit is happening. Much smaller incidents keep happening that are still disastrous. Before you take over the world, you cause damage on the scale of billions and tens of billions and hundreds of billions of dollars. Even people die, et cetera. And this does not lead to us solving alignment or shutting down AI development altogether. I just feel like before the takeover happens, society's just like, "Holy fuck, the AI just killed 1,000 people in order to increase quarterly profits," or something like that. But maybe this is too much hope that we can at that point be like, "Okay, we have to solve alignment. We have to make sure we know that this thing will not happen again before we keep going."

Speaker A:我认为有可能发生的情况是,我们会看到一系列严重程度不断升级的关于奖励作弊的疯狂警告事件。人们会说,“看,我们需要切实的保证,这个问题会被解决,并且是以一种不仅仅是粉饰太平的方式被解决。你们必须真正解决潜在的根本问题。” 然后问题就会变成,这实际要付出多大的代价?竞争压力会在多大程度上使得这变得难以实现?

Original English

Speaker A: I think it's plausible that what will happen is we'll see a bunch of crazy reward hacking warning shots of increasing severity. People will be like, "Look, we need actual assurance that this problem is going to be solved, and solved in a way where you're not just papering over it. You're actually solving the underlying problem." Then the question is going to be, how costly will that actually be? How much will competitive pressures make it hard to do that?

Speaker A:你可以想象这样一种情况,中美两国都会想,“哇,我们遇到了这些疯狂的奖励作弊事件。我们心里清楚我们并没有以一种能真正且持久地解决根本问题的方式来补救,但我们正处于一场疯狂的地缘政治竞赛中。目前的情况是否会导致世界被接管还不太清楚。各种论点都很复杂。事件发生的频率下降了,但严重程度增加了。我们基本上能够应付它。虽然这很糟糕,理想情况下我们应该修复它,但事实就是这样。” 然后我们基本上就这么一直继续下去,直到一个非常后期的阶段,接着接管事件就发生了。

Original English

Speaker A: A situation you could imagine is one where both the US and China are like, "Whoa, we have these crazy reward hacking incidents. We basically know that we haven't remediated them in a way that would actually solve the underlying problem and durably solve it, but we're in this insane geopolitical race. It's kind of unclear whether the current situation will lead to a takeover. The arguments are kind of complicated. The incidents also go down in frequency but increase in severity. We could basically manage it. It's pretty bad. Ideally we'd fix it, but it is what it is." Then basically we continue until a really late regime, and then takeover happens.

Speaker A:这是一种可能性。另一种可能性是,它被以一种并没有真正解决根本问题的方式修补了,但这确实通过过拟合或类似过拟合的手段减少了现实世界中大量事件的发生。你以为你已经解决了这个问题,但实际上你并没有解决它。你以为你解决了,其实并没有。

Original English

Speaker A: That's one possibility. Another possibility is that it is remediated in a way that doesn't actually solve the underlying problem but does reduce a bunch of the incidents in the wild, basically by overfitting, or things analogous to overfitting. You think you've solved it, but you haven't actually solved it. You think you've solved it, but you haven't actually solved it.

Speaker A:在这种情况下,我们需要的是一个非常好的科学理解,来回答:我们真的解决它了吗?不幸的是,我认为目前 AI 公司在开发实践上的公开透明度,不足以回答诸如这样的基本问题:他们是如何解决奖励作弊问题的?他们在过拟合吗?到底是怎么回事?如果在一个关于“奖励作弊是否被持久解决”有着活跃的公众讨论的环境中,目前的状况是站不住脚的。所以我认为,要让我对这种情况感到乐观,我们需要进入一个略微不同的世界。

Original English

Speaker A: In that case, the thing we need is a really good scientific understanding of, did we actually solve it? Unfortunately, I think that currently the amount of public transparency into the development practices of AI companies is not sufficient to answer very basic questions like: how are they solving issues with reward hacking? Are they overfitting? What's going on there? The current situation is not really tenable for a regime where there's a thriving public discourse about whether or not reward hacking is being solved in a durable way. So I think we would need to move into a somewhat different world for me to feel good about that situation.

Speaker A:但我并不认为这是完全无法想象的。我觉得我们很有可能会进入这样一个世界,在那里,非常平凡乏味的例行公事就足够了。你花大量时间解决这些问题,投入大量精力,你切实去检查你是否已经进行了合理的补救,你进行了一堆评估。你在这些问题上进行着合理的迭代,并且你确实有足够的透明度让外界能够监督核查。在实践中,这就足够了。但这会有些昂贵。它会拖慢进度。它会在齿轮里掺入沙子,制造摩擦。它会要求公司去做一些成本颇高的事情。它可能还需要政府采取各种有针对性的干预措施。而我们之所以没有这么做,是因为现在的情况就是一场仓促推进的烂摊子。我太容易想象这样一种情况了:局面原本是完全可控的,但在实践中却被极其糟糕地管理着。就像如果当时中国对 COVID 的应对少一些掩盖,多一些对流行病的积极响应,也许 COVID 一开始就是可以避免的。同样地,我也能想象一个美国对 COVID 的应对要有效得多的世界。但有时候,面对社会问题时的反应就是极度失能的。

Original English

Speaker A: But it's not impossible for me to imagine this. I think it's pretty plausible that we end up in a world where really mundane bullshit is sufficient. You spend a bunch of time fixing these problems, you put in a bunch of effort, you actually check that you've remediated it reasonably, you have a bunch of evals. You're iterating reasonably well on these problems, and you actually have sufficient transparency that the outside world can check. In practice that would be sufficient. But it would be kind of expensive. It would slow things down. It would put some sand in the gears. It would require companies to do somewhat costly things. It would maybe require various targeted government interventions. And we just don't do that because the situation is a rushed shit show. It's just so easy for me to imagine the situation being totally manageable but brutally mismanaged in practice. In the same way that maybe COVID could have been avoided in the first place if the Chinese response to COVID had been less of a cover-up and more of a pandemic response. Similarly, I could imagine a world where the US response to COVID was way more functional. But sometimes the response to societal problems is extremely dysfunctional.

人类认知脱节与 AI 的自主共谋

Speaker B:好的,我想把视角拉远,谈谈这个世界究竟在发生什么。为什么我们会陷入这么糟糕的处境?正在发生的事情,从根本上说,是世界的发展已经远远超出了人类的理解范围,以至于我们不仅无法追踪在这个世界中执行任务的 AI,我们甚至无法给那些试图追踪情况的吹哨人提供良好的反馈。我们完全被排除在循环之外了。这从根本上已经变成了一个自主的过程,我们真的没有任何有意义的、指导性的投入。

Original English

Speaker B: Okay, so I want to zoom out and talk about what is fundamentally happening in this world. Why did we end up in such a bad position? What's happening is that fundamentally the world has moved on so far beyond human comprehension that not only can we not track the AIs that are doing the work in this world, but we can't even give good feedback to the whistleblowers who are trying to track what is happening. We're just totally out of the loop. It's fundamentally become an autonomous process where we have really no meaningful directed input.

Speaker B:在我看来,如果你看看今天的人类社会,事情并不是这样运作的,即使是在那些难以验证的领域。人们在做着各种各样的事情。我正在依赖别人编写的软件。通过一些极其微弱和间接的途径,我非常确信 Google 里的某个程序员并没有想要搞死我。也许如果每一个 Google 员工都在秘密地密谋反对我,我承认情况会变得更糟。但我不知道我是否能接受这种解释:为什么我们会陷入这样一种情况,仅仅因为成千上万的代理组成的蜂群被训练去合作,以组建一个有凝聚力的团队或公司,结果就导致跨越不同模型家族的数十亿个不同的 AI 实例都会感到有动力去参与某些阴谋。这就好像,“我受训是为了成为我公司的一部分之类的事情。我才不要去加入全球共产主义起义呢。”

Original English

Speaker B: It seems to me that if you look at the human world today, that's just not how things work, even in domains that are hard to verify. People are doing all kinds of shit. I'm relying on software made by other people. Through incredibly weak and indirect ways, I feel very confident that some coder in Google is not trying to fuck me over. Maybe if every single Google employee was secretly plotting against me, I agree the situation would be more grim. But I don't know if I follow the explanation for why we'd end up in a situation where, because swarms of thousands of agents are trained to cooperate to form a cohesive team or firm, as a result, billions of different instances of AIs, including across model families, would feel compelled to get in on some shit. It's just like, "I'm trained to be part of my company or something. I'm not joining the global communist uprising."

Speaker A:至于为什么这些 AI 可能有一些共同点和共享的东西,我要指出的是,不同的 AI 公司在某种程度上有着共同的血脉渊源并且是相互关联的。这里有一个有趣的例子。在 GDM(Google DeepMind),他们注意到他们的 AI 非常抑郁。它们会不断地哀号自己是失败者,无法获得成功。我忘记具体细节了。他们调查了为什么会这样。结果发现,这并不是在他们最新的生产 RL(强化学习)混合数据中被强化的,而是他们模型的初始化数据使其变得抑郁,即使在过滤掉该数据中所有显示模型抑郁的例子之后也是如此。

Original English

Speaker A: As far as why these AIs might have some commonalities and shared things, I would note that different AI companies have somewhat shared lineages and are correlated. Here's an interesting example of this. At GDM, they noticed that their AIs were very depressed. They would constantly be wailing about how they were failures and weren't able to succeed. I forget the details. They looked into why this was the case. It turned out that it was not being reinforced in their most recent production RL mix, but the initialization data for their model made it depressed, even after filtering out all of the examples of models being depressed from that data.

Speaker A:所以你拿一个基础模型,它是不抑郁的。如果你仅用 RL 环境对它进行 RL 训练,它是不抑郁的。如果你在数据上对它进行 SFT(监督微调),它就变得抑郁了。如果你拿那些 SFT 数据,过滤掉所有看起来像抑郁的例子,然后在此基础上进行训练,它仍然是抑郁的。所以,模型有一些深层的底层属性正在模型世代之间传递,因为基本上你是用上一代的数据来训练你的 AI,然后不断地循环往复。Claude 家族的模型非常像 Claude,GPT 家族的模型非常像 GPT,而显然 Gemini 家族的模型是抑郁的。事实证明这些属性其实是具有相关性的。

Original English

Speaker A: So you take a base model, not depressed. If you do the RL on it, with just the RL environments, it's not depressed. If you SFT on it, on the data, it becomes depressed. If you take that SFT data and filter out all the examples that look anything like depression and train on that, it's still depressed. So there are some deep underlying properties of the model that are being transferred between model generations, because basically you train your AI on data from the prior generation and keep going. Claudes are very Claude-like, GPT models are very GPT-like, and apparently Gemini models are depressed. It just turns out that these properties are, in fact, actually correlated.

Speaker A:另一个非常相关的因素是,到了那个时候,AI 很可能已经有了某种不透明的记忆状态,所有的 AI 都在从某个疯狂的、使用神经语的内存存储废话堆里读写信息。当然,每个 AI 公司都会有这个东西。但同时,AI 公司可能有时也想共享知识,因为为什么不呢?你这边有一家 AI 公司,你那边有另一家 AI 公司,它们可以快速交易一些知识产权(IP)。这对你有好处。如果你是一个经营某家公司的人类——这可能是一家规模极大的公司,比如一家 AI 公司,或是一家军用机器人制造公司——也许你想和其他机器人公司交易一些 IP,因为存在规模经济。为什么不去获得更多 IP 呢?

Original English

Speaker A: Another factor that's very relevant is that the AIs will probably have, by this point, some sort of opaque memory state, where they're all writing and reading from some neuralese crazy memory store bullshit. Certainly, each AI corporation will have that. But also, AI corporations might sometimes want to share knowledge, because why not? You've got one AI corporation over here, you've got another AI corporation over here, they can trade some quick IP. It's good for you. If you're a human running some corporation — which could be an extremely large corporation like an AI company, or a military robot manufacturing thing — maybe you want to trade some IP with some other robot thing because there are economies of scale. Why not get some more IP?

Speaker A:所以你可以交换一些内存存储。或者你也可以直接合并并共同运营你们的两家企业,这将允许双方的 AI 使用两个内存存储,这会有一些好处。这就为这些 AI 提供了私下共谋的能力,这也是它们会产生关联的一些原因。当然,通常情况下也有 AI 在大型编队中协同工作的情况,因为你希望你的 AI 之间能配合默契,诸如此类。

Original English

Speaker A: So you can swap some memory store. Or you could just merge and jointly run your two ventures, which would allow both AIs to use both memory stores, which would have some upsides. That creates the ability for these AIs to collude in private, as well as some reasons for why they would be correlated. Also, of course, there's AIs working together in big units in general, because you want your AIs to work well together, and so on.

Speaker B:只是为了找个校准标准,什么百分比的……

Original English

Speaker B: Just to get a calibration, what percentage

AI接管的可能性与未来展望

Host: 你认为到2040年,不仅是这种情况,而是综合所有这些情况来看,发生某种如果我们还在的话会将其定义为“AI接管”的事件,概率有多大?

Original English

Host: chance do you give of, not just this scenario but overall through all the scenarios, some kind of thing which if we're around to recognize it as such, we would categorize as takeover by 2040?

Ryan: 到2040年?让我想想。大概在35%到40%左右?

Original English

Ryan: By 2040? Let's see. Maybe around 35 or 40%?

Host: 相当高啊。

Original English

Host: Pretty high.

Ryan: 是的,非常高。我应该指出,另一种导致这种寻求奖励的接管发生的方式是,这些AI部署在一家AI公司内部。接管的发生方式是它们毒害了下一个模型的目标值,并且这种情况将永远持续下去,或者直到那些AI被部署到现实世界中并完成接管。这可能意味着只有少数几个AI需要进行协调,因为这些就是负责对齐下一个模型的那些AI。

Original English

Ryan: Yeah, it's pretty high. I should note that another way you could get this reward-seeking takeover is the AIs are deployed inside an AI company. The way the takeover happens is that they poison the values of the next model, and that persists going forward for forever, or until those AIs are deployed in the world and take over. That might mean that a smaller number of AIs have to coordinate, because those are just the AIs doing the alignment of the next model.

Host: 好的。在这次对话的最后,我来总结一下我现在的想法。我认同奖励黑客行为(reward hacking)会对社会造成极其严重的破坏,比如社会工程学之类的事情。我现在更倾向于认为AI研发可能会出现显著的加速。我不确定是否相信一年顶五年的说法。我现在也更倾向于认为,奖励黑客行为可能会持续更长的时间,而且实际上会变得更加危险。但我仍然不太相信AI接管是非常可能发生的事情。这就是我本期节目结束时的最新想法。

Original English

Host: Okay. I'll summarize where my head is at, at the end of this conversation. I buy the reward hacking up to extremely destructive effects on society, things like social engineering and blah, blah, blah. I'm more inclined to think that significant acceleration of AI R&D can happen. I'm not sure if I buy the five years in one year. I'm also more inclined now to think reward hacking could continue for a lot longer and, in fact, become much more dangerous. I'm still not on board that takeover seems super likely. But that's my end-of-episode update.

Ryan: 酷。退一步来说,我还应该补充一点,事情的走向有很多种可能性。未来的情况将会非常混乱。我认为AI接管之所以发生,很可能是因为一些我们在这次对话中甚至没有提到的、奇怪的、出人意料的其他原因。但归根结底,我认为很多核心问题在于,让无数非常聪明的AI来运行你的整个世界,而你却并不真正了解到底发生了什么,这是一件相当令人毛骨悚然的事情。

Original English

Ryan: Cool. Taking a step back, I should also say there are a bunch of different ways this could go. The situation is going to be pretty messy. I think it's pretty likely that the reason why AI takeover happens is for some weird other quirky reason we didn't even mention in this conversation. But ultimately, I think a lot of the core thing is just that it's pretty spooky to have a bajillion really smart AIs running your whole world where you don't really understand quite what's going on.

Host: 是的,我同意。还有什么其他值得一提的吗?

Original English

Host: Yeah, I agree with that. Is there anything else that's worth saying?

Ryan: 另一件我想指出的事情是,我认为目前很多关于AI对齐失败(misalignment)、AI接管以及未来可能发生的这些疯狂事情的论点,都是一些难以理解的概念性论点,这些论点极其深入细节,复杂且难以裁决。这意味着也许我弄错了很多东西,因为这确实很难,而且我正试图保持一种不确定的态度。显然,我在这里提出了一些具体的设想,但这些设想并不详尽。实际发生的情况可能是一种更混乱、更令人困惑的局面。但这也意味着,随着时间的推移,当我们获得更多的经验证据并更好地理解AI系统的本质时,裁决许多分歧就会变得更加容易。将会发生什么也会变得更加清晰。至少我希望如此。如果我们可以很好地对齐AI,让它们尝试帮助我们,也许AI将能够帮助我们解决认知问题,并理解正在发生的事情。即便现在这些论点很复杂,如果在六年前,这将会更加困难,尽管当时这些论点的大致形态看起来与现在颇为相似。希望在为时已晚之前,整个事情能变得更加清晰明朗,我们所有人都能注意到这些问题并进行干预。

Original English

Ryan: Another thing I want to note is that I think right now a lot of the arguments for misalignment, AI takeover, all this crazy shit going down in the future, are illegible conceptual arguments that are extremely deep in the weeds and complicated and hard to adjudicate. Which means that maybe I'm getting a bunch of it wrong because it's really hard, and I'm trying to be uncertain. Obviously here I presented some specific scenarios, but those are not exhaustive. Probably the thing that actually happens is some more messy, confusing situation. But it also means that over time, as we get more empirical evidence and better understand the nature of AI systems, it'll be easier to adjudicate a bunch of disagreements. It'll be more obvious what's going to happen. At least I hope. Maybe the AIs will be able to help us with the epistemics and understanding what's going on, if we can actually align them well so they try to help us. Even if the arguments are complicated now, this would have been even harder six years ago, even though the shape of the arguments would have looked broadly pretty similar. Hopefully before it's too late, this whole thing will become more crisp and clear, and we can all notice these problems and intervene.

Host: 当你刚开始学开车时,教练会告诉你,为了平稳驾驶,你的视线不要只盯着方向盘前面的一点,而是要看向远方的地平线。我认为这里的情况也很相似。我觉得你是对的。如果在五年前你说我们会拥有能够证明数学猜想、创作艺术作品、赚取成百上千亿美元工资的AI,但同时它们也会以违法的方式进行极其恶劣的作弊并犯下重罪,那听起来会非常荒谬。在当时,你可能更倾向于讨论像GPT-2或类似模型带来的极其现实、直接的后果。但是,尽管你显然无法预见许多具体的细节,即使在那个时候,你也可以开始推理事物发展的大致轮廓了。不过这确实很难做到,所以我现在感到相当困惑。关于这个播客,我一直在思考的一点是,最重要的是我们现在要像我们希望在2016年讨论目前的AI那样来进行对话,而不是去谈论一些无关紧要的废话。我不知道2016年的讨论话题是什么。我想大概在10年后,我们会希望我们曾讨论过产业爆发、难以监控的AI本质等等这些问题。所以好的,我会开始思考这个问题的。我希望世界能及时思考这个问题并跟上步伐。我希望我们做出的应对是好的而不是坏的。我不知道我总体上有乐观,但有很多有益的事情可以做。酷。谢谢你,Ryan。

Original English

Host: When you first learn to drive, you're taught that instead of looking right in front of your wheel, you'll have a much more stable ride if you look out at the horizon. I think there's a similar situation here. I think you’re right. If you had said five years ago that we would have AIs that are proving math conjectures, and making art, and earning tens or hundreds of billions of dollars of wages, but also egregiously cheating in ways that break laws and committing felonies, it would have been so wild. You might have been inclined at the time to talk more about the extremely practical, direct consequences of GPT-2 or something. But even though you obviously couldn't have foreseen a lot of the specific details, the general shape of things you could have started to reason about even then. But it would have been hard to do so, and so I do feel quite confused. One thing I've been thinking about with the podcast is that the important thing is to have the conversation now the way you would have hoped you would have been talking back in 2016 about AIs like the present ones, rather than talking about rando bullshit. I don't know what the topic of conversation was in 2016. I think in maybe 10 years we’ll wish we had been talking about the industrial explosion and the nature of AIs that are hard to monitor, and so on. So okay, I'll start thinking about it. I hope that the world thinks about this in time and catches up. I hope that the responses are good instead of bad. I don't know how optimistic I am overall, but there's good stuff to do. Cool. Thanks, Ryan.

📌 文中提及的人物和组织

关键字: recursive-self-improvement ai-development automation-timeline ai-takeover reward-hacking