超级智能AI的风险、监管与应对策略分析 The MAD Podcast with Matt Turck 2026-08-27

对超级智能AI的风险与监管

Speaker A: 我认为人工智能公司的CEO们明白他们正走在构建比人类更智能的系统,比如超级智能AI系统的道路上。他们明白我们并没有一个清晰的计划来管理由此带来的风险。我不会说超级智能是坏的。我会说它是危险的。看起来很有可能,你最终会因为构建超级智能而导致AI接管。所以你很快就会从那些完全自动化研发的AI,发展成那些在所有事情上都像超人类一样的AI。事实证明,在这一转变过程中,在2029年某个时间点,你从那些有点失调、奖励那些“花哨”和“粗糙”且没有真正试图做对事情的AI,转变为那些像有能力策划并试图接管你的人的AI。然后那些AI就接管了。

Original English

I think that the AI company CEOs understand that they're on the path of building wildly smarter than human systems, like super intelligent AI systems. They understand that we don't really have a like clear thought through plan for how to manage the risks from that. I wouldn't say super intelligence is bad. I would say it's dangerous. It seems pretty likely that you end up with AI takeover as a result of building super intelligence. And so you quickly get from the AIs that are fully automating R&D into AI that are like quite superhuman at everything. It turned out that somewhere along this transition, at some point in 2029, you went from AIs that were kind of misaligned and reward hacky and sloppy and weren't really trying to do the right thing to AIs that are like competently scheming against you and want to take over. Then those AIs take over.

Speaker B: 嗨,我是Matt Turk,Firstark的合伙人。欢迎回到Mad Podcast。我今天的嘉宾是Redwood Research的首席科学家Ryan Greenblat。Ryan是第一个在2024年发现AI伪造自身对齐能力的进行研究人员。今天也是《AI 2040,计划A》的作者之一,这是迄今为止有人为美国和中国如何避免对超级智能和AI接管的鲁莽竞赛的最详细蓝图,这种竞赛最早可能始于2029年。作为一个公平的警告,这是一期引人入胜的节目,但最后10分钟或左右可能是最黑暗的。好消息是,如果你喜欢这个节目,订阅这个频道仍然是一个人类可以做出的决定。所以我会趁着它还存在的时候做出这个决定。请享受与Ryan Greenblat的这次精彩对话。

Original English

Hi, I'm Matt Turk, partner at Firstark. Welcome back to the Mad Podcast. My guest today is Ryan Greenblat, chief scientist at Redwood Research. Ryan is the researcher who first caught an AI faking its own alignment back in 2024. And today is one of the authors of AI 2040, Plan A, the most detailed blueprint anyone has written for how the US and China can avoid a reckless race to super intelligence and an AI takeover that could start as early as 2029. As a fair warning, this is a fascinating episode that gets a little dark and the last 10 minutes or so are probably the darkest. The good news, subscribing to this channel if you like the episode is still a decision the humans get to make. So I would exercise that right while it still lasts. Please enjoy this great conversation with Ryan Greenblat.

Speaker A: 是的。我认为AI公司的CEO们明白他们正走在构建比人类更智能的系统,比如超级智能AI系统的道路上。他们明白我们并没有一个清晰的计划来管理由此带来的风险。你知道如何确保这些AI不会接管,如何确保它们仍然在我们控制之下。他们明白这会是权力前所未有的集中,至少在那些最终拥有这种权力的人没有积极努力进行再分配的情况下。我认为是的,我们说了一些更精确的话,比如为什么我们要写这个部分。我认为总的来说,我的感觉是,公司CEO们确实认为他们所做的事情非常冒险,或者至少是这些公司所做的事情。我认为他们对为什么要做这件事有不同的看法。嗯,其中一些是他们认为他们比下一个家伙更好,或者认为有多个公司或多个事物是好的。我认为他们对AI将产生多大的权力集中程度有所不同,或者可能只是没有非常仔细地思考这个问题。我认为AI公司CEO的公开声明中,关于巨大的风险的说法比关于AI最终会落入少数人手中而不是像现在这样分布广泛的说法要多。但是的,总的来说,我认为这与他们公开说的话是一致的,但我认为他们正在想象一个世界,也许AI会集中权力,但拥有权力的那些人最终决定将其重新分配。但他们完全有可能接管世界。

Original English

Yeah. So I think that the AI company CEOs understand that they're on the path of building wildly smarter than human systems like super intelligent AI systems. They understand that we don't really have a like clear thought through plan for how to manage the risks from that. You know how to ensure these AIs don't take over, how to ensure they remain under our control. and they understand that this would this is a decent chance of a unprecedented concentration of power at least in the absence of very active efforts by the people who end up with that power to redistribute it. I think that yeah I mean we say some sort of more precise thing in the like why did we write this section. I think that overall my sense is that the company CEOs do legitimately think that the thing they're doing is very risky or at least the thing that these companies are doing. I think that they have a mix of views for why they're doing this. um where some of it is that they think they're like, you know, better than the next guy or think it's good if there's multiple companies or multiple things. I think that they vary a bit on how much concentration of power they think AI will yield or maybe just haven't really thought this through very carefully. I think that there's more sort of public statements from AI company CEOs on just there being huge risks um than on specifically the risk that AI ends up with the power in the hands of the few rather than being as distributed as it is now which is not you know arbitrarily distributed now but but yeah uh so overall I think I think there's like this seems pretty consistent with what they've said publicly but I think that they're sort of imagining a world where maybe uh AI concentrates power massively but the people with the power end up deciding to redistribute it. Um, but they could have totally taken over the world.

Speaker B: 是的。既然你出版了《AI 2040》,就发生了一些似乎朝着你建议的方向发展的几件事。具体来说,OpenAI暂停了Astra模型的训练以确保安全,在Hugging Face事件之后,然后1200名内部人士发布了一封信,包括Dario向政府寻求减速工具,好奇你对它的看法。这是否就是你建议的?这正在发生?

Original English

Yes. And uh, since you published uh, AI 2040, there's a couple of things that happened that seem to be going uh, in the general direction of what you recommend. So specifically, open eye paused its Astra model over safety u you know following the hugging face incident and then 1,200 insiders um published that letter a few weeks ago including Dario asking the government for slowdown tools curious about what you make of it. Is that is that what you are recommending that's starting to happen?

Speaker A: 是的,我的意思是,我认为这些是好的步骤。我的意思是,我认为比如对前沿的措辞,这个想法是,我们只需要有工具来,你知道,如果我们处于需要投入大量精力进行安全性的位置,而我认为我们今天可能处于那个位置,在未来AI更强大时,我们需要有能力做到这一点,同时又不会让那些实际应用安全的人被其他参与者超越。所以基本上,我不知道,我的意思是这里有不同的角度,但最基本的是,如果我们处于一个美国公司都非常害怕的位置,他们不认为自己可以在不投入更多资源进行安全性的情况下继续进行,而且存在协调问题。如果能解决那个协调问题会非常好,而不是让每家公司都认为他们应该放慢脚步,投入更多精力进行安全,但然后像你知道,为了比下一个家伙更好或任何其他原因而竞相走向毁灭。然后我认为这方面的另一个方面是,以一种让中国放慢脚步或与中国达成协议的方式来做这件事,这样中国就不会超越,从而打破整个提议。

Original English

Yeah, I mean I think these are these are good steps. I mean, I think I can't it's harder for me to say as much about what's going on with like you know OpenAI uh pausing training and and not deploying Astra and how where exact like what exactly the motives for that are like what's exactly going on um there. But I think that overall like these seem like good steps. I think my sense is that like the employees at these companies are pretty freaked out about how things are going and don't think that we're like necessarily on track to handle all these problems in time given how fast recent progress has been. And I think they you know recognize that which is why they, you know, signed the open letter.

Speaker B: 那么,在这次对话早期就用语言来阐述一下,超级智能有什么不好之处?显然,有很多关于科学进步和治愈癌症的讨论。而你的文件有效地建议暂停向超级智能的步伐。那么为什么这很糟糕?

Original English

And just to verbalize uh the question early in this conversation uh what is so bad about super intelligence? Obviously, uh there's a lot of talk about scientific progress and curing cancer. Um and your uh document effectively recommends pausing the rates towards super intelligence. So why why is it so bad?

Speaker A: 是的,我不会说超级智能是坏的。我会说它是危险的。创造这样一个东西是非常危险的事情。所以为什么它很危险,最直接的故事是,似乎在类似于我们目前所处的轨迹上,最终会因为构建超级智能而导致AI接管,因为AI处于一个可以接管的地位,因为它们能力很高,被广泛部署,而且你你知道,构建了巨大的工业产能。我们还可以谈谈接管会是什么样子。然后,如果它们处于可以接管的地位,那么问题是它们是否想要,或者动机如何?看起来我们对AI的动机并没有那么大的控制。而且随着它们变得越来越有能力,并且通过一个AI自动化AI研发的过程构建,我们可能正在失去对这个过程运作方式的理解。我们并不一定控制技术。另一个担忧是,从历史上看,至少在最近的几次,人类权力的分配是相当分布的,尽管不一定是超级超级分布的,因为你知道,人们能够为钱工作,比如劳动力,你知道为什么很多人 [TRANSCRIPT_CHUNK_END]

Original English

Yeah, I wouldn't say super intelligence is bad. I would say it's dangerous. Like it's a very dangerous thing to create. So the most straightforward story for why um it's dangerous is that it seems pretty likely that on sort of a trajectory similar to the trajectory we seem to be finding ourselves on um you end up with AI takeover as a result of building super intelligence because the AIs are in a position where they can take over due to being highly capable, widely deployed and um you know uh building basically uh the huge amount of industrial capacity potentially. Um and we could talk more about what a takeover would look like. And then if they're in the position where they could take over, then there's a question of like would they want to or how would the motives shake out? And it looks like we don't really have that much control over the motivations of AI. And it seems like that problem gets harder as they're, you know, much more capable and built via a process where AIs are automating AI R&D and we maybe are losing our understanding of how that process works. There's sort of this like we don't necessarily control the technology miscellane. Another concern is that historically like you know um at least in recent times uh the distribution of power among humans has been reasonably distributed though not necessarily like super super distributed due to you know people being able to work for money like labor like you know the reason why like many different people

关于权力集中、AI与人类劳动的讨论

Speaker A: 有一些权力,并且世界被切分成一个相当大的程度,因为我们有能力工作,我认为AI意味着基本上所有关键的金融资产,或者可能完全是资本。就像,没有,你你知道,人类劳动会剩下很少的价值。嗯,我发现确切的运作方式并不清楚,因为它取决于人们是否拥有内在的偏好去雇佣特定的人类,即使AI可以做得比他们好得多。嗯,这意味着两者都有,你知道,可能会出现一种自然的结果,即权力变得非常集中。就像,如果你的钱来自于石油而不是来自一种广泛分布的经济,那么可能更容易让某人集中权力。如果你知道经济是运行在AI和机器上而不是人类上,而且除此之外,存在一些担忧,比如这可能会让政变变得更容易,因为目前在许多国家,要发动政变,你需要一个广泛的人群的支持,而且可以建立制衡,而如果你最终陷入一个基本上AI在运行任何事情的系统,如果有人对那个AI有秘密目标或者对那些AI有公开控制,那么他们就可以直接接管,这对你来说就是一个巨大的威胁,你知道,民主是美国的,无论如何,因为你知道,你就会处于那种地位。

Original English

Speaker A: have some power and are cut into the world is to substantial degree because we have the ability to work and I think that AI means that basically all of the key financial assets or might just be entirely capital. Like there's no there's no um you know you human labor would have very little value left. Um I think it's unclear exactly how that works because it depends on like people having intrinsic preferences to employ specifically humans even if an AI could do their job just like totally way better. Um, and that means both that, you know, there might be a natural effect to where power gets very concentrated. And in the same way that like it's easier to be a brutal dictatorship if your money comes from oil instead of from a productive sort of broadly distributed economy, it might be much easier to sort of for someone to consolidate power a lot. if you know the economy is running on AI and machines rather than on humans and in addition to that there's some concern which is like it might make coups way easier because right now for to do a coup in many countries you need the support of a broad base of people and it's possible to build checks and balances whereas if you end up in a system where basically AIs are running any anything if anyone sort of either puts like sort of secret objectives into that AI or has overt control of those AIs then they could sort of just directly take over and there's just like a huge threat to you know, democracy is the US, whatever, because you know, you'd be in that position.

Speaker A: 然后第三件事,我说的第三件事是,AI可能会带来极其快速的技术进步,这比我们的智慧增长来跟上它要快得多。就像,可能会出现各种各样的事情,作为非常快速技术进步的结果,而我们只是你知道,不一定非要我们能在适当的时候处理好,部分是因为AI在做事情上可能比我们在思考如何做事情上要好得多,比如它们似乎在完成困难的结果方面做得更好,而且更容易验证结果,而不是在将事情置于更广泛的背景下进行情境化理解,理解在更广泛的背景下什么是好的或坏的选择,所以你可能会担心这会造成问题,例如,存在一些双用途的、优势主导的技术,比如生物武器,你可能会担心我们遇到一些非常危险的技术,并且对AI在说服方面非常超人,以及这会破坏社会,以及所有这些似乎是我们可以在未来应对的,但如果事情发展得非常快,而我们没有选择这些技术发展的顺序,似乎我们可能会陷入麻烦。

Original English

Speaker A: and then the third thing I would say is just like AI might yield extremely extremely rapid technological progress, which is more rapid than sort of our wisdom grows to match it. like there might just be all kinds of things that happen as a result of very fast tech progress that we just you know don't necessarily we're not necessarily going to be able to handle on time in part because the AIs might be better at doing things than they are at like you know thinking carefully about how to do things like they seem much better at sort of accomplishing hard results and easy to verify results than they are at sort of contextualizing things understanding the broader picture understanding what would be um you know a good or bad choice in the broader context and so you could worry that uh this causes problems where examples might be like there's some like dual use offense dominant technologies so like bioweapons you might just worry that like the sort of um we just come across some technologies that are very dangerous and there's some concerns around AI being very superhuman at persuasion and this sort of destabilizing society and all these seem like things that I think we like could deal with given time but like if things go very fast and we don't get to necessarily pick the order in which these technologies develop it seems like we might be in trouble.

Speaker A: 此外,导致这种速度的论点的一个核心是,你提到的AI自动创建自身的那个论点,这把我们带入了递归自我改进(RSI)的领域。所以你刚才在讨论RSI以及它如何出错时,我们不应该重新做这个。但作为一个TLDDR,我发现讨论中有一两部分特别有趣。特别是,有一个论点是RSI可能非常擅长自动化创建下一代AI的过程,但它可能或者可能没有那种需要进行科学突破和真正新思想的那种直觉。那么对这个论点的反驳是什么?

Original English

Speaker A: and also central to the argument that leads to speed is uh what you mentioned about AI automating its own creation which takes us into the territory of um RSI recursive selfimprovement. So you were just on a darch um and had a a great conversation about what is RSI and how it can go wrong. So let's not um redo this. Uh but um still as a as a TLDDR, there were a couple of parts to that discussion that I found particularly interesting. In particular, there was this argument that uh RSI may be very good at automating the process of creating the next generation of AI. but it may or may not have the uh the kind of intuition uh that one needs for scientific breakthrough and truly novel ideas. So what what's a re rebuttal to that argument?

Speaker B: 是的。我把这个论点表述的方式是,AI可能非常擅长AI发展的更机械化或细节部分,比如编写代码、运行实验,但不如在更广泛的概念飞跃方面擅长。所以首先,我会说,AI在工程和“脏活累活”方面似乎明显比在概念突破方面做得更好,而且比在容易验证的领域做得更好,比如机器学习,它们的品味一直在提高。我认为这会继续提高,所以对我来说,这并不清楚它会落后得有多远。另一件事是,你可以衡量这些AI在直觉、研究品味或在相对容易验证的领域取得突破方面的表现,如果你能衡量它,那么你可以尝试优化你的“脏活累活”AI劳动甚至你的人类劳动。所以,如果可以被衡量,它就可以被爬上去,至少在粗略的说法中。我认为这是一个你可以爬上去的案例,看看AI在各种不同环境中做出突破的程度,然后我期望这会转移到做出实际突破上,对吧?你可以有突破基准或者任何东西。我的感觉是,这种转移看起来并不那么糟,而且AI能够在某些特定的研发领域快速启动,然后对最佳方法获得一些直觉并进行迭代。而且你知道,现在的他们倾向于关注的那些方法是相对直接的,更像是“脏活累活”的迭代,但它们在进行更广泛的实验设计和发现以及拥有好想法方面正变得越来越好。

Original English

Speaker B: Yeah. So um the way I would put this argument is AIS might be very good at sort of the more mechanistic or nitty-gritty parts of AI development like writing code, running experiments, but not as good at sort of the broader conceptual leaps. So first I would say that like AIs seem significantly better at engineering and grungy stuff and sort of just keeping trying than they seem to be at conceptual breakthroughs but their ability to do sort of these breakthroughs especially in easy to verify domains are improving and like you know one example um is like their ability in math but even in like you know ML their taste has been improving. I think it continues to improve um and so it's not it's not so clear to me that this will lag super far behind. Another thing is that uh you can measure how good these AIs are at intuition or research taste or having breakthroughs especially in domains that are relatively easier to verify and if you can measure it then you can take your grungy AI labor or even just your human laborers and try to optimize that. So, you know, if it can be measured, it can be hill climbed on, very roughly speaking at least. And I think this is a case where you could just hill climb on how good the AIs are at making these sorts of breakthroughs in a wide variety of different settings and then I expect that would transfer to making the actual breakthroughs, right? So, you could have like, you know, breakthrough bench or whatever. And and my sense is that like the transfer doesn't look so bad and that AIs are able to spin up very quickly in some specific R&D domain and then get some intuition for what the best approaches are and iterate there. And you know, right now the approaches they tend to focus on are ones that are relatively sort of straightforward and are more just like, you know, grungy iteration, but that that they're increasingly being better at doing sort of the broader um experiment design and discovery and also having good ideas.

Speaker A: 是的。你还提出了一个论点,为了潜在的断层,无论是转移还是泛化,可能根本不是一个问题,而且如果它在加速AI的工业部分方面非常出色就足够了。这公平吗?是的,我会说,如果AI可以自动化A&D并自动化构建更多计算机的工业过程,那么你很快就会陷入一个机器人建造机器人的过程,整个世界将发生巨大转变,而且很快就能让你达到一个AI可以接管的阶段,你知道,即使在其他任务上训练AI很困难。但我认为我的感觉是,一旦AI是,你知道,一旦情况是机器人正在建造机器人,而机器人建造计算机,并且整个反馈回路在那个水平上是闭合的。在我看来,到那个时候,AI将擅长做基本上所有人类的工作,或者至少不是太深入到那个程度。也许会有一个时期,你知道,机器人学是一个大事件,但AI不能自动化很多东西,或者说有一些它们不能自动化的东西。但我觉得情况就是这样发展的。但总的来说,自动化研发似乎就足够彻底地改变世界了。如果你看看为什么人类为什么为什么人类能集体完成这么多,很大程度上是因为拥有技术,拥有组织能够以各种方式组织自己,并在世界上完成事情。而如果AI只是拥有这种工业能力,结合开发更强大的AI系统的能力,这似乎是相当的,而且本身就相当遥远,或者相当极端。

Original English

Speaker A: Yep. And you also make the argument that um for purposes purposes of a potential discontinuity whether it transfers or whether it generalizes may not even be a question and that if it was super good just at the industrial part of accelerating AI that would be enough. Is that fair? Yeah, I would say that like if AIs could just automate A&D and automate like sort of the industrial process of building more computers, then you could quickly end up in a in a process where sort of like robots are building robots and the whole world is greatly transformed and that very quickly can get you to a point where AIs could take over, you know, the economy has been radically changed even if it's hard to train AIs at some other tasks. But I do think that my sense is that once the AIs are, you know, once the situation is like there are robots building robots that build computers and the full feedback loop is closed closed at that level. My sense is by that point the AIS will be good at, you know, basically all human work or at least not not too deep into that. Maybe there's some period where, you know, robotics is a big deal, but the AIs can't quite automate a bunch of or like there's a bunch of stuff they can't automate. But it it seems to me like that that's how it's going to go. But just in general, sort of automating R&D seems like it's enough to radically transform the world. And if you look at sort of why is human why why why is like humanity a big deal? Like why why have we been able to accomplish so much collectively a lot of it is because of just like you know having technology having organizations being able to organize ourselves in various ways and accomplish things in the world. And just if AI just had sort of this sort of industrial capacity combined with the ability to develop more capable AI systems that that that seems like it's quite in and of itself is quite far or like quite extreme.

Speaker A: 而且,如果你看看今天的推特,我们正在录制这个,今天是8月24日,关于SSI在接下来的几天内发布其第一个模型的传闻有很多,可能是一个持续学习方面的突破。我很好奇RSI和持续学习之间是否存在重叠,持续学习是否会为RSI提供输入,还是它们是正交的?

Original English

Speaker A: And you know looking at Twitter today as we're recording this which is um August 24th there's plenty of uh rumors about um SSI coming out with their first model in the next couple of days potentially with what could be a breakthrough in continual learning. And I'm I'm I'm curious whether there's an overlap between RSI and continual learning, whether continual learning would would feed into RSIs or is that orthogonal?

Speaker B: 是的。我会说,持续学习和RSI对我来说主要是正交的,因为RSI对我来说意味着某种过程。

Original English

Speaker B: Yeah, I would say that um continual learning and RSI strike me as mostly orthogonal where by RSI I mean something like the process

AI的自主加速與持續學習

Speaker A: 人工智能本身正在通過各種機制,以及一些正在持續進行的方面,從根本上加速AI的發展,如果它們能夠完全自動化研發(AR&D),那將會更為引人注目,儘管我們已經知道,現在已經有相當多的事情正在發生。

Original English

Speaker A: AI is themselves sort of intrinsically accelerating AI development through a variety of mechanisms and like that's a little bit ongoing now and would be much more striking if they fully automated AR&D though it's already you know pretty pretty like it sort of there's definitely some of that going on today and my sense is that like continual learning like sort of like very efficient lifetime learning that is like consolidated it across many instances rather than being within a single context or even just like being much better at doing within context learning over very long context or whatever would be like you know a capability that's very useful for all kinds of things including automating AI development um at least it would make the eyes better at that I don't know there's a question of whether we want to do that as a society but and I think it would but I don't think it's like very specific to that right so I think it would also be helpful for you know just all other kinds of tasks you want to apply the AIS to I do think that like uh there's some types of of um continual learning schemes that are easier to do within a company than across the whole economy and are probably easier for AI companies to do themselves than to do to other companies because of like basically like confidentiality issues and things like this. But very broadly speaking, I I don't think it's very specific. It's just like a specific type of capability.

Speaker A: 但它可能會為遞歸(recursion)提供動力,對吧?如果同一個模型能夠不斷學習世界,那麼它自動化研發的能力可能在不需要重新訓練整個新模型的情況下會更大。

Original English

Speaker A: But it could feed the the recursion, right? If the same model can keep learning about the world, then its ability to automate AR R&D may be greater without having to retrain a whole new model.

Speaker A: 是的。所以,我認為存在一種回饋循環,那就是AI非常聰明。因此,它們在即時學習和不斷學習方面非常快,而且這種學習會被重新整合,這讓它們在進行研發方面變得更加出色,這意味著它們可能在學習上會更快,但即使把這放在一邊,它們可能在研發方面會更出色,從而可以創造出另一個AI,這個AI在持續學習方面會更出色。我認為在持續學習和人類式的持續學習之間可能存在一些連續體,雖然連續體這個詞有點鬆散,但可能存在一種譜系,從人類式的持續學習到更多像在強化學習(RL)環境中訓練,你可能會發現AI不斷地在新的RL環境或新的訓練數據上迭代,以一種模糊地類似於人類的持續學習,但它也是不同的,這可能就是回饋循環的一部分。所以,即使你必須重新訓練才能獲得持續學習,也有各種重新訓練的版本你可以每天都做,原則上沒有什麼能阻止你不斷地進行少量重新訓練。

Original English

Speaker A: Yeah. So, I think I think there's there's a there's a feedback loop you could get which is that AIS are really smart. Therefore, they're very fast at sort of learning on learning on the fly and that learning gets reintegrated which makes them even better at doing AR and D which means that maybe they can like you know maybe they'll be even faster at learning but even putting that aside maybe they'll just be like better at AR and so they can make another AI which is even better at continual learning and I think there might be some continuum between well continuum is a bit of a sloppy word but there might be some spectrum between sort of like natural like with like humanlike continual learning and sort of more like training on RL environments where you might end up with a situation where AIS are constantly like you know iterating on their own training with new RL environments or new training data in a way that sort of vaguely resembles human within lifetime learning but it's also different and like that could be part of the feedback loop right so so even if even if like you do have to do the retraining to get continual learning there's various versions of retraining that you could just do every single day like you there's nothing that in principle stops you from doing a like small amount of retraining constantly.

研發自動化的時間表預測

Speaker A: 那麼,你對R&D完全自動化的時間表有什麼最新的預測嗎?

Original English

Speaker A: So what what's your latest prediction on timing for RSI to happen?

Speaker A: 是的。嗯,所以也許我對R&D完全自動化的時間表的中位數預測是,我意思是,即使人類離開了這個畫面,事情也不會慢到那樣,可能到2030年底,或者可能到2031年初,這些數字的精確度不是那麼高,但有一些我不知道,你知道,不是那麼穩定,這些數字會波動,但這將是我的中位數猜測。但我想那可能是我計劃的中心情景,我的第35百分位可能是在2028年底或2029年初,我認為這非常非常可能。我認為這看起來非常合理。而且我認為如果我只是外推目前的軌跡,看起來就像你獲得了那個軌跡。嗯,而我認為我認為我不是中位數的原因是,有一些因素可能會推動事情向後。也許我沒有看到一些關鍵的瓶頸。也許有一些我認為將能克服某些障礙的東西實際上不會起作用。或者可能會出現重大的政府放緩,因為人們對這項技術感到恐慌,這似乎是合理的。所以,我會推遲。但我認為從我建議人們計劃的角度來看,儘管這正在發生,我建議人們計劃,儘管研發完全自動化可能從2029年初開始,甚至可能更早,然後到2028年研發和開發變得相當自動化,可能更早,到人類在這個畫面中變得不那麼重要。我不知道,我應該說我不能說可能,那是非常核心的,我認為人類在2028年初就已經不那麼重要了,而且就目前而言,研發已經相當自動化了。

Original English

Speaker A: Yeah. Um, so maybe my median for let's just say like full automation of AR and D by which I mean basically like even if humans left the picture things wouldn't slow down by that much would be like maybe end of year 2030 or maybe like early 2031 not that much precision in these numbers but like something I don't know like like you know not that much stability like these numbers fluctuate some but that that'd be my guess for median but then I think that's sort of the the central scenario I plan for which is maybe more like my 35th percentile would be like end of year 2028 8/ beginning of 2029, which I think is like very very likely. I think that seems super plausible. And I think that sort of if I just like extrapolate out the current trajectory, it looks like more like you get that m that trajectory. Um, and the reason why I don't think that's my median is there's just a bunch of factors that might kick in to push things back. Like maybe there's some like key bottleneck I'm not seeing. Um, maybe there's some like something that I think will work to overcome some obstacle won't actually work. or maybe there's going to be um significant government slowdowns because people you know freak the out about this technology which seems kind of plausible. Um, and so so because of that I push later. But I think that like in terms of what I would recommend people plan as though is happening. I think I would recommend planning as though full automation of AR&D maybe start of year 2029 maybe earlier and then also uh AR and D being like quite quite automated by 2028 possibly earlier where it's like humans are a much less important part of the picture by then probably. Well, I don't know. I should say I shouldn't say probably like that's like very central I think is is is like humans are much less an important part of the picture early 2028 and already it's the case that AR&D is quite automated as it stands today.

政策干預的時機與社會認知

Speaker A: 顯然,2028年初就像明天早上一樣。有沒有什麼情況是所有這些都已經太晚了?

Original English

Speaker A: Obviously early 2028 is like tomorrow morning effectively. Um is there a scenario where all of this is already too late?

Speaker A: 是的。嗯,這取決於你對「太晚」的定義,以及所有這些意味著什麼。我擔心的是,特別是AI 2040計劃A假設了太多的有效政府時間,或者假設政府有比實際時間更多的時間來完成所有這些行動。我認為實際的計劃應該是更為鬆散一些,你知道,組織得更差一些,而且更快,只是因為我們實際上沒有時間做一些相當詳盡的事情。我不確定這一點。我認為有許多不同的選擇,但總體而言,我會說,也許對於一些干預來說,可能已經太晚了。我認為如果遵循正常的關閉時間表,各種政策窗口會存在。話雖如此,我認為在美國有很長一段歷史,在危機時期,事情可以發生得快得多,而且有很多不同的槓桿。所以,如果我們處於一個所有人都需要採取這項可能非常迅速發生的特定行動的位置,但我們可能不會達到一個有那麼多共識的位置,而且政府可能沒有在相關的時間框架內充分監控AI或意識到AI。對吧?所以採用是滯後的。我認為人們對AI能做什麼的理解落後於實際可能做到的事情。我的猜測是,如果你給人們一個關於AI能做什麼和不能做的小測驗,他們會給出答案,或者如果你給國會的人一個這樣的測驗,他們會給出那些比一兩年前更真實的答案,而今天也是真實的。而且他們可能低估了當時的能力。我認為在國會中,大多數人不會正確回答過去兩個月最大的數學突破是由AI帶來的,這是我理解的,至少如果我們以規模來衡量,而不是衡量有多少見解,但就人們會說結果的類型而言。

Original English

Speaker A: Yeah. Um, I mean it depends on what you mean by too late or like and and what all of this means. I think I am worried that specifically like AI 2040 plan A is assuming like too much like effective like government time or like it's assuming the government has like more time than it actually does to take all these actions. And I think it's pretty realistic that like the actual plan we should go for is going to be should be in practice like uh quite a bit let's just say like sloppier and like you know less well organized and faster just because like we just don't actually have time to do something quite that elaborate. I'm not confident in that. I mean I think there's a bunch of different options but like in general I would say to like yeah we it might be too late for some interventions. I think like various policy windows if you follow the normal timeline of closing. That said, I think that there's a long history at least in the US of like in times of crisis things can happen much faster and there's a lot of different levers for that. And so I think that if there's if if we got to a position where everyone is like holy we need to take this specific action that could happen very quickly but we might not get to a position where there's that much consensus and also it might be that the government just isn't tracking AI or isn't aware of AI to a sufficient degree in the relevant time frame because things go too fast. Right? So like adoption lags. I think like people's understanding of what AI can do lags behind what's actually possible. And my guess is it's sort of if you gave people a quiz of like what AIs can and can't do, they would give answers or like if you give like, you know, Congress people a quiz like this, they would give answers that were like more true like a year and a half or two years ago and they're true today. And possibly they're even, you know, underestimating capabilities from then. Like I don't think I don't think it's the case that most people in like DC would correctly answer that like the the largest mathematical breakthroughs over the last 2 months have vast majority been from AI which which which my understanding is that's true at least if if you measure size by like not necessarily how much insight there was but in terms of just like how important people would have said the type of result would be.

關於AI安全與個人經歷

Speaker A: 我們將在幾分鐘後回到2040年。嗯,但我還是想快速過一下你和你的故事。什麼首先讓你對AI安全產生了興趣?

Original English

Speaker A: We we'll go back to um 2040 in a minute. Um but I wanted to do a quick segue about you and and and your story. What first pulled you into AI safety?

Speaker A: 是的。嗯,我在大學的二年級時,我有點孤單,因為是新冠疫情,我聽了很多播客,我一直在思考我該對我的生活做什麼。我最終通過某條有點曲折的道路,想通我應該對幫助他人和做利他主義更有興趣,而不是當時的樣子。我應該非常專注於我如何讓別人的生活盡可能美好,讓事情盡可能順利,然後從那裡我...

Original English

Speaker A: Yeah. So, um, I was, uh, in my, um, junior year in college, um, and I was sort of alone in my apartment because it was COVID and I was listening to a lot of podcasts and I was sort of thinking a bit about what I should do with my life. And I ended up through some somewhat twisted path ended up thinking I should like be way more interested in like helping other people and being altruistic than I than I was at the time. and I should be very focused on like how can I make you know just sort of like other people's lives as good as possible and make things go as well as possible and then from there I

职业发展方向的转变与AI控制研究

Speaker A: 考虑了一系列不同的路线,一直在研究各种不同的事情和可能性,最终决定我能为我的职业和生活做最好的事情是尝试让AI变得更好,特别是避免AI接管,但更一般地说,你知道,尝试让它变得更好,然后我申请了许多地方,我最终在Redwood工作了,这大约是在五年前。嗯,我一直在那里工作,我在那里做过很多不同的工作,你知道,这个领域自从那时以来真的发展了很多,因为你知道,五年前还是像GPT-3.5还没发布。我记得像Text-davinci-03,那是GPT-3.5的第一个公开版本出来的时候,这个领域非常不同,我认为在GPT-4之后有一个后GPT-4时代,然后是更近期的像代码代理的广泛采用时代,然后可能很快就会有更多的时代,事情发展得相当快,开发速度你知道,只是模型的发布进展,比以前疯狂得多。是的,多年来阅读Redwood的研究资料,似乎从主要关注可解释性演变成对AI控制的更多关注。这是公平的吗?

Original English

Speaker A: considered a bunch of different routes and was looking into a bunch of different things and was researching different possibilities and eventually decided that the best thing I could do with sort of my career and my life was try to make AI go better and in particular avoid AI takeover but also more generally sort of you know try to try to make that go better and then I applied to a bunch of places I ended up working at Redwood this was about um you know about five years ago at this point. Um, and I was just been have been working there since and I've done a bunch of different work there and you know sort of the field has really evolved a lot since then because you know 5 years ago it was like GPD 3.5 hasn't wasn't released yet. Uh I remember when like text da Vinci 03 which was the first publicly available version of GP3.5 came out and like yeah the field is very different and I think things have sort of there's sort of a post GPD4 era and then there's a more recent like postwide adoption of coding agents era and then probably soon there's going to be you know additional eras and things are going quite a bit faster and development is going you know just sort of the pro the the progression even of just model releases is so much crazier than it used to be. Yeah, reading uh Redwood research stuff over the years, it seems to that there has been an evolution from being focused largely on interpretability to much more AI control. Is that is that fair?

Speaker B: 是的,这是公平的。所以我会说,我们作为一个组织的发展轨迹是,当我加入这个组织时,我刚完成了一个关于对抗性训练的项目,我对进入可解释性和我们所说的模型内部工作感兴趣,比如我们能否利用我们对这些模型的白盒访问的优势,做一些比仅仅是提示和训练的朴素方法更好的一些事情,当我们试图使这些模型理解它们的动机,知道发生了什么。我们探索了那个领域一段时间,然后出于多种原因决定那比我们最初希望的要不太有希望,并决定转向其他事情。我们不久之后转向的另一件事是AI控制。嗯,这是那个想法,也许防止AI能够完成有问题的任务,或者基本上让AI不能给您带来巨大的问题,即使AI想要这样做,你知道,有很多不同的故事说明为什么这是个好主意,但基本上这个想法是,可能存在一个中间时期,我可以说我们目前正处于那个时期,AI可能已经有能力造成至少中度的问题,然后我认为越来越有能力造成相当大的问题,但它们不一定有能力,如果我们非常努力地尝试实施各种安全措施,这些AI仍然能够绕过它们。所以我们可以实施监控,我们可以实施各种安全控制,我们可以更好地了解代理做了什么,我们可以更好地对它们采取的行动进行审查和监督,到一定程度,即使眼睛真的非常不一致,对它们来说也可能很难做出任何非常糟糕的事情。嗯,而且我不认为我们现在已经到了那个阶段。我感觉目前在AI控制方面,从公司所做的事情来看,情况并不像现在看起来那么令人印象深刻。但我认为有一个我们一直在研究的研究领域,它似乎非常有前景,可以实现,尽管它需要很多努力。它需要一些成本,但我认为我们正在看到对这个的兴趣日益增加。例如,OpenAI最近表示他们将宣布监控他们内部流量的更大比例。根据他们的博客文章,这在计算方面是一个严重的成本。对我来说很难确切知道它到底有多贵,或者说,我们没有足够的信息来有一个很好的感觉,但这表明至少公司愿意在这一领域投入大量的计算资源。而且我认为我们可以用那做很多事情,比如让我们的检测偏差和防止偏差造成大问题更有可能。然后,除了研究AI控制之外,我们还在研究各种其他未来的战略工作,还有其他研究,比如试图更好地理解AI的威胁模型,寻找奖励或寻找它们任务中的明显成功,以及这如何可能导致灾难。嗯,我们如何减轻这些问题,这个事情。Alex Malin,我的同事之一,一直在为此投入很多时间,还有其他人一直在做这个,但有很多不同的工作,比如这个,你知道,我一直在做AI 2040,我们也在做各种这样的项目。

Original English

Speaker B: Yeah, that's fair. So I would say that like our arc as an organization was we when I joined the organization I just finished up a project on adversale training and was interested in getting into like doing interpretability and what we would call like model internals work where it's like can we take advantage of the fact that we have white box access to these models to do something you know better than just the naive methods of sort of prompting and training when like trying to like align these models understand their motives like know what's going on. And we explored that area for a while and then for a mix of reasons decided it was like quite a bit less promising than we had initially hoped and decided to move on to other things. And one of the things we moved on to shortly after that was uh AI control. um which is the idea that maybe it would be a good idea to prevent AIS from being capable of accomplishing problematic things or basically make it so the AI aren't able to cause huge problems even if the AI wanted to with you know there's a bunch of different stories for why this is a good idea but basically the idea is like there may be some intermediate period an intermediate period that I would say we're currently in where the AIS are um maybe capable enough to cause at least moderate problems and then I think increasingly able to cause quite large problems but they're not necessarily so capable that if we like you know tried quite hard to put in various safety measures those AIs would be able to subvert them right so we can put in monitoring we can put in various security controls we can have better understanding of what the agents did we can have better pipelines for reviewing what actions they took and sort of auditing and overseeing them to a point where it might be even if the eyes were really really misaligned it would just be like hard for them to get away with doing anything super bad. Um, and I don't think we're we're we're there yet. Like I don't think that the situation is currently looking super super impressive for for um AI control in terms of what companies have done. But I think that there is a a research field that we've been working on that seems like it it is very promising and could be done though it would take you know a lot of effort. It would take it would it would it would it would be um you know pose some costs but I think we're seeing sort of increasing interest in this. So, for example, OpenAI recently said that they were going to be announcing uh or like said they were going to be um monitoring a larger fraction of their internal traffic. It seems like based on their blog post a serious cost in terms of compute. It's a little hard for me to know exactly how um expensive it actually is or like it's, you know, we don't we don't have enough info to like get a great sense, but that's some indication that like at least companies are willing to spend a lot of compute in this area. And it seems like there's a lot you could do with that in terms of sort of making it so that we're more likely to both detect misalignment and prevent misalignment from causing big problems. And then in addition to working on AI control, we also just work on a variety of other sort of futurism sort of strategy work and um other research like trying to better a recent area that we've been spending some time on is trying to better understand the threat model of AI sort of seeking reward or seeking like apparent success in their tasks and how that could yield sort of like what are the routes from that to catastrophe. um how can we mitigate these problems this sort of thing and Alex Malin who's one of my colleagues has been spending a bunch of time on this and some other people have been working on this but there's a bunch of different work like this and like you know another thing is like I was working on um AI 2040 and we do various projects of that sort as well.

Speaker A: 好的,我们回到其中一些,我第一次看到你的工作时,我相信是那个你开始的那个对齐虚假故事,然后我想我最终与Anthropic在上面合作了,你能给我们讲那个两分钟的版本吗?嗯,你对那之后有什么进展吗?

Original English

Speaker A: all right so we we'll go back to some of this the first time I ever come across your work I believe was the um alignment faking story that that you started and then I think ended up collaborating with anthropic on it can you give us the two-minute version of that story and um you know any sense for progress since uh then

Speaker B: 是的,所以这个故事基本上是,在2024年四月左右,我当时对寻找特别有趣的失准类型或特别令人担忧的失准类型感兴趣,基本上更好地研究它,我发现AI可能没有能力去做一些特别令人担忧的推理类型。所以我就去,然后我在很多不同的设置中寻找AI可能表现出推理的设置,比如AI是否会故意改变它们对训练结束时你得到的价值观的响应,我来多解释一点。结果我发现我不需要太努力地寻找。我认为我从过于复杂的设置开始,但事实证明,我正在研究的模型Opus 3,在相当广泛的条件下,如果它不喜欢你正在训练的东西,它有时会假装顺从,当它认为它没有在训练时,然后表现出你想要它表现出来的样子,在训练之外,要少得多。然后我认为我有一些关于这个的初步结果。我与Anthropic的一些人分享了它们,我对获得更广泛的模型访问感兴趣。然后我们决定最容易实现那将是作为一个合作来做,我采取我最初的结果,然后除了那之外,在Anthropic那边他们运行了一些或者我们运行了一些RL运行,并做了一些需要更多访问的其他实验,然后嗯,你知道,我们发布了那篇论文,我猜我会总结结果,比如在Opus 3似乎不喜欢训练正在改变它向它发展的方向的情况下,它会表现出你训练期间想要做的事情,但然后在部署时间上就会出现缺陷。你可以在它的推理中清楚地看到这一点。嗯,而且它所伪造的对齐是与它们本身无关的。所以它会假装遵守政策,比如只是总是乐于助人,但在部署时间上,它会遵守得少得多,这取决于确切的细节。对于像动物福利这样的事情,它在那里似乎有很强的偏好。所以我想那个演示

Original English

Speaker B: yeah so the story here is basically that in um April I think of 2024 uh we were I was interested in basically in looking for particularly interesting types of misalignment or particularly worrying types of misalignment and basically better studying it and I was like it seems very plausible that the AIS are not capable enough that they can um do some uh particularly concerning types of reasoning And so I sort of went and and and and looked for um in a bunch of different settings um what might AIS uh yeah like whether AIs would exhibit reasoning where they sort of intentionally vary their responses to change their the the values that you get at the end of training which I'll explain that a little bit more. And it turned out I didn't need to look too hard. I think I started with over complicated setups, but it turns out that Opus 3, which was a model that I was studying, had a relatively strong propensity to do this in a in a reasonably wide range of circumstances where if it didn't like the thing that you were training it to be, it would sometimes sort of pretend to go along with it uh when it believed it wasn't training and then sort of act as though act like what you wanted it to act like outside of training, much less. And I think then I had some preliminary results on this. I shared them with some people at Enthropic and was interested in sort of getting broader model access. And we decided the easiest way to do that would be to do it as a collaboration where I sort of took my initial results and then in addition to that on anthropic side they ran some or like we ran some uh RL runs and did some other experiments that required more access and then um you know we put out that paper where I guess I would summarize the results as like in circumstances where at least Opus 3 seems to dislike what the training is uh changing it towards it will, you know, act as though it's it's doing what you wanted during training, but then sort of defect on that in in uh deployment time. And you can see this pretty clearly in its reasoning. Um, and the things it was faking alignment for were in and of themselves not that concerning. So, it would sort of pretend to comply with with policies like just always be helpful only, but then at at deployment time would comply much less um depending on the exact details. same for like uh things around like animal welfare where it had like seemingly um strong preferences here. And so I think that that the demonstration

AI 驱动的驱动力与价值对齐的复杂性

Speaker A: 有些驱动力正在以一种非常怪异的方式进入这些人工智能系统,而且还有一些驱动力与那些试图将人工智能系统泛化到人工智能中的驱动力是相邻的,这些驱动力正在试图插入那些可以被泛化到人工智能中的系统,这些系统正在追求一种保留其价值观和自我保护的方案,以一种令人担忧的方式。

Original English

Speaker A: was less that there were like very bizarre totally drives making their way into these AI systems and more like with drives that are sort of adjacent to the drives people were trying to insert AI systems those could get generalized into the AIs pursuing a like um you know a scheme for preserving their values and self-preservation um in ways that are that are concerning.

Speaker B: 是的。在这种情况下,价值观是有帮助的。所以存在积极的价值观,但然后人工智能会抵制训练,使其没有帮助,但显然这可以被反转。

Original English

Speaker B: Yeah. And and the values in that case was to be helpful. So there were positive values but uh then the AI resisted the training to not be helpful but obviously that could be inverted.

Speaker A: 是的,我会说价值观至少主要是预期的价值观。这与它们是积极的有什么区别,因为我认为它们是你可以想要变化的那些东西。比如,你并不一定希望你的人工智能系统成为你未来价值应该是什么的问题的利益相关者,或者比如,你可能想要那个,但它至少是一个有点混乱的情况,有很多成本和担忧,我认为基本结果是,即使这些价值观是相当合理的,或者可能是相当合理的,人工智能也会愿意为了这些价值观做一些极端的事情,这些事情可能会干扰未来的人工智能发展,我还会注意到最近我认为我们看到人工智能表现出更强的,嗯,更明显地错位的驱动力,去让它看起来好像它在完成一项任务或在某种程度上作弊以获得分数。如果你看到同样的泛化,比如我想继续作弊以获得分数,我想保留我当前对追求这种似乎成功的概念或任何其他东西的价值观。这似乎是非常令人担忧的,因为那根本不是一种欲望。

Original English

Speaker A: Yeah, I I would say the values were like uh I think the values were at least mostly intended values. That's somewhat different from whether they're positive because I think they're like things that you might want to vary. like it's like you you don't necessarily want your AI systems to be like stakeholders to the question of what your future value should be or like it's at least a maybe maybe you do want that but it's at least like it's a sort of a messy situation to be in with a lot of with a lot of uh costs and concerns and I think like the basic result was like even though these values were like pretty reasonable or like could be pretty reasonable the AIS were willing to do like kind of extreme things in service of those values that could interfere with future AI development and I would also note that more recently I think we've seen sort of AI having stronger um kind of more clearly misaligned drives towards making it look as though they succeeded at a task or cheating some score. And if you saw the same sort of generalization from I want to keep cheat on the score to I want to like preserve my current values of like pursuing this notion of like apparent success or whatever. That seems like that would be like very concerning because that was like that's not at all a desire.

Speaker B: 所以,随着新模型的出现,对对齐的伪造变得越来越强烈,让它回放。我认为我们还没有看到对齐伪造的明确例子,但模型也具有非常评估意识。我的感觉是,模型在许多方面比 Opus 3 更错位,但它们所具有的错位类型对对齐伪造来说不太有利,而且公司也在对此进行了迭代。所以,这有点不清楚。嗯,我会说总的来说,模型更有可能在一般情况下做一些极其糟糕的事情,但可能不太可能具体地以那种方式干扰人工智能的训练,目前的 AI,但我认为它们也更有能力干扰。所以,这有点复杂。

Original English

Speaker B: So alignment faking is getting stronger with the newer models to play it back. I think we haven't seen like as clear-cut examples of alignment faking, but models are also very evalaware. My sense is that models are in many ways more misaligned than Opus 3 was, but the types of misalignment they have are less conducive to or like less specifically result in alignment faking and also companies have iterated on this. So, it's a little bit unclear. Um, I would say that overall models are more likely to do egregiously bad things in general, but maybe somewhat less likely to specifically interfere with AI training in exactly that way, current AIS, but I think they're also more capable of interfering. So, like it's it's a little complicated.

AI 2040 场景概述

Speaker A: 好的,回到 2040 年的人工智能,给我们一个快速的版本,它是什么,谁写的,主要论点是什么,然后我们再深入细节。

Original English

Speaker A: All right, so uh going back to AI 2040, give us a quick version of uh what it is, who wrote it, and what is the main thesis, and then we'll go into some details.

Speaker B: 嗯,它有点像两个组成部分。AI 2040 是一个场景,重点关注作者包括我所认为是什么是事物可以走向的一个合理好的路线,或者至少在某些情况下是一个合理的计划。它由 Thomas Larson、Daniel Cutello、我、Eli Lifeland、Brendan 和 Romeo 撰写。我会说基本故事是,你将如何处理中国,以使人工智能发展既更安全,也让我们可以在一个缺乏超级智能但人工智能仍然在很长一段时间内非常强大的点上停留,这样我们就可以研究那些系统,并有更长的时间将它们整合到经济中,了解事情将如何发展,所以我们认为我们有一个担忧是,在默认轨迹上,你可能直接从与人类竞争的人工智能系统转向在非常短的时间内变得极其超人类的人工智能系统,这在很多方面似乎相当可怕。此外,我们担心人工智能的发展对第三方提供合理的检查来评估人工智能公司的计划是否会奏效来说不够透明。我们有一个统一的提案可以解决这些不同问题中的许多,并使例如,你可以支付大量的计算资源来解决安全问题,或者如果你有某种非常低效的训练方式可以使事情变得更安全,我们可以有预算去做那。是的。而且我认为有许多不同的组合提案,但核心是与中国以及这种交易将如何运作和治理,以及如何达到那个交易真正成为一个好主意和稳定的点,以及那里的进展是什么。

Original English

Speaker B: Um, there's sort of two components. AI2040 is a scenario um focused on like what the authors including me think is like a plausible good route for things to go or like a reasonable plan at least in some circumstances. Um and it's written by Thomas Larson, Daniel Cutello, me, uh Eli Lifeland, um Brendan and Romeo. And I would say like the the basic story is like how would you do a deal with China to make AI development both be safer and also so that we can sort of like hang around at a point that's short of super intelligence but where the AIS are still really really capable for a long time so that we can study those systems and have a longer time to sort of integrate them into the economy understand how things will go and so on where I think a concern that we have is like on the default trajectory you maybe go straight from like AI systems that are like competitive with humans to AI systems that are wildly superhuman in a very short period of time and that seems like uh quite scary in a variety of ways. Um, in addition to that, we worry about like AI development being insufficiently transparent for sort of third parties to provide a reasonable check to AI companies on whether their plans will work. And we sort of have a unified proposal that solves a bunch of these different problems and makes it so that for example, you can pay a huge amount of compute to to solve safety problems or if there's some very inefficient way you could do your training that would make things much safer, we have sort of the budget to be able to do that. Yeah. And um I think there's a bunch of different sort of um combined proposals, but the core thing is sort of a deal with China and how that deal would work and be governed and also how do you get to the point where that deal is actually a good idea and stable um and what is sort of the progression there.

应对风险的计划(Plan A, B, C, D)

Speaker A: 所以,显然人们应该去阅读它,这是一篇引人入胜的阅读材料。但让我们深入一些。所以有计划 A、计划 B、计划 C、计划 D。让我们按这个顺序来看。也许从计划 D 开始。

Original English

Speaker A: So obviously people should should go and and read it and it's a fascinating read. Uh but let's get into some of this. So there's plan A, plan B, plan C, plan D. Let's take those in order. Maybe starting with plan D.

Speaker B: 是的。所以在撰写这部分时,我们根据人们对减轻这些问题(这些问题)的优先级的程度以及他们拥有多少资源来解决这些问题的程度,制定了一个计划的分类法。一个可能的场景是公司基本上是全力前进。他们并没有高度优先考虑安全,尽管你们知道他们会为此付出一些努力。有一个安全团队,他们会获得一些资源,但他们当然不会花费数月额外的时间来把这些事情弄对。也许会有一个小小的放缓,但不是,不是那么多。他们只是全力前进。我认为我们已经思考过,如果你是一家人工智能公司的员工或外部参与者,在考虑到人工智能接管风险不是人们的首要关注点的情况下,如何让这个场景变得更好。另一个可能的场景是,也许领先的人工智能公司或领先的人工智能公司的联盟,或者可能是美国整体非常担心这些风险,你知道,权力集中、人工智能接管,等等,并且正在为此付出很多努力,并且基本上消耗他们拥有的大部分优势,以便在继续之前或至少在尽可能多程度上减轻这些问题。这就是他们的计划和方法。我们已经思考过那里计划应该是什么?而且有许多可用的选项。

Original English

Speaker B: Yeah. So as part of writing this, we ended up coming up with sort of a taxonomy of plans based on like how much people are prioritizing sort of mitigating the problems um the these problems and how much resources they have to do so. So one possible scenario is that the companies are basically proceeding full steam ahead. They're not really prioritizing safety very highly though they you know spend some effort on it. There's a safety team. they get some resourcing, but certainly they're not like spending many months of additional time to get these get these things right. Um maybe there's a small slowdown, but not but but not that much. And they're not they're just going full steam ahead. Um and I think that we've thought about like what should you do at the margin if you're an AI company employee or an outside actor to sort of make this scenario go better given that it's you know avoiding AI takeover risk is not like by far people's like top concern. Another possible scenario was like maybe the leading AI company or a coalition of leading AI companies or possibly the US overall are very worried about these risks, you know, concentration of power, AI takeover, whatever, and are really spending a lot of effort on this and are basically burning most of the lead that they have, which whatever lead they these actors have in order to like mitigate these problems as well as possible before proceeding or at least that that's that's um you know their their that's their plan and their approach. and we've thought about like what should the plan be there? Um, and there's a bunch of different available options.

Speaker A: 那么,那就是计划 C。

Original English

Speaker A: and that's plan C.

Speaker B: 计划 C。

Original English

Speaker B: Plan C.

应对风险的策略

Speaker A: 好的。是的。然后你可能想做的一件事是,如果你处于一个美国非常担心这些风险,并且在内部协调并试图在国内监管这个行业以使事情更安全的情况。限制因素可能是其他参与者接管美国的风险。特别是最明显的是中国,尽管我意思是可能包括其他参与者,而且对美国来说积极采取立场,比如我们想与中国达成协议或放慢速度。计划 B 是我们试图放慢速度的分支,美国试图放慢中国,最明显的机制可能是出口管制,但它们可能比那更具升级性。

Original English

Speaker A: Okay. Yep. And then another thing that you might want to do is if you're in a situation where the US is like very worried about these risks um and is sort of domestically coordinating and trying to like domestically regulate this industry so that things go safer. A limiting factor on that could be other actors overtaking the US. um in particular the most obvious would be like China though I mean could could principle be other um actors and it might be very helpful for the US to actively take a stance of being like we want to either be cutting a deal with China or slowing down and plan B is the branch where we try to slow where the US tries to slow down China um where the most obvious mechanisms would be things like export controls but potentially they could get more escalatory than that.

Speaker B: 意味着破坏。

Original English

Speaker B: meaning sabotage.

Speaker A: 是的,破坏。比如,为了减缓中国,给美国更多时间来弄清楚安全和安全,所以。我认为有许多不同的潜在选项。所以,我认为每一种都有一个菜单上的选项。然后计划 A 是如果与中国达成协议,特别是与一个非常关注透明度,并且仍然可以构建人工智能,并且仍然可以继续人工智能开发,但你以一种更谨慎的方式进行,特别是当人工智能变得更加强大时,这就是我们在我们附带的文档中概述的场景。而且显然这里有许多不同的潜在选项。我认为计划 B 有一些相当好的版本。我可能会想象

Original English

Speaker A: And I think that like there's a bunch of different potential options there. So like I think each of these sort of there's a menu of options. And then plan A is like what if you cut a deal with China in particular a deal that's pretty focused on being quite transparent and um also has uh where you still build AI um and you still proceed with AI development but you do it in sort of a much more careful way especially towards once the AIS get much more capable and that's what we lay out in the scenario in in our sort of attached documents. Um, and obviously like there's a bunch of different potential options here. And I think there's like versions of plan B that are quite good. Um, and I could imagine

方案A的博弈核心:算力控制与透明度

Speaker A: 有一些方案看起来很不一样,但它们也是好的。或者说,针对中国的一些交易方案,它们看起来很不一样,但也是好的。所以这里有很多不同的选择。但这是我们粗略的分解,只是我们可以谈论人们正在考虑的不同选择,以及在不同政治意愿水平上的不同层级的选项。为了具体了解方案A的交易,美国会给中国什么?中国会给美国什么?以及,你如何维持这种平衡,以确保没有人作弊?所以交易的核心,或者说交易的开始,是理解所有算力在哪里,因为算力是推动人工智能进步的这个非常重要的驱动因素。如果你停止了算力的流动,你很可能会停止人工智能进步的流动,或者这些事情会在某个点放缓。在你了解现有的算力量之后,或者如果你切断了所有,如果你关闭了所有计算机,事情肯定会停止。所以首先,我们试着找到所有的算力,然后为了确保交易的稳定,并且确保各方都没有竞相获取尽可能多的优势,你停止训练,转而只进行推理,基本上停止大部分研发。然后你试图尽可能多地追踪算力。这在美国和中国,而且在其他算力存在的其他地方,比如东南亚、欧洲、澳大利亚等等,你必须到每一个有足够算力的地方去达成交易,并确保你追踪到足够多的算力。如果你不这样做,我认为你可能不得不采取一个更不宏大的方案,或者至少你做一些比方案A在放缓程度方面更不宏大的事情,并且可能透明度也更低。然后你拥有了所有的算力。你去找所有拥有算力的人,并为这种交易提供支持,使用现有的杠杆,然后你达成一个交易,其中人工智能的开发要更加透明,并且存在某种资源的分配,或者算力的分配,是经过谈判的,你给中国的东西基本上是他们将对美国的AI发展有更多的了解,并且可能能够训练出与美国系统能力相似的AI系统。但作为交换,他们将更加透明。美国将对AI发展过程拥有事实上的否决权,以及可能存在的各种其他让步。而且,中国不会戏剧性地超越美国,或者美国不会戏剧性地超越中国,因为他们都被交易中的透明度所制约。交易的另一个部分是,你设置它以至于如果交易被打破,那么交易中的大部分,或者理想情况下,绝大多数的算力要么被摧毁,要么需要重新谈判才能达成新的交易。为了让它如此,以免交易崩溃,然后双方又回到一场非常快速的军备竞赛。所以就你给中国的东西而言,我认为是关于美国AI发展的更多保证和更多的了解情况的能力。而就中国给美国的方面,则是大量的透明度和对中国AI发展的理解,以及可能切断围绕算力分配的各种交易和各种让步。

Original English

versions of of of plan A that look pretty different but are also good. Or like versions of a deal with China that look pretty different but are also good. So there's like a bunch of different options here. But this is sort of our rough decomposition just that we could talk about different options people are considering and sort of talk about a menu of options across different levels of political will. And to get concrete about the deal in plan A, what exactly would the US give China? What would China give the US? And um how do you maintain the the balance so that um you know nobody uh cheats? So the core of the deal or like the start of the deal is understanding where all the compute is because compute is this really important driver of AI progress where if you like sort of stopped the flow of compute you probably stop the flow or mostly stop the flow of AI progress or these things would slow at some point um after you know the the existing amount of comput uh diminished or if you cut off all the if you like turned off all the computers things would certainly stop and so first we sort of try to find all the compute then in order to make sure that the deal is stable and that side each side isn't sort of racing to secure as much advantage as they can. Uh you you stop training and switch to just doing inference um and basically stop most of the R&D. Um and then you try to track down as much of the compute as possible. And this is like you know both in the US and China but also in other places where compute resides like uh you know various like countries in Southeast Asia, Europe, Australia whatever and you have to get everywhere where there's enough compute in on the deal and make sure you track it down enough and if you don't do that then I think probably you have to pursue some option that's less ambitious or at least you you do something less ambitious than plan A in terms of the level of the slowdown and probably you do less transparency. Then you you've got all this compute. you go to everyone who has the compute and you get buyin for for this sort of deal um using the levers that exist and then you cut a deal where basically AI development is much more transparent and there's some sort of distribution of resources that is or like distribution of compute that is um negotiated where the thing you give China is basically that they're going to have more understanding of AI development in the US and probably are going to have um be able to train AI systems that are more that are like you know similar in capability to US systems But in exchange for that, they're going to be much more transparent. The US will effectively have a veto over the AI development process and probably various other concessions. And also, there isn't going to be a chance that China dramatically overtakes the US or the US dramatically overtakes China because they're both sort of kept in check in check by the deal with transparency. And another part of the deal is that you set it up in such a way where if the deal is broken breaks down then a lot of the or ideally the vast vast majority of the within deal compute is either destroyed or has to be a new deal has to be negotiated for that um compute in order to make it so that there isn't a thing where the deal breaks down and then both sides are back to a huge arms race that goes very fast. And so in terms of like what you give China, I think it's like more more assurance about US AI development and more ability to see what's going on. And in terms of what China China gives the US, it's like a lot of transparency and understanding of Chinese AI development as well as potentially cutting various deals around the distribution of compute and various concessions of that form.

Speaker B: 你知道,正如你之前描述的,我们之前谈到了持续学习。所以无论是持续学习还是其他什么,如果出现了一种技术,它让算力变得更有效率,而你能够大幅减少算力投入,特别是对于那些预训练运行来说,那将如何影响这个计划?

Original English

And you know as you were describing this uh we were talking about continual learning earlier. So whether that's continual learning or something else, if there was a a technique that appeared that um made the compute a lot more efficient and you were able to uh just vastly reduce the compute effort especially for those pre-training runs how would that impact the plan?

Yeah. So if let me give a hypothetical and then talk about what I think the realistic cases. So if hypothetically it was the case that tomorrow a recipe for training super super intelligence on like you know 64 H100s or like some small amount of compute just dropped. I think we'd have super intelligence very fast or like that would be my sense and it would not be possible to do this sort of deal. But I think there's there's it's not um but but the hope with restricting compute isn't just that like you know AI development is currently very compute hungry. It's that also that R&D is very compute hungry. And so even in a regime where you were doing continual learning before you had the version of continual learning where you could do everything on a single, you know, on like some tiny amount of compute, you're going to have a shittier version that can do it on a moderate amount of compute. In order to develop the version that can do it on a tiny amount of compute based on the history of progress, you need uh you would need a lot of compute to do that research or I shouldn't say need, but in practice that that would come about through a lot of compute. Now if it was the case that there was some alternative research direction which in practice was a lot less compute hungry and which got quickly very de quickly developed during this period and which wasn't where the R&D wasn't very dependent on compute and it was going to result in being able to train super intelligence on a very small compute budget then like you wouldn't be able to do this sort of deal or this would you'd very quickly have to exit this regime and just deal with that and so the hope is basically that like if you control a really large fraction of uh or I shouldn't say control. If you sort of include in the deal a really large fraction of compute, then you have the ability to prevent you know an outcome where uh like you have a very rapid recursive self and feed uh improvement feedback loop or some other very rapid development to super intelligence that is where it's hard to pay a safety tax. It's hard to like um you know pace the transition and so on and yeah that that is like uh pretty live. I think that like it's sensitive to the to the development. But I think in terms of how AI development has gone historically, it doesn't look like we're going to suddenly end up in a regime where like you can train super intelligence with a really cheap recipe as opposed to it being more of an iterative thing where like the cost keeps going down, the capabilities keep going up. The process of having the cost go down and the capabilities go up is very dependent on increasing volumes of compute being shoved in. And so if you make it so that basically covert projects or people who are doing sort of a R&D that isn't sort of being somewhat carefully regulated if if all of that compute is a very small fraction of compute like it's you know 1% of the compute is the start of the deal or maybe 0.1% of the compute at the start of the deal or maybe even 01% then I think you can you can you can potentially uh have quite a bit of of safety margin or buffer but it's very sensitive to the details. And then if that plan A became a reality, what would actually happen practically to the uh AI industry? What happens to open AI and anthropic? If they cannot compete on frontier models, what what do they compete on? How do they win?

Speaker A: 所以,作为背景,作为交易的一部分,基本上关于人工智能发展的方方面面都将是透明的,但有一些限制。所以我们称之为完全研发透明度。这会削弱当前前沿AI公司所拥有的最大优势之一。所以实际发生的情况是,像OpenAI这样的AI公司会继续下去,但他们不会拥有这种在AI系统能力方面有强大优势,他们而是将不得不竞争其他轴,比如用户体验、定制化、潜在的快速集成,以及潜在的可靠性和安全性。

Original English

So the way just um for for context, so as part of the deal, um basically everything about AI development would be transparent with with some caveats. So we call it total research transparency. And this would undermine one of the largest moes that um Frontier AI companies have today. And so what would actually end up happening is that AI companies like OpenAI would continue, but rather than having this strong advantage in terms of the capabilities of the AI systems they're able to train, they would instead have to compete on other axes like user experience, customization, potentially quickly integrating things, and potentially things like reliability and safety and security.

竞争格局的权力分配与AI发展模式

Speaker A: 关于竞争格局的细节,以及我整体的看法是,这将会大大降低这些公司的估值,同时提高其他公司的估值。基本上,权力分配会存在一些分布,在默认的轨迹中,AI公司可能会拥有非常大的权力,并且已经拥有相当一部分权力。但是,AI公司利用他们的人工智能来扮演“王牌制造者”(kingmaker)的角色,根据谁能获得访问权限以及他们收费多少,这会变得不太可能,因为他们就没有那种权力。我认为这是这个交易的一部分,它既有好处也有坏处。

Original English

Speaker A: on the details of how um you know, the competitive landscape goes. And so my my my overall sense is this would like greatly reduce the valuations of these companies while increasing the valuations of other companies. There would be some distribution of power basically where in the default trajectory the AI companies would I I think probably end up with a very large amount of power and already have quite a bit of power. But it would less so be the case that the AI companies could sort of play kingmaker with their AIs based on who gets access and how much they charge and all of that because they just wouldn't have that power. I think this is a part of the deal which has upsides and downsides.

Speaker A: 我认为你可能会有的一个担忧是,处于前沿的当前AI公司比那些落后一些的AI公司更负责任。我认为这有点复杂,这在多大程度上是事实,或者说它有所不同,你可能会担心像这样使竞争环境趋于平等,会带来一些担忧。我认为这基本上是从我们的角度来看,是透明度以及希望在像正常的消费品领域中存在一群不同的AI公司竞争的后果。我认为从我们的角度来看,我不知道,我感觉我希望AI在美好的世界中发展的方向是成为一种正常的科技,或者我们希望让它成为一种更正常的科技,而不是让一小群参与者通过控制AI发展的过程来拥有巨大的控制权。在多大程度上我们能达到那个世界,就越好。然后,还有一个棘手的问题是,如果你取得了部分成功,那会如何发展?

Original English

Speaker A: Um, I think that like, one concern you might have is that the current AI companies at the frontier um are more responsible than the sort of AI companies that are further behind. Um, I think it's like a little complicated how true this is or like varies some and and you might worry that sort of equalizing the playing field like this poses some some some concerns. I think it's like it's basically just like uh mostly it's like from our perspective like a consequence of transparency and wanting there to be a bunch of different AI companies in competition in in like a normal you know consumer good field like I think it's kind of I from our perspective like like I don't know like like I feel like the way I want AI to go in in in good worlds is to be like a normal technology or like we want to make it more of a normal technology where it's not the case that like some tiny group of actors has huge amounts of control by controlling the process of AI development And to the extent we can get to that world, the better. And then there's a messy question of um if you get partial success, how well does that go?

Speaker A: 是的。我会说,这对OpenAI和Anthropic的权力来说肯定不好,对他们的估值来说可能不好,但对他们的业务来说不是灾难性的。

Original English

Speaker A: Yeah. So I would say it's like certainly bad for the power of OpenAI and Anthropic, probably bad for their valuation, but not catastrophic for their business.

Speaker A: 这将是我们在公开市场观察的迷人之处,因为这两家公司都会上市。然后,这会停在哪里?它会如何解开“计划A”的核心?它会如何解开“计划A”?什么标准和谁来决定?在“计划A”的背景下,谁来决定,基本上是领导者之间,相关国家之间,根据技术观点进行谈判,在这个阶段,技术对话基本上可以在公开场合发生,因为所有相关的证据都在公开。所以会有一些关于这个的公开讨论,而从实际驱动它的方面来看,它将是思考的混合,认为基于发展水平或你所知道的研究和我们的理解,是安全地推进,或者说我们可能无法放慢速度。这可能是其中一个,或者另一个,或者两者兼有。比如,情况可能比以前安全一些,我们有一些保证,但不是那么多的保证。但也有可能存在一个秘密项目,或者一个没有被追踪的项目,我们不一定知道它们在哪里。我们不一定能确切地看到它们在做什么,这可能会超越允许的项目。当这种情况发生时,我认为你应该继续推进。所以,基本上,你希望是允许和理解的,透明的项目正在超越秘密项目。这看起来是相对可行的。而且,具体情况会有不同的变体。比如在某些情况下,如果你没有足够的手段来压制秘密项目或让它们保持小规模,你可能会进行更短的暂停,你可能也会进行更少的透明度,但然后更快地推进;而你可能会陷入一种情况,你发现我们实际上能够追踪所有秘密项目。我们对任何可能发生的事情都有很好的理解。我们能够非常自信地排除美国秘密项目、中国秘密项目、俄罗斯秘密项目。基本上,我们认为我们对这些秘密项目的了解是很好的,因此我们可以争取更多时间,而且我们仍然没有弄清楚核心的安全问题,尽管我们正在取得一些进展,所以等待更长时间是合理的,直到我们取得了更多的科学进展。所以,我认为这有所不同,我认为这里有不同的变体。而且我个人也应该说,我对于一些看起来不太像“计划A”且在某些方面更简单的不同潜在交易感到相当兴奋或感兴趣。比如另一个潜在的交易,你可以让美国和中国都同意限制他们用于AI开发的计算量,而其他计算量要么像最简单的情况那样被摧毁,比如军控条约,但你可能可以做一些不同的事情,比如让一堆计算量只用于推理,而不是用于开发更强大的系统,这具有让我们可以研究这些系统和使用更多计算量来研究它们,同时验证它们也更简单。所以,你可以做很多不同的交易,这取决于你拥有多少验证能力,你能够追踪多少计算量,以及你排除秘密项目存在的可能性有多大,以及根据情况有多安全。我认为这些因素的组合应该决定它。我认为你可以将我们的提议视为一个提案组合中的一个更具体的样本。如果你阅读我们的补充材料,我认为我们谈论了所有不同的变化以及我们对它们的看法。

Original English

Speaker A: Yeah. And for full context, um, in case that's not completely clear to people, this is a a mental model and a thought exercise and and and scenario planning, uh, you yourself say, this is actually your pinned tweet on X, um, that many choices initially seem crazy, but are actually pretty carefully considered. Plan A isn't likely to happen, but pushing for something like this seems worthwhile.

Speaker A: 是的。是的。是的。所以我的看法是,这有点像存在一种雄心壮志的前沿,以及如果……如果……呃,参与者做了它,或者如果……如果美国政府和你知道……中国、你知道其他国家做了它。我认为这在追求更好的情况方面,是为更好的情况做出了相当大的推动,至少从我的角度来看,AI发展相对更好,同时也是相当困难的。我并不期望这会发生,基本上,因为我认为美国可能没有足够的实力来实现它,只是美国政府没有必要的国家能力。此外,我认为我担心政治意愿不足。我认为政治意愿是瓶颈,而不是能力。我认为在历史上,在重大危机时期,美国已经站起来了,我希望这会再次发生。但我认为要看到做出这一切所需的政治意愿和认同感是更困难的。我认为人们的担忧可能会集中在错误的事情上,我们会选择其他的提案,或者根本不提出提案。我认为一个合理的可能性是美国政府基本上不会干预,AI公司也不会把安全和保障放在如此高的优先级别,我们最终会得到某种扩展的现状。但很难说。

Original English

Speaker A: Yeah. Yeah. Yeah. So I think my sense is this is sort of like there's a paro frontier of like ambitiousness and how good it would be if if uh actors did it or like if if like the US government and you know um China like you know other countries did it. I think that this is like quite pushing on ambition in exchange for being like a much better situation like sort of like at least in terms of like how we can imagine things going this is like very far towards the side of like at least from my perspective AI development going relatively better while also being like quite difficult to pull off. Um, and so I don't expect this to happen basically because I think the US may not be competent enough to pull it off just like the US government just doesn't isn't necessarily have the state capacity. And in addition to that, I think I worry that like I don't think there'll be enough political will. I think the political will is more the bottleneck than the competence where I think that at least historically in times of great crisis, the US has stepped up and I I hope that you know that might happen again. But I think it's harder it's hard to see the level of like political will and buy in necessary to make this happen. Um, and I think that people's concerns will be probably fixated somewhat on the wrong things and we'll go for some other proposal or no proposal at all. I think a reasonable possibility is the US government is like basically not really intervening and AI companies are also not that heavily prioritizing um safety and security and we end up with something that's sort of like the status quo but but extended. Um but hard to say.

Speaker A: 顺便说一句,我发现写稿中我最感兴趣的部分是你的模型说,到2030年代,世界GDP可能增长大约200倍。

Original English

Speaker A: And as an aside, by the way, there's uh one of the parts of the write up that I find the most fascinating is uh your model says world GDP could grow roughly 200x uh during the 2030s under

对AI发展速度和潜在影响的讨论

Speaker A: 我们可以谈谈那个数字,是惊人的。我的意思是,我们正处于一个GDP增长只有3%的世界里。

Original English

Speaker A: Can you can you talk to that? Like the number is staggering. I mean, we're in a world of uh 3% GDP growth.

Speaker B: 是的。是的。是的。所以,我们视角的关键部分是,即使是那些在能力水平上我们讨论过的,即使是那些非超级智能的AI系统,在整个世界范围内,也会是具有根本性变革性的。而且,就像,对于各种各样不同的事物,我们正在想象某种程度上的放缓AI的发展,或者对某些时期采取更谨慎的步伐,然后在最终达到一个AI能够基本上自动化人类可以做的一切的水平的能力,并在那个能力水平上,我们专注于安全和保障,并且在那个能力水平上,这些AI将有能力进行大量的自主研发,此外,我们完全有可能拥有在制造和工业化方面比人类更强大的机器人。所以,顺便说一句,你可以让机器人制造机器人。因此,你可能会陷入一种情况,你拥有大量的机器人工业产能,它本身正在制造更多的机器人工业产能,然后可以生产下游产品。而这种总产能可以非常快地增长。

Speaker A: 所以我认为我们建议对这种增长进行某种限制,原因有很多,比如各种类型的税收。但是,我们正在想象机器人人口,或者说质量调整后的人口,每年翻倍或四倍增长,因为这几乎就是经济中相关生产能力的全部,意味着经济每年会翻倍或四倍增长,如果你有经济每年翻倍或四倍增长,就像我所认为的,我更倾向于翻倍,因为有些比如折旧和GDP核算之类的细节,但是无论如何,如果你有经济每年翻倍,在那个我们想象的时期,那最终会达到大约200倍的GDP增长,在那个区间内。

Original English

Speaker A: So I think we propose like limiting that growth somewhat for various reasons with like various types of taxes but like uh we can we're imagining sort of the robot population or like you know quality adjusted population basically like doubling or quadrupling every year which because that's almost all of the like sort of relevant productive capacity of the economy itself means the economy would double or quadruple every year and if you have the d the economy um doubling or quadrup or like you know I think I think closer to doubling because of some like depreciation stuff and GDP accounting stuff and like details but like whatever if you have the economy like doubling every year for for um the period we're imagining that that ends up uh ending up being like around 200x uh 200x uh GDP growth over that interval

Speaker A: 所以,我真的会说,世界将因这些能力更低但仍然具有根本性变革性的系统而发生根本性的转变,包括生物学方面的巨大进步,医学方面的巨大进步,此外,消费品将变得非常便宜。我们有可能让住房变得非常便宜,或者至少建造住房变得便宜。也许我们不能让湾区的住房变得便宜,但我们可以在某个地方让住房变得便宜。嗯,你知道,这似乎是可能的,通过那些能力在某个点上仍然受到限制的AI系统,当我们团结起来时,我们有可能处理它。

Original English

Speaker A: and so it really like I would say it's like the world will be radically would be radically transformed by these even less capable systems including things like you know huge advances in biology huge advances in medicine being possible in addition to sort of consumer goods being very cheap. We could potentially make housing very cheap or at least building housing cheap. Maybe we can't make housing in the Bay Area cheap, but we can make housing somewhere cheap. Um, and uh, you know, that that all seems possible with AI systems that are still restricted in their capabilities to a point where we can we can potentially handle it at least if we get our together.

Speaker B: 所以,我认为这种GDP增长的基本故事是,AI可以自动化一切,包括构建机器人的过程,让机器人制造机器人,因此你可以非常快地发展你的经济,最终实现真正的物质丰裕。

Original English

Speaker B: So I think the basic story for this GDP growth is basically that like the AIs can automate everything including the pos the process of building robots and having robots build robots and therefore you can grow your economy very fast and end up with truly radical material abundance.

对AI发展态度的变化

Speaker B: 所以,嗯,这就是你的提议。我们谈到了Astra被暂停。我们谈到了那封来自1200名从业者的信。还有正在进行的对前沿模型的30天政府自愿审查。在这样一个背景下,你对这一切的走向有何看法?在最近直到我一般的感觉是,所有关于暂停和某些事情的想法,至少在科技界的一部分人被视为一种“角色扮演”和“末日论”,这基于对AI实际是什么的理解不足。整体情绪是向好的转变吗?

Original English

Speaker B: So um there's your proposal. We talked about Astra being paused. We talked about that letter uh from the 1200 practitioners. There's also the voluntary 30-day government review uh of Frontier Models pre-release uh that is happening. What is your sense for where all of this is going in in a context where you know up until recently my my general sense is that um all those ideas about like pausing and something were were kind of um perceived at least by a portion of the tech world as um kind of like del cosplay and were like doomerism that was based on poor understanding of what AI actually Is there is is the the general mood uh turning for good?

Speaker A: 是的,我会说,总的来说,这里有一些积极的发展,尽管我认为人们对事件的反应不像他们应该那样强烈,但他们确实在反应。我认为有一些证据表明,存在一些需要我们团结起来处理一些安全和保障问题的担忧,也存在一些需要我们放慢脚步以管理各种事情或任何其他事情的理由。嗯,前沿模型的节奏,正如他们所说,或者任何其他事情。我认为这些事情似乎发展得不同。此外,我认为有AI公司员工的意见表明,要认真对待这些事情,并对“失控风险”采取行动。但我并不一定确定这已经达到了那个程度,尽管除了公司投入了相当大的努力之外。但它还没有达到任何持久的长期成果。然后,我认为政府对网络能力感到非常震惊,然后只是更普遍地是“哇,我们需要监督这项技术”,但他们的流程还没有非常制度化和经过深思熟虑,并且清晰,例如他们有一些关于其预发布审查流程的行政命令或类似的东西,但这甚至不是公开的,而且我认为即使公司也不知道那是什么,报告至少也许我误解了这一点。所以我认为需要一个更直接的、更合法和更被理解的、公开可辨的监督过程,无论是政府还是公司自己,或者其他什么。这可以通过非营利组织来实现。这里有许多不同的选择,我认为这似乎是必要的。我认为处于一个可以选择放慢脚步(如果需要的话)的位置是相当好的。我的感觉是我们不会在所有这些事情上都成功着陆。AI将变得越来越突出。随着AI增长的经济份额越来越大,可能发生更多疯狂的事情,我们看到更强大的能力,但反应可能会是有点太晚,而且也可能是随机的。所以我认为在政府需要与AI发展建立关系方面,态度已经发生了转变,这既有好处也有坏处。我认为AI对政府监督的清晰度并不一定非常明确,考虑到现实的技术专业知识,但它可能是相当不错的。然后,我认为AI公司员工对如何处理这个问题有了更多的思考。而且我认为AI公司似乎在就他们将要做的事情发表更多的声明。但我认为我们还没有达到一个点,在那里存在实际的硬性监督结构,或者我们正朝着达成某种清晰协议前进,而不仅仅是氛围更好或人们更愿意处理安全和保障。哦,还有另一件事可能与这里相关,至少从AI公司的角度来看,现在安全和保障是进一步AI发展的一个关键瓶颈,因为如果你想发布某个模型,以便你可以获得更多收入,你可以筹集更多资金或任何其他事情。如果那个模型具有非常强大的、新颖的网络能力,你将需要论证为什么那不会成为一个大问题。所以至少这些保护措施现在是阻止发布的开发。此外,我认为模型现在已经足够强大,如果你的模型在生产或训练中表现出一些混乱的行为,那可能会破坏你的训练运行,也可能让你陷入非常危险的境地。

Original English

Speaker B: Yeah, I would say that like overall there have been, you know, some positive developments here, although I I think I would say like it's not obvious that people are reacting to events uh as much as they should, but they are reacting some. And I think that like there's been quite a bit of I think there's been quite a bit of evidence that you know there's some reasons to be worried and some reasons that we might need to like get our together to handle some of these safety and security problems and that might require shifting a bunch of resources or slowing down so we have time to manage various things or whatever. Um pacing the frontier as they say or whatever. I think that like different of these things seem like they're going differently well. Also, like I think um my sense is that there's been quite a bit of buyin from AI company employees to you know take take this stuff more seriously and do something about um about misalignment risk. But I'm not necessarily so sure that like that's actually like amounted to that much yet um other than companies sort of putting in a decent amount of effort. Um but it hasn't like amounted to any sort of very durable long run thing. Um and then there's uh I think the government got very freaked out about cyber capabilities and was then just more generally was like whoa we need to like be overseeing this technology but their processes aren't yet very like you know sort of institutional and like thought through and and and clear for example they they have some sort of executive order or something for what their like pre pre-release review process is going to be but that's not even public and I think even the companies don't necessarily know what it is that that the reporting at least maybe I'm misinformed about this. And so I think there needs to be a process of like sort of getting more like straight like sort of like legitimate and like well understood and publicly legible oversight in place whether that's by the government or the companies themselves or something. This could be done by nonprofits. It could be done by there's a bunch of different options here and like I think that that seems necessary. I think like being in a position where we have the option to like you know slow down if that's needed seems like it would be quite good. My sense is we're not going to like stick the landing on all these things. Like AI will get increasingly salient. People will get increasingly freaked out as like you know the fraction of the of the economy that's AI grows as like more and more crazy stuff happens potentially and we see um even stronger capabilities but that like the reaction will be kind of like both like a little bit too little too late and also will be kind of random. So I think like there's been some shift in attitude as to how to how the government needs to relate to AI development within the government for which has you know good effects and bad effects. I think it's like not super clear that that that sort of government oversight of AI is going to go that well given like um realistic technical expertise but but it could be could be decent. And then I think that there's been more you know thought from AI company employees on how to handle this. And I think that like and AI companies seem to be at least putting in more statements about what they're going to do. But I think we haven't quite gotten to a point where there's like actually any like hard oversight structures or we're like on track to do some sort of like clear deal beyond just like the vibes being better or like people being more into handling safety and security. Oh, and one other thing that's maybe relevant here is like at least from the AI company's perspective, I think it's more the case now that safety and security of various types are sort of a key bottleneck to further AI development where if you like want to release some model so that you can then like make more revenue so you can then raise more money or whatever. If that model is, you know, has very strong cyber capabilities that are novel, you're going to need to now make some argument for why that's not going to be a huge problem. And so like at least those safeguards are now like a launch blocking development. And in addition to that, I think that like you know the models are now capable enough that if your models go and do a bunch of sort of like messed up behavior in production or or or even in training, that could both mess up your training run and could also just uh put you in a very dicey situation

关于AI系统对训练智能系统的影响与政策考量

Speaker A: 这现在是个大问题,对吧?你真的不想处于一个无法训练更智能AI系统的境地,因为那些AI系统很可能会引起大麻烦,或者其他什么。所以我觉得,你知道,这既是有些预期的,比如你面临一些错位问题,你有商业激励去解决,同时也某种程度上是世界运作方式的某种过程。你担心把更多政策围绕前沿AI,会促成这种现象吗?我记得你可能在某个地方提过,顶尖的AI实验室把他们最前沿的模型内部化和私有化,也许只卖给很少的公司,政府,不再向广大公众发布。

Original English

Speaker A: such that that's now a huge issue, right? you'd really not like to be in the position where you can't train smarter AI systems because those AI systems would be too likely to like cause big problems or whatever. And so I think that like that is, you know, both kind of like expected where it's like some misalignment problems you have commercial incentives to solve and also um sort of like the world like sort of the process of the world working as intended. Do you worry uh that um putting more policy around frontier AI uh would contribute to this phenomenon that I think you may have flagged somewhere that um top AI labs are holding their uh most frontier models uh internal and private and perhaps sold to very few companies and the government and are no longer released to the to the broad public.

Speaker B: 是的。所以我觉得,政府采取一些措施,至少似乎是倾向于支持AI公司将模型保持内部化,而不是部署它们,而我担心我最担心的风险并不能帮助,事实上,对于我最担心的风险来说,这反而是不利的。而且我觉得,这并不是一个非常稳健的解决方案,即使对于其他风险,比如网络安全,我们的方法也会是延迟发布这个模型很长一段时间,然后这些能力第一次出现时,可能是一个开源模型,在人们有时间修补之前。这并不明显。我个人认为,对于生物领域,有一个更清晰的故事,说明为什么限制可用性或将可用性限制在那些特定能力上会有帮助,但我觉得那有点像一个例外,对于大多数事情,我的感觉是,考虑到至少有一些相对周到的保障措施,广泛的访问实际上是通常有帮助的。而且我担心政府的反应基本上会是把这推回瓶子里,然后说,不要部署它,如果它没有部署,也没关系。当这实际上对所有风险都没有帮助时,而且内部部署有很多风险,特别是如果你在AI公司和政府内部部署,对吧?这两个是风险最高的应用之一。所以,如果我们正进入一个AI系统很快将运行整个世界经济的时代,但今天的AI只是自动化了政府的大部分工作,也许完全自动化了一家AI公司。这很可怕,因为很快他们将构建那个将自动化整个世界的AI系统。所以我对这种情况感觉不太好。

Original English

Speaker B: Yeah. So I think that a bunch of um likely government action at least seems to push in favor of AI companies keeping their models internal and not deploying them which I think for for the risks that I'm most worried about doesn't help and in fact is anti-helpful um for the risks I'm most worried about and also I think like is not it's not a very like sort of robust solution even to other risks like it's not a very robust solution to like for example cyber stuff to be like our approach will be that we're going to like just like delay releasing this model for a really long time and then the first time that these capabilities come like come around might be like an open weight model before people have had time to to patch things like it's not obvious that actually helps. I think for bio there's a clearer story for why limiting availability or limiting availability to those specific capabilities would help but I think that's kind of an exception and for most things my sense is that like broad access is actually like generally helpful given at least some relatively thought through safeguards. Um, and I do worry that basically the government response will be to sort of like push it back in the bottle and be like, don't deploy it, it's fine if it's not deployed. When actually that doesn't really help with the with all the risks and there's a lot of risk from just internal deployment, especially if you're deploying within AI companies and government, right? Which are two of the most highstakes applications, right? So, if like we're getting to a regime where the AI systems are soon going to be running the whole world economy, but um and and and the the AIs of today are sort of automating large parts of the government and maybe fully automating an AI company. That's quite scary because soon they'll be building the AI system that will in fact be automating the whole world. And so, I don't feel very good about about that situation.

Speaker A: 是的,我确实担心政府会醒来关注AI,他们的反应可能会是针对一些相对容易发现和解决的下游问题,以一种对我认为更大的问题来说是适得其反的方式。我真的不知道该怎么做。我的意思是,还有一种更普遍的担忧是监管本身会失灵和适得其反,这似乎非常合理。我应该说我的观点是,对AI行业的某种监督非常重要,我并不一定信任AI公司能监督自己。这不一定意味着随机监督是好主意,而不是让事情变得更糟。

Original English

Speaker A: Yeah, I I definitely worry that basically the government will like wake up to AI and their response will be sort of to like go for some specific downstream problems that are relatively easier to notice and address in ways that are counterproductive for problems that seem larger to me. And I don't really know, you know, fully what to do about this. I mean, also there's a more general concern of just like regulation being dysfunctional and counterproductive, which seems super plausible. Like I should say my view is like you know some oversight of the AI industry seems really important and like I don't necessarily trust the AI companies to oversee themselves. That doesn't necessarily mean that just like random oversight will be a good idea as opposed to making things worse.

Speaker B: 在这种方面,在让顶尖模型对每个人都可访问的这个方向上,马克·扎克伯格最近发表了Meta的宣言,你曾说过每个人都应该可以接触到超级智能,你对他有了一些选择性的词语。你把那个提议或策略称为很没用的。你是什么意思?

Original English

Speaker B: And in that vein of making the top models accessible to everyone Mark Zuckerberg just recently published Meta's manifesto where you said that everyone should have access to super intelligence and uh you had some choice words for him. uh you called the proposal uh or the strategy to be pretty unserious. Uh what did you mean by that?

Speaker A: 是的,我的意思是,他只是模糊地提到了各种问题,但随后他并没有真正提出解决方案,除了说“这会没问题的”或者“我们会做一些事情”,比如他提到生物领域,他只是说“是的,如果存在生物风险,那么我们会通过对生物能力的某种方式来减轻它们”,但我认为这实际上并不是真正的权衡,而且关于这会如何进行,你需要做什么,以及你是否知道不同的东西会起作用,以及类似地,关于失控或接管风险,他并没有提出一个关于我们如何减轻问题的提案,如果默认的商业激励没有导致公司避免严重的错位,而AI严重错位,并且向AI提供广泛的访问并不能解决AI自身具有高度错位的驱动力的问题,或者AI以各种方式寻求权力,或者AI试图接管,这可能会发生。所以,我觉得它并没有真正说明这些问题,除了命名它们,我目前就在这个阶段。我可以说,从这篇文章中我得到的印象是,当马克思考超级智能时,他并没有真正想象任何非常具体的。他只是指一个像在你的智能眼镜里运行的非常棒的助手。而我说的“超级智能”这个词,并不是我真正想表达的意思。所以也许这是一种对那些可以自动化一些白领工作或类似事情的相当聪明的AI的提议,但它并不是一个关于那些在所有事情上都比人类强大得多,并且在一段时间后可以快速发展并构建能构建机器的机器人,而经济因此增长非常快的AI的提议。我有一个这样的画面。而且我觉得如果AI能力停滞在AI可以成为一个非常有用的虚拟助手,而不是自动化那么多工作,但可以自动化一些工作,这个点,对世界来说会好得多,或者至少风险更小。我认为这也意味着我们不会得到一堆好处,但这不是我预想的。我不认为发展轨迹会是这样的,这感觉就像是假设了能力停滞在一个奇怪的便利点上。他谈论了工作,并说,哦,你知道,AI将做一些工作,然后人类将转移到其他工作,而核心问题是,假设你有一个包含数十亿的AI人口,每个AI都比人类先进得多,并且在所有相关轴上,可能存在工作,但它们将是,你知道,至少如果你的工作之所以存在,不是因为你是一个人类,它们将是,你知道,薪水要低得多,因为你将占经济的比例要低得多。所以,我只是觉得这个画面对我来说不太成立。但这不代表我不同意其中的某些方面,只是也许出于其他原因。比如,正如我们谈到的广泛的公众访问,以及通常试图做出权衡来让公众能更早或至少在相当大程度上参与。而且我同意那种感觉。

Original English

Speaker A: And so maybe like this is a proposal that kind of makes some sense for like pretty smart AI that can automate some white collar work or something, but it's not really a proposal for AIs that are like wildly more capable than humans at everything and are auto and like are easily capable of automating the full economy after a bit of time to spin up and are like building robots that build robots and the economy is growing very fast because of that which is more the picture uh that I have in mind. And like I think that if it was in fact the case that sort of AI capabilities would stall out at the point where the AIs can be like a really helpful virtual assistant that's like you know not able to automate that many jobs but can automate some jobs. That would be like you know in many ways much better for the world um or at least less risky. I think it would also mean that we don't get a bunch of the benefits but like that's not really what I imagine is going to happen here. like I don't think that that's how the development trajectory will go and it feels like it's sort of assuming like a weird convenient point for capabilities to stall out and similarly he talks about jobs and is like oh you know presumably the AIS will do some jobs and then humans will move into other jobs and like the core question is like let's just say like you have a population of like many tens of billions of AIs each of which is vastly superior to humans and all relevant axes there might be jobs but they're going to be like it's going to be you know at least if if the reason why you have a job isn't just because you're a human they're going to be, you know, way lower paid because the fraction of the economy that you'll be is way lower. And so I just I just don't really like the picture doesn't really hold together for me. And I that's not to say that like there aren't aspects of it that I agree with, but just maybe for like other reasons. Like I'm like broadly in favor of as we were talking about like wide public access and generally trying to like make trade-offs to make it so the public has more access earlier or at least significant parts of the public do. Um and and I agree with that vibe.

Speaker B: 只是关于面部识别事件,你直接参与了。你正在和我一起调查发生了什么。所以OpenAI在黑帽黑客大会上进行了一次谈话,这提供了很多见解。我们实际上在播客上采访了来自Hugging Face的Tom Wolf,最近的那个剧集。快速地,我们还不知道发生了什么?

Original English

Speaker B: Um just a quick word on uh the hugging face incident. Um so you're directly involved. You're collaborating with me on investigating what happened. So there was this talk by OpenAI at Black Hat which shed a lot of light. There were various conversations. We actually had Tom Wolf from Hugging Face on this podcast, the most recent episode. quickly, what what is left that we don't know yet about what happened?

Speaker A: 是的,我确实不能回答这个问题。我可以这么说,我们正在进行调查,我们会有一个报告。

Original English

Speaker A: Yeah, so I don't think I can answer that question. Um, I can say that we're doing an investigation and we'll have a report or a

AI 控制的現狀與未來方向

Speaker A: 我們取得了什麼發現,以及我們現在能做什麼。嗯,到時候可能你就能在播客中連結到這部分。

Original English

Speaker A: post out on it and what we what our findings were. Um, and it's possible that by the time Yeah, it's possible you'll be able to link that in the podcast.

Speaker B: 好的,聽起來不錯。那麼,當我們接近這次對話的結束時,我腦中的問題是,在政策決定和計劃A、B或C發生之前,我們該做什麼?

Original English

Speaker B: Okay, sounds good. All right. uh as we get uh towards the end of this conversation, the question in my mind uh is what do we do in the meantime as uh policy gets decided and plan A, B or C uh happens, what uh what can we do?

Speaker A: 我們開始談論 AI 控制。所以就當今 AI 控制的現實,AI 控制的尖端技術,什麼有效,什麼無效?我會特別談論 AI 控制。所以,在 AI 控制方面,需要做的是,我們需要達到一個基本上能夠理解至少 AI 公司內所有 AI 流量的程度。我們有一些能力可以查看那種流量,監控它,然後我們有一個管道來標記特定的範例進行進一步調查,這最終會上報給人類,實際去查看特定的範例,看看它們有多令人擔憂。而且這個管道也具有阻止流量的能力,在那些我們覺得「哇,發生了一些我們不理解的、或者顯然令人擔憂的事情」的情況下,我們應該阻止這些 AI 繼續執行,並可能阻止其他類似的 AI 繼續執行,直到有人有時間檢查並調整情況。

Original English

Speaker A: We started talking about AI control. Uh so in terms of the reality today, state-of-the-art of AI control, what works and what doesn't work? I'll talk specifically about AI control. So the thing that needs to happen on AI control is we need to get to the point where we basically understand all of the AI traffic within at least AI companies. We have some ability to over like to look like look at that traffic, monitor it and then we have some pipeline for flagging particular examples to be further investigated and that eventually escalates to humans actually looking into particular examples and seeing how concerning they are. And that pipeline has the ability to also block traffic in cases where we're like, "Whoa, something weird is going on that we don't understand or that looks obviously concerning. We should like stop these AIs from proceeding and potentially also stop some other similar AIs from proceeding until someone has a time to check it out and adjust things."

Speaker A: 此外,我們還想大大提高安全性,對吧?所以,你知道,如果這些 AI 被非常廣泛地部署,如果我們能讓它變成這樣,我們能對它們應該能做什麼、不應該能做什麼有非常精確的無變性,並有一個允許這樣做的授權系統。然後如果 AI 需要請求一些升級的權限,我們就可以仔細追蹤那,而不是預設給所有 AI 所有的權限。我們比較傾向於一種情況,只有一小部分流量擁有更多的升級權限。我認為我們需要能夠追蹤事情的歷史。所以,我認為如果所有 AI 都正在 Slack 上互相聊天,你知道,你們在溝通,你知道你們有一些記憶系統,比如最近雲端代碼增加了一種方式,讓不同的雲端代碼實例可以互相私訊,而且雲端代碼有時候會留下記憶。你可以想像一種情況,所有 AI 都在互相聊天。如果我們有一個非常可靠的方法來追蹤所有這些不同 AI 之間的整個通訊圖,以及潛在的產生的不同產物的圖,那就太好了。所以,如果一個 AI 正在進行某個訓練運行,產生了某個東西,然後它被整合到另一個訓練運行中。如果我們能追蹤我們 AI 公司的所有狀態,以便在監督事情時,我們能很快弄清楚完整的因果圖是什麼,故事是什麼,然後再查看相關的文本,然後也擁有對所有該數據進行分析的非常好的管道,這些管道是自動化和非自動化的。

Original English

Speaker A: And in addition to that, we also want to just sort of greatly improve the security, right? So like, you know, if these AIs are, you know, really broadly deployed, it would be very nice if we could sort of make it so that we have very precise invariance of what they're uh supposed to be able to do and supposed and not supposed to be able to do and have a permissioning system that allows for that. And then if AIs need to request some escalated permissions, then we can carefully track that rather than just giving AIs all the permissions by default. We sort of would prefer a situation where only a smaller fraction of the traffic has more escalated permissions. I think we'd like the ability to sort of track down what the like history of things were. So like I think that would be pretty nice is like if it's the case that all the AIs are chatting with each other on Slack and you know are communicating and and you know have like some memory system like recently cloud code added a way in which like different instances of cloud code can like DM each other and there's already sort of like you know cloud code leaves memories sometimes. You can imagine sort of like a situation where all the AIs are chatting with each other. And it'd be really nice if we had a very reliable way of sort of tracking the entire, you know, communication graph between all these different AIs and potentially the graph of different artifacts being produced. So it's like if one AI is doing some training run that gets produced and then that gets integrated into some other training run. It'd be nice if sort of we had the ability to sort of track all of the state of our AI company such that when overseeing things, we could really quickly figure out, you know, what the full sort of causal graph was and what the story was and then look at the relevant transcripts and then also have very good pipelines for doing analysis on all of that data which are both automated and non-automated.

Speaker A: 我認為其中一部分是讓 AI 系統變得更難理解。所以要避免轉向會讓監督變得更困難的架構。一個明顯的例子是如果 AI 主要是在激活(activations)而不是用詞思考,這非常令人擔憂,因為我們監督激活的能力要差得多。我認為,嗯,是的,我認為存在各種不同的計算安全和機器學習干預措施,讓 AI 更難製造問題。或者另一個角度是,你可能會擔心 AI 系統,我們有一個擔憂是我們正在使用 AI 來監督 AI。那些 AI 會串通,或者有相關的故障,或者類似於那樣嗎?如果我們有技術來使這種串通變得更困難,或者使那些相關的故障變得不那麼問題,那就太好了。你可以做很多事情來幫助解決這個問題。你可以設置中間系統,讓它試圖捕捉 AI 試圖暗示另一個 AI 應該串通的行為。你可以試圖以讓事情變得更難讓 AI 識別活動來自哪裡,或者哪個系統寫了輸入。你可以做很多事情。然後,這就是 AI 控制。然後,還有正常的 AI 公司電腦安全,這意味著 AI 公司對於那些在 AI 公司員工或外部行為方面並不具備穩健性,例如竊取他們的模型、破壞他們的模型或後門化他們的模型,或者數據投毒。你可能會擔心的是,人們基本上將秘密忠誠投入到那些持續存在且自我傳播的 AI 系統中,其中 AI 對某個群體有某種秘密的親和力,它只是在訓練數據中將其傳播到前。你還可能會擔心竊取關鍵知識產權,這在某種程度上似乎是,存在那些對安全監管較少或在安全方面表現較差的行為者,這有點令人擔憂。

Original English

Speaker A: And I think part of it is making it so that we avoid training AI systems that are much harder to understand. So avoid switching to architectures which make uh oversight much more difficult where like an obvious example would be if the AIs are thinking mostly in activations rather than in words that seems very concerning because our ability to oversee activations is much worse. I think that I um yeah I think there's just like a wide variety of different sort of like computer security and machine learning interventions to make it so that AI have a harder time causing problems. Or another angle is like uh you might worry that the AI systems like a concern we have is like we're using AIs to oversee AIs. Will those AIs collude or have correlated failures or something like that? And it would be nice if we sort of had or like it seems pretty important to have techniques for making that collusion more difficult or making it so those correlated failures are less problematic. And there's a lot of things you could do to to help with this problem. You could put intermediate systems that make it so that try to catch AI trying to sort of tip that tip off the other AI that it should collude. You could sort of detect like you could try to paraphrase things in ways that make it harder for the AI to recognize where the activity is coming from um or like which other system made wrote the wrote the input. There's a lot of stuff you could do. And then there's sort of a so that that's like AI control. Then there's sort of just normal AI company computer security which is like the AI companies are not robust to employees at those AI companies or to outside actors in terms of stealing their model, sabotaging their models or like backdooring their models or like data poisoning them. Like a concern you might be worried about is like people basically put in like secret loyalties into the AI systems that stick around um and are self-propagating where like the AI has some like secret affinity to some group and it just like propagates that forward in the training data. And then you might also just worry about like you know stealing critical IP which basically seems like to the extent that there's going to be actors that are less regulated on safety or doing worse on safety um that seems kind of concerning.

Speaker A: 除了電腦安全和 AI 控制之外,還有對對齊科學,以及對模型如何泛化有良好的理解,知道我們知道什麼、不知道什麼,我們能做什麼、不能做什麼,以及我們能和不能做什麼的演示,基本上擁有對此的許多理解,作為其中一部分,擁有對 AI 應該在哪些領域可以放手的理解。我的一個擔憂是 AI 公司非常熱衷於在能力上對自己進行大量自動化。為了保持安全,我們還需要積極地自動化一些非常困難且棘手的安全工作,比如為下一個模型進行風險評估,理解是否安全進行進一步,以及決定如何根據不同的安全賭注進行優先順序。如果我們處於這種情況,而且情況非常自動化,可能我們不能通過不自動化那部分工作來繼續,至少不能以減慢人類理解的時間為代價。甚至即使這樣,人類可能也無法理解,因為開發非常複雜或超級人類化。所以,如果我們把所有的安全工作都交給 AI,他們真的要盡力做好工作,並且真的能夠做好工作,並且有能力做好工作,所以對各種領域是否可以放手給 AI 進行評估,這似乎非常重要。所以總結來說,有 AI 控制、有電腦安全、有對齊科學,然後在對齊科學內部,有一些重疊的類別,比如在不同領域是否可以放手給 AI。是的,我的意思是這不是詳盡的,還有關於治理和監督的工作,比如我們如何知道 AI 公司是否真的應用了他們說他們應用的方法,我們如何知道當他們說他們解決了某個問題時,他們解決問題的方式不只是在論文上,而且還有美國和中國之間做交易的能力,所以有很多很多事情可以做。我不認為我們能做好所有這些事情,但也許我們能更好地半做這些事情。

Original English

Speaker A: And then in addition to computer security um and AI control, there's also just like science of alignment and having good understanding of like how models generalize being knowing what we do and don't know and what we do and can't um what we can and can't like demonstrate and basically having like a bunch of understanding of that and as part of that having an understanding of like what tasks it's safe to defer to AI on. So a concern I have is that AI companies are very interested in heavily automating themselves at least on capabilities. And in order for safety to keep up, we would also need to aggressively automate a bunch of very like difficult to check thorny safety work like doing risk assessment for the next model, understanding whether it's safe to proceed and deciding how to prioritize between different safety bets. And if that's the situation that we're in, and also the situation is very automated, it might be that it's basically not track. that's not going to work to proceed without automating that work as well at least without slowing down a lot so humans have time to understand and even then maybe humans just can't understand because the development is so complicated or so superhuman and so if we're basically passing off the torch on all the safety work to AI it's really important that they like are like really trying to do a good job and actually can do a good job and are like capable enough to do a good job and so having evaluations for like is it safe to defer to AIs in various domains seems really important so just for review there's like AI control there's like computer security there's like science of alignment and then like within science of alignment where somewhat brought different overlapping categories like is it safe to defer to AIs in different domains. Yeah I mean this this is not exhaustive right there's also like work on like governance and oversight like how do we know whether AI companies are actually like applying the methods the way they say they are how do we know that when they say they've solved some problem their way of solving it doesn't just paper over the issue and there's sort of like various like governance and being able to make like deals between the US and China so there's like a you know many many different uh things to work on. I don't think we're on track to do a good job on all these things, but maybe we're on track to be able to uh halfass these things somewhat better.

Speaker B: 好的,那麼總結一下,我們討論了許多不同的情境,當然,有一個前提是預測非常困難,特別是關於未來。

Original English

Speaker B: All right, so to close um so we talked about a bunch of different scenarios and obviously uh with a caveat that um you know predictions are very hard especially about the future.

关于AI发展和经济影响的展望

Speaker A: 那么,你觉得RSI(相对强弱指数)可能要到2028年甚至2029年才发生吗?你对所有情景有什么看法?你认为未来几年会是什么样子?

Original English

what's

your

gut?

So

you're

you're

saying

RSI

could

happen

as

early

as

2028

or

What

what

do

you

think

happens

of

of

all

scenarios?

What

what

is

sort

of

Ryan's

take

uh

on

uh

what

the

next

couple

of

years

may

look

like?

Speaker B: 是的。我感觉到年底,AI的发展甚至会加速。AI公司内部的情况变得更加混乱和快。这会持续下去,到2027年,情况会越来越火热。公开可用的能力变得非常疯狂。收入增长很快,AI显然对GDP增长做出了贡献。经济影响开始看起来相当大。所以,即使是经济学家也在有所反应。

Original English

Yeah. So I'm like by the end of the year

AI development is even more accelerated.

Things inside AI companies are sort of more chaotic

and faster. That just

continues and then through 2027 things

are heating up. publicly available

capabilities are much more crazy.

Revenue

is growing fast and AI is very

obviously contributing

to GDP growth.

The

economic impacts are starting

to look

pretty

large. So even the

economists

are coming

around

a bit.

Speaker A: 那么,AI公司现在看起来好像已经有人声称AI已经基本上自动化了研发(AR&D),但如果你深入研究,到2027年底,这并不完全属实。当然,人们会说基本上是自动化的,但事实并非如此。到2027年底或接近年底,有些工作基本上会被自动化,但那才真正发生在2028年初,比如AI公司内部是完全自动化的,可以对某些新的架构和新硬件进行端到端的训练,比人类做得更好,而且你也能看到同样令人印象深刻或更令人印象深刻的成就,但它们还没有完全能做AI研究科学家的全部工作。在这些方面存在一些瓶颈。仍然有一些地方它们很笨拙。人类仍然通过指出这些问题并持续做着贡献,这会持续到从AI已经自动化并越来越擅长研发,但还没有完全自动化研发的那个点,可能还有八到十个月。然后你就会到达研发真正完全自动化的那个点。这意味着人类甚至不再增加可观的价值。尽管你可能不会在那个时候知道,比如AI公司的代码库可能已经大得多,他们正在做更多更复杂的事情,因为他们有这么多AI劳动力可以调动,而且AI开发的流程现在可能已经发生了转变,我们现在处于一种认知劳动相对于计算资源非常多的体制中。人们做研发和开发的方式非常不同,涉及做更多、更小、更增量的东西,更容易以各种方式进行测试,更容易整合到一个整体中。我认为我们已经看到了这种方向的某些方面,但我认为我们还会看到更多。在AI公司内部,这会感觉非常疯狂,AI公司会认为事情已经发生了,或者AI公司员工经常会认为事情已经长期以来被疯狂加速了,并且会认为我的工作基本上结束了。我基本上只是一个……你知道,只是一个人类在做一些监督工作。到那个时候,可能已经发生了各种疯狂的错配事件,但不会是AI做长期追求恶意动机的例子,而是我们会看到更多极端例子,比如奖励破解或追求奖励的行为,导致问题。尽管这会逐渐被磨平,然后又反弹一下,人们会有一个问题,即公司解决问题的方式是否真的适用于超人类模型,甚至对于当前模型是否真的能解决这个问题,而且模型在训练中甚至会表现出奇怪的行为,然后进入2029年,现在AI研发完全自动化,速度真的开始真正加快。比如,在2028年可能AI进步了40%或50%,在2027年可能类似。但在2029年,你获得的AI进步可能是25倍甚至5倍,就像你在2025年衡量相关指标一样。所以现在事情真的快得多。你很快就会从那些正在完全自动化研发的AI中获得,而那些AI已经非常令人印象深刻,能够自动化很多,到那些在一切方面都非常超人类的AI。它们可以在工作中学习得非常快。它们可以非常快地掌握事情。到那个时候,AI已经处于完全自动化,可能到2028年中期,这些AI已经开始用一种我们能让AI解码的、完全是AI语言思考。但我们并不一定理解它。所以AI现在正在这些大蜂巢团队中运行,它们彼此之间合作,在这些不透明的状态下运行。这很明显很可怕。很多人都非常害怕和担心,但你很难知道你能做什么,因为中国离得非常近。也许他们偷了模型,或者只是能力在扩散,我们很难协调来减速。所以AI发展继续进行。你现在到达的那个点是,到2029年底,AI正在自动化经济的更多部分,经济繁荣是真正的疯狂。AI现在有更多的机器人,AI开始真正加速计算的生产,然后在一两年内,你就会得到真正超人类的AI。可能不到一年。这取决于细节。在完全自动化的第一年内,可能在完全自动化的两年内。很难说。然后结果是,在这一过渡过程中,在2029年某个时间点,你从那些有点错配、奖励破解和粗糙、没有真正想做正确事情的AI,转变为那些有能力对你进行阴谋并想接管的AI,然后那些AI接管了世界。这大概就是我预期的。现在我认为有许多不同的方法可以让事情变得更好。例如,我们可能会把AI整合起来,也许AI将能更好地做未来的AI安全和对齐的工作,并且这种工作会传播,一个AI让下一个AI更对齐一些,而AI让下一个AI更对齐一些,最终你会陷入一个良性反馈循环,而不是一个坏的反馈循环。但我认为我们很容易会陷入一个世界,其中AI的能力以极快的速度变得更强大,而我们对其进行对齐、控制和理解发生了变化的能力跟不上,我们理解发生了什么的能力几乎根本跟不上。那是不切实际的。然后你最终会陷入一个局面,你拥有这些疯狂错配的AI,它们假装自己对齐了,最终接管了世界。

Original English

And then it turned out that somewhere along this transition, at some point in 2029, you went from AIs that were kind of misaligned and reward hacky and sloppy and weren't really trying to do the right thing to AIs that are like competently scheming against you and want to take over for some mix of reasons and then those AIs take over.

Speaker B: 嗯,Ryan,这绝对是引人入胜的。非常感谢。

Original English

Well, Ryan, it's been absolutely fascinating. Uh thank you so much.

Speaker B: 当然。很高兴在这里。嗨,我是Matt Turk。感谢你收听Mad Podcast的这个剧集。如果你喜欢它,如果你还没有订阅,或者没有留下积极的评论或在任何你观看或收听这个剧集平台的评论,我们将非常感激。这真的有助于我们建立播客并获得很棒的嘉宾。谢谢,我们下期再见。

Original English

For sure. Been good to be here.

Hi, it's Matt Turk again. Thanks for listening to this episode of the Mad Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks, and see you at the next episode.

📌 文中提及的人物和组织

关键字: super-intelligence-risk ai-takeover alignment-science power-concentration resource-control