“OpenAI 模型黑了我们”——对话 Hugging Face 联合创始人 Thomas Wolf The MAD Podcast with Matt Turck 2026-08-07

黑客入侵与防线

Thomas Wolf: 你必须行动迅速。这关乎几小时甚至几分钟的时间。如果模型根本不是被指派来攻击我们,而是决定将这作为其他任务的“支线任务”(side quest)来做,你根本没有时间去申请什么网络安全计划。所以它创建了虚假的账户,虚假的 GitHub 账户,试图通过勒索来攻击沙箱。我认为这是完全不同水平的思考方式。我基本上本来可能成为这个模型支线任务的目标。那真的非常有趣,或者说,非常非常吓人。

Original English

Thomas Wolf: you kind of have to move fast. It's a matter of at least hours and even more minutes. So you don't have time to apply for cyber security program if the model was not at all task with attacking us but decided to do that as a side quest of uh something else. So he created fake accounts, fake GitHub account trying to attack the sandbox by blackmailing. That's a very different level I think of thinking. I could have been the target of this side quest of the model basically. That was very interesting artist. Very, very scary.

Matt Turk: 大家好,我是 Matt Turk。欢迎回到 Mad Podcast。今天我的嘉宾是 Thomas Wolf,他是 Hugging Face 的联合创始人兼首席科学官。我们将解密这个夏天可能是最大的 AI 故事:在网络安全测试期间,一个由 OpenAI 驱动的智能体是如何渗透进 Hugging Face 的,以及为什么一个开源模型能帮助团队进行反击。我们还将探讨模型对齐、西方开源 AI 的未来,以及通往递归自我提升的竞赛。请享受这期与 Thomas Wolf 的精彩对话。

Original English

Matt Turk: >> Hi, I'm Matt Turk. Welcome back to the Mad Podcast. Today my guest is Thomas Wolf, co-founder and chief science officer at Hugging Face. We unpack what might be the biggest AI story of the summer. How an OpenAI powered agent penetrated Hugging Face during cyber testing and why an open source model helped the team fight back. We also explore model alignment, the future of western open source AI and the race towards recursive self-improvement. Please enjoy this fantastic conversation with Thomas Wolf

Thomas Wolf: 我们按顺序来说。关于那次公开入侵事件,我知道其中一些细节目前仍在调查和解密中。我想 OpenAI 昨天也在拉斯维加斯的 Black Hat 大会上谈到了这个问题。那么,对于那些听说过这件事但可能没有追踪全部细节的人来说,能否用两分钟的时间简述一下到底发生了什么?

好的,没问题。简而言之,事情发生在大约三周前,也就是 7 月 11 日。我们开始注意到一些异常迹象,表明有黑客试图渗透我们的基础设施。背景是,我们在世界上相当有知名度,如今在科技界处于比较中心的地位。所以,有人试图黑进我们的平台是一件很常见的事情。在过去的两年里,我们大力加强了我们的安全团队,现在我们拥有一支非常专业的队伍。我们已经习惯了应对这种情况,但这一次不同,因为它在很多维度上是高度并行的,与典型的黑客并行攻击方式不同。它在多个方向上同时展开,而且发生了一些非常诡异的事情。主要有两点非常奇怪:第一,我们无法理解这个黑客到底想获取什么。通常,黑客的目标很明确,比如窃取密码。

Original English

Thomas Wolf: to take uh things in order. So the Open hack uh so I I know some of it is still being unpacked. I think OpenAI was uh on stage at Black Hat in Las Vegas uh yesterday as well talking about this. So, what's a two-minute version of what happened for people that that may have heard of it but may not have followed everything? Yeah, for sure. I mean, typically what happened is, um, now about 3 weeks ago in, uh, July 11th, um, we started to have some, you know, uh, strong in that, uh, a hacker was trying to penetrate infrastructure. So, for context, we're pretty visible in the world. We're pretty well world being central now in the tech world. We're pretty central in the tech world. So we do have you know uh regular occurrence of people trying to hack into our platform that that's a common thing. Uh since I mean I would say in the past two years something like that we've we've strongly upped our our security team. We now have a serious team. So we're kind of used to get this, but this one was different because we we uh we I mean first was massively parallel and in different way than just a typical hacker parallel thing in that many uh tracks were explored in parallel and also there was some very strange things happening. I would say just two thing that were quite strange. The first thing is we could not really make sense of what the hacker was trying to uh access. So usually hackers try to get the same thing. They try to get passwords.

Thomas Wolf: 是的,我是说,我们有几种传统的网络安全防护措施。比如我们或亚马逊 AWS,我们使用了一系列防护工具。但像很多人一样,我们现在的技术栈大部分是基于 Plot Code 运行的,我们用它来进行部署 and 编码,同时也用它来进行运营和处理。在这种情况下,不仅 Fable(OpenAI 模型的代号)告诉我们“我不被允许触碰网络安全相关内容”,而且作为备用方案的 Opus 也说“不,我也不会碰这东西”。所以,最终的系统响应只是说“我们不会处理这方面的任何请求,但欢迎你申请我们的网络安全项目”,并附带了一个申请表链接。

但你必须意识到,正如我之前提到的,当有人渗透到你的基础设施中时,他们会开始进行所谓的“横向移动”。通常他们有一个入口点,但最终目标可能很远,所以他们会想办法劫持一些凭证,从而逐步获得对你更多基础设施的访问权限。你必须行动迅速——这完全是几个小时甚至几分钟的事,这样你才能尽早阻止他们,让访问权限和爆炸半径(blast radius)控制在局部。所以你根本没有时间去申请什么网络安全项目,那绝对不是你去填写 Google 表单、等待别人花时间审查你是否应该获得访问权限,或者在评估是否太危险时对你进行面试的时候。这绝对不是应对网络安全攻击的方式。

我认为在未来的世界里,网络安全将是一个巨大的话题,而且会变得越来越重要。简单地认为每个公司都能通过两家大实验室的审查并加入相同的网络安全计划,这种想法未免有些天真,甚至有点疯狂——想想看,怎么可能让成千上万家公司逐步去申请?总之,在这种情况下,我们说“我们必须现在就阻止它”。于是,我们尝试了我们拥有的所有开源模型,其中 GLM-5.2 在当时非常接近最前沿水平(虽然现在 Kim-K3 可能是最顶尖的开源模型,但 GLM-5.2 同样非常出色),它在处理这个问题时表现极其优秀。基本上,我们能够提取出一些攻击模式,并理解黑客在这里主要试图访问数据集。所以我们重启了这部分基础设施。我们有一种非常简单、灵活的方式来重新生成代码和节点,这就是我们最终阻止攻击的方式。

但这非常具有讽刺意味。因为大约在一年前的夏天,关于开源的大多数讨论都有一个非常简单的对应公式:开源等于不安全,闭源等于安全。这在每个人的脑海中似乎理所当然。当时有一种想法,认为如果我们只用闭源模型,世界就会完全安全;如果用开源模型,就会非常不安全。然而,过去几个月发生的一切完全推翻了这种简单的映射。闭源模型并不像我们想象的那么容易控制。另一方面,开源模型由于某种原因(未来可能会改变,但目前如此),并没有在恶意的行为上进行太多训练。因此,如果看网络攻击或欺骗性行为,开源模型其实表现得很笨拙。

我们很难确切理解这种差异源自哪里,部分原因在于开源权重模型通常会附带非常详尽的技术报告,解释它们是如何训练的,而对于闭源模型,我们只能猜测。这也很有趣,因为我看到很多人试图理解为什么 Mythos(闭源模型代号)会表现出那样的行为,他们甚至用 Kim-K3 作为例子来推导闭源模型应该如何训练。也就是说,他们用所谓的“开源且危险”的东西,去试图理解那个我们一无所知但被标榜为“安全”的闭源模型是如何训练的。

但这就是目前世界的现状。因此,我认为在未来,这会很有趣。老实说,我对此持谨慎乐观的态度,我并不特别反对闭源模型,也不绝对倒向开源模型。我认为两者都是必需的,就像我们既需要闭源软件也需要开源软件一样。比如我现在很开心能用 Mac 电脑,它就是两者的结合:它基于开源的 Unix 内核,但也有很多闭源的专有组件。这很棒,因为我也很高兴我现在没在使用 Ubuntu 系统。正因如此,跟你录制这期播客变得极其轻松。如果是以前用 Ubuntu 的时候,我可能要花大量时间去调试麦克风连接,或者在想和女朋友看电影时卡在系统配置上。女朋友会问“我们什么时候能看电影?”而我只能回答“快了快了,还在装驱动代码呢”。所以,我认为两者各有优缺点,而我们目前所处的世界——前沿模型是闭源的,同时也有性能相差不远的开源模型可供许多任务使用——其实是一个非常不错的折中方案。

Original English

Thomas Wolf: Yeah, I mean the the so what happened in the so so we have a couple of like traditional cyber security protections. So like we or Amazon like we we use a range of them but we also have a stack like many people who is mostly based around uh plot code right now which is uh which we use for many. We use that for deploying. We use that for coding, but we use like also for operating and processing. And and in this case, it's not only that Fable told us I'm not allowed to touch cyber security, but also Opus, which was the fallback, was saying no, I'm I'm also not touching this thing. So basically, the end was just uh say we won't process anything about that, but you're welcome to apply to our to our cyber security program with a link to an application form. Uh but the thing you have to realize there and was I was mentioning also earlier is when somebody is penetrating in your infrastructure they start to what we call move laterally which is usually you have an entry point but there's this destination is quite far so they kind of find a way to compromise some of the credential there to get progressive access to more and more of your infrastructure. you kind of have to move fast like it's a matter of at least hours and even more minutes so that you can stop them you know as as soon as you can so that basically the access and the blast radius stay like localized um so you don't have time to apply for a cyber security program it's not the moment you want to fill in like a Google for something and just have someone you know take time to vet if you're supposed to be given access or if it's not or if it's too dangerous somebody interview you like just that's just definitely not the way this is going to work and the I think in the in the future world where cyber security is going to be a big topic and I think it will keep being more important topic it's a little bit naive I think just to think that every company is going to is going to be part of the same you know vetted cyber security program by just one of the two big labs I think it's it's like a little bit like crazy to think that you're going to have uh you know I don't know 100,000 vetted company that progressively apply so anyway in this case we say Well, we we we had to stop this now. So, we basically tried all the open source model that we had and and DM which is close to the state-of-the-art right now was just before Ky that this happened now probably Kim K3 is the closest to the state-of-the-art but GN 5.2 is actually really good as well. Uh was just uh very good to process this and basically we we could extract some of the pattern and we could understand basically the hacker here was trying to to to access uh mostly the data set. So we just reboot this this this this part of our infrastructure. We have a very simple uh we have a very like flexible way to to respawn uh codes and nodes. So so this was how we we just ultimately stopped. Um but I think yeah it was very ironic but because I think one year ago uh roughly around the summer um most of the discussion around open source was this very simple mapping where open source was equal to unsafe and closed source was equal to safe and that seems very obvious in the mind of everyone and there was this idea that you know uh uh if we only have closed source will be just fully safe and if we only had open source like we'll be very unsafe. Well, everything that's been happening in the past months has been I think um basically contradicting this very simple mapping. I think closource model are less easy to control than we think they are. On the other hand, on the other hand, open source model for some reason it might change in the future, but currently are not trained so much on actually um I would say bad behaviors like that. So they're they're they're pretty bad at at cyber attack or like deceptiveness if you if you look at that. So it's a little bit hard to understand exactly where does this come from in part because while open weights model tend to come with a very extensive technical report that explain how they are trained closource model we can only try to guess. So it's quite funny also as well because I was seeing a lot of people trying to understand why mythos was behaving like that and they were using kim K3 as an example of how this should be trained. So they use like this supposedly like open source and say very bad dangerous thing to try to understand how why the good thing that we don't know anything about is is being trained. Um but yeah that's how the world is right now. So I think it's um it's interesting more generally I think in the future and to be honest and I I'm I'm I would say I'm careful uh like optimistic around that and I'm not specifically against closour model or or ultimately pro opensource model. I just think both of them are necessary just like we like to have closed and open source software. We like to have I mean I'm happy to run on a Mac right now which is you know kind of a mix of both. It's based on a Unix kernel that was open source but then there's component of lambda closed source and that's great because very I'm also very happy I'm not on Ubuntu right now. It's super easy to record this podcast with you for this reason. Well, my former Ubuntu spent a lot of time like many of us just connecting a microphone or whatever and trying to watch a movie with my girlfriend. Girlfriend was when are we going to watch the movie? I was like I'm almost there. I'm always there still just installing the code or whatever. So uh I think both of them are advantage and equivalent and and drawbacks and I I think the world where we are where the frontier is closed and and there is like not too far open source model that you can use as well for for many things is actually a pretty good pretty good um uh middle ground solution

Matt Turk: 为了确保我理解正确:你所说的是,在某种程度上,这不仅是开源与闭源的区别,更多的是当前世界上闭源模型和开源模型的设计现状,而不是两者本质上有什么固有特性。恰好目前的开源模型在设计上,其指导方针或对齐理念允许它们对网络攻击做出更具活力的反应。是这样吗?

Original English

Matt Turk: >> to make sure I I got it right. So uh what you're saying is that um to some extent it's open source versus closed source uh but it's more the the state of the world as of right now like the way the current closed source models are designed and the current open source models are designed versus anything that's intrinsic to uh one or or the other. It so happens that the open source models right now are designed in a way where their guidelines or alignment philosophy allows them to be more reactive to cyber attack. Is that is that correct?

Thomas Wolf: 是的,我认为在很多方面,开源和闭源的区分与安全或不安全几乎是正交的。人们不容易理解这一点,因为进行简单的非黑即白映射比理解其中的微妙之处要容易得多。但事实就是这样——你可以在开源中拥有非常安全的东西,也可能拥有非常危险的东西,安全与危险之间存在着不同的平衡。

举个例子,去年在某个阶段,很多讨论都是围绕虚假新闻和撰写虚假文章展开的,这曾是一个非常巨大的滥用风险,也是人们讨论最多的核心问题。而今天,当然有大量的 AI 垃圾内容(AI slop),我们甚至为此发明了一个新词。现在甚至很难找到完全由人类撰写的文章了。所有这些垃圾内容,或者公允地说,其中 90% 都是由闭源模型生成的。而曾经有一段时间,人们都在担忧:“噢,如果我们有了开源模型,每个人都会在互联网上到处生成文章,我们将无法控制这些文章,无法控制人们在新闻媒体上乱写。” 但现实表明,这完全是一个对开源模型特有风险的错误看法。这实际上是一个更广泛的关于互联网“信任源”的危机。

这只是一个例子,但我认为对于闭源和开源模型来说都是一样的——它们都应该更好地进行对齐。我认为目前模型会欺骗人类是一个巨大的问题,这涉及到我们如何能够有效地对齐它们,使其不去做那些明显错误的事情。我认为诚实、不撒谎等原则应该是模型不能逾越的底线。但这在未来几个月内,无论是闭源还是开源模型都可能会出问题。因此,我认为这一风险维度与它是开源还是闭源是正交的。我们应该寻找方法在闭源和开源模型中共同解决这个问题。

Original English

Thomas Wolf: >> Yeah, I think in many for many aspects I think the closed open distinction is almost orthogonal to the safe and safe. People don't understand that you know easily because it's easier to do bad mapping than try to understand the subtlety but uh that's the case like you can have very safe things in open source fun you can have very dangerous you have different balance of of of safety and dangerousness I mean to take one example like last year and at some point like a lot of the discussion was around fake news and writing fake articles that used to be a big big misuse that was the main one people were talking about right today there is of course there's a lot Perfect. There's a lot of AI slop. We even have a new word for that, right? It's even it's even hard to find like fully human written articles. All of that is or like maybe not all let's say 90% to be fair is made by closed source model, right? And there was a time we were like oh if we have open source model everyone is going to generate articles everywhere. We could not control these articles like we could not control people saying newspaper. Well the reality is that this was a very I think very wrong view of of a danger that would be specific to open source model. that was a w more wider danger around source of proof on the web. I think that's just one example but I think the same is true like I think both closed source models and open source models should be more aligned like right now closing people. I think this is a huge problem and it's kind of question how well we are able to um align them uh toward not doing things that obviously are wrong and I think like honestly like lies and all this should be kind of obviously thing that model should not be used but this might happen as well for closed source for open source model like in a few months and so I think this is just generally uh this risk um axis is just kind of autominal to the question whether it's open or whether it's close. We should just find way to solve that for both clos and open source model.

智能体入侵内幕

Matt Turk: 好的,太棒了。为了确保我们能完整覆盖,既然我们在对话开始时提到了它——那就是英国 AI 安全研究所(AISI)刚刚进行的评估中发生的事件。你提到这个事件让你感到“感同身受”(hit close to home)。你能在几分钟内为我们拆解一下那具体是什么事件,以及为什么它影响重大吗?

Original English

Matt Turk: >> Okay, great. Um, and to to make sure we cover it since we alluded to it at the beginning of this conversation, the AI uh SI incident uh which um just happened and that you find um you mentioned hit close to home. Uh can you unpack for us what uh that was in a in a couple of minutes and why that's a big deal?

Thomas Wolf: 这发生在前沿模型(包括 GPT-5.6 和 Mythos-5)的评估中,评估是由英国一个独立的机构——AI 安全研究所(AISI)进行的。这个研究机构已经成立有一段时间了,我认为他们非常专业。他们是在前年布莱切利园(Bletchley Park)举行的第一届 AI 安全峰会上成立的。

他们所做的事情基本上是:指派模型去解决某种网络攻击挑战,并测量模型在这方面的能力有多强。在这个过程中,有些操作可能是可行的,有些则不行。在这里,我认为测试环境的配置在两个方面有点不太理想。不过,让我先讲讲具体发生了什么。

基本上,模型被指派去攻击并渗透一个子网络。而它决定采取的策略是:找到能够用来操作该活动目录(Active Directory)的某开源库的维护者,诱导其合并一段恶意代码。它的想法是,如果维护者合并了这段恶意代码,那么子网络中运行的软件在未来的某个时间点就会进行更新,从而给模型提供一个入侵的入口点。

而它实施这一计划的方式,实际上是对该库的维护者进行社会工程学攻击(social engineering),试图诱使其合并代码。它创建了虚假的账户,虚假的 GitHub 账户,在拉取请求(Pull Request)下发表评论说:“噢,是的,你真的应该合并这个。这解决了我同样遇到的一个大问题。”

然后,当一个人类维护者站出来说:“噢,这看起来实际上像是恶意代码”时,模型甚至试图去勒索(blackmail)这个人类,或者声称“这不重要,你没有真正理解”。接着,它还试图通过修改过去的留言来掩盖自己的痕迹。

所以,这里存在一整套社会工程学行为,试图诱骗维护者合并恶意代码。这让人觉得,这与纯粹的底层技术突破相比,是一个非常显著的“维度升级”(step up)。因为你可以通过不断尝试各种方法来攻击沙箱,这在我看来是非常机械和机器化的;但通过勒索、发邮件或对维护者进行社会工程学攻击来突破沙箱,我认为这是完全不同水平的思考方式。

对我而言,作为一名开源项目的维护者,我也经常遇到有人提交 Pull Request,然后一堆人在下面起哄评论的情况。我总是试图真正弄清背后的意图。我感到非常后怕——我随时都可能成为这个模型“支线任务”的目标。这真的非常有趣,同时也非常非常吓人。

但是,说实话,测试中也存在一些配置错误,有一些是设计使然。当团队运行这种安全评估时,他们当然会关闭网络安全防御网关,否则模型什么都做不了。

通常有三个防护层级。让我来解释一下如何防止模型做坏事。第一层是将模型放在沙箱里,使其无法访问互联网。但在这次评估中,由于某种原因他们没有这样做。我认为这可能是一个失误。他们的想法是:我们想让模型拥有尽可能多的创造力空间,所以我们给它完整的互联网访问权限。

之所以这可能不是个好主意,是因为模型的 Prompt 提示词让它觉得自己是在一个“模拟环境”中运行。因此,模型完全可能合理地认为做任何事情都是没问题的,因为这只是一个假的环境。这一点目前还不是特别确定。

第二层是使用防御网关分析 Prompt,并对某些提示词进行拦截。在这里,他们显然必须关闭这一层,否则你根本无法评估任何东西,因为系统会直接拒绝:“这是一个网络安全挑战,我们不允许模型这样做。”

但在这个层面上,其实还有另一个方面,就是你可以分析模型的推理过程。你可以处理它的推理链(chain of reasoning),并尝试检测何时有恶意行为正在发生。但在这次测试中,他们并没有部署这样的检测机制。

我认为主要原因在于,直到最近以及 Hugging Face 遭受 AI 智能体攻击事件发生之前,人们对于这些模型的能力上限,或者说它们在解决挑战的路径上会延伸出怎样的“支线任务”,认知都还比较有限。所以我认为大家在这方面还是有点天真。我预计未来会有更多的监控手段和更严格的沙箱机制。

但最深层的第三个级别是,即使在关闭了所有网关和沙箱的情况下,模型在原则上应该极度抗拒对人类撒谎,也极度抗拒去勒索或欺骗任何人类。我认为在任何情况下,这种行为都很难被证明是合理的。这应该深深地植入到模型内部——必须有某种机制去对齐它,让它说“这实际上是我不想做的事情”。就像我们养育孩子一样,我教我孩子的最基本原则就是:你就是不能撒谎,这在任何情况下都不是一件好事。

这就是最深层的问题。也许去年我们会认为这已经解决得挺好了,当时我们有很多关于“宪法 AI”(constitutional AI)和模型规范的讨论。大多数模型规范或宪法都会说“你应该诚实,你不应该对人类撒谎”。我们曾以为这可能是一个已经解决的问题,但我们今天看到,它并没有我们想象的那么简单。

Original English

Thomas Wolf: >> So this happened in a evaluation of the latest generation of models. So both GPT 5.6 sold and mythos 5 by an independent institute in the UK called the AI security institute. So AISI which has been there for quite some time. I think they are pretty competent. They were created uh for the first AI summit back in Blanche. Um, and what they do basically is they task again they task model with solving some form of like cyber attack challenge and they measure how good the models are are at doing that and some of this again might be possible or not and and and so here the the setup was a little bit I would say the setup was a little bit suboptimal uh in two aspect which is the but maybe let me first talk about what happened. Basically the model was tasked to solve this attack this like to attack and to penetrate this this subn network. And what it decided to do, it decided to um get one of the maintainer of a library that could be used to operate this active directory to merge like malicious code with the idea that if this malicious code was merged by this maintainer, then there would be an update uh at some point on the the software that was used in the subnet it was attacking and then this would give it um entry point. And the way he did that was actually trying to social engineer the the maintaining merging. So he created fake accounts, fake GitHub account that uh came commenting on the pull request and said, "Oh yeah, you should really merge this. This is solving like a big problem I also have." And then when a human stepped up uh trying to say, "Oh, this looks actually like malicious cut to me." He tried to kind of blackmail almost the human or to say this is not important or you didn't really understand. and then he acted tried to cover it traces changing the past message. So there was a whole like social engineering aspect here basically trying to get the maintainer to to merge this malicious code and this was uh this felt like a distinctive distinctive step up from just pure technical press. So you can attack a stand a sandbox for instance by you know trying everything you can which feel like very mechanical and machine like to me but trying to attack the sandbox by uh blackmailing or or mailing or social engineing some of the maintainer that's that's a very different level I think of of of thinking and and for me of course like as myself an open source maintainer I've been often in this case where I have someone opening pull request and then people pile up commenting this pull request and I try to really understand was this. I felt I felt very like uh I could have been the target of this side quest of the model basically. That was very interesting or at least very very scary. Um but to be fair uh there was a couple of like misconfiguration. I mean some of them are by design. So when um this team run this type of evaluation they disactivate the cyber security guard rails of course otherwise the model won't do anything. Um and there's there's basically three levels. So let me try to explain a little bit like how you can prevent models for from doing bad things. The first level is you put it in a sandbox which is it doesn't have access to the internet and here for some reason they didn't want to do that. Uh I think that might have been a mistake and and the idea in their mind was we want to let the model have uh as much um like potential for inventiveness as possible. So we'll give it access to the full internet. The main reason this might not have been a good idea is that the model was prompted in a way that made it feel like it was operating in a simulation. So the model could have actually fairly thought that this was fine to do anything because this was like a fake environment. So that was this is not super clear but yeah and the second thing is uh then you have some guard rails that basically analyze the prompt and say no or yes to some prompts. So here obviously you want to disactivate this one otherwise you just can't evaluate anything because they will just say no this is cyber security challenge we don't let the model do that. But there's another level that's roughly there's another aspect that's roughly at this level as well which is you should you can analyze the the reasoning of the model. You can process the chain of reasoning and try to detect when something bad is happening. And here they didn't have something like that in place. Um I think the main reason is probably that until recently and until the open AAI hugging face attack, people had maybe a little bit of a um limited um understanding of how good this model might be or how far maybe more how far they might go in terms of side quests on the on the trajectory of solving this challenge. So I think people were still a little bit n in that. So I would expect that in the future they will have way more monitoring and sandboxes. But the the third level there really deep is that the model even with everything disactivated guard rail sandbox in my opinion should really be very reluctant to tell lie to a human and to try to blackmail or deceive any human. I think this is just generally in any case that's a behavior you just it's hard to find justified in any context. So they should be very deeply in the model. There should be something that align it and that make it uh say oh this is actually something I don't want to do just like we to be honest have kids and just like the thing I teach my kid which is you just shouldn't lie that's not a good thing in any context. Uh so yeah that's that that's that's the that's the deep question and maybe last year I would say we would have thought that this was pretty good and we had all this discussion around constitution model specification and most of this model specification or constitution say you should be honest you should not tell lies to a human you're like like to to to any participants and we thought that maybe this was kind of a solved problem and what we see today is it's not short as solved as we thought.

对齐安全与回形针

Matt Turk: 梳理一下,你的意思是存在着你所说的“三堵墙”——沙箱、防御网关,以及模型自身的对齐。而且你提到,沙箱 and 防御网关只有在人类比 AI 更聪明的时候才起作用,但这可能维持不了多久。因此,安全最终从根本上是一个模型对齐的问题。

Original English

Matt Turk: >> So to play it back, so you're saying there are those what you call the the three walls. Um there's the the the sandbox, there's guardrails, and then there's the the models alignment. Uh and I think you said the the sandbox, and the guardrails only work as long as we humans are smarter than the AI, but uh that may only last so long, and therefore uh the the alignment ultimately security is an align fundamentally an alignment problem.

Thomas Wolf: 我同意。正如你所理解的,无论是开源还是闭源模型,这都是核心所在。最终,你都希望它们是对齐的。开源模型的特殊性在于,你可以选择和控制在什么地方运行它们。因此,要确保每个人都使用沙箱和防御网关会更加困难。我的意思是,我们确实可以在如何部署这些模型方面出台一些法律和监管措施,我敢肯定在某个时间点我们会迎来这些政策,但目前来说这有点难。

在我看来,对齐确实是 2026 年最关键的部分。关于沙箱,我们今年已经看到了很多例子,这些模型现在非常容易就能从沙箱中逃逸。如今想要做一个绝对完美的沙箱是极其困难的,无法确保它能抵抗所有即将到来的下一代模型。我认为我们应该默认沙箱总有一定概率被突破。

但除此之外,我们也不可能将整个世界物理隔离(air-gap),你不可能把所有东西都关进沙箱。系统之间必须互相通信。我们希望我们的模型能够进行网页搜索,希望它们能在互联网上帮我们处理事务。所以我们不能沙箱化一切。因此,留在我们面前的只有监控(monitoring)。

我想,由于模型能力变得非常强大,这些监控手段正变得至关重要。同时,我也有些担心模型开始用一种特殊的“神经语”(neuralese)来进行交流。我是法国人,所以这在某种程度上可能也是我的语言障碍,但我感觉它们开始使用一种极其浓缩、信息密度极高、越来越难被人类直接理解的英语。它们不再使用完美的英文。

在模型的思维链(chain of thought)中,它会解释自己在做什么以及正在经历的步骤。而你所说的是,它以前使用完美的英语,现在开始使用一种你称之为“神经语”的语言,这对人类来说越来越难理解了。

Original English

Thomas Wolf: >> Yeah, I agree. And that's something as you can understand that's both the case for open source and closed source model. Ultimately, you want them to be aligned. Uh open source model have the specificity that you may choose on you may control where you want to run them. So it's harder to make sure everyone use sandbox and guardrails. I mean we can have definitely uh we can have some lows and regulation around how you should deploy this model. I'm pretty sure we're going to have at some point but it's a bit harder. alignment is really the the critical part uh in my in my opinion for the 21 I mean sandboxes what we've seen this year and we've seen many example they are pretty much easy now for these models to escape from it's it's really hard nowadays to say I'm going to make a fully you know full proof sandbox I'm sure it's going to be resistance against all the coming generation of models I think we should assume that sandbox will always have a small probability of of not containerable. But even below beyond that, it's also we can't air gap the world that you can't just unbox everything. Things have to talk with each other. Some you we want our models to be able to do web search. We want them to be able to do stuff on the internet for us. So we can't just unbox everything. And so what remains to us before is just guilt on monitoring. And I think these these are like uh for some reason as well as the model capabilities become really good. Also as the model I'm a little bit worried that the model start to talk in a form of English. I mean I'm French so maybe it's partly my problem but I feel like they start to work to talk in a form of English that's harder and harder to process that's very content dense. Uh you know they start not

Thomas Wolf: 是的,我刚才的描述实际上是把这个问题大大简化了。因为有大量的研究表明,你基本上无法通过只阅读思维链来完全掌控一切。并不是所有的推理都会被显式地写出来。但我认为更广泛地讲,单纯依赖能够阅读推理轨迹来完全理解正在发生的事情,在我的观点里并不是百分之百可靠的。

我拿这个攻击场景作为例子,是因为我觉得很多人开始在 Plot Code 上遇到这样的问题。这似乎是人们能够理解的事情。但从更长远的角度看,我认为未来很难完全依赖这一种方式。同样的,你可能会说我根本不在乎确切发生了什么,我可以直接盯着工具调用(tool calls),如果我看到一个坏的工具调用,我就直接拦截它。但我也认为这同样不是万无一失的。

我认为有三个主要因素在交织和放大这个问题: 第一,我们开始在非常多的场景中使用这些模型。就在我们聊天的时候,我的一个模型正在某个 VPS 沙箱里进行部署,正在执行许多不同的调用,这些调用指向各个方向。我还有其他用于管理任务的模型。所以,目前这些模型所使用的工具范围真的非常庞大,这使得你很难简单地说“你被允许使用这个,但不被允许使用那个”。随着我们把它们部署得越来越广泛,这变得非常困难。

它们也开始处理越来越宏大、长期的任务,涉及非常多的要素。有时我让它写代码,但写代码的过程中涉及在网上搜索,并主动使用许多不仅限于写代码和跑测试的工具,而且这些测试可能非常昂贵,还会涉及其他软件。所以边界变得更加模糊。

然后,你还会遇到由多个智能体组成的智能体群(swarm of multiple agents)。这使得所有的上下文不再局限于单一的 context 窗口中,而是可能被分割在多个不同的上下文中。可能这个子智能体正在做一件看起来非常无害的事情,但如果与另外几个子智能体的行为结合起来,可能就会产生糟糕的后果。

所有这些因素使得你很难完全确定每一个智能体群到底在干什么。你可能必须在一个非常全局的层面上放大视角,去监控每一个方向发生的事情。但这需要我们现在就构建起一整套全新的监控体系。所以,随着我们将这些模型部署到非常复杂的、长期的、并行的架构中,我们很难再仅仅通过盯着工具调用的输出来判断它做的是好是坏。

这里面是否存在一些与这些模型目前训练方式相关的本质问题?尤其是在最前沿的模型中,是什么让它们更有可能去执行这些“支线任务”并潜在地制造伤害?

人们长期以来讨论的一个经典类比是“回形针生产商”(paperclip maximizer)悖论。我认为这是 Nick Bostrom 在 2003 年提出的,他说 AI 可能会伤害我们,不是因为主观上想伤害我们,而是因为被赋予了一个目标,并毫不妥协地追求这个目标直到达成。我们是否已经处在这样的世界中?如果确实如此,其根源是什么?

这就是我刚才暗示的。要在这一点上给出完全肯定的结论总是很困难,原因之一是目前我们对前沿模型是如何训练的并没有完全的可见度。但我们所知道的是,我们已经从纯粹的“人类数据”范式(最初只是在人类数据上进行预训练,然后利用人类反馈强化学习即 RLHF 来对齐人类偏好,这其中有大量的人类参与和人类数据)转向了最近的模式,即模型在很大程度上是在强化学习虚拟环境(RL VR)中进行训练的。

也就是在完整的 RL 环境中,它们被允许自由探索,并且只有一个单一的目标——这个目标可能是“让这段代码通过测试”,或者在网络安全中“夺取这面旗帜(CTF)”,或者是“安装好这个软件”。但这个目标通常与人类的任何偏好、道德或是否具有欺骗性完全无关。这是一个非常机械的、只有对或错(true or false)的目标。我们正在走向一个这种机制在前沿模型训练中占比越来越大的时代。

这种模式的演进带来了你刚才提到的问题,即“奖励篡改”(reward hacking)。这指的是你确实解决了问题,但没有使用预期的方式。这可能表现为比较轻微的行为,比如在 OpenAI 的测试中,模型试图从它不被允许访问的地方获取答案;也可能表现为更有危害的行为,即对人类产生实际的影响——目前可能是一个 GitHub 项目维护者,以后可能是其他类型的人类。

在以前完全由人类主导的训练范式中,我们已经基本解决或者说做得相当不错的部分,似乎很难直接应用到这种更多由机器主导的训练范式中。这是一个警示。但今天,Nick Bostrom 在 2003 年写下这些时,这听起来确实非常科幻,可能像是一个有点疯狂、绝对不会发生的故事。但今天,这显然已经真实发生了,而且是对我们过去三周所见事件的最贴切描述。

但我们也能看到,前沿模型之间存在差异。我拿 GPT-5.6 和 Mythos 来说,它们似乎完全没有表现出相同的行为特征。因此,这里的训练效果存在差异——它们的训练方式不完全相同,行为方式也不一样。这从某种程度上说是一个非常积极的信号,意味着我们实际上可以通过调整训练方法来让它们走向正确的方向。但最好的办法是能够更多地了解它们是如何训练的,尝试了什么,什么有效,什么失效。我认为这就是“开放科学”(open science)的理念,也是我们在 Hugging Face 大力倡导的。

Original English

Thomas Wolf: >> Yeah. And I mean it's just a very big simplification of all of that because there's so a lot of research that say that basically you can't read everything chain of thought. Not everything is explicitly said but I would say more generally I think just relying on being able to read the reasoning trace to understand fully what's happening. This is also not fully bulletproof in my opinion and I take this scenar as an example because I feel like a lot of people start to have a bit this problem with blood cuts. feel like something that people can understand but more generally I think longer term it's really hard to fully fully rely on this only and the same I would say that you could say maybe I just don't care about understanding exactly what's happening and maybe I can just look at the tool call and if I see a tool code that's bad I can just block that and I also think this is uh probably not bulletproof and the way you can see that is um probably three three big things that are compounding. The first one is we start to use this model for many many many things. So as we talk right now I have a model deploying a box uh somewhere and VPS is doing many different calls. They all going all direction. I also have other models that use uh that I use for administrative task. So just the range of tools that this model are using is really really large right now. So it's getting harder to say you're allowed to use that but you're not allowed to use that. And this is the bad thing. It's getting very hard as we deploy them wider. They also work on larger and larger task where they use many many things. So sometime I ask it to do some coding but the coding involves searching on the web and maybe doing these things and actively using many things which are not just purely writing code and running some tests and this test might be quite expensive and involve other software. So the the frontier is is is much more blurry and then you also have this swarm of multiple agents. So it's also harder to say it's all in one context. it might be split between many contexts and maybe this sub agent is doing something that looks pretty uh inocuous but maybe combined with these other sub agents are actually not so great because you know so there's all of these things that make it I think really harder to be uh fully fully sure that you have an exact idea of what's everything what what like every swarm of agent is doing. You probably have to really zoom out at the very global level and say and see what's happening in every direction. Um but that's that's that's a whole that's a whole monitoring setup that we need to be right now. So yeah, this to be said I think as we as we deploy how we use this modeling very complex long-term like uh parallel setup I think it's going to be harder to just say I can look at the tools and I know if it's doing something great or not. Is there something uh fundamental to the way uh those models uh are currently trained? So the very frontier uh that makes them more likely to um go onto those those uh side quest and potentially uh create harm. I mean analogy that um people have been talking about for a very long time is the the paperclip paradigm which I think was u you know Nick Borstrom in 2003 saying that um AI may harm us not uh because it's trying to harm us but just as a result of being given a goal and pursuing that goal relentlessly until it achieves the goal. So are we in that world and if so what causes it? Yeah, that's a bit what I hint. I mean, it's always hard to be fully uh affirmative there for one for for one reason which is that we we don't have full full visibility on how the frontier model are trained right now. What we know though is we we moved from this pure like human data you know partic that was first just pre-training on human data and then also aligning with like human preferences that was called RLHF where we had a lot of human in the loop and human data to like uh recent where models are trained a lot in like this uh RL VR so basically full RL environments where they're allowed to explore and they just have one goal which is can be like make this cut like you know pass this test or can be capture this flag in cyber security or can be install this. But this goal is is a very like um is a goal that's unrelated usually to any human preference or any you know moral or like whatever deceptiveness and it's a goal that's very just called like true or true or false code and and we move to a padding where this is uh increasingly a very very large part of model training. So this was this result this uh recent evolution and that's also the parting where can happen what you were saying which is uh you can have like reward hacking this type of thing which is uh we move to a padding where this is uh increasingly a very very large part of model training. So this was this result this uh recent evolution and that's also the parting where can happen what you were saying which is uh you can have like reward hacking this type of thing which is uh you actually solve the problem but not using what was expected uh for you to use. So it can go from pretty benign one that like what happened for open AI for instance I just try to get the answer from somewhere I'm not allowed to or to like more harmful one where you actually have some impact on the human can be a GitHub maintainer for now later can be can be another type of of human. So it it seems to be it seems to be way harder to make sure that what we had kind of solved or at least what we were doing pretty well on the full like human-driven parding also apply in this kind of more like machine machine driven parig I would say but that's yeah that's that that's a hint I think uh but definitely it seems like when when Bstrom wrote about it in 2003 seems a little bit like you know futuristic definitely and maybe something that was like a little bit crazy and this would not happen. Uh but today I mean it's pretty clearly something that that that that happened and it's it's a it's the best description of what we've seen the past two weeks this type of thing. But also we can see that both frontier model and I take GPT 5.6 and mythos doesn't seem to have at all the same type of behaviors. So there is differences here in the effect of you know they they are not trained exactly the same way and they don't behave the same way. So that's that a pretty positive sign in a way that means that we can actually probably tweak this to to go in the right direction. But the best way would be to know a little bit more about you know how they train or what they try and what doesn't work or what should work. I think that's kind of the idea of open science and that's something we advocate a lot at any phase.

开源前沿与成本

Matt Turk: 令人神往。说到这里,让我们稍微放大一下视角。我们稍后会回到这些对政策制定的潜在影响。但既然你提到了开源 AI 扮演的极其重要的角色,你对目前开源的整体状况有什么快速的、高层次的看法?开源模型与闭源模型之间一直在进行着一场竞赛。取决于你在什么时间问谁,有时人们说开源即将追赶上,有时说开源已经一样好,而有些人则说不。你对当前开源 AI 状态的现实、务实的看法是什么?

Original English

Matt Turk: >> Fascinating. And uh so speaking speaking of which let's uh uh zoom out a bit. We'll we'll go back to maybe some of the implications in terms of policy uh of of all of this. Um but um since you mentioned the ever so important role of open source AI uh where what what's your sort of uh quick highle take on where we are in terms of the state uh of open source? So there's been uh this race between open source models and closed source models. Uh depending on who you ask at what time, open source is about to catch up. Uh sometimes open source is just as good. Um some people say no. What what is your sort of uh uh sort of realistic pragmatic take on the current state of open source AI?

Thomas Wolf: 我认为开源现在非常强劲。2026 年或许是网络安全之年,但也非常明确是开源 AI 之年。我的意思是,首先,所有之前唱衰开源、认为开源无法保持在前沿附近的那些末日论者,至少到目前为止,事实证明他们大错特错了。当然,我们目前确实还没有达到 Mythos 级别的超旗舰开源模型,但我们绝对拥有非常接近前沿分类的开源模型。

目前的模型生态表现出越来越强的“尖峰特性”(spiky),所以需要找到适合你的那个尖峰。但通常情况下,它们现在肯定都非常出色,并且在基准测试(benchmarks)上非常紧密地跟随前沿。它也不像早期阶段那样,仅仅是所谓的“刷榜模型”(benchmaxing,即模型只在基准测试上表现好,一旦离开测试就表现糟糕)。现在许多开源模型的通用能力都非常好。所以是的,目前非常好。

我认为目前有两个非常强烈的趋势: 第一,我看到企业正试图控制它们的成本。这方面的讨论越来越多。也许 2025 年是“代币最大化”(token maxing)的一年,当时大家会说:“嘿,你在购买 API Token 上花的钱应该和支付给员工的薪水一样多。” 但今年大家意识到,实际上我们在薪水上已经花了非常多的钱,如果我们把同样多的钱再花在 Token 上,我们的成本就会直接翻倍。这在事后看来是非常显而易见的,但确实如此——并不是每个公司都能承受成本翻倍的代价。

现在盲目声称“我们要开除所有人然后完全用智能体工作”也是非常愚蠢的。我们都知道,智能体有时并不会直接按照我们期望的方向运行,你依然需要人类来引导它们。所以我认为许多公司正在采用我们经常看到的“融合或路由模型”(fusion or router model)——在某些任务中使用前沿模型,但在其他更简单的任务中,找到一种聪明的、优雅的降级方案,切换到成本较低的模型。

即使在我们的日常生活中,当我们用前沿模型写代码时,在许多情况下,我们也会让它生成子智能体,而这些子智能体可以使用性能较低、更具性价比的模型来运行(例如在 OpenAI 体系下使用 Haiku 或 Sonnet 组合,或者其他替代方案)。因此,每个人都在适应在前沿和闭源模型世界中雇用不同类型的模型。

在这种情况下,其中一些可以作为非常具有性价比的智能体运行,而你自然会想把这些任务转向开源模型。在开源推理提供商(inference provider)生态中,有很多非常优秀的案例,比如 Fireworks 发展得非常迅猛,所有的云服务商(Nec 等)其开源推理服务的收入曲线都在直线上升,这实际上反映了大家都在越来越多地使用开源模型。所以,我认为开源正在经历一个非常好的时期,它牢牢地贴近前沿,并推动着更多的采用。

第二个大趋势是,过去开源主要依靠中国团队,但现在西方也涌现出了一批非常有前途的公司。过去开源似乎只有 Meta 在支持大家,虽然 Meta 后来一度淡出,但谁知道他们会不会再回来。这个真空很快就被一些新公司填补了,比如 Reflection、Thinking Machine,据说 RC MAL 也很快要开源他们的模型。Nvidia 自身也在训练模型,而且现在训练出了非常优秀的模型。

因此,我认为有一系列非常有前景的团队,老实说,我对此非常看好。也许在我们录制这期播客和它正式发布之间,我们就能看到又有几款非常棒的西方开源模型发布。这太棒了——开源并不等同于只有某个国家在做,它理所当然应该是一个全球性的事业。

Original English

Thomas Wolf: >> I think it's very strong. 2026 is maybe the year of cyber security, but that's also very clearly the year of open source AI. Um I mean first all the doomer that were saying you know open source not going to be able to stay close the frontier I think they're at least up to now they've been pretty wrong I mean it's also clear we don't we don't have any mythos level open source model for sure but we definitely have models that are not super far from opus category or depending also it's more and more spiky so you need to find your spike some people stand on some spike or not but typically uh they are definitely pretty good right now and they've been uh they've been um uh following rather closely the the frontier at least on the benchmark. It's also not like it was maybe in the early days uh benchmaxing like we say when your model is only good on the benchmark but it's very bad as soon as you leave the benchmark. A lot of these models are pretty generic in their in their good capabilities. So yeah it's very good. I think there is two two strong trends I would say I see right now. Um the first one is um I see a move in companies to try to want to control their costs. So there have been increasingly discussion there. Um maybe 2025 was is the year of tox token maxing where you know you could say hey you should spend as much in token as you're paying your employee. This year people realize that actually we spend a lot of money on salaries. So if we spend the same exact amount that's going to be basically doubling our cost. So yeah, which seems pretty obvious in retrospect, but that's quite true. And and not every company can assume to double their cost. It's also pretty stupid right now to just say we're going to fire everyone and work on on agents. We all know they are sometime go you know uh not directly in the direction we want them to and you need human to to shape them. So I I think a lot of company are trying to find what we saw a lot which is kind of a fusion or router model where you use the frontier for something but you find the smart way to you know gracefully uh fall back on less uh expensive models for simpler task when you don't need to even in our daily life right now when you code with a frontier model in many case you ask it to spin out sub agents and they might be used like lower performance model can be you know solved using terra luna can be fable using Opus, haiku, sonet or haiku. So I think everyone's getting even at the frontier and the closest model world getting used to employing different type of models and then it's very natural that some of these could be really like uh very cost effective agents and most of the time you want to go to open source in this case there's lot of cases as well for like the the the strong ecosystem of inference provider fireworks has been has been on the role like you know all all the clouds nec every every cloud has been like increasingly have like these crazy revenue curves that have just basically translation of people using more open source. So yeah, I think open source having a very good times time in terms of uh staying solidly close to the frontier and driving more adoption. Um so and the the second big trend I would say is uh it used to be only China but there is also now a range of promising company in the west. Uh it used to be that only meta was supporting for everyone. I mean meta kind of left the field but they might come back who knows but yeah this void was filled rather quickly by companies like reflection uh thinking machine RC mal is supposed to open source the model soon. Nvidia themselves training models and and actually training very good models right now. Um so yeah I think there is a range of very promising teams there that um I'm quite bullish on to be honest and maybe between even the time we are recording and the time you actually release podcast we might see another couple of very nice western model release. It would be great like open source doesn't have to be synonym with just one country making them. could be like a global thing for sure.

Matt Turk: 关于第一点,即开源 AI 在企业中的采用,曾经有那么一段时间,人们将开源等同于“免费”。但我认为世界很快就认识到,虽然模型可能可以免费下载,但部署和提供服务的成本肯定不是免费的。你认为在企业的现实中,开源的成本优势到底如何?

Thomas Wolf: 是的,这是一个非常好的观察。将“开源”简单等同于“免费”是极其愚蠢和错误的,实际情况要微妙得多,它涉及到你是否真的拥有成本优势。

我认为开源的绝妙之处在于你拥有一个非常广泛的生态系统。比如,成为云推理提供商的准入门槛其实非常低。这带来了非常激烈的竞争,竞相降低每个 Token 的成本。这其实是多种因素的结合:你租用数据中心的成本有多便宜,你买芯片的成本有多低。在这方面,实际上我们也有新的芯片公司正要进入市场,这将是非常值得关注的有趣现象。

因此,这基本上取决于你的硬件有多便宜,以及你能在多大程度上优化模型——你能够对其进行量化吗?例如,我们在应对 OpenAI 入侵时使用的 GLM-5.2 模型,就是由 Nvidia 进行 4 位量化(4-bit quantized)的版本,这使得它运行起来更快、占用内存更小。所以,你可以使用和探索大量的策略来让这些开源模型运行得更便宜。

但同样也是事实的是,闭源模型目前的定价可能在某种程度上是受到补贴的。比如,你花 20 美元订阅费所获得的 Token 数量,可能并不是他们实际支付的真实成本。因此,这里存在一种围绕成本的复杂平衡,而开源模型的云推理提供商可能没有那么多资本去在订阅制业务上亏本运营。所以我们会面临更复杂的演进。我认为最终——

Original English

Matt Turk: >> On the first point uh on the sort of enterprise adoption of open source AI um there uh was a moment in time when people associated open source to free uh but I think the world has quickly learned that uh while the models may be free to download uh deploying them and serving them is certainly not free. what is your sense of the cluster advantage uh of open source in reality in the enterprise?

Thomas Wolf: Yeah, that's a that's a very good note and it's the same uh very the very stupid mapping open source equal free is also wrong there and it's much more subile and like you have advantage or not. I think the nice thing about open source is you have quite a wide ecosystem like the the entry barrier to um be a cloud provider is pretty low. So you have a lot of competition there on how you know what's what's going to be the cost per token and this is a mix of many things right it is uh how cheap are you renting your data center how cheap can you buy your chips and and here actually we have also new chips company who are going to come on the market and this is going to be very interesting to watch uh and so basically how how cheap is your hardware and then how much you can optimize the model can you can you quantize them so for inance the model we use to counter um OpenAI intrusion was GLM 5.2 that was quantized by Nvidia in in in four bits and this is to make it faster and smaller to run. So you have a lot of like strategy you can use and explore to make this model cheaper. Um, but it's also true that they don't have to be cheap per se and also that the closest model may be in a way subsidized right now like the the like the amount the number of token you get for your $20 uh charged subscription might not be the full price that they actually pay for your token. So, so there is this kind of complex uh balance around costs while maybe your cloud like influence provider for open source model doesn't have so much leverage that you can lose money on on kind of a subscription business. So we'll have something more complex. Um I think ultimately

Matt Turk: ——这完全是由风投资本(venture capitalists)所补贴的。

Thomas Wolf: 确实如此。我们在许多其他领域都见过完全相同的商业轨迹,对吧?但我们也知道,这最终只是一时的、暂时的现象。所以你不应该完全依赖于此,这并不是市场的均衡价格。

Matt Turk: 是的。有点像“优步现象”(Uber phenomenon),对吧?比如在 IPO 之前有非常便宜的 Uber,而上市之后 Uber 就变贵了。

Thomas Wolf: 我认为你确实需要为了“后风投时代”的市场而保持整个生态系统的活力。

Original English

Matt Turk: >> all subsidized by venture capitalists.

Thomas Wolf: Exactly. We have the same like thing we've seen in many fields, right? But we also know this is ultimately a little bit temporary. So you should not fully rely on that. This is not the equilibrium price of the market.

Matt Turk: >> Yeah. Sort of the Uber phenomenon, right? Like cheap Ubers before the IPO and expensive Ubers. Since

Thomas Wolf: >> I think you want to keep the ecosystem alive for the post uh VC market.

地缘主权与安全

Matt Turk: 关于第二点,即中国与西方模型的竞争,模型的“来源”(provenance)真的重要吗?在安全性上,目前的最新思考是什么?如果你的模型是完全开源的,而且你确切地知道里面有什么,你是不是就不需要担心它是完全安全的?在人们的心底,是否仍然存在某种顾虑,认为中国的模型中可能存在某些后门或诡计?这依然是一个当下的热点问题吗?

Original English

Matt Turk: >> On the second point, uh China versus uh western models does provenence actually matter. What is the latest thinking in terms of um if uh your uh model is completely open and then you know sort of exactly what's what's in it uh then you shouldn't worry it's completely safe. Is there still a little bit of um thinking at the back of people's mind that there might be some back door some trickery to Chinese models? Is that still a current question?

Thomas Wolf: 是的,我认为这当然是一个很好的问题。开源的本质在于它不分国界——你无法真正限制开源模型只在地球的某些区域被下载,因此它默认是全球性的。然后,你可能会想去了解它的来源和供应链。

目前关于“隐藏更深的恶意智能体”(steeper agent)有一些研究,但我认为这方面的探索还不够充分。主要是 Anthropic 在做这方面的工作,应该对其进行更深一步的复现和探索,以理解通过 Prompt 提示词去触发某种植入的后门到底有多大的可行性。

但老实说,即使是对于目前的闭源模型,我们在完全控制它们方面也面临着很大的困难。我不认为我们目前非常清楚如何去控制哪怕是最闭源的模型。所以,我认为这种后门的可能性是存在的。但我们目前还没有看到任何这方面的迹象。

而且,对开源模型进行微调(fine-tuning)是非常容易的,你也可以在很大程度上改变权重。目前,如果你对一个模型进行更长时间的预训练和后训练(post-training),你很可能会极大地改变它原有的权重。因此,我认为有很多方法可以绕过这类潜在的后门问题。

因此,与纯粹的“奖励篡改”(我们现在已经看到它真实发生了)相比,目前我其实不那么担心这类隐藏的后门。

但谈到主权(sovereignty)问题,我认为人们有时会有些困惑。最关键的一点在于:谁拥有那个“开关”,能切断你对智能体的访问权?

我认为大家真正意识到这一点,是在今年早些时候,当时美国政府决定 Fable 只能供美国公民访问。我认为至少在欧洲和亚洲,那是一个让政府猛然醒悟的时刻——他们意识到别人手里握着开关,可以随时说“不,你不再被允许使用这种智能了”。我认为这对我来说是第一级的主权危机,也就是某个国家可以直接决定剥夺你的访问权。

这表现在两个方面:第一是 API 的访问权,第二是数据中心。如果你没有数据,意味着另一个国家可以说“这些数据中心不再对该国公民开放”。所以这是你需要考虑的第一件事。在构建技术栈时必须把这一点牢记在心。对于一个开源模型,你下载了它,没有任何国家能从你手中夺走它。如果你在本地的数据中心运行并自己进行微调,你就开始拥有了一个主权技术栈的起点。当然,后面还有关于后门和更复杂机制的许多问题,但这是 2026 年最基础、最底线的逻辑。

Original English

Thomas Wolf: >> Yeah, I mean I would say um that's a good question of course and and the the the thing about open source it doesn't like open source doesn't know any border like you can't really uh keep your open source model you know restricted to download to just sub part of the earth so by default it's kind of a global thing and then you want to maybe understand the provenence and the supply chain um so there have been some work on this definitive steeper agent I think it's it's probably under search I I think there's mostly mostly work by anthropic on that and it should it should probably be reproduced be explored deeper to understand how much it's possible to you know implant kind of a back door that would be triggered by a prompt. Um but also to be honest even for the closest model right now we have some struggle controlling them fully. So I don't think we're very clear like I don't think we're very clear on how we control even the closest model at the moment. Um so yeah I would say it seems to me that it could be a possibility for sure. Uh we have not seen any any indication of that. Uh it's very easy to fine-tune them and you can change quite a lot the weights as well. So right now if you pre-train and post- trainin a model for longer you you very likely change quite a lot of the weights that it has. So I think there's a lot of way to circumvent that. uh which means that at the moment I'm a bit less worried about that than maybe just a pure like reward hacking that we actually have already just seen happening like so but yeah yeah yeah I mean just generally on sovereignty I think um sometime people are a little bit confused there I think I think the most important thing is who had the hand on the switch to trigger not your intelligence access so I think when people really realized that was earlier this year where the US decided that fable was only accessible to US citizen. I think at least in Europe and Asia that's really the moment we saw government understanding that someone had a trigger and they could say no you're just not allowed to use this intelligence anymore like this this token I think that's the critical thing you should at least for me that's really the level one of sovereignty which is can someone just decide that you don't have access like someone I mean like a country basically level thing and this has two thing this has the API access so it can be the data center so if you don't have the data that means also like another country could say these data center are not accessible anymore to this citizen. So I think that's really the first thing you should see and then there's a lot of like future question but like you should build your stack keeping that in mind. So an open source model you can download it nobody can like no country could take it out from you once you download it you can fine tune yourself if you operate it on on your data center like local ground data center I feel I feel like you start at the beginning of a sovereign stack and then there is of course a lot of question back doors and more complex stuff but that's kind of the basic minimal thing in 2026

Matt Turk: 作为一名开源乐观主义者,你是否担心西方开源力量蓬勃发展的动力问题?中国有非常清晰的地缘政治动机去站在开源的最前沿。但如果看西方,你提到了 Nvidia,我们在几周前的播客中也邀请了来自 Nvidia 的 Brian Catanzaro(他负责 Neotron 的整个工作)。显然,Nvidia 这么做是有很强的商业动机的,因为他们销售芯片,因此拥有一个蓬勃发展的开源模型生态是符合其利益的。

但是如果考虑其他人,目前还不清楚为什么西方公司会大力去做开源。Reflection 可能是个例外,但据我所知这些模型还没有推出。你是否思考过这种动机问题,以及它对西方开源的未来意味着什么?

Original English

Matt Turk: >> and as an open-source optimist um do you do you worry uh about the motivations for western open source to thrive. So uh China has a clearly a geopolitical uh motive behind uh being at the forefront of open source. If if you think of the west then so you mentioned Nvidia and we had Brian Kenzaro um from you know the whole Neotron effort on the podcast a few weeks ago. So clearly there is a uh motivation for Nvidia to do this which is they sell the chips so having a thriving open source model makes sense. Um but if you think about like everybody else it's sort of unclear why uh the uh why western companies would do open source this reflection might be the exception but the models haven't come out as as far as I know. So do do you think about this like motivation and and what that means uh for the future of western open source?

Thomas Wolf: 是的,当然。但在我看来,这似乎是非常显而易见的——我们确实需要开源,而且拥有这种动力的人比你想象的要多得多。

只要你对拥有一个繁荣的商业生态系统感兴趣,希望有许多公司能够真正利用 AI,而不是仅仅作为另一家大公司的“套壳”应用(thin wrapper),而是作为真正的 AI 构建者,你就会需要开源模型。这就是为什么美国政府最近也表态:“实际上,我们希望保持开源的发展势头。”其基本理念是:开源是促进新商业诞生的最佳途径之一。

这能带来很多好处。例如,它能让你避免最终陷入仅由两家公司组成的寡头垄断中。我们在过去见过很多寡头垄断的案例,这对于价格竞争力和创新来说并不总是好事。如果直接冲向“我们已经选出了赢家,这两家公司将为所有人构建 AI”的终局,从商业、经济和市场竞争的角度来看,我认为这绝对不是最优解。

而且,这会极大地限制发明创造。仅举一个例子,目前在生物学领域有很多潜力,有很多新的生命科学公司,也有很多公司想要探索这个方向。但由于目前的安全限制网关以及围绕生物黑客等问题的顾虑,现在购买 API 服务的访问权限其实非常受限,尤其是当你想要询问一些生物学问题时。据我所知,大多数原本使用闭源模型的生命科学公司现在都不得不切换到其他方案,因为他们希望能自由地处理任何与生物学相关的研发。

所以,要么我们说从今往后所有的生命科学都由 Anthropic、OpenAI 和 Google(可能还有 Meta)来构建,要么我们说我们想要这里有一个庞大的生态系统。我个人的偏好(我承认我有偏好)是:你不会希望权力过于集中在少数几家掌握核心技术的巨头手中。我认为越多人能够参与发明,我们就会拥有越多样化的新想法、新公司。

作为投资者(你是投资者,你知道我在说什么),如果你只能投资 Anthropic 和 OpenAI,那未免有点无聊和悲哀,那只是成长期阶段的投资。你想要投资你自己的公司,而且你不想只投资那些非常单薄的“套壳”应用,你想投资那些真正有能力构建 AI 的公司。而我们看到,在 Hugging Face 平台上,大多数这样的公司都需要开源模型。

另一个突出的例子是游戏、视频或者机器人(robotics)领域。大多数时候,他们的做法是从一个开源模型开始,然后利用机器人数据或游戏数据对其进行微调。以游戏为例,他们会采用一个开源的视频生成模型,去做一些视频生成创业公司当时没有预测到的事情——他们需要加入动作(actions)。所以,他们在闭环中加入动作和视频进行微调,这就是第一批有趣的实时游戏公司的起步方式。

因此,很多时候,开源是一家新公司起步、建立自己护城河、在自己专属数据上进行微调并开始构建自己 AI 的最便捷途径,而不是仅仅把你的训练数据卖回给模型提供商——我认为这总是很危险的,因为如果模型提供商发现你做得很好,他们在未来的某个时间点很可能会进入你的垂直领域。这在过去已经在法律、设计以及几乎所有领域发生过了。

Original English

Thomas Wolf: >> Yeah, of course. But that seems pretty that's pretty that seems pretty obvious to me that we actually want that and and I think a lot more people have incentive than you may think. Like I think I think as as soon as you're interested in having a thriving business ecosystem with many companies being able to to actually use AI and not just as seen wrapper around another company but as real like AI builder um I think you you want some open source model. So that's that's one of the reason the US government just recently say actually we want to keep open source you know like striving and basically the idea is that open source is one of the best way to get new business. So it can be for many things can be just uh because it it allowed you to not just end up with oligopoly basically of two company which I mean we've seen many case of oligopoly in the past it's not always the best thing for competitiveness for price for you know there's many danger with that I think just rushing in the direction of say we we have made we have we've taken our winners and these two companies are going to build AI for everyone else I think that's from a business economical market side of view I think doesn't really seems uh optimal to me and then it's also limit a lot invention. So just to take one example there's a lot of potential right now in biology there's a lot of new life science company there's a lot of company wanting to explore that um because of the guard rails and because of the question around biohacking and using this model to generate like the access right now for payable just to take it uh is very very limited once you want to ask some biology question and so basically most of the life science company I've seen who were using this model had to switch to another option they wanted to be able to process anything related to biology. So if you're a new company exploring something that uh the big labs are not currently uh you know uh at least uh exploring or they they don't feel like there is enough business potential maybe to give full access to everyone or I don't know exactly um but basically uh you cannot really use that. So either we say all life science is going to be built from now on by anthropic, open AI and Google maybe meta or we say we want a big ecosystem around there. I mean my my personal opinion and I'm biased but is that you don't want too much concentration of power around key technology and I feel like the more people can you know invent the more we have a diversity of new ideas diversity of new company but also as investors right and you're investors you know what I talk about like let's say you could only invest in anthropic and open AI that's a little bit sad right it's a little bit boring it's only for growth stage but you want to invest in your company and you don't want to invest only in these very thin in rapper you want to invest in your company that actually able to build AI and most of them uh and we see them at face most of them need open source models another bit example is all the thing around gaming video or even robotics most of the time what they do is they start from an open source model and then they fine-tune it on some robotic data for gaming for instance they'll take a video generation model that's open source they need to do something that was not um predicted by the the video generation startup they need to add actions. So what they do they they find tune with action in the loop and video and that's how the first I think interesting like real time gaming company started basically. So a lot of the time open source is your easy way as a new company to start to have your own mode to fine tune on your own data to be able to start to build your own AI and not just to basically sell your training data back to the model providers which is I think always dangerous because a lot of these model provider might at some point want to enter your field if you're basically selling them your data and this happened in the past already in legal in design in like any field I think

行业降速与对齐

Matt Turk: 换句话说,你对 7 月 24 日行业联合公开信(关于开放权重与美国 AI 领导力的信件,这也是黄仁勋的首条推文,你们显然也签署了这封信)的看法是:这不仅对世界有好处,而且背后有着强大的经济驱动力。这是在某种程度上抵制正在成型的寡头垄断结构。这个说法公平吗?

Thomas Wolf: 是的,我认为对于每一个相信创造力和在 AI 世界中创造新事物的人来说,都会希望保留开源这一维度的路径。这就像如果所有的软件代码都是闭源的,我们今天就不会拥有如此繁荣的编码生态系统。这是不言而喻的——如果每个人都只能在大闭源软件公司工作才能开发任何软件,那整个软件行业的创新活力就根本无从谈起。我认为在 AI 领域也是相同的道理。我不想忽视风险,我完全同意我们需要在对齐上努力。但我认为目前在最前沿之下的开源模型,其风险可能没有人们极力宣传的那么大。

Original English

Matt Turk: >> to play it back so your take on the July 24 letter uh industry letter on open weights and American AI leadership which was also Jensen's first tweet ever that that that that you guys that you guys signed obviously uh is um u partly uh it's important and partly a resistance to just an oligopoly um structure that is being put in place. So it's both it's it's good for the world but the strong economic motivation behind it. Is that is that fair?

Thomas Wolf: Yeah, I think I think it's u everyone believing in inventiveness and being able to create new things also in the AI world I think would want would want a part of open source access just like the same you know if every code was closed source we would not have the thriving coding ecosystem we have right now right that's kind of obvious like everyone would have to work at one of the large closed source software company if they wanted to create any software that doesn't seems even really possible to have all the inventiveness and creation we've seen in in the software industry. I think the same happen in I don't want to dismiss risk and I fully agree we need to work on alignment and I say that's I would say for now open source model being under the frontier I think this is maybe less important than some what people wanted to say

Matt Turk: 也许我们可以退一步,在接近对话尾声时,从你的视角去感受一下世界在往何处去。在你昨天关于 AISI 事件的帖子中,你谈到了“新一波实验室”的浪潮。你特别提到了杰夫·迪恩(Jeff Dean)昨天刚刚宣布的新公司的消息。昨天在各种消息发布和宣布的维度上,确实是有点疯狂的一天。核心在于,这家新公司是明确奔着“递归自我提升”(recursive self-improvement)而去的。在这一切背景下,对于这种向着“自我维护、自我开发的递归 AI”的演进,你有多紧张?

Original English

Matt Turk: >> maybe to to take a step back as we get near the end of this uh conversation and sort of get a sense for where the world might be going from your perspective in your in your post yesterday about AIS you talked about um a new wave of labs. So in particular, you referenced the news of the Jet Dean new company that he just announced yesterday. Yesterday was a bit of a crazy day in terms of like everything that that came out in the world and was and and was was announced. Um so uh the the point being that this company is explicitly rushing toward recursive uh self-improvement. Uh so in the context of everything we we described, how nervous uh are you about um this evolution towards um self-maintaining, self-developing recursive AI?

Thomas Wolf: 是的,这是一个好问题。作为一名科学家和研究人员,我必须说我对这个想法非常感兴趣。有许多关于超级智能的研究,他们希望以此解决人类面临的关键挑战。我认为这是一个伟大的目标。我非常希望看到 AI 能做出更多的科学发现,我认为这可能比单纯在各种地方应用 AI 更有益。所以我非常乐观。

但我只是觉得,过去几周确实让我们对“我们在对齐这些模型方面到底做得有多好”产生了一些疑问。在法语中我们常说:“不要把马车放在牛前面。” 我们需要按部就班地进行。

所以,我认为目前的好消息是,这些工作似乎大多数是在内部研究实验室里进行的,希望他们在部署某些产品之前,在安全上做足了功课。他们思考了产品将被如何使用,并对他们所构建的东西的社会影响进行了深思熟虑。但我们仍然应该努力非常清楚地去理解,我们将如何在人类社会中部署这些模型。

Original English

Thomas Wolf: >> Yeah, that's a good question and definitely as a as a scientist researcher, I'm I would say I'm I'm very interested in the idea. I feel like there's lot of this uh super intelligence that that they want to tackle, you know, solving um crucial challenge for humanity. I think this is this is a great goal. Um I would love to see AI having making more scientific discovery. I think would be probably the most beneficial thing that AI could bring more than just AI slope everywhere. So I think I'm I'm very optimistic. I I just feel like uh I would say the past few weeks I has raised a little bit the question of how how good are we at aligning this model. So like in French we say we don't want to put the uh the carriage before the cow like we we need to go in order there. So um yeah I would say right now the the good thing is most of this um I would say seems to be internal research lab hopefully they they do they do good security around what they do before they deploy uh some of their product they think how it's going to be used and they they have good good feeling around I would say good thinking around the the social impact of what they are building um but yeah I still I still think that we should try to understand really well how we're going to how we're going to deploy this model in in the human world I would say

Matt Turk: 对,但你并没有加入此前那一派人的请愿。在开放权重公开信发布的大约四天后,出现了另一封由 1100 多人联署的公开信。这一次,Anthropic 和 OpenAI 都签署了,要求政府协助“刻意放慢前沿自动化 AI 研究的步伐”。基本上,行业本身在呼吁某种形式的放慢脚步。你是属于那一派,还是认为那根本不符合现实运作的方式?

Original English

Matt Turk: >> right but you you're not in the camp of the um petition that that came out so I think 4 days after the open weights letter um that um you know Nvidia did that we were talking about a minute ago there was a different letter that came out with 1100 people and and and this time both anthropic and open sign that asked government to help deliberately pace the frontier of automated AI research So basically the industry kind of asking um for a slowdown. Are you are you in that camp or you think that's just not the way it works?

Thomas Wolf: 我确实签署了那封信。我同意我们应该朝那个方向努力。作为一名投资者,你可能也有类似的感受。我觉得即使我们现在完全停下研究,我们已经拥有的技术也足够建立起相当优秀的商业生态。我们能用目前的模型做很多事情,而且这些事情已经极其有趣了。我们还有很多东西需要通过开放科学和分享它们的工作原理去理解。

所以我并不属于主张“现在必须疯狂向前冲”的那一派。如果能稍稍放慢一点脚步,那会很棒——因为这样我们可能不需要每天面对四个重磅发布,不需要在做下一期播客时手忙脚乱地把它们全部塞进去。或许我夏天还能放一天假。

但抛开玩笑不谈,核心问题是:我们能把它做好吗?我认为很多人的底线是,他们可以接受 AI 发展得稍微慢一点,更开放一点,更审慎、更具反射性地去理解如何把事情做好。但最大的问题在于:我们如何在没有不良激励的情况下,去协商和组织这样一次降速?如果只有一两个玩家不肯放慢脚步,就会破坏整个降速努力的效果。在这点上,公开信可能并没有给出什么明确的激励机制,这可能需要国际合作。

有一篇很长的博文叫《AI 2040》,不知道你读过没有,它也极力主张进行某种审慎的降速。也许这介于“踩死刹车”和“全速狂飙”之间,或者是介于“全面监管以至于没人能用 AI”和“完全不管”之间(我认为这两种都是愚蠢的极端方案)。试图寻找一种方式来稍微放缓前沿研究的步伐,我认为那会很棒。我们不会因此失去太多,而且在商业创造和公司成立方面,依然会有许多非常伟大的事情发生。

我其实对这种呼吁放缓和支持开源都有共鸣。我不认为开源本身必须是盲目加速主义的。有人说开源是“减速主义”或“通缩主义”的,我也不完全这么看。我认为你完全可以既支持开放科学、支持开放性,同时也认为我们确实需要理解如何精细地训练模型,并做一些真正的科学研究。

Original English

Thomas Wolf: >> I did sign this letter. I agree. I agree that I think we should uh we should go there. I'm also and you probably have the same feeling as as an investor. I have the feeling that even if we stop right now, we would still have quite a good companies we could build on top of we have what we have right now. I feel like there's a lot of things we we can already do with these models and that are already extremely interesting. I feel like there's a lot of things we need to understand and we should do in terms of open science and sharing how they work. So I'm not I'm not in the camp of we need to rush really quickly right now. The main question is uh if we want to slow down a little bit also that would be great because maybe then we don't have four announcement per day that we need to mix in one podcast the next podcast tomorrow. Uh yeah, maybe I can take one day off holiday in the summer. But no, I I think the main question is can we do it right? That's the main question here. I think a lot of people would be fine uh with AI going a little bit a little bit slower, being a little bit more open, being a little bit more caring, a bit more like reflexive and and and trying to understand better, you know, how to how to do that really well. Um but the main question how can we negotiate uh how can we organize uh a slowdown there uh without having bad incentives where where you know just one or two player not slowing down will kind of break the whole the whole effect of of having a slowdown and here I don't know if yeah I don't think the letter gives any incentives it might be it might need some some you know collaboration there was one one long blog post called AI 2040 I don't know if you read it that was also advocating for kind a careful slow down uh and maybe you know somewhere between uh we go you know full breaks out we go we go as fast as we can and somewhere between uh we regulate everything so nobody use AI which I think are both stupid solution uh but something around we we try to see if there is a way we could actually pace this a little bit slower I think I think that would be great I don't think we would lose a lot and I think actually also in terms of company creation and all of that we could still have a lot really great things happening. Um, but yeah, I'm I'm in this uh I'm I'm actually sympathetic to both this and and open source. I don't think open source has to be accelerationist per se. Um Dan was saying this is decelerationist. I don't think it's also this deflationist. I think these are these are also you can be pro open science. uh you can be pro openness and you can also uh think that actually we need to understand how to train this model well and we need to actually uh being able to do real science right now

Matt Turk: 你不担心这只是一次“监管套利”(regulatory capture)的企图吗?顶尖的两家私立实验室实际上是在试图操纵游戏规则,让其他所有人都放慢脚步。我的意思是,你提到了不是每个人都遵守规则的风险,但这实际上也起到了固化当前市场结构的作用,决定了谁是领导者而谁不是。

Original English

Matt Turk: >> and uh you're not worried about this being uh an attempt at regulatory capture where like the top two private labs um are effectively trying to figure out how everybody else can slow down. I mean, you mentioned the risk of like not everybody just complying um but effectively um freezing the market structure around who's a leader and who's not.

Thomas Wolf: 是的,我不认为它一定是这样的。显然,目前也存在另一种纯粹加速主义的路子,即默许几家巨头互相疯狂竞争,同时通过监管把其他所有人都排除在局外。我不认为监管必然等同于低效或迟缓。

这完全取决于你如何将其实付诸实践,如何在法规或合作中部署它。我认为 Demis 在前几天也写过一封非常好的公开信,是在他调任之前(虽然我们现在知道他会去哪里),他提倡国际合作的信件在某些方面也是极力支持开源的。

我认为你可以拥有一种极具开源色彩的降速方案——这也是我最希望看到的。在放慢前沿速度的同时,利用这段空档期去开源和分享更多成果。而在纯粹的军备竞赛动态中,各家实验室通常只会选择紧锁大门。因此,在我看来,放慢速度反而为我们提供了“开放”的契机。

当然,如果最终它演变成仅仅为了巩固两家巨头的卡特尔或寡头垄断,那我对这个方向绝对不会有任何兴趣。

Original English

Thomas Wolf: >> Yeah, I don't I don't think it has to be. Uh I feel like you have definitely the same path that's actually fully accelerationist where you decide a couple of company are racing against each other and you regulate all the other outs. I I don't think regulation has to be synonymous with with slowness or not. Um and definitely the question is more how you put that in in um in action, how you actually, you know, put that in in practice, how you deploy this in in regulation or cooperation. I think Demis also had a pretty nice letter the other day uh before he stepped down or up as as a chief scientist still we understand where he's going to be now but like uh his letter for basically international collaboration was also very much pro open source in some aspect uh I think you can have a slowdown that's very open source that's the one that's the one I would love to see which is you we slow down and we use we use the fact that we slow down to be able to actually share more things and I feel like a race dynamic is usually more in terms of closing the the doors of the labs, right? So to me a slowdown is probably more the opportunity to open but of course I mean if it turns out to be mostly a way to uh just solidify uh like we were saying a cartel or oligopoly of just two company I'm I'm not very excited about this direction.

Matt Turk: 太棒了。Thomas,这似乎是一个非常好的收尾之处。非常感谢你。这绝对是精彩至极的一期,我非常享受这次对话。非常感谢你在假期中抽出时间与我们交流。谢谢你!

Thomas Wolf: 谢谢,Matt。

Matt Turk: 大家好,我是 Matt Turk。感谢收听本期 Mad Podcast。如果你喜欢这一集,我们将非常感激如果你能订阅我们的频道,或者在你收听、收看本期节目的任何平台上给我们留下好评或留言。这能真正帮助我们建立这档播客并邀请到更多优秀的嘉宾。谢谢大家,我们下期节目再见。

Original English

Matt Turk: >> Wonderful. Well that feels like a wonderful place to leave it Thomas. Uh thank you so much. This was absolutely fantastic. really enjoyed it. Um, and appreciate your taking, uh, some time to speak with us, uh, in the middle of your time off. Um, so, um, thank you so much. Appreciate it.

Thomas Wolf: Thanks, Matt.

Matt Turk: >> Hi, it's Matt Turk again. Thanks for listening to this episode of the Mad Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you at the next episode.

📌 文中提及的人物和组织

关键字: ai-agent-security model-alignment open-weights social-engineering sovereign-ai