数据驱动的通用智能体范式与前沿架构演进 Latent Space 2026-06-24

引言与节目介绍

Matei Zaharia: 我们所拥有的一个突破口是,一旦你能把数据放到正确的位置,AI模型就会变得相当不错。通用的智能体已经相当成熟了[音乐声]。我的意思是,Ali在这里已经谈到过AGI了。它们具有相当好的推理能力。实际上,我认为许多传统软件都会被这种新范式所重写,也就是只要把数据准备好,然后在上面加一层AGI。奇迹就会出现。

Original English

Matei Zaharia: One of the fuses we have is actually once you can get the data in the right place, the AI models are becoming pretty good. The generic agents are fairly [music] I mean Ali talked about AGI already here. They have pretty good reasoning capabilities. Actually, I think many of the traditional software will be sort of rewritten with this new paradigm which is just get the data to be there and then let's slap some AGI on top. Magic will come out.

Host: 是的。

Original English

Host: >> Yeah.

Matei Zaharia: 嗯,但是如果没有正确的数据,你真的无法做到这一点。

Original English

Matei Zaharia: >> Um but without the right data, you can't really do that.

Host: 在我们进入今天的节目之前,我只想给听众朋友们留个简短的信息。谢谢你们。如果你们没有选择点击并收听我们的内容,我们就不可能为您带来您如此明确想要的那些人工智能工程科学和娱乐内容。几乎每天都有赞助商主动联系我们,但幸运的是,有足够多的听众实际订阅了我们,使这一切能够在没有广告的情况下维持下去,而且我们希望保持这种状态。但我只求大家帮个忙。你能做的最有效且完全免费的一件事,就是点击那个订阅按钮。这是我唯一会向你们提出的要求,这对我和我的团队来说意味着一切,他们如此努力地工作,只为了每周把《In Space》节目带给你们。如果你们这样做了,我向你们保证,我们绝不会停止努力,我们会把这个节目做得更好。现在让我们进入正题。来自Databricks的Matei Zaharia。欢迎来到《In Space》。

Original English

Host: >> Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering science and entertainment contents that you so clearly want if you didn't choose to also click in and tune into our content. We've been approached by sponsors on an almost daily basis, but fortunately enough of you actually subscribe to us to keep all this sustainable without ads and we want to keep it that way. But I just have one favor to ask all of you. The single most powerful completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you and it means absolutely everything to me and my team that works so hard to bring the In Space to you each and every week. If you do it, I promise you we'll never stop working to make this show even better. Now let's get into it. Matei Zaharia from Databricks. Welcome to In Space.

Databricks 峰会与早期回忆

Matei Zaharia: 感谢邀请。

Original English

Matei Zaharia: >> Thanks for having us.

Host: 是的,非常感谢。呃,感谢您抽出宝贵时间。你们的 Databricks Data AI 峰会正在进行中。您刚才还在跟我说,你们举办的第一届峰会只有 50 个人。

Original English

Host: >> Yeah, thanks so much. >> Uh thanks for taking time out. You you have your Databricks Data AI Summit going on. You were just telling me how the first summit that you guys ran was just 50 people.

Matei Zaharia: 嗯哼。是的,我想那是在伯克利的一个小型聚会。我们组织了一个……是的,我们做了一些教程,是的,就是教人们使用 Spark。

Original English

Matei Zaharia: >> Mhm. >> Yeah, it was a little meet-up at Berkeley, I think. We put together a Yeah. >> We did these tutorials and yeah, just teach people Spark.

Host: 是的。你知道,显然现在的规模,我想我看到的头条数字大概是全球有 10 万人参与,其中 3 万人是现场参加。这真是一个疯狂的社区。嗯,我是说我刚刚看了主题演讲。[笑声] Ali 真是……你们以前就知道,那时候能明显看出 Ali 会成为这么棒的首席执行官吗?他是个极好的演讲者。你怎么看?

Original English

Host: >> Yeah. You know, obviously now it's like I think I had like a the headline number is like 100,000 people around the world, 30,000 in person. Uh it's a crazy community. Well, I mean I just saw the keynote. >> [laughter] >> Ali is just Did you know that was it obvious that back when that Ali would be like such a great like CEO? Like he's a great presenter. >> What do you think?

Matei Zaharia: 呃,我的意思是,我认为在我们的创始人团队中,很明显我认为他会是做这个最合适的人选。而且,是的,结果非常棒。他……我的意思是,为了运营一家公司,他恶补了非常多的课题。他会直接钻研,比如去学习,并且,你知道,去和所有的专家交谈。就像,即使你无法雇佣到那个人,你也能去学习足够多的,比如财务、销售,或者其他任何知识。嗯,并且你知道,我会从那里继续推进。是的。

Original English

Matei Zaharia: >> Uh I mean, I think among our group of founders, it was clear that I think he'd be the best at this. And and yeah, it turned out great. And he's I mean, he's ramped up on so many topics going a company. He would just go in and like study at and, you know, become talked all the experts. Like, even if you can't hire the person, you know, learn enough about like finance and sales and whatever it was. Um, and you know, I'd go from there. Yeah.

Host: 我的意思是,他显然智商很高,情商也很高,但这并不像……今天的 Ali 和十年前的 Ali [笑声] 已经大不相同了。我认为他……为了达到今天这个地步,他投入了大量的工作。

Original English

Host: >> I mean, he's obviously very high IQ and have very high EQ, but it wasn't like Ali today is quite different from Ali from like 10 years [laughter] ago. I think he there's a lot of work that he put in to get to this point.

Matei Zaharia: 是的。我的意思是,不。我的意思是,对我来说,他最吸引人的地方在于他很幽默。而且就像……你知道,很难去拿,你知道的,严肃话题中的数据、安全之类的事情来开玩笑。

Original English

Matei Zaharia: >> Yeah. I mean, no. I mean, to me the the most appealing thing about him is that he's funny. And like it You know, it's >> It's true. Yeah. >> It's hard to make jokes about, you know, data about serious topics and security and what have you.

Host: 哦,是的。那是肯定的。

Original English

Host: >> Oh, yeah. That's for sure.

Omnigents 与代理架构的演进

Host: 是的。所以,你们推出了一大堆新产品。我只简单点一下名,因为我们不可能涵盖所有内容。Omnigents,你的你的宝贝。El Tap,你的宝贝。你的 Dream Engine。呃,我们还要报道 Genie,报道 Customer League。你们收购了 Panther,Open Sharing,还有 Unity AI Gateway。我认为其中很多东西是你会期望 Databricks 去做的。这就好像是路线图的一部分。在你们这个类别里的每个人都有类似的东西。但我认为,可能你们两位正在引领这一领域中最独特、最具差异化的两项倡议。也许我们将从……从 Omnigents 开始,然后我们将……我们将深入探讨它。我确实认为有很多人在探索这种元框架 (meta harness) 的概念。是什么引导你想到这个的?

Original English

Host: >> Yeah. So, you you guys launched a whole bunch of things. I'll just go to name check briefly the stuff because we were not going to cover everything. Omnigents, your your baby. El Tap, your baby. Your Dream Engine. Uh, we're also going to cover Genie, cover Customer League. You acquired Panther, Open Sharing, and there's Unity AI Gateway. A lot of these I think like are things that you would expect a Databricks to do. It's like part of the the road map. Everyone in your category has has similar things. But I think probably the two of you are leading the two most unique and differentiated initiatives in the landscape. Maybe we'll start with uh with Omnigents and then we'll we'll go into it. >> I do think that a lot of people are exploring this sort of meta harness concept. What led you to it?

Matei Zaharia: 是的。实际上有几条相互交汇的线索,我认为这是一个很好的迹象,表明你需要一些新的东西。所以,一方面,在内部有所有的编码智能体。我们有非常棒的开发基础设施团队。他们构建了一个叫做 Isaac 的东西,这基本上就像是云代码 (cloud code) 和 code X 上的一个包装器,让你可以在网络上的沙盒里,或者就在你的开发机器上,或者你的笔记本电脑上,或者是任何地方使用它们。然后,你知道,他们在那里添加了各种各样的东西,我们看到所有……这种更高级的工程师们,比如正在构建他们自己的工作流,里面有成吨的智能体,并且他们还在其上,甚至在那之上构建他们自己的 UI 和其他东西。然后另一点是,像我们构建智能体一样,我们在研究团队中发布了这个名为 Genie 的数据科学智能体,基本上这个团队是我共同领导的。我们也为各种不同的事情构建了许多内部智能体,然后我们还有所有的客户智能体,所有这些都遇到了这样一个问题:哦,我需要切换模型和测试框架等等,你知道,每隔几个月就得换一次。此外,如果你不能与某人共享会话、拥有历史记录、拥有搜索功能,以及所有这些建立在其上用于协作的层,那么这个智能体就像是完全无用的。我从这两种背景中思考了一下,起初人们认为这很奇怪,心想你为什么要把编码智能体和自定义智能体放在同一个东西里?但我说这是……这基本上是相同的问题,你只是想构建那种能让你交付智能体的东西,如果你关心安全性的话,或许还可以控制它,并让它能在不同事物之间移植。然后我们做了一些实验性的原型。我们说,是的,其实我们可以让它发挥作用,然后我们,你知道,我们就真的把它构建出来了。

Original English

Matei Zaharia: >> Yeah. There were actually a couple of like converging lines, which is I think is a good sign that you need something new. So, on the one hand, there's all the coding agent in for internally. We have really great dev infra team. They built something called Isaac that's basically like a wrapper on cloud code and and code X and lets you use them either on the web in like sandboxes or just on your dev machine or on your laptop or whatever. And then you know, they were adding all kinds of stuff there and we saw all the the sort of more advanced engineers like were building their own workflows with tons of agents and they were building their own UI's and stuff on top of even on top of that. And then the other one was that like us building agents, we shipped this like data science agent called Genie on the research team which I co-lead basically. We also build a lot of internal ones for various things and then we have all the customer ones and all of them were running into this thing of like, oh I need to switch model and harness and so on you know, every few months. Plus the agent is like completely useless if you can't share sessions with someone and have history and have search and all this like layer on top of it for collaboration. I thought a bit about it from both contexts and at first people thought it was weird and like why are you doing coding agents and custom agents in the same thing but I said it's it's it's basically the same problems and you you just want to build the stuff that lets you deliver the agent maybe control it if you care about security and make it portable across things. And then we prototyped some things as experiments. We said, yeah, actually we can make it work and then we you know, we sort of built it for real.

Host: 我在想,这种我们称之为架构的东西,是否能与你们过去职业生涯中的任何事物相对应。你知道,就像我总是会思考,很多事物实际上都可以追溯到操作系统。

Original English

Host: >> I'm wondering if this kind of let's call it architecture >> Yeah. >> maps to anything in your careers in the past. You know, like I always think about how a lot of things actually just tie back to operating systems.

Matei Zaharia: 嗯哼。[笑声] 许多操作系统又可以追溯到数据库,或者是反过来。所以,我确实认为它和网络协议有很多相似之处,你知道,互联网协议。我们也做过……是的,我们也做过有关数据共享方面的工作,这可能是大多数观众不会了解的,除非他们……

Original English

Matei Zaharia: >> Mhm. >> [laughter] >> A lot of operating systems tie back to databases so or the other way around. >> So the thing I do think it ties a lot to like network protocols, you know, internet protocol. We also did Yeah, we did stuff with like data sharing also which is probably most viewers probably won't know unless they

Host: 是的,开放协议(open protocol)是用于共享的术语。开放共享(Open sharing)。

Original English

Host: >> Yeah, open protocol is the term for sharing. Open sharing.

Matei Zaharia: 开放共享。是的,就好比你有一家公司,你维护着某种数据表,比方说像沃尔玛或者什么公司。他们拥有,你知道的,库存信息以及每家店售出了什么。然后你还有供应商,他们非常乐意生产更多的商品,并在你正好需要的那一刻把货物运送给你。所以他们非常希望能实时访问你的数据表。因此,与其到处发送电子邮件、Excel 表格或者打电话,为什么不能与他们实时共享该表格的视图呢?然后他们进行查询,他们,你知道的,将它与他们自己的数据结合起来,然后决定要发送什么。所以这属于这样一种情况,你会……就像你可能会问,既然今天我们能如此快速地凭借直觉编写出任何代码(vibe code anything),那为什么我们还需要去设计诸如协议、API 或者软件呢?为什么你不能就根据需求随心所欲地去写代码呢?但实际上,对于这种互操作性,即多个以不同速度推进的参与方正在构建各自的东西,而你仍然需要在顶层有一层来进行协调,这时候你确实需要去设计并构建它。所以这让我想起了那种感觉,就像是智能体之间互相交流,以及用户与智能体和工具进行交流。

Original English

Matei Zaharia: >> Open sharing. Yeah, so it's like you have a company, you maintain some kind of table like like let's say like Walmart or something. They have like the you know, inventory and what's been sold in each store. And then you also have suppliers and they would love to produce more things and ship them like exactly the moment you need them. So they would love like real-time access to your table. So instead of like sending emails around or Excel sheets or phone calls, why can't you share like a view of that table in real-time with them? Then they query, they you know, join it with their data and they decide what to send. So it's it's one of these things where you you like that you you might ask like today since we can vibe code anything so fast, why do we even need to design like protocols or APIs or software? Why can't you just vibe code things on demand? But actually for this type of interoperability where multiple parties that are moving at different speeds are building stuff and you still want some layer on top to coordinate, you do want to design it and build it. So it reminds me of that like agents talking to each other and users talking to agents and and tools.

Host: 我们知道有任何其他评论或不同的观点吗?

Original English

Host: >> Do we know of any other comments or alternative viewpoints?

Guest 2: 顺便说一下,我想我们就我们所说的 Matter 的好处进行了很多次辩论,而且我……我想就在我们决定要做这件事的时候,我正在告诉 Matei,“嘿,碰巧有这么一个特定的星期,我从醒来的那一刻起就不停地敲代码,直到我上床睡觉。我当时就在看我的云端会话,我的 Codex 会话,其中一件特别烦人的事情就是必须保持我的笔记本电脑处于打开状态。” 嗯哼。我当时实际上正开车去看医生,我记得因为我想确保整个工作继续运行。顺便说一句,听到你这么说真是太令人欣慰了。

Original English

Guest 2: >> I think by the way, we had a debate on exactly we said the benefits with Matter a lot and I I think around the time we decided to do this thing, I was telling Matei, "Hey, it just happened to be there's a particular week that I was coding non-stop from the moment I woke up to like the moment I went to bed. I was like looking at my cloud sessions, my Codex sessions and and one of the things particularly annoying was having to keep my laptop open." >> Mhm. >> I was actually driving to a doctor's appointment and I remember because I want to make sure the whole thing continues working. By the way, it's so comforting to hear you say

远古的编程体验与云端沙盒的诞生

Speaker A: 因为我当时在想,我不知道自己是不是像个小丑一样在做这件事,或者说,是的,说实话,我当时在开车,并且把笔记本电脑连到了手机的热点上。我就把它放在旁边,每当遇到红灯时,我就开始看笔记本电脑上发生了什么。我只是觉得这太荒谬了。

Original English

Speaker A: that because I'm like I don't know if I'm a clown and I'm doing this or like Yeah, yeah, like honestly, I I was driving and I was tethering my laptop to my phone. Keeping on the side, whenever I hit a red light, I started looking at what's going on on my laptop. And I just felt that was ridiculous.

Speaker B: 是的。

Original English

Speaker B: Yeah.

Speaker A: 感觉就像我们回到了编程的黑暗时代。

Original English

Speaker A: It felt like we went back to the dark ages of programming.

Speaker B: 我的意思是,你从这个编程时代获得的所有生产力都是惊人的,但是,呃,比如,你听说过云吗?

Original English

Speaker B: I mean, the productivity you gain from all this coding age is amazing, but uh Like, have you heard of cloud?

Speaker A: [笑声]

Original English

Speaker A: [laughter]

Speaker B: 这对我来说太疯狂了。

Original English

Speaker B: It was crazy to me.

Speaker A: 我们当时在做的是沙盒吗,还是在那之前?

Original English

Speaker A: Was it the thing we were working on was the sandboxes or was this before that?

Speaker B: 那是一个沙盒。

Original English

Speaker B: It was a sandbox.

Speaker A: 好的。所以,你当时是——

Original English

Speaker A: Okay. So, you you were

Speaker B: 所以,我是从一个非常不同的角度来处理的。我当时想,“嘿,我们要有实际上不会关闭的云沙盒。你可以非常快速地获取一个,而且不仅仅是为了运行智能体(Agentic)会话。它实际上也是为了运行开发环境。”所以,我那周其实亲自在构建那个东西,而在构建的过程中,我遇到了所有这些问题。然后,我实际上为我的情况写了一份文档,说明“这是我对实际环境应该做什么的愿望清单。”而且我认为他最后实际上几乎实现了其中的每一项。

Original English

Speaker B: And so, I was approaching from very different angle. I wanted to like, "Hey, we're going to have cloud sandboxes that actually doesn't shut down. You can get one very quickly, but not just for running agentic sessions. It's actually also for running development." So, I was actually personally building that that week and through building that, I ran into all these issues. And then, I wrote actually a document for my case that, "Here's my wish list of what the actual environment should do." And I think he actually ended up almost implementing every single one of them.

Speaker A: 是的,我记得 Reynold 曾经说过,因为我最初的这个原型仅仅是与你的智能体聊天,然后他说:“我必须能够打开一个像我自己那样的 Shell,比如列出文件,还有跟踪日志(tail)等等。”所以,我当时——

Original English

Speaker A: Yeah, I remember Reynold saying cuz my first prototype of this had just chats with your agent and he said, "I have to be able to open a a shell like my own shell and like list files and like tail them and stuff." So, I was

Speaker B: 这是 SSH 连接到大型机吗?

Original English

Speaker B: Is this an SSH into a mainframe?

Speaker A: 是的,实际上它有 [笑声] 那个功能。

Original English

Speaker A: Yeah, actually it has [laughter] that.

Speaker B: 可以跟踪(tail)我的日志。

Original English

Speaker B: Telling my logs.

Speaker A: 是的。是的。

Original English

Speaker A: Yeah. Yeah.

Speaker B: 而且,我认为我提出的另一件事是,呃,我仍然只为了渲染 Markdown 文件这一个唯一目的而在使用 Cursor。

Original English

Speaker B: And also, another thing I think I asked was uh I I I had I still use cursor for the sole purpose of rendering markdown files.

Speaker A: 嗯哼。是的。

Original English

Speaker A: Uh-huh. Yes.

Speaker B: “给我一个能查看我的 Markdown 文件并正确渲染它们的方法。我再也不需要一个单独的工具了。”是的。我想你们也把这个功能加进去了。

Original English

Speaker B: give me a way to see my markdown files and render them properly. I don't need a separate tool anymore." Yeah. I think You also built that in.

Speaker A: 我们确实做了。是的。是的,我们有很多工程师构建了,你知道的,他们自己那种随性(vibe)的编程设置。但是后来,他们说的另一件事是,“嘿,我构建了一些对我来说很棒的东西,但团队里没有其他人可以使用它,因为我没有一个可以协作的服务器。”这也正是,这正是为什么我们试图建立 Omnigen,这样你就可以有一个服务器,并在里面建立好安全机制。所以,你知道的,比如用 Google 登录或者什么的,而且能够真正安全地共享东西。这就是为什么我们看到许多其他的智能体会遇到一些问题,比如人们认为他们开发了一个很棒的智能体原型,但你知道的,因为安全团队的原因,它不被允许连接到某些非常重要的数据或任何东西。所以,是的。

Original English

Speaker A: we did that. Yeah. Yeah, we had a lot of engineers building, you know, their own vibe coding setup. But then, the other thing they all said is like, "Hey, I built something that's amazing for me, but like no one else on the team can use that cuz I I don't have a server to collaborate." And And this is This is why we tried to set up Omnigen so you can have a server and have the security uh set up in there. So, you know, like log in with Google or whatever and like actually securely share stuff. Would And that's why we've seen a lot of other agents like hit things like people think they prototyped an awesome agent, but you know, it's not allowed to connect to like some really important data or whatever because of the security team. So, yeah.

Agent Cloud 架构与开源策略

Speaker B: 是的,在这一点上,所以,对于那些在 YouTube 上观看的人,我们将在这里调出一张关于这个结构的图片,我们可以稍微讨论一下这个架构。我想我只是,我只是想让大家理解,因为当我们谈论软件时,它可能会非常抽象,而这实际上就是我们所谈论的。你已经在开源中基本做出了整个平台,这有一个运行器(runner)组件和服务器组件,并带有一个你已经弄清楚的统一的 API。任何其他类型的元素,显然你可以插入所有这些持久层和计算层。这整个就是一片云。这就是 Agent Cloud。

Original English

Speaker B: Yeah, at this point So, for those watching along on YouTube, we're going to bring up a image of the structure here and we can talk through a little bit of the architecture. I think I just I just want to have people understand cuz like when we're talking about software, it can be very abstract and like like here's actually what we're talking about. You've worked out in open source this entire platform basically and there's a runner component and server component with a a sort of uniform API that you've you figured out. Any other sort of element and obviously you can plug in all these persistence layers and and compute layers. This is a whole cloud. This is Agent Cloud.

Speaker A: 是的,我的意思是,它有这些组件来协同工作。你知道的,很多动作都发生在你部署智能体的那台机器上。所以,无论你在上面有什么,你都可以运行。但是,是的,我认为这几乎是你想托管这类协作型智能体、并拥有那个服务器所需要的最小限度的东西。我们将其开源的原因之一是,这给任何构建智能体的人提供了一个可以作为起点的应用程序,并且可以进行定制,这也正是我们在 Databricks 中看到的,比如有人可能会制作一个很棒的,你知道的,智能体应用程序,然后其他团队就会问,“哦,我能直接把你的应用用在我的智能体上吗?”

Original English

Speaker A: Yeah, it's I mean it's got these components to work with it. The you know, a lot of the action happens like on the machine where you deploy your agent to. So, whatever you've got on there, you can run. But yeah, it's I think it's sort of the minimal thing you you want to have hosted like collaborative agents and to have that server. And one of the reasons we open sourced it is anyone building agents, this gives them an app they can start with and customize, which we were seeing in Databricks too like someone would make a nice, you know, agent app and then other teams would ask, "Oh, can I just use yours for my agent?"

Speaker B: 是的,我想我们大概有五到六个由每个不同团队构建的不同的智能体框架。它们所做的事情多多少少都是一样的。

Original English

Speaker B: Yeah, I think we had like five or six different Agentic frameworks built by every different team. They do all do more or less the same thing.

Speaker A: 是的,你基本上需要,人们想要采用在 Forkit 中能够运行的东西,而你最好也有一些开源的东西。是的,这也引出了另一个问题,这对于像 Databricks 这样的公司来说很有趣,那就是你选择将什么开源,你又选择将什么做成专有产品?这,这,我的意思是,这要追溯到 Spark,对吧?

Original English

Speaker A: Yeah, you need to basically people want to take something that works in Forkit and you might as well have something open source. Yeah, which which also was another question that which is interesting for a Databricks like what do you choose to open source, what do you choose to make it proprietary? It's It's I mean this goes back to Spark, right?

Speaker B: 是的。

Original English

Speaker B: Yeah.

Speaker A: [笑声]

Original English

Speaker A: [laughter]

Speaker B: 所以,我的意思是,开源某个东西的原因之一是,如果你认为这一层实际上会产生一些网络效应。它将受益于许多人的共同协作。所以,比如在 Spark 方面,我不知道你是否知道,当,当 Spark 问世时,我们也把很多重点放在让你能在上面构建库(libraries)上。所以,比如过去有不同的分布式计算引擎,用于机器学习和图计算。我们说它们都应该成为你可以组合使用的库,而且我们还让连接数据源变得超级容易。然后我们就从中受益,因为,你知道,我们没有时间去为比如一千个不同的数据库和文件格式编写连接器。但是我们可以直接使用别人做出来的那些,当然他们也从加入,你知道的,这种类似生态的事情中受益。所以,这算是其中之一。另一种思考方式是,我可以想象,你知道,如果我们的东西不开源。我们有某种托管智能体的东西,但它不开源,然后市面上有一个开源的。如果是你,从长远来看谁会赢?所以,像在这里,因为确实能从人们编写的集成中获益,所以一定会是这样。然后还有一些其他东西,比如你根本无法以开源的形式交付的内容,那些是公司所做的事情。例如,你如何确保你的像流处理,你知道的,任务或者你的基于湖的数据库晚上不会,你知道,丢失所有数据?好吧,这就需要一个会坐在那里的运营团队。没有别的办法,它必须是一项服务。所以,比如我们希望确保作为一家公司,我们非常擅长那些基础设施服务,然后在你在其上构建的内容方面,我们尽可能做到开放。

Original English

Speaker B: One So, I mean one of the reasons to open source something is if you think it's a layer that will actually there'll be some network effect. It'll benefit from many people collaborating on it. So, for example, with Spark, I don't know if you if you know what when when Spark came out, we we also focused a lot on letting you have libraries on top. So, like they used to be different distributed computing engines for like machine learning and graph computation. We said they should all be libraries that you can compose, and we made it super easy to add connectors to data sources, too. And then we benefit because, you know, we we don't have the time to write like connectors to like, you know, a thousand like different databases and and file formats. But we can just use the ones people make, and of course they benefit from joining, you know, kind of this this thing. So, that's like one of these as a Another way to think about it is I can imagine, you know, we our thing wasn't open. We had some kind of agent hosting thing, but it's not open, and then there's an open one. If you're Which one's going to win in the long run? So, like here, because there is this benefit from like people writing integrations, it'll be it'll be that. And then there are other things that like you just can't even deliver as open source that are things the company does. Like, for example, how do you make sure your like streaming, you know, jobs or your your lake-based database doesn't like, you know, lose all your data at night? Well, that requires an an operational team that's going to sit there. There's no way it has to be a service. So, like we want to make sure as a company we're really good at those infra services, and then we're as open as as we can in terms of like what you build on top.

开源生态的初步反馈与现代 AI 堆栈

Speaker A: 我的意思是,从收益的角度来说,我想我们已经看到了 pull request 还有生态系统的集成,尽管它是在星期六才发布的。

Original English

Speaker A: I mean, speaking from a benefits, I think we're already seeing pull request and the ecosystem integration, even though it was only released on Saturday.

Speaker B: 是的,星期六。是的,所以有人——

Original English

Speaker B: Yeah, Saturday. Yeah, so someone

Speaker A: 让我们看看发生了什么。是的。

Original English

Speaker A: Let's see what's going on. Yeah.

Speaker B: 是的,你可以看一下合并的代码行数。今天早上我实际上问了一位传奇人物关于——

Original English

Speaker B: Yeah, you can look at the merge lines. I actually asked some legend this this morning about the

Speaker A: 已经有 400 次合并了?

Original English

Speaker A: 400 merge already?

Speaker B: 是的,我想,很可能,我猜大约有一半不是来自我的团队。但举例来说,有人添加了在 Kubernetes 上运行它的支持。人们添加了许多云沙盒。所以,这个可以启动一个云沙盒并在里面运行你的智能体,这对于共享来说非常棒,因为这不像在你的笔记本电脑上,别人在上面运行一些临空(sky)代码。嗯,所以,是的,很多初创公司已经把这些放进去了,我们期待看到更多。我们已经有了更多的智能体测试框架(harnesses),包括 Cursor、CLI 以及 Anti-gravity 等等。

Original English

Speaker B: Yeah, I I think quite I would guess around half are not from my team. But for example, someone added support for running it on Kubernetes. People added many cloud sandboxes. So, this can launch a cloud sandbox and run your agent in there, which is great for sharing, too, cuz it's not like on your laptop and someone's like running sky code on there. Um so, yeah, many startups have put those in and we expect to see more of them. We also have more agent harnesses already, cursor, CLI, and anti-gravity, also.

Speaker A: 是的。

Original English

Speaker A: Yeah.

Speaker B: 这些都很,呃,完美。我,你懂的,我感觉上一次发生这种情况的时候,还是现代数据堆栈(modern data stack)兴起之际。我不知道它是否真的那么有用。我,我实际上对你关于这个的复盘挺好奇的。我,我想大多数人 [清嗓子] 会同意,它终于算是死了。呃,但或许,这正在催生一个新的现代 AI 堆栈,它实际上是在做相同的事情。

Original English

Speaker B: That's all uh beautiful. I I you know, I I I feel like the last time this happens, there was the rise of the modern data stack. I don't know if it was that useful. I'm I'm actually kind of curious in your your postmortem. I I think most people [clears throat] will agree that it is finally dead. Uh but it maybe this arises to a new modern AI stack that like does the same thing.

Speaker A: [笑声]

Original English

Speaker A: [laughter]

Speaker B: 我不知道。

Original English

Speaker B: I don't know.

Speaker A: 我的意思是,我认为现代数据堆栈曾是一个非常有用的东西,可能一直到今天都是如此。我想对于那些并不真正了解这段历史的观众来说,我想现代数据堆栈实际上被解构为:你需要一层来将数据摄取(ingest)进来。你需要一层来转换你的数据。呃,然后所有这些都在运行,之后你需要一层来可视化你的数据。而所有这些都运行在某种数据仓库上,或者后来我们在做数据仓库时,也出现了数据湖仓(lakehouse)。我认为这些概念都非常强大而且非常有用。它在某种程度上促成了许多工作负载。人们最终遇到的,算是一种关于统一和整合的问题。它,

Original English

Speaker A: I mean, I think the modern data stack was a pretty useful thing, probably even up until this day. I I think what maybe for the audience who don't actually understand the history. I think the modern data stack is effectively decomposed into you need a layer to ingest the data in. You need a layer to transform your data. Uh then all of this are run and then you need a layer to maybe visualize your data. And all of this runs on some sort of data warehouse or later on as we're doing data warehouse, also lakehouse. I think that concepts are all very powerful and very useful. It's sort of enabled a lot of workloads. What people eventually run into is kind of a question of unification and consolidation. It's,

统一平台的演进与 API 抽象

Speaker A: “嘿,你真的需要把所有这些东西切成不同的碎片,然后和这么多不同的供应商和平台合作,就为了完成一个非常简单的可视化吗?”对吧?所以,我认为随着时间的推移,每个人都开始意识到客户在推动我们。我们开始意识到这一点,所以我们开始构建越来越多的功能,并试图进行整合。归根结底,现在客户不需要担心为了生成一个图表而必须连接五个不同的系统。但我认为,老实说,类似的事情可能正在发生,嗯,为了生成一个非常简单的智能体,你想要把多少个不同的框架连接在一起。

Original English

Speaker A: "Hey, do you really need to chop all of this into different pieces and work with so many different vendors and platforms in order to get like a very simple visualization done." Right? So, I think like over time everybody start realizing that customers are pushing us. We start to realize that, so we start building more and more capabilities and trying to consolidate. And at the end of the day now, customers don't have to worry about having to hook up five different systems in order to produce a chart. But the I think I honestly something like this is probably happening um in how many different frameworks do you want to hook up together in order to produce like do a very simple agent.

Speaker B: 澄清一下,我会说这个的核心是所有这些外壳之上的这个通用 API。所以这个 API 基本上就是你有一个智能体会话,你可以发送一条消息或者比如一个文件。基本上,这就是你可以发送的内容,然后你输出,你知道,这些流,当它流式传输文本或者当它进行工具调用时。而且,嗯,或者你可以发送的另一件事是,你可以告诉它取消或转向。所以这就是 API。现在,我们所做的是,我们可以在像 Claude code(在终端中运行)、Codex、你知道的 Phi、OpenAI SDK 等所有这些东西之上为你提供它。我们将它们全部映射到同一个接口。所以,如果你构建自己的比如智能体编排器,你将不得不自己维护这个。然后,每当 Claude 更改其 API 时,你必须,你知道,调整你的代码,否则它会丢失一些消息。所以那是维护起来很有价值的东西。然后在此之上,就像我们构建了一些应用程序。我认为我们构建了一个非常酷的用户界面之类的东西,但那是……而且我们构建了安全和控制部分,这让我感到很兴奋。但正是那个通用接口。所以我们不……它并没有试图成为一个技术栈。事实上,你可以在这个服务器之上插入你自己的用户界面。这是我们非常关心的一个用例,因为我们想在自己的产品中使用它。

Original English

Speaker B: Just to be clear, I would say the core of this is this common API on top of all the harnesses. So the API is basically like you've got an agent session and you can send in a message or like a file. Basically, that's what you can send in and then you get out, you know, these streams as it's streaming text or as it's doing tool calls. And uh or the other thing you can send in is you can like tell it to cancel or turn. So that's the API. Now, the thing we did is we we could get you that on top of like Claude code, running in a terminal, Codex, you know, Phi, OpenAI SDK, all that stuff. We map them all to that same interface. So that is something that you'd have to maintain yourself if you built your own like agent orchestrator. And then whenever Claude changes its API, you got to, you know, tweak your thing or it's going to lose some messages. So that's the thing that's valuable to maintain. Then on top of that, like we built a few apps. I think we we built a pretty cool UI and stuff, but that's um and and we built the security and control piece which which I'm excited about. But it's that common interface. So we don't we it doesn't try to be a stack. And in fact, you could plug in your your own UI on top of this uh server. That That's one of the use cases we care a lot about cuz we want to use this in in our own products.

Speaker C: 是的,它应该无处不在。

Original English

Speaker C: Yeah, it should be everywhere.

Speaker B: 是的。

Original English

Speaker B: Yeah.

计算沙盒与数据库的演进

Speaker C: 我认为其中让我觉得非常有趣的一件事是,好吧,首先,我会我会努力做所有事情,而不是把它叫做现代 AI 技术栈,因为我认为 [笑声] 我们有一个名字了。但是像,是的,像你知道的,嗯,所以最早告诉我计算沙盒的人之一是 Neon 的 Nikita。因为很多人认为 Neon 就像是,嗯,它是无服务器的 Postgres,将计算和存储分离,并且,嗯,你知道,即时分支以及所有这些东西。但实际上,每个数据库公司也是一家计算公司。

Original English

Speaker C: I think one of those things that is is really interesting to me is well, first of all, I'll I'll endeavor to do everything and not call it the modern AI stack because I think [laughter] we have a name. But like yes, like you know, uh so one of the first people that told me about compute uh sandboxing was Nikita from Neon. Cuz a lot of people think about uh Neon as like, well, it's serverless Postgres with like the separation of the compute and storage and uh you know, instant branching and all those things. But actually, every database company is also a compute company.

Speaker A: 是的。

Original English

Speaker A: Yeah.

Speaker C: 所以他当时实际上在向我展示他的整个他的沙盒解决方案。我认为他从来没有发布过它。

Original English

Speaker C: And so he was actually showing to me his whole his sandboxing solution. I don't think he ever launched it.

Speaker A: 所以我们的沙盒解决方案,我们之所以能这么快地构建它,是因为我们意识到如果你只采用实际的数据湖基础架构,并从中移除数据库,顺便说一句,这是你的主意。现在有一些区别。例如,在支持这种特定工作负载的架构中,拥有本地持久性很重要。嗯,因为你希望你的状态持久化。你的库,你不需要每次都安装你的库。对吧?而 Neon 的架构,因为存储与计算分离,你不需要持久性的本地磁盘。所以有一些区别。但归根结底,是的,它是,它是,嗯

Original English

Speaker A: So our sandbox solution, the reason we could have built it so quickly was because we realized if you just take the actual lake base architecture and remove the database from it, by the way it was coming from you. Now there are some differences. For example, in the ones that support this particular workload is important to have local persistence. Um because you want your state to persist. Your libraries you don't have to install your library every time. Right? Whereas the Neon architecture, because of the separation of storage from compute, you don't need persistent local disk. So there's some differences. But the at the end of the day, yeah, it's it's uh

Speaker C: 是的,所以这是当你运行像一个编程沙盒的时候。就像如果我使用那个我们在 Databricks 内部的开发基础设施,那里有像几十 GB 的数据,仅仅是用来存储所有源代码和我构建的工件之类的东西,我希望下次它还能回来。所以,但是是的。

Original English

Speaker C: Yeah, so this is when you run like a a coding sandbox. Like if I use that we had the dev infra internally at Databricks, there's like many many like tens of gigabytes of data just for like all the source code and like artifacts and stuff that I built and I want that to come back next time. So but yeah.

云计算和数据库的规模与采用

Speaker B: 节目前,我们谈到了一些可能会让人对采用率感到惊讶的统计数据。可能是内部的,也可能是外部的,想到的任何数据都可以。只是为了向人们展示这一切发生的规模。

Original English

Speaker B: Before the show, we was talking about some statistics that might be surprising at the adoption. It could be internal, could be external, whatever comes to mind. Just to impress people the scale this is happening.

Speaker A: 所以我们在分析方面,我认为我们每天在我们的三个云上可能启动 5000 或 6000 万个虚拟机。所以我们是那里最大的计算编排器之一。这在 CPU 计算方面是肯定的。

Original English

Speaker A: So we on the analytics side, I think we launch maybe 50 or 60 million virtual machines a day across our three clouds. So we're one of the biggest compute orchestrator out there. So that's for sure for CPU compute.

Speaker B: 是的。

Original English

Speaker B: Yeah.

Speaker A: 嗯,而所有这些过程,我认为处理了 EB 级的数据。我曾开玩笑说,取决于你所在的时区,通常在你吃早餐之前,Databricks 在那一天就已经处理了 EB 级的数据。嗯,而在 Neon 上,它真的很有趣,它现在每天启动我想是 1300 万个数据库队列。

Original English

Speaker A: Um the and all of those process, I think exabytes of data. I joked about depending on which time zone you are, typically before you have breakfast, Databricks would have processed processed exabytes of data already on that day. Um and on Neon it's it's actually pretty interesting queue it's launching I think 13 million databases a day now.

Speaker C: 是的,这对我来说就像是一个巨大的……

Original English

Speaker C: Yeah, that to me that was like a big

Speaker A: 而且那就像……

Original English

Speaker A: And that's just like

Speaker C: 你是什么……你这是什么意思?

Original English

Speaker C: What do you what do you mean?

Speaker A: [笑声]

Original English

Speaker A: [laughter]

Speaker C: 后来很大程度上得益于智能体以及分支实验。

Original English

Speaker C: And then a lot of those were thanks to agent agents and branching experimentation.

Speaker A: 是的。

Original English

Speaker A: Yeah.

Speaker C: 因为我们让它变得如此简单和快捷,并且非常感谢 Nikita 的团队启动数据库。这是,嗯,那个,所以它正在改变人们使用数据库的方式。

[BODY_START]

统一平台的演进与 API 抽象

Speaker A: “嘿,你真的需要把所有这些东西切成不同的碎片,然后和这么多不同的供应商和平台合作,就为了完成一个非常简单的可视化吗?”对吧?所以,我认为随着时间的推移,每个人都开始意识到客户在推动我们。我们开始意识到这一点,所以我们开始构建越来越多的功能,并试图进行整合。归根结底,现在客户不需要担心为了生成一个图表而必须连接五个不同的系统。但我认为,老实说,类似的事情可能正在发生,嗯,为了生成一个非常简单的智能体,你想要把多少个不同的框架连接在一起。

Original English

Speaker A: "Hey, do you really need to chop all of this into different pieces and work with so many different vendors and platforms in order to get like a very simple visualization done." Right? So, I think like over time everybody start realizing that customers are pushing us. We start to realize that, so we start building more and more capabilities and trying to consolidate. And at the end of the day now, customers don't have to worry about having to hook up five different systems in order to produce a chart. But the I think I honestly something like this is probably happening um in how many different frameworks do you want to hook up together in order to produce like do a very simple agent.

Speaker B: 澄清一下,我会说这个的核心是所有这些外壳之上的这个通用 API。所以这个 API 基本上就是你有一个智能体会话,你可以发送一条消息或者比如一个文件。基本上,这就是你可以发送的内容,然后你输出,你知道,这些流,当它流式传输文本或者当它进行工具调用时。而且,嗯,或者你可以发送的另一件事是,你可以告诉它取消或转向。所以这就是 API。现在,我们所做的是,我们可以在像 Claude code(在终端中运行)、Codex、你知道的 Phi、OpenAI SDK 等所有这些东西之上为你提供它。我们将它们全部映射到同一个接口。所以,如果你构建自己的比如智能体编排器,你将不得不自己维护这个。然后,每当 Claude 更改其 API 时,你必须,你知道,调整你的代码,否则它会丢失一些消息。所以那是维护起来很有价值的东西。然后在此之上,就像我们构建了一些应用程序。我认为我们构建了一个非常酷的用户界面之类的东西,但那是……而且我们构建了安全和控制部分,这让我感到很兴奋。但正是那个通用接口。所以我们不……它并没有试图成为一个技术栈。事实上,你可以在这个服务器之上插入你自己的用户界面。这是我们非常关心的一个用例,因为我们想在自己的产品中使用它。

Original English

Speaker B: Just to be clear, I would say the core of this is this common API on top of all the harnesses. So the API is basically like you've got an agent session and you can send in a message or like a file. Basically, that's what you can send in and then you get out, you know, these streams as it's streaming text or as it's doing tool calls. And uh or the other thing you can send in is you can like tell it to cancel or turn. So that's the API. Now, the thing we did is we we could get you that on top of like Claude code, running in a terminal, Codex, you know, Phi, OpenAI SDK, all that stuff. We map them all to that same interface. So that is something that you'd have to maintain yourself if you built your own like agent orchestrator. And then whenever Claude changes its API, you got to, you know, tweak your thing or it's going to lose some messages. So that's the thing that's valuable to maintain. Then on top of that, like we built a few apps. I think we we built a pretty cool UI and stuff, but that's um and and we built the security and control piece which which I'm excited about. But it's that common interface. So we don't we it doesn't try to be a stack. And in fact, you could plug in your your own UI on top of this uh server. That That's one of the use cases we care a lot about cuz we want to use this in in our own products.

Speaker C: 是的,它应该无处不在。

Original English

Speaker C: Yeah, it should be everywhere.

Speaker B: 是的。

Original English

Speaker B: Yeah.

计算沙盒与数据库的演进

Speaker C: 我认为其中让我觉得非常有趣的一件事是,好吧,首先,我会我会努力做所有事情,而不是把它叫做现代 AI 技术栈,因为我认为 [笑声] 我们有一个名字了。但是像,是的,像你知道的,嗯,所以最早告诉我计算沙盒的人之一是 Neon 的 Nikita。因为很多人认为 Neon 就像是,嗯,它是无服务器的 Postgres,将计算和存储分离,并且,嗯,你知道,即时分支以及所有这些东西。但实际上,每个数据库公司也是一家计算公司。

Original English

Speaker C: I think one of those things that is is really interesting to me is well, first of all, I'll I'll endeavor to do everything and not call it the modern AI stack because I think [laughter] we have a name. But like yes, like you know, uh so one of the first people that told me about compute uh sandboxing was Nikita from Neon. Cuz a lot of people think about uh Neon as like, well, it's serverless Postgres with like the separation of the compute and storage and uh you know, instant branching and all those things. But actually, every database company is also a compute company.

Speaker A: 是的。

Original English

Speaker A: Yeah.

Speaker C: 所以他当时实际上在向我展示他的整个他的沙盒解决方案。我认为他从来没有发布过它。

Original English

Speaker C: And so he was actually showing to me his whole his sandboxing solution. I don't think he ever launched it.

Speaker A: 所以我们的沙盒解决方案,我们之所以能这么快地构建它,是因为我们意识到如果你只采用实际的数据湖基础架构,并从中移除数据库,顺便说一句,这是你的主意。现在有一些区别。例如,在支持这种特定工作负载的架构中,拥有本地持久性很重要。嗯,因为你希望你的状态持久化。你的库,你不需要每次都安装你的库。对吧?而 Neon 的架构,因为存储与计算分离,你不需要持久性的本地磁盘。所以有一些区别。但归根结底,是的,它是,它是,嗯

Original English

Speaker A: So our sandbox solution, the reason we could have built it so quickly was because we realized if you just take the actual lake base architecture and remove the database from it, by the way it was coming from you. Now there are some differences. For example, in the ones that support this particular workload is important to have local persistence. Um because you want your state to persist. Your libraries you don't have to install your library every time. Right? Whereas the Neon architecture, because of the separation of storage from compute, you don't need persistent local disk. So there's some differences. But the at the end of the day, yeah, it's it's uh

Speaker C: 是的,所以这是当你运行像一个编程沙盒的时候。就像如果我使用那个我们在 Databricks 内部的开发基础设施,那里有像几十 GB 的数据,仅仅是用来存储所有源代码和我构建的工件之类的东西,我希望下次它还能回来。所以,但是是的。

Original English

Speaker C: Yeah, so this is when you run like a a coding sandbox. Like if I use that we had the dev infra internally at Databricks, there's like many many like tens of gigabytes of data just for like all the source code and like artifacts and stuff that I built and I want that to come back next time. So but yeah.

云计算和数据库的规模与采用

Speaker B: 节目前,我们谈到了一些可能会让人对采用率感到惊讶的统计数据。可能是内部的,也可能是外部的,想到的任何数据都可以。只是为了向人们展示这一切发生的规模。

Original English

Speaker B: Before the show, we was talking about some statistics that might be surprising at the adoption. It could be internal, could be external, whatever comes to mind. Just to impress people the scale this is happening.

Speaker A: 所以我们在分析方面,我认为我们每天在我们的三个云上可能启动 5000 或 6000 万个虚拟机。所以我们是那里最大的计算编排器之一。这在 CPU 计算方面是肯定的。

Original English

Speaker A: So we on the analytics side, I think we launch maybe 50 or 60 million virtual machines a day across our three clouds. So we're one of the biggest compute orchestrator out there. So that's for sure for CPU compute.

Speaker B: 是的。

Original English

Speaker B: Yeah.

Speaker A: 嗯,而所有这些过程,我认为处理了 EB 级的数据。我曾开玩笑说,取决于你所在的时区,通常在你吃早餐之前,Databricks 在那一天就已经处理了 EB 级的数据。嗯,而在 Neon 上,它真的很有趣,它现在每天启动我想是 1300 万个数据库队列。

Original English

Speaker A: Um the and all of those process, I think exabytes of data. I joked about depending on which time zone you are, typically before you have breakfast, Databricks would have processed processed exabytes of data already on that day. Um and on Neon it's it's actually pretty interesting queue it's launching I think 13 million databases a day now.

Speaker C: 是的,这对我来说就像是一个巨大的……

Original English

Speaker C: Yeah, that to me that was like a big

Speaker A: 而且那就像……

Original English

Speaker A: And that's just like

Speaker C: 你是什么……你这是什么意思?

Original English

Speaker C: What do you what do you mean?

Speaker A: [笑声]

Original English

Speaker A: [laughter]

Speaker C: 后来很大程度上得益于智能体以及分支实验。

Original English

Speaker C: And then a lot of those were thanks to agent agents and branching experimentation.

Speaker A: 是的。

Original English

Speaker A: Yeah.

Speaker C: 因为我们让它变得如此简单和快捷,并且非常感谢 Nikita 的团队启动数据库。这是,嗯,那个,所以它正在改变人们使用数据库的方式。

Original English

Speaker C: Because we made it so easy and so quickly and thanks a lot to Nikita's team to launch databases. It's uh that So it's changing the way people use databases.

智能体安全策略:状态化和上下文控制

Speaker B: 是的。好的,我们稍后会深入探讨更多关于数据库的话题,但我想确保我们了结任何关于 Omnigens 的事情。嗯,你提到了,嗯,你对安全和控制方面感到兴奋。嗯,许多公司现在都在弄清楚这一点,以及支出方面。嗯,你在那里发现了什么?

Original English

Speaker B: Yeah. Okay, we're going to go into more database talk in a bit, but I want to make sure we close up anything on Omnigens. Uh you mentioned uh you're excited about the security and control side. Uh a lot of companies are figuring that out right now, as well as the spend side. Um what have you found there?

Speaker A: 是的,所以我花了不少时间与内部用户交谈,嗯,开发者,安全团队,你知道的,嗯,经理,以及很多客户。有几件事。比如首先,一件事,你知道,显而易见的是对于安全性,嗯,你知道,在可用性和安全性之间存在这种紧张关系。而且,嗯,人们做事的方式,比如今天很多编程智能体都有非常基本的功能,比如你可以告诉我哪些工具模式是允许的或不允许的之类的。就像是简单的“是”或“否”。但这让你处于一个非常艰难的境地。所以,仅举个例子,比如我的智能体是否应该能够阅读,你知道的,一些机密文档?或者,假设它是否应该能够从 NPM 安装新包,你知道,也许它是一个被攻陷的包。“是”或“否”。像也许,也许我想允许它。我的智能体是否应该能够将东西发布到公司网站?好吧,如果我在网站上使用代码,是的。但它应该能够两者兼做吗?所以,它可以像抓取机密文档并被提示注入然后泄露它吗?可能不行。所以,我们决定我们需要的是状态化的或我们称之为上下文策略的东西,你在其中跟踪该会话的状态。这不像它是否被允许推送到营销网站?而是比如,嘿,如果它做了一件危险的事情,比如它安装了,你知道的,一个才发布一天的 NPM 包,或者它,它读取了比如一千份机密文档,那么不行。那就不要不要这么做。否则,也许可以。这是比如平衡那种权衡的一个例子。所以,通过拥有一个更强大的引擎,它基本上既更安全又更有用。这需要跟踪会话。另一个有趣的部分是,它正在执行这些非常底层的事件,你需要在此之上有一些库来解析它们。比如举个例子,我们在 Google Drive 内部有一个 MCP 服务器。它有 60 个 API 调用。比如我怎么知道其中哪些,比如会将文档与互联网上的内容共享,哪些不会?这这很烦人。所以我们在 Omnigen 中设计了策略层,这样它就是一些函数,并且你可以拥有库。比如有人可以构建一个将底层事件映射到高级事件的东西,然后你针对出来的高级事件编写策略。所以,而且那与 Panther 相关,并且

Original English

Speaker A: Yeah, so I spent quite a bit of time talking to internal users, uh developers, security team, you know, uh managers, and also lots of customers. And there's a few things. Like first of all, one thing, you know, that immediately was became obvious as for security, uh you know, there's this tension between like usability and security. And um the way people do like a lot of coding agents today have very basic things like you can tell me which tool patterns are allowed or disallowed or whatever. It's like yes or no. But that puts you in a very tough spot. So, just as an example, like should my agent be able to read, you know, some confidential documents? Or or let's say should it be able to install new packages from NPM, which you know, maybe it's a it's compromised. Yes or no. Like maybe maybe I want to allow it. Should my agent be able to publish stuff to the company website? Well, if I'm using at the code on the website, yes. But should it be able to do both? So, it can like grab a confidential document and be prompt injected and leak it? Probably not. So, the thing we decided we need is stateful or what we call contextual policies, where you keep track of the state of that session. It's not like is it allowed to push to the marketing site or not? But like hey, if it did a risky thing, like it installed, you know, a one-day-old package from NPM, or it it read like a thousand confidential docs, then no. Then don't don't do it. Otherwise, maybe it's okay. That's one example of like moving that trade-off. So, it's both more secure and more useful by having a more powerful engine essentially. This requires tracking sessions. The other piece that was interesting there is like there are these very low-level events it's doing and you want some libraries on top that parse them. Like for example, we have a MCP server on Google Drive internally. It's got 60 API calls. Like how do I know which of those like will share a document with stuff on the internet and which ones won't? It's it's annoying. So we designed an Omnigen the policy layer so that it's it's functions and you can have libraries. Like someone can make something that maps the low-level events to high-level ones and then you write a policy about the high-level things that came out. So and that was related to the Panther and

Speaker B: 是的,Panther 会有所帮助。Panther 在事件处理方面是一种类似的想法,它是基于 Python 的,而不是一种奇怪的自定义语言。这更多是实时的。

Original English

Speaker B: Yeah, Panther is will help with that. Panther is kind of a similar idea on the event processing side and it's Python based versus a weird custom language. This is sort of more as in real time.

Speaker A: [笑声]

Original English

Speaker A: [laughter]

Speaker B: 这些事情正在发生。是的。

Original English

Speaker B: Those things are happening. Yeah.

Speaker A: 所以,是的,但但是这些是那些很酷的事情。我认为上下文的或状态化的部分,然后是,它可以作为库的方式,这也是开源它的另一个原因,因为其他人会编写库,就像我们和我们的客户可以使用它们一样。最后一件事,因为它是状态化的,我们跟踪的状态之一是你在那次会话中花了多少钱。所以,我曾经有过比如,我我让一个智能体调试一些东西,结果它花了 500 美元,因为它是

Original English

Speaker A: So yeah, but but these are the cool things. I think the contextual or stateful part and then the the way it can be libraries and that was another reason to make it open source because others will write libraries and like we and our customers can use them. And the final thing because it's stateful, one of the states we track is how much you spent in that session. So I can I've had like I I asked an agent to debug something and it spent $500 because it

安全与成本控制

Matei:……决定去读取大量日志文件,从而消耗大量 token。但我实际上可以直接说:“启动一个子代理来做这件事,并把成本限制在 5 美元以内。如果需要更多资金,请先征求我的许可。”而且因为我们在该会话中计算了成本,它会弹出来提示我:“好的,你已经花了 5 美元。你想继续吗?”

Original English

Matei: decided to read a lot of log files and burn a lot of tokens. But I can literally say, "Okay, launch a sub agent to do this and cap it to spending $5. Like ask me for permission if it needs more." And because we're counting that within that session, it'll pop up and tell me, "Okay, you spent five $5. Do you want to go on?"

Host:我在这里补充一点背景信息。Matei 在过去五年中,将大量时间花在了构建 Databricks 的 Unity Catalog 上,这是数据的治理层。

Original English

Host: So I'm more than context here. Matei spent the last five years, a lot of his time was architecting Unity Catalog at Databricks, which is the governance layer for data.

Matei:没错。是的,是的。

Original English

Matei: That's right. Yeah. Yeah.

Host:这实际上是将那一层的专业知识与这里所有的 AI 治理结合在了一起。

Original English

Host: And it's sort of combining expertise at that layer together with all the AI governance here.

Matei:是的。不过,我也花了很多时间,因为那些编程代理让我感到很烦,有时还会弄出些“定时炸弹”。而且作为首席技术官,我可不想因为安装了某个奇怪的 NPM 包导致所有代码泄露而登上新闻头条。所以我格外偏执,但同时我的时间又非常有限。因此,我不想坐在那里不断批准诸如“你想运行这个 20 行的 bash 脚本吗?是还是否?”这类请求。这就是为什么我花了很多时间去弄清楚:我怎样才能让它尽可能安全,同时又不那么烦人。

Original English

Matei: Yes. Yeah, but I also spent a lot of time being annoyed by coding agents and and getting bombs. And also as the CTO, I don't want to end up on the front page as like I installed some weird NPM package and leaked all the code. So, I'm especially paranoid, but also I have very little time. So, I don't want to sit there approving like, "Do you want to run a 20-line, you know, bash script? Yes or no?" So, that's why I spend a lot of time figuring out like, how can I make it as safe as possible and not annoying.

Host:是的。安全,或者我们称之为安全性(security),是不是比追求最大的 token 数量或 token 预算更让你担忧?哪一个更像是……

Original English

Host: Yeah. Is safety and let's call it security a bigger concern than token maxing or token budgets, you know, which one is like

Matei:哦,是的,这两个问题都存在。我的意思是,我不知道,我想这取决于你是什么类型的公司。所以,我认为一些公司,比如预算有限的,他们真的非常在意这一点。或者,你可能像 Uber 那样(规模庞大),但也依然会有同样的顾虑。

Original English

Matei: Oh, yeah, they're both there. I mean, I don't know. I I guess it depends on the type of company you are. So, I think some companies like the budget is is limited and you know, they they really care about that. Or I mean, you can be Uber and still be concerned, you know.

Host:是的,哦是的,完全正确。

Original English

Host: Yeah, oh yeah, totally. Yeah, [laughter] yeah, yeah.

Matei:是的,对我们来说,作为一家云服务提供商,安全性绝对是至关重要的,它是最重要的事情。至于 token 消耗上限,你知道,我们现在还不太担心它。但我见过这样的例子:我曾和一些咨询公司交谈过。他们有大约 10 万名员工,都在为客户编写代码。如果他们每个人每个月多花 1000 美元,那可不是闹着玩的。你知道,我们只有几千名工程师。

Original English

Matei: If you have Yeah. Yeah, for for us security is is absolutely critical as a as a, you know, cloud provider. It's it's the most important thing. And token maxing, you know, we we we're not so worried about it yet, but but I've seen the like for example, I talked to some consulting companies. They have like 100,000 employees who are all coding for customers. If those each spend like an extra $1,000 a month, that's that's not fun. You know, we have like only a few thousand engineers.

Host:那么 Databricks 的政策是什么?是没有限制,还是说……

Original English

Host: What's the policy in Databricks? Is it is it just unlimited or

Matei:是没有限制,但你知道,我们会使用自己的产品来分析使用痕迹之类的东西。我们有一个团队专门负责优化,并检查是否有人在做些奇怪的事情。我们实际上仅通过分析当前的使用痕迹就获得了一些非常酷的洞察,比如哪些模型在处理 Rust 语言时比处理 TypeScript 或其他语言表现得更好。所以,至少在我们的代码库中是这样。

Original English

Matei: unlimited, but we do you know, we use our own product to like analyze the traces and stuff and we have a team that's you know, looking to optimize and and to see if anyone's doing something weird. And we actually had some really cool insights just from analyzing current traces like which models are better at say Rust versus like, you know, TypeScript or whatever. So, yeah, at least in our code base.

OmniGen 与代理控制层

Host:非常棒。显然,我不得不问关于 token 上限的问题。我显然认为这是一个关键问题,但安全性与控制凌驾于此之上。我们需要找出一个合理的层级,让你能在拥有一定自主权的同时,又不至于失去控制。

Original English

Host: Yeah, amazing. Obviously, I have to ask the token maxing question. Obviously, I think it's a it's a key thing, but but yes, security and control above that and figuring out a sane layer that you can have some autonomy, but not too much.

Matei:是的。[笑] 是的,而且我们想让这件事变得超级简单。作为一名工程师,你应该能够自己设定这些规则。所以在 OmniGen 中,你可以要求你的代理为你自己设定一个执行策略。

Original English

Matei: Yeah. [laughter] Yeah, and we want to make it super easy. As an engineer, you should set the thing. So, in Omnigen, you can ask your agent set up a policy on yourself to do this. So, it

Host:如果有什么我应该展示的内容,我在 GitHub 上没有看到,但你知道,这是……

Original English

Host: If there's anything I should be showing, I I I don't I don't see it on the GitHub, but you know, this is

Matei:把它放进那里的文档里了。所以,你可以晚点再看。如果你想看的话,只需在文档中查找关于上下文策略(contextual policies)的内容。嗯……

Original English

Matei: put in the docs there. So, you can look at it later. Just look in the docs on contextual policies if you want to see. Um

Host:我只是想向人们展示一下策略。

Original English

Host: I just like to show people policies.

Host:是的,如果你想跟进了解这一点,这就是你需要查看的确切位置,对吧?是的。

Original English

Host: Yeah, if you want to, you know, follow up on this, this is exactly where to look, right? Yeah.

Matei:是的,是的。嗯,关于这些的背景故事是这样的,你知道,就像我在你们开发这些功能之前,写了一份包含 10 个功能想法的文档。这就像是我整理的用户需求愿望清单,我告诉团队说:“嘿,在发布时,你们能至少实现其中的五个吗?”结果他们全部都完成了。所以,

Original English

Matei: Yeah. Yeah. Uh yeah, and the story of these is like I just wrote uh you know, like I wrote a doc with like 10 ideas for things before as you were working on them. Well, it did That was like my wish list of things people asked, and I told the team like, "Hey, can you do like at least five of these for the launch?" And then they just got back with all of them. So,

Host:哇。

Original English

Host: Oh, wow.

Matei:嗯,所以你可以想出更多点子,但这其中有些仅仅是为了作为示例。实际上,你可以拦截代理发出的任何事件,然后你可以选择阻止它,或者强制它询问用户,又或者允许执行,并且你可以更新状态来记录相关信息。

Original English

Matei: Um so, you can come up with with more, but then some of them are just meant to be examples. Really, you can intercept like any event the agent is making, and you can then either block or force it to ask the user or like allow, and you can update state to keep track of stuff.

Host:是的,因为你知道,说到底,我认为你就像是一个系统设计师,你让人们可以随时接入,对吧?这就是你们所做事情的整个操作模式。

Original English

Host: Yeah, cuz you know, ultimately, I think I think of you as like a systems designer, you let people plug in, right? That's the whole modus operandi of what you do.

Matei:是的,是的。我们也非常看重可组合性。比如,其他人能否编写一个库供别人使用,这正是这个工具的初衷……

Original English

Matei: Yeah. Yeah, and we and we care a lot about also composability. Like, can someone else write a library that others use, which this is meant to

Host:这里也有一种“自带电池”(开箱即用)的哲学,这可能与你当年做 Spark 时非常相似,也就是你可以直接开始使用。是的。

Original English

Host: There's also a batteries included philosophy here, probably very similar to how you did Spark, which is you could just start using. Yeah.

Matei:是的,没错。它必须在某些特定功能上开箱即用就能有很好的表现,然后你可以用它来构建你自己的东西。正如你所知,在 Spark 中,如果你只是想读取一个表格或者进行数据聚合,它在开箱状态下就应该表现得非常出色。

Original English

Matei: Yeah, that's right. It has to be good out of the box at certain things, and then you can build your own things on top you know, we don't want to do but you know, in Spark, if you just want to like, I don't know, like read a table or do like aggregation, it should be awesome out of that out of the box.

开源贡献与创业机会

Host:是的。如果人们想了解 OmniGen 的最新进展,他们应该去看看你的主题演讲,去查看 GitHub 上的文档。如果他们想做出贡献,或者想在这个生态系统上构建应用,你会指出哪些地方是投入精力后杠杆率最高、最值得参与的地方?

Original English

Host: Yeah. People want to catch up on Omnigen, they should watch your keynote, uh they should go through the GitHub in the docs. If they wanted to contribute or they want to build on this ecosystem, where would you call out as the most high-leverage places to get involved?

Matei:是的,一定要加入 Discord 和 GitHub,我们的团队在那里会进行监控。人们要求的一些东西,我们就自己直接做出来了。还有些东西,你知道,我们正在和他们合作构建。同时也请告诉我们,你希望如何使用它。因为我认为特别是对于开发者来说,每个人都希望工具能按照自己的方式工作。而一个真正优秀的开发者工具,你必须听取关于各种使用方式的反馈,然后弄清楚如何进行抽象化设计,以及如何让用户进行自定义。所以我们很乐意倾听,如果你觉得“嘿,我不希望它这样工作”,请告诉我们。我们真的很想在不同代理之间建立起那个兼容层,然后让你能在上面做各种事情。

Original English

Matei: Yeah, do get involved in the Discord and then get hub our team is is there is monitoring and some of the things people ask for we just built ourselves. Some of them, you know, we're we're collaborating with with them to build that and also tell us like how you would like to use that cuz I think especially for developers like everyone wants it to work their own way and a really good developer tool like you have to hear the feedback on all the ways and figure out the abstractions and how to let people customize. So we love to hear like if you think hey I you know I don't want it to work this way uh tell us. We really just want to get that compatibility layer across agents and then let you do stuff on top.

Host:是的。那么在初创公司方面呢?比如,我是一名创始人,我看到了一个机会,我想向你展示。如果有一家初创公司在做某些事情,你会提出什么要求,会让你觉得“我真希望有人在做这个”?

Original English

Host: Yeah. Is there any you know in terms of like the startup side I'm a founder I want to I see an opportunity I want to get in front of you. What's your request for like a startup that like you know I wish someone someone was working on this.

Matei:哦,对于一家初创公司来说。

Original English

Matei: Oh for a startup.

Host:是的。比如,你的初创公司发展得很好,但如果你现在没有在做自己的公司,有什么显而易见的事情是你认为应该去做的?

Original English

Host: Yeah. Like you know you got your own startup it's doing well but like you know if you weren't working on your own startup what what is like obvious that you should

Matei:[笑]

Original English

Matei: [laughter]

Host:显然,你也为许多初创公司提供咨询。

Original English

Host: You advise many startups too obviously.

Matei:我的意思是,我认为作为一家拥有大量工程师的公司,任何能帮我理清人们如何使用编程代理、了解成本支出以及代码质量的工具都是有价值的。或者,你能告诉我“你应该编写或添加这个技能”、“你应该写这个东西”,又或者指出“你的代理在处理涉及此服务的任务时表现非常糟糕”或“你需要去花点时间解决这个问题”。那将会非常棒。是的。

Original English

Matei: I mean I do think just as a company with a lot of engineers like anything that helps me make sense of how people are using coding agents and and spend but also quality or like you should write you know you should add this skill or you should write this thing or your agents are really horrible at tasks involving this service or like go spend time. That would be nice. Yeah.

Host:是的,我发现最接近这个概念的是一个叫 get AI 的团队。

Original English

Host: Yeah the closest I found is this team get AI.

Matei:嗯。哦,酷,是的。

Original English

Matei: Mhm. Oh cool yeah.

Host:他们起初的想法是“我们就只做代码和人工归因追踪”,但他们现在基本上正在这之上构建分析层。我确实觉得,比如“人工分析”(artificial analysis)这样的机构在他们那块领域做得非常好。所以我认为肯定会有人去做这个,这最初可能是顾问们的领域,但后来人们肯定会真正开发出软件,就像为编程代理提供一个管理控制台一样。

Original English

Host: They they started with like we will just do code and human attribution but they're basically building the analytics layer on top of that. I do think like there are a bunch of like artificial analysis is obviously doing super well with with their stuff. So there's there will be people I think I think this is like the domain of consultants first but then people will actually build software that is like they have like the management plane for coding agents.

Matei:是的,我想那里会有很多深刻的洞见。在其他领域你也见过类似的情况。

Original English

Matei: Yeah I think there'll be a lot of insights there. You have it in other areas.

数据库系统基础:OLTP 与 HTAP

Host:好的,那么另一个重要的事情就是你的梦想引擎(dream engine)了。

Original English

Host: Okay, well, and then the other big thing is your dream engine.

Matei:[笑]

Original English

Matei: [laughter]

Host:如果你想讲讲关于 OLTP 的故事……

Original English

Host: If you want to tell the the story of of you know, OLTP.

Host:我们这么说的背景是,我会让大家去听听我们那期有 Ankur Goyal 参与的播客,我们在那里讨论了 SingleStore、HTAP 以及所有的历史背景。

Original English

Host: And our background was I'm going to make people listen to our Ankur Goyal episode where we talked about single store, HTAP, and all that history.

Matei:是的,是的。OLTP 的概念其实很简单。所以,如果人们听过 Ankur 关于 HTAP 的演讲,那实际上就是数据库的世界。抱歉,这里可能需要补充很多背景信息。数据库的世界……

Original English

Matei: Yeah, yeah. The OLTP idea is actually pretty simple. So, people have heard of the Ankur's talk about HTAP, it's effectively the world of databases. Sorry, there's like maybe a lot of context that needs to be injected here. The world of databases

Host:伙计们,这将会成为我强迫大家去学习你们数据库知识的那期播客。你不能只靠看 markdown 文件就能搞懂复杂的代码逻辑。

Original English

Host: to be the database podcast that I'm forcing people to like learn your databases, guys. You cannot vibe code with just markdown files.

Matei:这是现有系统技术中最重要的一项基础知识。[笑] 不过,数据库的世界实际上被划分为大致两半。有一半被我们称为 OLTP 数据库,它们是事务型的,想想你的 Postgres、你的 MySQL……

Original English

Matei: It's one of the most important fundamentals of [laughter] systems technologies out there. But, the world of databases is effectively split into roughly two halves. There's what we call OLTP databases, which are transactional, and think of your Postgres, your MySQL,

数据库架构演进与CDC的痛点

Speaker A: 你的Oracle数据库。另一方面是我们称之为分析的部分,有时也可能被称为OLAP。区别在于,对于OLTP,你通常可能是运行一些事件的事务,这些事务会查找特定的某一行,然后我们更新这一行,对吧?这是一种非常面向行的数据结构。而在分析方面,你试图对数据进行推理,你试图计算:“嘿,我每个门店的收入是多少?我的网站每天表现如何?”然后你最终可能要在上面运行机器学习,来预测,“嘿,未来我的销售额可能会怎样?”它们的架构非常不同,大家都是从OLTP数据库开始的,因为每一个应用,当它发展到一定规模、需要的不止是Markdown文件时,你就需要有一个数据库。你不想丢失你的数据。你想要有一定的事务(轻哼)一致性。但是,一旦你想对数据进行推理分析,如果你只有比如100行数据,那在你的Postgres或MySQL数据库上运行可能没问题。可是,一旦你有了更多数据,并且想要运行更复杂的分析,这种分析本身可能就会压垮你原本用于保存数据的数据库。所以你开始考虑将数据移出……

Original English

Speaker A: your Oracle databases. And the other side is what we call analytics, and sometime might refer to the term OLAP. And the differences is on OLTP, you typically have maybe run some transactions on event that looks up at one specific row, we update that row, right? It's a very row-oriented data structure. And on analytics, you're trying to reason on the data, you're trying to compute, "Hey, what's my revenue per store? What's my How's my website doing every day?" And then you eventually want to probably end up running machine learning on it to predict, "Hey, how will my maybe sales be going in the future?" They are so very different architecture, and everybody start with OLTP databases cuz every app, when it becomes serious enough, that needs more than markdown files, you need to have a database. You don't want to lose your data. You want to have some transactional [snorts] consistency. But, once you want to reason on the data, if you only have like 100 rows, it's probably okay to run it on your Postgres or your on your MySQL database. But, once you have more data and want to run more complicated analysis, the very analysis might crush your save post crash database. So you start doing getting data out of the

Speaker B: 也就是做复制同步,将它们复制到分析系统中。

Original English

Speaker B: Replication replicate them into the analytic systems.

Speaker A: 是的,对于人们来说,Elasticsearch 就像是一个巨大的……

Original English

Speaker A: Yeah, which is for people elastic search is like a big

Speaker B: 是的,所以其中一些数据实际上进入了Elasticsearch,比如用于日志分析。显然,我们的很多客户会把数据导入Databricks来运行更复杂的操作。这里有一个术语叫做CDC。

Original English

Speaker B: Yeah, so some of them actually get into elastic search for like log analysis. A lot of our customers obviously get into data bricks to run more sophisticated things. And there's this term called CDC.

Speaker A: 也就是变更数据捕获(Change Data Capture)。

Original English

Speaker A: CDC change data capture.

Speaker B: 嗯,它的作用是读取数据库的binlog。如果你不明白什么是binlog,没关系。但它实际上是数据的一点点增量,然后根据这些增量,在分析端重建数据库的状态。但是CDC就像是一件非常痛苦的事情。它基本上是业界的标准。每个人都在使用它。但是……它最终会导致,我想许多数据工程师最终都会在凌晨3点被叫醒……因为某些数据管道(pipeline)的问题。

Original English

Speaker B: Um and what it does it reads the bin log of the database. And if you don't understand what bin log is fine. The but it's a little delta of the data and then reconstruct based on the delta the state of the database um on the analytic side. But CDC is like a very painful thing. It it's how basically standard in the industry. Everybody uses it. But um it ends up being so if I think many data engineers ends up being woken up at like 3:00 a.m. um because of some pipeline thing.

Speaker A: 你知道,我的解释是,大家都像是,你知道,仅仅靠做CDC就成了一家价值50亿美元的公司。

Original English

Speaker A: You know, my my explanation is like everybody is like a you know, became a $5 billion company just doing CDC.

Speaker B: 是的,完全正确。CDC就像是一个非常……它是最枯燥但也是支撑现代社会最基本的操作之一。但它是如此脆弱,以至于我们开玩笑说,它应该被称为“持续数据损坏(continuous data corruption)”,因为你可能会在你的OLTP数据库上改变了架构(schema),然后CDC管道就无法处理这种架构变更了。接着所有东西就崩溃了。

Original English

Speaker B: Yeah, exactly. CDC is like a very it's one of the most boring but one of the most fundamental operations like powering modern society. But it's so brittle that uh we joke that it's should be called continuous data corruption because you might change your schema on your OLTP database and then the CDC pipeline fails to handle the schema change. And then everything goes out.

Speaker A: 我的意思是,你可以使用各种各样的技巧,比如你可以加入一些版本控制或者之类的东西。但是……

Original English

Speaker A: I mean there's all sorts of tricks that you can do like you add in like some versioning or whatever. But

Speaker B: 是的,但这通常来说非常复杂。比如,我想在我的主题演讲中,我问观众,如果他们喜欢自己的CDC管道请举手。可能只有大概两三个人举了手。所以如果像SingleStore在大概十年前……我想业界就有这样一个想法:“嘿,如果我构建一个能够同时处理这两种工作负载的单一数据库会怎样?”

Original English

Speaker B: Yeah, but it's a very in general very complicated. Like I think at my keynote I asked the audience put up their hand if they love their CDC pipeline. Only like maybe two people put it up. So if single store like about maybe a decade ago I think the industry had this idea, "Hey, what if I built a single database that can handle both workloads?"

Speaker A: 顺便说一句,这是每一个做数据库的人一直梦寐以求的。

Original English

Speaker A: Which like by the way every database person ever has ever always dreamed about this.

LTAP架构与统一存储的探索

Speaker B: 是的,是的,这是数据库工程的圣杯。为什么不构建一个能够同时做这两件事的单一系统呢?但这最终只是导致了很多妥协。嗯,其中一个,我认为第一个问题是,嘿,他们说Postgres有一个庞大的生态系统。对吧?你希望使用为Postgres构建的那些工具。而Spark,举个例子,也有一个庞大的生态系统。有许多你想要使用的库。如果你现在要创建一个新的东西,你就没有生态系统。你往往会创建一个新的、更小的专有API,而这样你两边都会缺失。而且在性能上,也很难使其在两端都具有可比性。所以结果往往是两边都不讨好。

我们关于LTab的整个想法,显然是在“HTAP”这个词上玩的一点文字游戏,也就是我们认为这才是做对了的HTAP。HTAP想要为两者构建一个单一的引擎。而我们认为,通过统一存储,你可以获得99%你所需要的东西。只需拥有一个单一的存储层。一旦你有了这个单一存储层,如果你的Postgres数据库正在以列式的格式写入数据,所有的分析系统就可以直接去读取那些数据,没有任何延迟。对吧?中间没有数据管道,所以所有的数据都可以立即用于推理分析。

我想我之前告诉过一些客户,嘿,当我们谈论这个的时候,我说这对AI智能体(agents)将会超级有用。我一开始其实自己也不太相信,尽管我们写了那样的定位。但是昨晚我和一位澳大利亚的客户共进晚餐,他们真的告诉我,哦,嘿,我们遇到的一个大问题是,我们从我们的服务中获取了所有的这些日志,我们看到了SLA的下降并想要去调查。但是,对于那些智能体来说,它们甚至没有办法了解在实际数据库本身内部到底发生了什么。我们所能看到的只是数据库和服务的类似产品遥测数据。如果你了解,举个例子,实际上是谁在下那些订单,你实际上能让那些智能体变得强大10倍。嗯,发生了什么?他们具体在做什么?所以现在我实际上被我们自己的理念说服了。

Original English

Speaker B: Yes, yes, this is the the holy grail of database engineering. Why not build a single system that can do both of this? But it ends up just being a lot of compromises. Um one, I think one of the first issue is that hey, each they say Postgres has a massive ecosystem. Right? You want to be using the tools that's built for Postgres. And Spark, for example, had a massive ecosystem. There's a lot of libraries you want to use. If you were to create now a new thing, you don't have an ecosystem. You tend to create a new smaller proprietary API and you're lacking both. And it's also very difficult to make it performance-wise to be a sort of comparable on either side. So it ends up being actually sucking up both. And our whole idea of LTab is kind of obviously a word play on the term HTab, is that we think this is HTab done right. HTab wants to build a single engine for both. We think you can get 99% of what you need by unifying the storage. And just have a single storage layer. And once you have the single storage layer, if your Postgres databases are writing data in a column-oriented format, everything analytics can just go read that data directly without any delay. Right? There's no pipeline in between, so all the data would immediately be available for reasoning analytics. I think I was telling some customers earlier, hey, when we talked about this is going to be super useful for agents. I actually at first didn't really believe in it myself, even though we wrote that positioning. But then last night I was having dinner with Australian customer and they actually told me, oh hey, one of the big issue we have is we have all these logs from our services and we see SLA dips and want to investigate. But then there's no way for those agents to even understand what's going on in the actual databases themselves. All we see is just like product telemetry of the database and the services. You would actually make those agent 10 times more powerful if you understand, for example, who's actually placing those orders. Um what is happening? What exactly are they doing? So now I'm actually sold on our own message.

Speaker A: 是的。

Original English

Speaker A: Yeah.

Speaker B: 嗯,我认为它真的是让你基本获得了HTAP圣杯的几乎所有好处,那就是,嘿,让数据立即可用于推理分析。

Original English

Speaker B: Um I think it's really kind of it gets you basically the almost all of the benefits of the HTAP holy grail, which is hey, make the data available immediate for reasoning analytics.

Speaker A: 是的,我想你知道,就像人类普遍具有智能并且希望有能力和权限去查询任何东西一样。即使在他们做工作的同时,他们也需要历史记录,他们需要上下文,以及比如,他们还能从哪里获得上下文呢?那就是一种分析型的工作负载。

Original English

Speaker A: Yeah, I think you know in the way that humans are generally intelligent and want to have the ability and access to to query anything. Even while they do the work, they also need history and they need context and like where else do they get context? That's an analytical workload.

Speaker B: 完全正确。

Original English

Speaker B: Exactly.

Speaker A: [笑声]

Original English

Speaker A: [laughter]

Speaker B: 是的,我记得当我们的数据库出现故障时,工程师说,好吧,我不能直接在上面运行一个巨大的查询来看看发生了什么,因为那会把数据库弄崩溃,并且会让它受损更严重。像这样的事情就可以通过这个方法消除,因为你会启动一个完全独立的机器集群来做分析。你不会让那个仍在尝试提供服务的主数据库超载。

Original English

Speaker B: Yeah, and I remember when we had incidents with with our databases and the engineer said, well, I can't just run a giant query on it to see what's going on because that's going to bring down the database and it hurt it even more. Like that's the kind of stuff that this gets rid of because you spin up a whole separate fleet of machines that's doing the analytics. You're not overloading like the main database that's still trying to serve stuff.

Speaker A: 是的。

Original English

Speaker A: Yeah.

Speaker B: 所以,这一直是一段时间以来的梦想。为了达到今天这个状态,必须完成什么呢?就像,你知道,是的。我觉得你已经宣布过好几次这个概念的变体了。但当时都不如 L-TAP 这么清晰。

Original English

Speaker B: So, this has been a dream for a while. What had to get done in order to get to today? Like, you know, yeah. I feel like you have announced variants of this several times. But it wasn't as clear as L-TAP.

Speaker A: 是的。

Original English

Speaker A: Yeah.

Speaker B: 我觉得 L-TAP 就像是一个宣告,好吧,伙计们,我们搞定了。

Original English

Speaker B: I think L-TAP is like a like, okay, we've got it guys.

Speaker A: 我当时在和Meta的一个人聊天,然后他问我,嘿,这其中的玄机是什么?为什么现在这是可能的了?我认为现实是,我们花了很多时间真正致力于研究数据湖基础架构。我的意思是,显然很大一部分来自Neon团队,也就是存储与计算的分离。事实证明,距离目标只有一步之遥。从那里走向这个L-TAP的理念,也就是,嘿,我们在Neon的架构和数据湖基础架构中,我们将数据以面向行的格式写入开放的数据湖。但在那里,我们是以Postgres页的形式写入的。

实际上,Ali和我花了很多时间讨论,嘿,我们能不能直接把写入格式改成面向列的?我们就在那里讨论,然后有一天,我们的一个工程师,他其实超级聪明,走进来说,嘿,我刚做了一个原型,能行。

Original English

Speaker A: I was talking to somebody at Meta and then he was asking me, hey, what's the catch? Why is it possible now? And I think the reality is we took a lot of time to actually worked on the lake base architecture. I mean, obviously a lot of it came from the Neon team, which is hey, separation of storage from compute. And it turned out it was just a tiny little step away. Going from that to this L-TAP idea, which is hey, we just in the Neon architecture and in the lake base architecture, we're writing data in row oriented format to the open data lake. But in there, we're writing in Postgres pages. Actually, Ali and I were spending a lot of time debating, hey, can we actually just change that to write in column oriented format? And we're just debating and then one day, one of our engineers, who's actually super smart, came in and said, hey, I just prototyped it, it works.

Speaker B: 等等,原型的功能是什么?

Original English

Speaker B: Wait, prototype what?

Speaker A: 原型就是,不在数据湖中以面向行的格式存储数据。

Original English

Speaker A: Prototype instead of storing the data in the data lake in the row-oriented format.

Speaker B: 比如Postgres的数据,将它们写成Parquet格式。

Original English

Speaker B: Like Postgres payments, write them in parquet.

Speaker A: 是的。嗯,他只是观察到,嘿,我们的存储集群里有很多额外的空闲CPU。我们可以利用这些CPU来进行从行到列的转码(transcoding),其中行格式适合OLTP,而列格式适合分析。嗯,所以我们就利用那个时候来做转码。事实上,一旦你对数据进行转码,数据的压缩效果会更好。所以从那些写入到,比如说,S3或者其他类似数据湖的对象存储服务中,你实际上可以写入得更快,因为它们现在变小了。

Original English

Speaker A: Yeah. Um and he just make the observation that hey, our storage fleet is has a lot of extra idle CPUs. And we could use those CPUs to do the transcoding from row to column, where row is good for OLTP, but column's good for analytics. Um so let's do the transcoding at that time. And as a matter of fact, once you transcode the data, the data compresses better. So from those services writing to, for example, S3 or other data lake like object stores, you can actually write them faster cuz now they are now smaller.

Speaker B: 是的。

Original English

Speaker B: Yeah.

Speaker A: 所以这里没有开销,在性能上没有任何妥协。

Original English

Speaker A: So there's no overhead, it's no compromise in performance.

Speaker B: 没有开销。

Original English

Speaker B: overhead.

Speaker A: 是的,不过是因为反正我们有额外的CPU。

Original English

Speaker A: Yeah, but because but we had extra CPUs anyway.

Speaker B: 反正有闲置的机器集群,是的。

Original English

Speaker B: fleet anyway, yeah.

Speaker A: 所以,争论就结束了。我的意思是,这是一个经典的关于科技问题的……

Original English

Speaker A: So the the debate ended. I mean, it's one of the classics of the tech issue of a

创新文化与执行力

Speaker A: 尽管有很多的争论,但后来有人真的就付诸行动了,只是试着去制作一个原型,而且竟然成功了。

Original English

Speaker A: lot of debate, but then somebody actually went ahead and just tried to prototype it and it worked.

Speaker B: 但是,像对公司如此具有战略意义且如此重要的事情,我原本预期会有一个类似于启动会议那样的环节,或者是设计文档,结果竟然什么都没有。

Original English

Speaker B: But like something this strategic and important to the company, I expect there to be like a kickoff thing, like a design doc, nothing like that.

Speaker A: 什么都没有。他……他只是……我们在许多许多次的会议中一直在争论 [笑声],然后我们就只是一直从第一性原理出发,争论它是否具有可行性,然后某个人就直接去把它做出来了。

Original English

Speaker A: Nothing like that. He He just We We were debating [laughter] in many many meetings, and then we're just debating whether it's possible or not from first principle, and then somebody just did it.

Speaker C: 是的,我的意思是,如果你能创造出一种环境,让人们愿意那样去做,那将会是非常棒的。而且这在 Omni John 项目中也发生过一些。我想,如果如果我只是拿出一份文档,然后说像“我们可以一起做这些东西”,每个人都会,你知道的,都会想“哦,这个方面怎么办?那个方面怎么处理?” 但是,接下来如果你亲自去尝试一下,它就会有所帮助。然后如果你有了真实的用户,他们会去猛烈地测试它,然后像……它仍然能够正常运作,或者在这种情况下,如果你有了工作负载,你……你就能知道这个工作负载到底是什么样子的,你可以直接测试相同的模式,然后,是的。抛开技术层面不谈——虽然技术层面非常酷——但这就像是最重要的事情,那就是创新的文化。而且你不需要来请求我的许可,你也不需要去走一个类似于完整正式流程的东西,就是去做,你知道吗。

Original English

Speaker C: Yeah, I mean, if you set yourself up so people do that, that would be great. And that happened a bit with Omni John too. I think if if I just had a doc and like we can make these together, everyone would you know, would think oh, what about this? What about this? But then you if you try it out, it helps. And then if you have real users and they bash it and like it's still working or in this case, if you have the workload, you you know what the workload looks like, you can just test the same pattern and yeah. Tech aside, which is very cool, this is like the most important thing, the culture of innovation. And you don't have to ask my permission, you don't have to like do a whole form formal process, just do it, you know.

Speaker D: 嗯,特别是在当今这个时代,我认为有了人工智能的加持,去那样做实际上变得更加容易了。

Original English

Speaker D: Well, especially these days, I think with AI, it's actually easier to do that.

Speaker B: 我觉得你展现得非常真实,我的意思是,我见过很多诸如大型公司的高管团队(C-suite),并且就像……我认为在达到了那种规模的时候,事情是会慢下来的,而且我确定你肯定已经感受到了,但不知何故,你拥有一批这样核心的成员,他们好像是不受这种规模法则限制的,是可以豁免的。

Original English

Speaker B: I think you are very real I mean, I've met a lot of C-suite of like large companies and like I think that at scale is things slow down and I'm sure you felt it already but somehow you have this core of people that like are exempt.

Speaker C: 这是如何做到的?

Original English

Speaker C: How?

Speaker A: 我认为我们雇佣了并且我们在与非常非常优秀的人才一起共事,这不仅是其中极其重要的一部分,并且我们要赋予他们权力。但同时,也许我们自己在战壕里(在一线)花费大量的时间也是至关重要的。

Original English

Speaker A: I think we hire and we work with really really good people and that's a very important part of it and empowering them but also spending a lot of time maybe us in the trenches matter a lot also.

产品策略与聚焦

Speaker C: 是的,我想我的意思是,我认为首先人们可以某种程度上适应身处在一个更大的公司里,所以这会有所帮助。并且我们想要确保他们知道自己可以去尝试新事物,可以去平息争论,并且拥有大量过去是如何完成工作的先例,或者在测试阶段推出一个产品,诸如此类。然后另一件事我确实认为,作为一个公司,尽管规模庞大,我们并没有推出那么多,比如,那么多的产品。我们努力保持产品线的相当连贯性,那……那实际上是公司的整个理论基础,就像是……与其拥有像 20 个亚马逊服务那样,你需要去搭建比如一个分析和机器学习的技术栈,你不如只拥有一个平台,而且它有着像相同的 API,在所有的服务中有着相同的语义,拥有同一份数据的副本。所以这需要类似于统一化的过程,然后我们基本上是一次只增加一样新东西。比如我们通过 Delta Lake 增加了存储功能——我们以前从来不做任何存储的——然后我们增加了 SQL,你知道的,我们增加了机器学习平台方面的东西。所以,但是是的,不要……不要做太多,但是要把那些事情做好。而且……而且这也同样能帮助,你知道的,帮助保持它的可管理性。

Original English

Speaker C: Yeah, I think I mean I think first people could kind of adapt to being in the larger company so that helps and we we want to make sure they know that they can try stuff and and settle debates and and have a lot of examples of how it was done before or launch a thing in beta or whatever and then the other thing I do think as a company like despite the size we don't launch that many like products we try to keep it pretty coherent that's that was actually the whole like sort of theory of the company was like instead of having like 20 Amazon services you need to set up like a analytics and machine learning stack you just have one and it's like the same API, the same semantics across all of them, the same copy of the data so that requires like unification and then we basically added one more thing at a time like we added storage with Delta Lake we didn't used to do any storage then we added sequel you know we added machine learning platform stuff so but yeah don't don't do too many but do those things well and and that also helps you know helps keep it manageable.

Speaker B: 是的。

Original English

Speaker B: Yeah.

Speaker C: 另一件我们经常十分鼓励的事情是,与其为了所有事情去进行大包大揽的构建(boil the ocean),不如让我们弄清楚,我们如何能渐进式地去完成它,我们如何……如何才能非常快速地完成它?就像我们的许多产品,它们是在几周的跨度内构建完成的。然后我们会去问,嘿,就像……通常我向任何正在构建产品的团队提出的第一个问题就是:目标客户是谁?你在和谁合作?你和他们是以名字直呼其名地熟识吗?你们会互相发短信吗?我认为拥有那种非常紧密的反馈循环是极其关键的。

Original English

Speaker C: The other thing we kind of encourage a lot is instead of building to boil the ocean for everything let's figure out how do we do it incrementally how how do we do it very quickly like many of our products they're built in the span of weeks and then we go to hey like usually my first question to whoever team is building is who's the target customer who are you working with are you on a first name basis with them are you texting with them? I think having that very tight loop.

Speaker B: 你能提及另一个在脑海中浮现出来的,在这种模式下的产品发布案例吗?我只是为了……我只是想让你就那方面提供一些背景信息。

Original English

Speaker B: Can you bring up another launch that comes to mind with in this kind of thing I just to I just want to give you some background on that way.

Speaker C: 那个……呃……

Original English

Speaker C: The uh

Speaker A: 客户是谁?

Original English

Speaker A: Who's the customer?

Speaker C: 是的,我甚至觉得……这实际上更多的是一个内部的项目。因为我们会把它用作我们的开发者……就像,是的,基本上整个像人工智能团队都获得了使用它的权限并且一直在使用它,并且我们确保它从一开始就能在我们内部的代码库上正常运行,那是一个单一的代码仓库(mono repo),规模极其庞大。我们给了他们一些类似的基础设施。我们给了他们大量类似 token 额度的资源。所以那全部是开发者。是的。我们也有其他的例子。我不……这算是……我认为这是一个公开的故事了,但是……

Original English

Speaker C: Yeah, I'm even was more of an internal thing actually because we would use that for our developer like yeah, basically the whole like AI team got access to it and was using it and we made sure it works from the beginning with our internal code base which is a mono repo that's like enormous. We gave them like some infrastructure. We gave them lots of like token capacity. So it's all the developers. Yeah. We had others. I don't This is I think a public story but

Speaker B: 我正想问一下关于交易市场(marketplace)、开放共享平台,它们所有这些都有……我只是不太清楚地记得具体是哪几个公开提及过这个情况。

Original English

Speaker B: I was going to ask marketplace open sharing all of them had I I just don't remember exactly which ones publicly referenced it.

与早期核心客户的合作

Speaker C: 是的,他们也有其他的……嗯,在公司非常非常早期的时候,有像 Delta Lake 这样的项目,那是我们开发的事务性存储层。当时我们拥有的最大客户提出,说类似“好的,我需要一些……我想要在云端有一些东西。因为你知道的,我……我……如果我们网络的其余部分被攻破了,像这个东西需要独立开来,以便存储和查询事件。”然后他来找我们谈。他说,“好的,这是每秒事件的速率。这是我想要的类似于新鲜度的数据。你们能做到吗?”所以那个规模比我们当时拥有的任何工作负载都要大得多。然后我们有一位负责那个项目的工程师,我的同事 Michael Armbrust,他就是一直致力于让这个项目运作起来,一旦它对他们有效了,你知道的,它对其他所有人也就有效了。是的,这是在公司早期阶段,大概成立四年或者更早一点的时候。

Original English

Speaker C: Yeah, they had other Well, very very early in the company there was like Delta Lake which is the transactional storage layer we did. We had our largest customer at the time said like okay, I need some I want something in the cloud cuz you know, I I if the rest of our network is compromised like this thing needs to be separate to store and query the events and then talked to us. He said okay, this is the rate of events per second. This is like the freshness I want. Can you do it? So that was like way larger than any workload we had and we had our engineer working on that my Michael Armbrust and he worked just to make this work and once it worked for them, you know, it worked for everyone else. Yeah, this was early in the company probably like four years in or so.

Speaker B: 20……2018年?

Original English

Speaker B: 20 2018?

Speaker C: 是的,2017年、2018年的样子。

Original English

Speaker C: Yeah, 17 18.

Speaker B: [笑声]

Original English

Speaker B: [laughter]

Speaker A: 是的,数据无尘室(clean room)功能,这基本上就是一种让你在不共享底层数据的情况下,却允许特定操作的数据共享方式。最初那些功能实际上只为两个客户做了开发。我认为整个行业有一种感觉,嘿,也许如果你过度拟合于(overfit)像一到两个客户的需求,这对你来说将会是非常糟糕的,但我认为过度拟合的缺点远远小于它带来的好处本身。而且如果你试图变得过于雄心勃勃并且想要一次性解决所有问题(大包大揽),那反而是一个大得多的问题。

Original English

Speaker A: Yeah, clean room which is basically how you share data in a way without sharing underlying data but you allow specific operations. Those were done effectively initially just for two customers. I think the industry has a sense of hey, maybe if you overfit to like one or two customers going to be really bad for you but I think the downside of overfitting is much smaller than me the upside itself. And if you sort of try to be too ambitious and boil the ocean it's a much bigger problem.

Speaker B: 是的。

Original English

Speaker B: Yeah.

Speaker A: 因为你最终可能实际上连一个客户都没有。

Original English

Speaker A: Cuz you might end up actually having no customer.

Speaker C: 是的,那更是……那是更有可能出现的结果。然后……然后你可以在此基础之上进行一些调整转向(pivot)。我确实认为会存在像糟糕的客户这样的事情,有时候你是应该解雇他们的。

Original English

Speaker C: Yeah, that's more that's the more likely outcome. Then then you can sort of pivot from there. I do think there is such a thing as a bad customer that sometimes you should fire.

科技公司与传统企业的差异

Speaker A: 他们有时候可能是存在的。如果你推动……嗯,我认为我们……我们可能会看到,并且也许很多新一代的人工智能公司也正在看到的一个挑战是:所以说,科技公司与非科技公司,或者是传统的企业是非常非常不同的。并且如果你仅仅只是为了科技公司去优化一切,当你把它们扩展到科技公司之外的领域时,你可能会面临非常大的挑战。

Original English

Speaker A: They could exist sometimes if you drive Well, one of the challenge I think we we probably see and maybe many AI so newer generation companies are seeing is so tech companies are very very different from non-tech companies or traditional enterprises. And if you optimize everything just for tech companies, you might have very challenges scaling them outside of tech companies.

Speaker B: 好的,那像是……那像是你经常思考的前三大区别是什么呢?

Original English

Speaker B: Okay, what are like what are like top three differences that you always think about?

Speaker A: 治理能力(Governance)是一个很大的区别。

Original English

Speaker A: Governance is a big

Speaker C: 我认为,是的,一个很大的区别就像是,是的,安全性。你知道的,数据隐私、治理,所有这些东西。所以通常如果你在构建某种像 B2B 或者开发者工具类的产品,像你最大的市场将会是企业级客户。但是这非常不同。一个已经存在了例如,你知道的,它拥有某种形式的 IT 系统长达类似 30 年的公司,他们有那么多遗留的系统,或者他们在受监管的领域运营。而一个初创公司,甚至像是……你知道的,就像是一种更近期的科技公司,所有的、一切都是崭新的、一种原始未被污染的状态。所以是的,这只是不同而已。如果你从未与企业打过交道或者身在其中工作过,你……你就不会了解这一点。

Original English

Speaker C: I think yeah, a big one is like yeah, security, you know, data privacy, governance, all that stuff. So usually if you're building some kind of like B2B or developer tool, like your biggest market is going to be enterprises, but it's just very different a company that's existed for like, you know, it's had some form of IT for like 30 years. They have so many legacy systems or they operate in a regulated space. Whereas a startup or even like you know, like sort more recent tech company, all the everything is new and sort of pristine. So yeah, it's just different and if you've never worked with enterprises or been in one, you you just won't know about it.

Speaker B: 而且采购流程可能也是非常不同的。实际上会有多得多的利益相关者。

Original English

Speaker B: And the procurement process is probably quite different. There's actually far more stakeholders.

Speaker C: 这是一个方面。是的。另一个很有意思的部分是,我认为一些科技公司……你知道的,人们会说:“哦,我能自己去搭建那个东西”,对吧?我……我……我……我……我……我直接自己搭建就行了。所以……所以然后你就会去……

Original English

Speaker C: is one. Yeah. Another piece that's interesting is I think some tech companies you know, people will say oh I can build that myself, right? I I I I I I'll just build that myself. So so then you go

Speaker B: 我不认为人们会对 Databricks 说那样的话。

Original English

Speaker B: I don't think people say that about Databricks.

Speaker A: 呃……

Original English

Speaker A: Uh

Speaker B: [笑声]

Original English

Speaker B: [laughter]

Speaker C: 是的,他们确实会。他们会的。他们会的。他们会的。他们确实会那么说。他们会的。

Original English

Speaker C: Yeah, they do. They do. They do. They do. They do. They do.

Speaker A: 是的,我的意思是……是的,并且这取决于具体的……团队以及各种情况。但另一方面,像许多大型企业会说:“事实上,我不要自己搭建。我从来不想涉足去构建那种系统的业务中。”就像是我不想我的……你知道,无论我是个零售商还是别的什么,我绝不希望因为某个奇怪的类似于……书呆子一样的家伙没能把流处理管道弄好而导致系统宕机。那就是我想表达的……

Original English

Speaker A: Yeah, I mean yeah, and it depends on the the teams and things. But on the other hand like many of the enterprises say actually I don't. I never want to be in the business of building that. Like I don't want my you know, whatever I'm a retailer or something, I never want to be down because like some weird like nerd like couldn't get streaming pipelines working. That's is what I'm

Speaker B: 是的,老实说这让他们成为了非常棒的客户,对吧?

Original English

Speaker B: Yeah, this makes them great customers to be honest, right?

Speaker A: 但是你必须得理解这一点,如果没有在那里面工作过的经验和类似的背景,这是很难体会到的。像你可能无法真正领会其中的意味。

Original English

Speaker A: But you have to understand that it's hard without having worked there and stuff. Like you may not appreciate.

Speaker C: 听着,我觉得他们都很棒。别误会我的意思。他们有着不同的挑战。但是许多的科技公司,很肯定的是那里有大量多得多的 DIY(自己动手搭建)文化。

Original English

Speaker C: Look, I think they're all great. Don't get me wrong. They have different challenges. But the many of the tech companies, for sure there's a lot far more DIY.

Speaker A: 在另一方面,你有一群人,他们是……他们在各自的领域内绝对算得上是专家。比如,他们正在制造飞机。他们正在,你知道的,设计药物,或者其他什么工作。并且他们只想要一座通往知识的桥梁,在那个层面上,像他们根本不想去学习,你知道的,数据库之类的东西或者是别的什么,不管我们觉得它们有多酷。

Original English

Speaker A: On the flip side, you have people who are they're very much experts in their domain. Like they're building airplanes. They're, you know, designing medicines, whatever. And they just want a bridge to the knowledge where like they don't want to learn, you know, databases or whatever as cool as we

Databricks 的下一代数据引擎愿景

Speaker B: 我觉得稍微读一读的话,这甚至比一般的软件工程师想象的还要有趣。比如他们只是从来不想去了解。他们只会说,你知道,我有海量的,像矩阵一样的东西,或者任何包含我临床数据的东西。比如,我该如何,你知道,我该如何去对其进行聚类或者做点什么。所以,是的。

Original English

Speaker B: I think it is even as interesting as the average software engineer might think it is to read a little bit. Like they just never want to know. They just say I have you know, giant like, you know, matrix or whatever with my clinical data. Like how do I, you know, how do I like cluster it or whatever. So, yeah.

Speaker A: 是的,没错,确实如此。好的,接下来我想实际构建出那种,可以说是一种梦想中的引擎愿景。

Original English

Speaker A: Yeah. Yeah, that's true. Okay, so and then I wanted to actually build out the the sort of dream engine vision.

Speaker B: [笑声]

Original English

Speaker B: [laughter]

Speaker A: 这到底会通向哪里?

Original English

Speaker A: Where does this all lead?

Speaker B: 所以,我们在大概几年前意识到的一件事是,实际上现有的每一个数据库引擎,特别是在分析领域,差不多都有十年的历史了。基本上所有有一定吸引力的系统都有大约十年的历史。它们一开始都只针对一些非常具体、狭窄的用例。然后随着时间推移,它们变得越来越成功,野心也越来越大。接着它们尝试支持越来越多的用例。但是,支持这些用例最快的方法往往是在最初创建的基础上进行修修补补的破解(hack)。它们最初并不是为这些用例设计的。但后来,你或多或少还是能够凑合支持它们。不知不觉中,经过10年这样有机演进之后,它变成了一堆庞大且臃肿的代码堆。嗯,这其中也包括了 Databricks。我认为,很少很少有公司,或者说很少有系统有这个胆量说:我们从头开始吧。让我们回到原点,利用我们今天经过十年的工作负载积累以及可能高达数十亿收入所了解到的一切知识,尝试重新从头编写它,并切实确保它能运转良好且能支持所有这些用例。所以,我们开始这么做了。但是,这是一个非常雄心勃勃的项目。顺便说一句,你可以在维基百科上搜索一个叫“第二系统综合征”的词。

Original English

Speaker B: So one of the thing we realized maybe a couple of years back is that actually every single database engine out there, especially on the analytic side, are kind of a decade old. Um pretty much everything that have reasonable traction are about a decade old. And they all started targeting some very specific narrow use cases. And then over time it's become more and more successful. They've grown in their ambition. And then they try to support more and more use cases. But the fastest way to support those use cases tend to be hack around the that were initially created. They were not for those use cases. And then but you can kind of support them more or less okay. And before you know it, after 10 years of organic evolution that way, it becomes a gigantic pile of Um the and but then that includes Databricks. And very, very few company or very few systems I think have the gut to say let's go start from scratch. Let's go back to the drawing board and design knowing everything we know today after a decade of workflows and probably billions in revenue, let's attempt to rewrite it from scratch and actually make sure it will work and it can support all these use cases. So, we started doing that. But, it's a very ambitious project. Uh by the way, you can search on Wikipedia this thing called second system syndrome.

Speaker A: 是的,我知道这个。

Original English

Speaker A: Yeah, I know that.

Speaker B: 或者叫第二系统效应。

Original English

Speaker B: Or second system effect.

Speaker A: 每一个开发者都必须知道什么是第二系统。

Original English

Speaker A: Every developer must know what a second system

Speaker B: 基本上就是你构建了你的第一个东西,它运行得非常好,而第二个则注定要失败,因为它太过野心勃勃。然后你去问别人。

Original English

Speaker B: It's basically you build your first thing and it works out great and the second one's bound to fail because too ambitious. And then you ask someone

Speaker A: 你知道的,你以为自己无所不知,然后你就会觉得,我这次要设计一个完美的系统。

Original English

Speaker A: you know, you think you know everything and then you're like I'm going to design the perfect system this time.

Speaker B: 是的。结果证明它并不完美,然后开始失败,而你因为太好高骛远永远无法发布,嗯,然后你就完蛋了。实际启动这个项目的工程团队,他们非常聪明。我想我们为 Databricks 聘请了地球上最优秀的一批数据库工程师,他们非常出色。谢天谢地,这不是他们的第二系统。他们中的许多人过去都已经构建过不止两个系统了。

Original English

Speaker B: Yeah. And it turned out it's not perfect and then it start failing and you're too ambitious never launch um and you get killed. The um the engineering team that actually started this, they were brilliant. I think we hired some of the best database engineers um on the planet into Databricks and they were brilliant. Thank god it's not their second system. Many of them have built more than two in the past.

Speaker A: 很好。

Original English

Speaker A: Nice.

Speaker B: 但是,他们仍然对此感到担忧。“嘿,从头开始构建一个数据库引擎,我认为传统的观念是它大概需要 5 年时间才能成熟。这将是一个非常长期的项目。它可能会失败。”嗯,我想其中一个工程师半开玩笑地说:“嘿,也许我们直接叫它‘Stream 引擎项目’(Project Stream Engine)吧。”如果我们用联合创始人的名字来命名,也许我们就不至于被取消或扼杀了。不过,我认为他们构建的东西相当了不起。他们算是从范式(paradigm)的角度改变了构建数据库引擎的方式。通常,当你构建一个数据库引擎时,你会阅读大量的学术论文,试图了解最新的算法和数据结构是什么,然后把它们拼凑在一起,看看它们是否有效。这里面也有很高的失败风险,因为那些在纸面上看起来非常棒的东西,可能在 70% 的工作负载中表现确实很好,但在另外 30% 的工作负载中却会适得其反。嗯,他们实际上更像是在建造一个用来构建数据库的工厂。所以,他们花了更多的时间来构建这个工厂,而这个工厂采用了我们拥有的十年追踪数据(traces)。我想他们追踪表里的数据点算下来可能达到了一千万亿(quadrillion)个。

Original English

Speaker B: But, they were still worried about this. "Hey, building a database engine from scratch, I think the conventional wisdom is going to take like 5 years to mature. This will be a very long-term project. It could fail." Um I think one of the engineers kind of joking and say, "Hey, maybe we just call it Project Stream Engine." If we name after co-founder, maybe we don't get cancelled or killed. But, I I think they built something pretty remarkable. Um they went back to they they kind of changed the way the database engines were built from a paradigm point of view. Usually when you build a database engine, you read a lot of academic papers, you try to understand what the latest algorithms and data structures and you put them together and see if they work or not. And there's a high risk of failure there also because whatever that looks really good on paper might work out might might actually look really good in 70% of the workloads, but then it backfires on the other 30%. Um they actually went built a more of a factory for building the databases. So, they spent more time building this factory and the factory takes the decade of traces we have. I think they count as a quadrillion data points in the trace table.

Speaker A: 你们不丢弃任何数据?还是你们会进行采样?

Original English

Speaker A: You don't drop anything? Or you see sample?

Speaker B: 我们当然会采样,但是存在海量的数据。而且他们用这些数据来建立一个模型。嗯,比如一个机器学习模型。不是大语言模型(LM),就是机器学习模型。机器学习模型基本上能非常非常快速地告诉我们,对于任何特定类型的查询,任何算法和实现方式的性能表现如何,而且保真度非常高。基于此,他们就能挑选出最有可能真正在应对不同类型工作负载时发挥作用的算法和数据结构。

Original English

Speaker B: We for sure sample, but the there's like massive amount of things. And the and they use that to build a model. Um like a machine learning model. Not an LM, machine learning model. Machine learning model basically it can very very quickly tell us how any algorithm and how any implementation will perform for any specific type of queries with very very high fidelity. And based on that they can uh pick the most likely algorithm and data structure that will actually help with the different kinds of workloads.

Speaker A: 嗯哼。

Original English

Speaker A: Mhm.

Speaker B: 既在运行时,也在实现时。

Original English

Speaker B: Both at runtime as well as at implementation time.

Speaker A: 嗯哼。

Original English

Speaker A: Mhm.

Speaker B: 因为有无数的……

Original English

Speaker B: Because there's like unlimited number of

Speaker A: 我的意思是,听起来你们想要像路由一样分发给不同的数据结构。

Original English

Speaker A: I mean it sounds like you want to like route to different data structures.

Speaker B: 是的,我是说,如果你仔细想想,一个数据库内部将许多东西实现在一起,但你想确保它们都能彼此良好地协同工作,并且对于任何给定的操作,可能存在不止一种实现。所以我们让它实际成为一个非常非常……现实情况是,某些算法例如在应对极低延迟需求时表现极佳,但在扫描 PB 级数据时可能就不太好用。是的,对吧。其实,吞吐量和延迟之间经常存在这种权衡。

Original English

Speaker B: Yeah, I mean if you think about it a single database has many things implemented together, but you want to make sure they all work well with each other and then for any given operation there might be more than one implementation. So we make it actually one really really Reality is things algorithm that works super well for example for very very low latency might not work very well for say scanning through petabytes of data. Yeah. Right. Actually most often there's a trade-off there between throughput and latency.

Speaker A: 关键维度有哪些呢?比如规模(scale)、吞吐量、延迟,还有别的什么吗?

Original English

Speaker A: What are the key dimensions like scale, throughput, latency, what what what anything else?

Speaker B: 还有数据的分布。

Original English

Speaker B: and and the distribution of data.

Speaker A: 是的。

Original English

Speaker A: Yeah.

Speaker B: 没错。数据的稀疏程度。那非常重要。嗯,你命中相同数据的频率有多高?

Original English

Speaker B: Right. How sparse the data is. That matters very a lot. Um how frequently do you hit the same data?

Speaker A: 对。有多少不同(distinct)的值之类的情况呢?

Original English

Speaker A: Yeah. How many distinct values and stuff like that?

Speaker B: 那些因素影响很大。比如唯一值的数量基本就决定了你执行聚合(aggregation)操作时的内存消耗。你的哈希,比如在某些时候会涉及哈希表。

Original English

Speaker B: Those things matter a lot. Like number of distinct value basically impacts the memory consumption of your aggregation. Your hash like at some point there's a hash table.

Speaker A: 那个,我打算在我的报告里试着把这些都列出来,因为我真的很想要一个分类法(taxonomy)。对我来说,分类体系太有帮助了,因为它涵盖了所有他们需要考虑的方面。

Original English

Speaker A: That somebody I'm going to in my write-up I'm going to try to list all these out because I really want a taxonomy. To me taxonomies are so helpful because it covers everything they should think about.

Speaker B: 我觉得如果你真的试图把它列出来,可能得有上百万种不同的特征。

Original English

Speaker B: I think if you actually try to list it out probably like a million different features.

Speaker A: [笑声]

Original English

Speaker A: [laughter]

Speaker B: 我总是在想,好吧,给我 12 个就行。

Original English

Speaker B: I always want like okay give me like 12.

Speaker A: [笑声]

Original English

Speaker A: [laughter]

Speaker B: 呃,你知道的。

Original English

Speaker B: Uh you know

Speaker A: 你知道,就像有人做过的……我想大约 40 年前 Oracle 的一篇论文就总结过类似于“分布式系统的八大谬误”(the eight fallacies of distributed systems)。那种东西超级有用。

Original English

Speaker A: you know like a someone did like I think a Oracle paper in like 40 years ago did like the This is the eight fallacies of distributed systems. That kind of thing is super useful.

Speaker B: 对,就是那样。

Original English

Speaker B: Yeah, that's it.

Speaker A: 就像是,好吧,仔细考虑这八点。

Original English

Speaker A: It's like, okay, think through these eight.

Speaker B: 但让我给你举一个非常奇怪的例子,但它实际上对性能有着深远的影响,那就是比如:你的字符串是 ASCII 的,还是里面包含 Unicode?我应该如何对其编码?

Original English

Speaker B: But let me give you a very weird example, but it actually has a profound implication on performance, which is like it's your string is ASCII or does it have Unicode in it? How should I encode this?

Speaker A: 我的意思是字符串是最复杂的数据类型。

Original English

Speaker A: I mean strings are the most complex data types.

Speaker B: [笑声]

Original English

Speaker B: [laughter]

Speaker A: 所以说,那个就像是……例如,如果字符串极其密集,你甚至可以把每个字符串转换成……想象一下我要做聚合,你不是用哈希表,而是实际上可以用一个数组,因为如果你的字符串足够密集,如果你只有 256 个选项,你就不需要哈希表。你可以直接做数组查找。

Original English

Speaker A: So the that like for example if strings are super dense, you could actually convert every string into a like imagine I have to do a aggregation instead of having a hash table, you could actually have an array because if your string is dense enough, if you only have 256 options, you don't need a hash table. You can just do array

Speaker B: 是的。

Original English

Speaker B: Yeah.

Speaker A: 查找。

Original English

Speaker A: lookup.

Speaker B: 是的。

Original English

Speaker B: Yeah.

Speaker A: 嗯,那样会快得多。

Original English

Speaker A: Um and that that'll be far faster.

Speaker B: 比如国家代码或者类似的东西。

Original English

Speaker B: like a country code or something.

Speaker A: 没错。

Original English

Speaker A: Yeah.

Speaker B: 没错。

Original English

Speaker B: Yeah.

Speaker A: 所以实际上那个模型里可能有数以百万计的特征。但是利用它,他们首先可以优先考虑那些在实践中可能会真正产生影响的各种算法。而且其中许多非常反直觉。你在实践中发现,那些你认为可能会运行得非常好的东西实际上表现并没有那么好。但更重要的一点是,在运行时,你可以调度最合适的算法和数据结构。

Original English

Speaker A: So it's actually like probably millions of uh features in that model. But using that, they can one basically prioritize the different algorithms that might actually impact in practice. And many of them are very counterintuitive. It isn't actually things that you think it might work super well actually don't work that well in practice. But also more importantly at runtime, you can dispatch the right algorithm and structure.

Speaker B: 我正在倾听这伟大的愿景。我觉得 Databricks 在渐进式演进方面做得非常好。你们是否必须在某个时候强制切换到一个全新的系统,还是说……

Original English

Speaker B: I'm listening to to the the dream. I feel like Databricks is is doing a really good job with the incremental evolution. Do you have to hard cut to a new system at any point or like

Speaker A: 我们的设计方式使它可以是渐进的。所以首先我们发布了一个新的端点(endpoint)。呃,但这是面向更广阔的大环境而言,而不是……我们想做的是,一方面从设计上来说,这个新引擎应该能够做到我们以前能做到的所有事情,并且做得更好。对吧?特别是“更好”这部分是指能够在几十毫秒内完成的极低延迟工作负载。但我们想利用增量的功能模块来渐进式地推出它,这样就不需要花上 5 年时间才能真正看到隧道尽头的曙光。

Original English

Speaker A: We designed it in a way that it can be incremental. So first we're releasing a new endpoint. Uh but but this goes to the broader ocean versus What we wanted to do is one to the by design, this new engine should be able to do everything we're able to do before and better. Right? It's been particular the better part refers to very low latency low latency workloads that can finish in tens of milliseconds. But we want to roll it out incrementally with incremental capabilities so it doesn't take like 5 years to actually see the light at the end of the tunnel.

Speaker B: 我认为那是一项英雄般的壮举。我不知道该用什么别的方式来形容。我真的对任何类型的新工作负载和新数据库都非常感兴趣。我的意思是,很明显,我认为如果我算是证明了……

Original English

Speaker B: I think that's a heroic task. I don't know what what what other way to say it. I am really interested in any any sort of new workload and new databases. I mean, obviously, I think if I've maybe established that

数据库类别的演进与融合

Speaker A: 我算是个小小的数据库书呆子。那些事务型数据库,抱歉,是记账型数据库,比如 TigerBeetle。不知道你有没有见过。

Original English

Speaker A: I'm a little bit of a database nerd. The transactional databases, sorry, the accounting databases, like the TigerBeetles. I don't know if you've seen those.

Speaker B: 它们是做什么的?

Original English

Speaker B: What do they do?

Speaker A: 复式记账数据库。它其实就是专门用来对财务账户和信用系统之类进行建模的。

Original English

Speaker A: Dual entry accounting database. Like it's just meant to really model like financial accounts and credit systems and

Speaker B: 听起来是一个非常具体的领域。

Original English

Speaker B: It's like a very specific

Speaker A: 吞吐量非常高,是的。所以当你提到大家一开始都只做一个功能,然后规模扩大后就开始添加其他功能时,确实就是这样。我最近采访了 TurboFifo 的 Simon,情况也一样。还有 Chroma 也是。2023 年的所有向量数据库公司,现在突然间都变成了提供通用 Blob 存储的公司。

Original English

Speaker A: very high throughput, yeah. Yeah. No, it's so when you were talking about how everyone like starts with a thing and then they they scale up and then they tack on other things. It's exactly that. And then I use recently interviewed Simon from TurboFifo, same thing. Like and Chroma as well. Like they all the vector database companies of 2023 all are suddenly now just we're just generally general storage of blob storage.

Speaker B: 特别是数据库本来就不应该成为一个单独的类别。

Original English

Speaker B: Especially a database should have never been a separate category.

Speaker A: [笑声]

Original English

Speaker A: [laughter]

Speaker B: 我觉得这在以前可能算是个大胆的观点,但现在它已经成了普遍的共识。什么才应该是一个单独的类别呢?如果你知道,如果一切都变成了 ELT,那什么是……

Original English

Speaker B: I think that used to be a hot take. Now it's like the conventional wisdom nowadays. What should be a separate category? You know, if everything becomes ELT, like what's

Speaker A: [笑声]

Original English

Speaker A: [laughter]

Speaker B: 我认为 ELT 的核心论点是,我们并不是在实际的查询层折叠数据库,我们只是在存储层进行折叠。是的。我认为这是非常重要的一部分。而且我们其实觉得,把查询层折叠成一个单一的类似于 HTAP 风格的数据库是说不通的。顺便说一下,我认为很多人都有另一种想法,那就是,如果我只需要担心一种查询语言,而不是同时操心 Postgres SQL 和 Spark SQL,那该多好?为什么不能只有一种呢?但我认为对于 AI 智能体来说这不是问题。智能体能够非常流利地使用 Postgres SQL 或 Spark SQL。它永远不会搞混。只要数据在那里并且可以访问,智能体就能做得很好。

Original English

Speaker B: I think the thesis of ELT is we're not collapsing the databases at the actual query layer. We're just collapsing the storage layer. Yeah. And that's a I think a very important part. And we actually don't think it makes sense to collapse the query layer into a single like HTAP style database. And part of it By the way, the other thing I think a lot of people had is hey, it would be nice if there was only one query language I have to worry about instead of worrying about Postgres SQL and maybe Spark SQL, why not just one? But I don't think that's an issue for agents. Agents are very eloquent in Postgres SQL or Spark SQL. It's never going to get confused. As long as the data is there and it's accessible, agents will do fine. That might have been So,

Speaker A: 是的,而且……这在5年前对人类来说可能是一个问题。

Original English

Speaker A: Yeah, and the 5 years ago might have been a problem for humans.

Speaker B: 随着时间的推移,这个问题也可能出现,但这引导了我们应该如何循序渐进地做事,对吧?我们意识到你现在并不需要它。我们不需要非得解决那个问题,才能从当前的 Delta 应用中获取巨大的价值。

Original English

Speaker B: That could arise over time also, but it should and this is leads to how to do things incrementally, right? Like you we realize you don't need it right now. We don't need to solve that problem to have a lot of value from from the current Delta app.

Databricks 与 Snowflake 的发展之路

Speaker A: 是的,好吧。我打算用一些更尖锐的话题来结束这期播客。

Original English

Speaker A: Yeah, okay. I'm going to end the pod with a little bit of more of sort of spicier things.

Speaker B: [笑声]

Original English

Speaker B: [laughter]

Speaker A: 大家都接受了存储和计算分离的概念,并试图构建,你知道的,各种云。我听过 Snowflake 同样的推销。那么,在他们失败的地方,你们是如何取得成功的呢?

Original English

Speaker A: Everyone has like had the receive within a separation of storage and compute and try to build you know the clouds. I had the same pitches from Snowflake. How have you succeeded where they failed?

Speaker B: [笑声] 太尖锐了。

Original English

Speaker B: [laughter] Rough.

Speaker A: 嗯,我尊重他们作为竞争对手。但客观地说,你们的步伐已经超过了他们。在你们看来,核心的洞察是什么,导致你们走上了截然不同的方向?

Original English

Speaker A: Well, I mean respect that they are a competitor. Objectively, you have outpaced them. What is the core insight from your point of view that you guys just went different directions?

Speaker B: [哼笑] 可能最大的根本区别是,两家公司差不多是同时起步的。两家都上了云,两家都关注计算与存储分离的架构。但最大的区别是,一个是开放的。比如,Databricks 从来没有过专有的数据格式,对吧?我们是从开放的生态系统开始的。最初是 Parquet,然后演变成 Delta 和 Iceberg 之类的。它就像一个巨大的整体。我认为这非常重要。另一点是 AI。我是说,在 2022 年 10 月 ChatGPT 问世之前,我们一直将 Databricks 定位为机器学习加数据。平台的很多部分在构建时都考虑了机器学习的用例。很明显,AI 有点不同,而且 Matei 在这方面花的时间比我多得多。但整个平台就是这样,我们从来没有觉得我们只是一个数据基础设施平台。

Original English

Speaker B: Probably [snorts] the biggest fundamental difference. Both companies started around the same time. Both went to the cloud. Both focused on storage from compute architecture. But the biggest difference one is open. Like Databricks had never had a proprietary format, right? We started with the open sort of ecosystem. Started with Parquet and then evolved into Delta and Iceberg and all that. It's like one big thing. I think that matters a lot. The other one is AI. Um I mean before 2022 October 2022 when ChatGPT came out, we had always pitched Databricks as a machine learning plus data. And a lot of the platform were built with machine learning use cases in mind. And obviously AI is a little bit different. And Matei is like spent far more time there than I do. But the whole platform was we never felt hey we're just a data infrastructure platform.

Speaker A: 就像,不仅仅是 Databricks。

Original English

Speaker A: Like Databricks only yeah.

Speaker B: 我想他们一开始的想法是,“好吧,我们就管理最有价值的数据,并努力让它变得非常快。为此,我们将拥有自己的存储,它是与引擎优化结合的。” 然后目标就只对准那些少部分数据,也就是管理层、财务人员等会查看的数据,并且让这些数据的查询服务变得超级快。这是一个不同的领域。相比之下,我们的起点是批量处理和数据摄取。比如你有一大堆 JSON 日志文件之类的,我们来做这种超大规模的事情,因为这就是 Spark 擅长的大规模批处理。然后,我们将数据保存在一种开放格式中。它可能会慢一点,但数据已经在那儿了,你可以让下游去消费。后来事实证明,从那种在规模和摄取方面非常出色且成本极低的批处理方案,去创建具备商业用户所需速度和特性的较小数据版本,是要容易得多的。

Original English

Speaker B: I think that they started with a like they thought okay, we'll just manage the most valuable data and try to make it really fast. For that, we'll have our own storage, you know, which is optimized with the engine. And then we'll just target like the small amount of data that like the managers and whatever, you know, finance people and so on look at and make that super fast to serve. And you know, it was a different space. Whereas we started with like we'll do the bulk processing and ingest. Like you got a bunch of, you know, JSON log files, you got whatever, we do that very large-scale stuff cuz that's what Spark was was for, the large-scale batch processing like stuff. And then, we'll keep the data in an open format. Might be slower, but like it's already out there, you can consume it downstream. And it turned out that, you know, it's easier to to go from that batch thing that's really good at the scale and ingesting and super low cost and create versions in it that have the speed and features of the, you know, super easy-to-use like smaller data for business users thing.

Speaker A: 然后再去优化。

Original English

Speaker A: And then there's then optimize.

Speaker B: 是的,从开放开始,从大规模开始。在某种意义上,我们算是从他们的上游开始的。实际上,曾经有一段时间我们还互称合作伙伴,因为你说如果把两个解决方案结合起来用,用 Databricks 做数据摄取和计算,然后通过 Snowflake 提供数据表服务,你就能获得所有数据可视化和速度极快的功能。这很棒。但后来我们俩都意识到,客户在跟我们说,“我为什么要用这个额外的东西呢?我为什么不能直接查询你们的表?” 我们的回答是,“不,我们在这方面做得太糟糕了。请使用我们合作伙伴的 SQL 数据仓库产品。” 然后他们也意识到,“等一下,大量的计算正向上游转移到这个额外的东西里了。”

Original English

Speaker B: Yeah, start open and start large. Like in some sense, we started upstream of them. And there was a time actually when we both like sort of listed each other as partners because you said if you use both solutions together, use Databricks for like your ingest and compute and then serve the tables out of Snowflake, you get all the visualization, all the way fast stuff. Like that's great. And then, you know, we both realized like customers were telling us like, "Why do I need this other thing? Why can't I just query your tables?" And we said, "No, we're horrible at that. Like please use our partner for the SQL warehouse stuff." And then they realized like, "Wait a minute, so much of the compute is moving upstream into this this other thing."

Speaker A: 你们不得不进入彼此的领地。

Original English

Speaker A: You have to go into each other's territory, yeah.

Speaker B: 但我认为我们的起点确实是更广阔的范围,并且基于开放的理念。这其实很重要。对于很多大型企业来说,如果你的公司已经存在了比如 30 年,你肯定经历过被甲骨文之类锁定、面临各种疯狂状况的阶段。如果你是那里的首席技术官,要为公司规划未来的架构,你肯定希望能选择一个开放的基础。而且理想情况下,你只希望公司内有一种管理数据的方式。你绝不想要七个不同的系统。

Original English

Speaker B: But I think we did start with like the bigger scope and with the open thing. And that's important actually. Like as a kind of ghost enterprises, like if your company's existed for like 30 years, you've experienced, you know, being locked into Oracle and like all kinds of like crazy things. And if you're the CTO there and you're setting up the architecture for the future for your company, you're going to want to pick a foundation that's open. And you only want like one way to manage data in your company ideally. You don't want like seven different systems.

Speaker A: 我认为数据格式已经赢了。我想现在每家企业都想把数据放在开放的数据格式中。但在当时,这其实是非常有争议的。大概五、六年前,Snowflake 的一位联合创始人实际上还写过一篇名为《明智地选择开放》(Choosing Open Wisely)的博客文章,这基本上是在反驳开放理念。

Original English

Speaker A: But I think the data format have won. Like, I think now every enterprise wants to put data in open data format. But, uh it was actually very controversial like back then. I think five, six When exactly as one of the Snowflake co-founders actually wrote a blog called Choosing Open Wisely, which basically argued against

Speaker B: [笑声] 是的,是的。

Original English

Speaker B: [laughter] Yeah, yeah.

Speaker A: 我想他们可能已经把它撤下了。现在你得去找网页存档了。

Original English

Speaker A: I think they might have taken it down. You have to find the archive now.

Speaker B: 噢,我是说,这文章现在是永远抹不掉了。不,它还在那里。我很喜欢这种视角,因为显然这只有你们这些经营公司的人才会拥有。非常感谢你满足我们的好奇心。这真的是非常不可思议的视角。

Original English

Speaker B: Oh, I mean, it's never going away now. Uh no, no, it's still there. I love the sort of perspective of that only you guys will have because obviously you run the company. Uh and I Thank you for indulging this. It's a incredible uh perspective.

Ali Ghodsi 与 Databricks 的领导团队

Speaker A: 也许这是最后一个问题。听你刚才所说,我觉得我必须给 Ali 极高的评价。

Original English

Speaker A: Maybe one last one. Um as you were talking, I think I have to give Ali a lot of credit.

Speaker B: 嗯哼。

Original English

Speaker B: Mhm.

Speaker A: 他是一位不可思议的 CEO。我认为他是智商、情商、对技术的痴迷、执行力和商业头脑的完美结合。

Original English

Speaker A: He's an incredible CEO. I think he's the perfect combination IQ, EQ, technology obsession, execution, business acumen.

Speaker B: 嗯哼。

Original English

Speaker B: Mhm.

Speaker A: 而且他也是一位创始人,这让他更容易去动员和执行。我认为那是……

Original English

Speaker A: Um and he's also a founder, which makes a lot may him a lot easier for him to mobilize and execute. Um I think that's uh

Speaker B: 哦,就是这样。所以,你们有 Ali,他们没有……好吧。

Original English

Speaker B: Oh, that was that was it. So, did uh you have Ali and he they don't like okay.

Speaker A: 嗯,这里面有……

Original English

Speaker A: Well, there's

Speaker B: [笑声] 还有很多其他的事情,但我认为 Ali 在其中发挥了相当大的作用。

Original English

Speaker B: [laughter] a whole lot of other things, but I think Ali played a pretty big role in the

Speaker A: 我还以为他会主导某些技术上的决策。

Original English

Speaker A: I was I was I was I was I thought he was there was like going to be some technical uh choice that he he contributed to.

Speaker B: 他确实推动了其中的很多事情。比如,在我们面临岔路口的时候,他会坚持走某一条路,后来事实证明那确实是一条正确的道路。是的。

Original English

Speaker B: I he he pushed for a lot of these. Like, there were sort of forks in the road where he pushed for like one way and then it became clear that like that was the right way. Uh yeah.

Databricks 的 AI 模型战略与 Mosaic

Speaker A: 我的意思是,关于你们八个人是如何在一起合作的故事,真的值得写一整本书。我想人们已经做过一些相关的报道了。第二个问题,也不算是个明确的问题。关于 Mosaic。我们社区里有很多人都很好奇 Databricks 在模型方面的发展故事。比如,当你们收购 Mosaic 时,当时的想法是,“好吧,我们可以做微调。我们要打造内部模型。” 因为他们有 Mosaic 的模型。但现在看来你们似乎并没有这么做。而且你们似乎正朝着……

Original English

Speaker A: Uh I mean, there's a whole book that needs to be written about how like the eight of you like, you know, work together and all that. I think there's been profiles that people have done. Second one, not a clear questioning again. Uh Mosaic. Mosaic. A lot of people in our community are in are curious on like what's the sort of the model story of Databricks, right? Like, when you guys bought Mosaic, like the thing was like, "Okay, well, we can do fine-tuning. We're going to do in-house model." cuz they had uh the Mosaic models. And it seems like you're not doing that. And it seems like you're going towards more of

关于 Mosaic 与专有模型训练

Speaker A:关于 Alt App 还有 Harness 那些东西。背后的故事是怎样的?你知道,就是……

Original English

Speaker A: the Alt App and the Harness stuff. What's the story there? You know, just

Speaker B:是的,我想当 Mosaic 刚起步的时候,它最出名的是早期发布开源的大语言模型 (LLM),而且那些都是通用模型。实际上在那之前他们还在做别的事情。他们基本上是在优化训练系统。所以他们拥有世界上最快的图像模型训练栈之类的东西。然后他们决定做 LLM,这是个明智之举。他们在 ChatGPT 之前就进入了这个领域,所以他们拥有首批开源 LLM 之一。

Original English

Speaker B: Yeah, I guess when Mosaic started I think it was well known or became most well known for releasing open source LLMs early on and they were general models. Actually before that they were doing other things. They were about optimizing training systems basically. So they had the fastest image model training stack in the world and stuff like that. And then they decided to do LLMs which was smart. They moved into it before ChatGPT. So they had some of the first open source LLMs.

Speaker A:对。我们采访过 John Franco 和 Abby,聊了 MPC-7B。

Original English

Speaker A: Yeah. We interviewed John Franco and Abby for MPC-7B.

Speaker B:对,没错。哦是的,非常酷。是的。所以我们决定,虽然我们确实发布了开源模型 DBRX,并且达到了超过 Llama 3 的规模,但我们决定要真正把重心放在——未来会有很多人发布模型,与其去开发那些很大程度上仅仅依赖堆砌算力和扩大规模的通用模型,我们更想关注下一步:假设你已经拥有了非常聪明的模型,你该如何让它变得有用?对我们来说,很大一部分工作在于自动化,比如让它非常擅长查询数据。这就是我们称之为 Genie 的第一方智能体。它就像一个虚拟的数据科学家。想象一下,有这么一个人,他对你公司里所有的东西都了如指掌,熟悉所有的机器学习库、所有的数据库、网上所有的资料,你可以直接向他们提问。这就是我们首先想做的事情。这意味着我们不必那么专注于训练某种前沿模型,而是可以使用外部模型或微调过的定制组件来构建一个系统。

不过,我们仍在进行大量的模型训练工作。事实上,我们总是在不断采购大量 GPU 等设备来做这件事。我们在几个方面进行了此类工作。其一,在许多高吞吐量的用例中,如果你拥有一个专门优化的模型,它的表现会比你拿到的任何通用模型都要好得多。一个很好的例子就是理解文档,比如解析 PDF 或 Word 文档。如果你曾经尝试过这么做,你会感到很沮丧,因为你把它发给像 Claude 或其他什么模型,它差不多能理解,但总会弄错一些东西。而且这非常昂贵,你只是把一张图片扔进去,就消耗了大量的 Token。因此,我们的团队构建了这个用于文档处理的视觉模型,它接收一页文档,然后返回一个包含了所有组件的完美 JSON。它非常有竞争力,可能比那些前沿模型便宜 100 倍,而且效果更好。这实际上是由一位来自 DeepMind 的研究员完成的,他也是 Adept 的联合创始人,是非常早期的 LLM 扩展 (scaling) 专家,现在专注于这个领域。

同样地,我们也在为编程智能体的一部分功能开发专用的子智能体 (sub-agents)。如果你看过 Harvey 开发的顾问模型 (advisor models) 相关的内容,还有来自……

Original English

Speaker B: Yeah, exactly. Oh yeah, very cool. Yeah. So we decided, you know, even though we did launch a open source model DBRX and went up to sort of above the Llama 3 scale, we decided that we really want to focus on—there'll be so many people releasing models and instead of doing the general model where, you know, a big part of the recipe is just throwing a lot of compute and just scale, we want to focus on the next step also of let's say you have the very smart model. How do you make it useful? For us it was a lot about automating making it very good at querying data. That's the first party agents we have called Genie. So it's like a virtual data scientist. Imagine there's someone who already knows all the stuff in your company inside out and knows all the machine learning libraries, all the data libraries, all the stuff on the web and you can ask them questions. That's what we wanted to do first. So that meant let's not focus as much on just training some kind of frontier model, but let's build a system using either external models or fine-tuned customized components.

We're still doing quite a bit of model training though and in fact we're always procuring lots of GPUs and stuff all the time to do it. And there's a few places where we're doing that. One is, there are many high volume use cases where if you have a specialized model, it's just so much better than any of the general models you get. A nice example of that is understanding documents, like PDF, Word documents, parsing them. If you've ever tried to do that, it's frustrating cuz you send it to like Claude or whatever, it almost gets it, but it gets some things wrong. And it's super expensive. You just burned a huge amount of tokens plopping an image into there. So, our team built this document vision model that takes a page and gives you back a nice JSON with all the components. And it's very competitive. It's probably 100x cheaper than those frontier models and still better. And that's actually done by one of the researchers who came from DeepMind, was a co-founder of Adapt, a very early LLM scaling person, but focused on this.

Likewise, we're doing specialized sub-agents for part of what the coding agent does. And if you've seen the stuff on advisor models from Harvey, also from...

Speaker A:还有 Anthropic,以及 Cognition(口误),对。

Original English

Speaker A: And Anthropic, Commission also, yeah.

Speaker B:还有加州大学伯克利分校 (UC Berkeley),实际上,我在那里的一个研究生写过一篇关于顾问模型的论文,我想那是在这些模型问世之前。我的意思是,我确信其他人也同时有这个想法,但这确实能帮上大忙。所以,是的,我们今天就在主题演讲中展示了一些东西,关于……

Original English

Speaker B: And UC Berkeley, actually, one of my grad students there wrote a paper called advisor models. I think before those came out. I mean, I'm sure others had the idea at the same time, but that's something that helps a ton. So, yeah, we actually showed some stuff just today at the keynote on...

Speaker A:是 Parth 吗?哦,你认识 Parth?

Original English

Speaker A: Is it Parf? Oh, you know Parf?

Speaker B:Parth,对,对,Parth 正在……

Original English

Speaker B: Parf, yeah, yeah, Parf is...

Speaker A:正在我的活动上演讲。他要在持续学习基准测试 (Continual Learning Bench) 上演讲。

Original English

Speaker A: speaking at my thing. He's speaking at Continual Learning Bench.

Speaker B:是的,没错。我是他在 Adept 的顾问之一,对。

Original English

Speaker B: Yes, yes, yeah. I'm one of his advisors at Adapt, yeah.

Speaker A:我还采访过他兄弟 Chai,因为他也在 Adept。

Original English

Speaker A: interviewed his brother, Chai, cuz he's also at Adapt.

Speaker B:是的,没错。

Original English

Speaker B: Yeah, yeah, yeah.

Speaker A:那一家子都非常聪明。

Original English

Speaker A: That family's very smart.

模型定制化的未来与复合智能体

Speaker B:对,对。[笑] 他们非常棒。所以,我们在做这方面的工作,随着我们在第一方智能体上积累了经验,我们也在和客户一起做这些事。我的感觉是,随着时间的推移,定制模型实际上会变得越来越容易。这也是我们所发现的,因为基础模型变得更聪明了,所以它们在强化学习 (RL) 中已经能生成更好的轨迹 (traces),而强化学习本质上就是从你自己过去的轨迹中学习。然后,合成数据生成现在变得更好、更容易了。我们有完全使用开源模型的管道 (pipelines)。比如用同一个模型生成训练环境,然后训练自己,并在某项任务上击败了 Opus 和 GPT 5.5 之类的模型。所以,我确实认为这会流行起来。关键在于,训练算法的便利性只会随着时间不断提升。现在的问题是,它什么时候能成为主流。比如,我们不需要一个硬核的 LLM 研究员来做我们搞的那种专门的文档解析,什么时候它能变得足够简单,让任何人都能把一些资料丢进去并描述一下任务就能完成。

Original English

Speaker B: Yeah, yeah, yeah. [laughter] They're awesome, yeah. So, yeah, so we're doing some of that and as we get experience with these in the first-party agents, we're also doing them with customers. So, my feeling is like customizing models is actually going to get way easier over time. That's what we're finding cuz the base models are smarter, so they generate better traces in RL already, and then RL is about learning from your own past traces. And then synthetic data generation is way better, way easier now. We have pipelines just using open-source models. Like the same model generates training environments and trains itself and beats Opus and GPT 5.5 and stuff at a task. So, I do think it's going to pick up. The thing is the ease of training the algorithms is only going to go up over time. There's a question of when it crosses into mainstream. Like, instead of specialized document parsing thing we did, where you need a hardcore LLM researcher, when does it get easy enough that anyone can sort of plop in some stuff and describe a task.

Speaker A:对。

Original English

Speaker A: Yeah.

Speaker B:对。

Original English

Speaker B: Yeah.

Speaker A:你知道什么能让它变得简单吗?界面。还有统一的 API。因为显然,如果它不可互操作,你就无法切换。

Original English

Speaker A: Well, you know what makes it easy? Interfaces. And unified APIs. Cuz obviously if it's not interoperable, then you cannot switch.

Speaker B:这也是我们在这个复合智能体 (composable agents) 和 Omnigent 身上看到的趋势。你可以拥有搭载专业模型的子智能体,然后你可以对整个系统进行训练。我认为那也会有很大帮助。

Original English

Speaker B: That's what we're seeing with Omnigent and the composable agents. Like you can have sub-agents with specialized models, and then you can train the whole thing. I think that'll help a lot, too. Yeah.

数据与上下文:新的石油

Speaker A:我想留到最后讨论的一件事——实际上,这是我精心安排的,所以我还有点为自己感到骄傲。Satya 也在谈论这个。几周前我在微软 Build 大会上采访了他。然后他写了这篇文章,我确信你看过,是关于整个移动构建前沿生态系统的。当我和他交谈时,他听起来,呃,更像是一个 Databricks 的 CEO。

Original English

Speaker A: The last thing I was going to leave, actually, I'm sequencing this, so I'm actually kind of proud of myself. Satya is talking about this. I interviewed him at Microsoft Build a couple weeks ago. And then he wrote this essay, which I'm sure you've seen, which is the whole mobile building frontier ecosystem. He sounded, when I was talking to him, more like a Databricks CEO.

Speaker B:[笑]

Original English

Speaker B: [laughter]

Speaker A:嗯哼。我想问的是,这东西在我的圈子里大概率已经火了。我不知道在你们的圈子里火没火。关于将 Token 视为知识产权 (IP)、建立上下文,背后的理论是什么?他基本上是在说,除了数据之外,没有什么是新的石油,或者说上下文 (context) 才是新的石油。你们以前肯定听过类似的版本。

Original English

Speaker A: Uh-huh. Is there a I mean, this thing presumably went viral in my circles. I don't know if it did in your circles. What's the sort of theory of like, I guess tokens as IP, building up the context, you know, he basically said everything but data is the new oil or context is the new oil. Some version of that you guys have heard before.

Speaker B:对,我同意。我认为随着你围绕数据拥有更好的技术,你就可以在你的领域内利用你所拥有的数据做更多的事情。这甚至不仅仅是关于 AI。即使在人们开始实时收集数据的时候——比如我记得所有电力公司都安装了智能电表,所有汽车制造商都开始安装传感器和摄像头。任何技术都会让数据变得更有价值,并能给你带来一些优势。任何能帮助你利用数据并做出决策的东西都是如此。AI 也是一样。以前你有一堆数据就放在那里。现在你可以让一个智能体自动告诉你。例如,以前我发现产品里的某个功能坏了,是因为客户抱怨;现在智能体直接告诉我:“我注意到没人再上传文件了,因为他们遇到了报错”或者之类的情况。正如你在 Raiden 看到的那样,作为一家数据库公司,因为我们掌握了所有查询的历史记录、表布局以及它们的工作方式,我们可以非常快地构建一个全新的引擎,而且效果很好,我们也有信心它会表现出色。所以,我认为这个观点是对的。问题只在于它究竟将如何落地,但我确实认为,正如 Satya 所说的那样,自定义模型的定制化随着时间推移会变得越来越容易。

Original English

Speaker B: Yeah, I agree. I think that the data you have, as you get better technology around that, you can just do more in your domain with it. It's not even just about AI. Even when people started collecting stuff in real time, like I remember all the power companies put the smart meters and stuff, and all the car manufacturers started putting sensors and cameras and stuff. Any technology makes data more valuable and can give you some advantage. Anything that helps you do something with it and make some decisions. And AI is the same way. You had all this stuff that's just sitting there. Now you can have an agent automatically tell you. Like for example, instead of I discovered a feature in my product is broken cuz a customer complained, the agent tells me, "I noticed no one is uploading files anymore cuz they got errors or whatever." And as you saw with Raiden, as a database company, because we have the history of all the queries and all the table layouts and how they work, we can build a new engine very quickly that actually is good and we're confident that it's going to be good. So, I think this is right. I think the question is exactly how it will land, but I do think custom model customization which Satya talked about is going to get easier over time.

Speaker A:顺便说一下,这也是为什么我提到模型这个话题,因为他们有他们的 MEI,而你们没有。这才是最初我想问的问题。

Original English

Speaker A: Which is why, by the way, I brought up the model thing cuz they have their MEI things and you guys don't. That was the mental question.

Speaker B:对,[笑] 我们确实有,我们在做像强化学习微调即服务 (RL fine-tuning as a service) 这样的事情,服务了一批客户。基本上我们有预览版客户,而且我们还有一个通用的被称为 AI runtime 的东西,就像是我们为你提供按需分配的 GPU 集群,里面有软件栈,让你很容易就能进行训练。所以,我们并没有……但那东西已经存在一段时间了。我们提供 GPU 算力已经有一段时间了。

Original English

Speaker B: Yeah, [laughter] we do have, we're doing RL fine-tuning as a service with a bunch of customers. We have preview customers and we have a general something called AI runtime that's like we got you GPU clusters on demand with software stack in there that makes it easy to do training. So, we didn't like sort of but that's existed for a while. We've had GPU compute for a

Scaling the Mosaic Stack and Customer Engagements

Speaker A: 而这正是大量 Mosaic 技术栈被用来帮助扩展的地方。

Original English

while and that's where a lot of the Mosaic, uh, stack went to to help scale that.

Sweaks: 是的。

Original English

Yeah.

Speaker A: 但是,没错,我们发现这种合作参与,像是有两种类型的客户。有一些客户只是想要 GPU 和相关库,以便能够把数据输入、输出并进行监控。这就是 AI runtime(AI 运行时)的意义所在。然后还有一些客户会说:“嘿,你知道吗,你能实际和我一起合作,构建评估体系,构建合成数据吗?”

Original English

But yeah, we found that the engagements, like some of the There's two types of customers. There's some who just want GPUs and libraries to like get data in and out and monitors. So that's what AI runtime is. And then there's some that say, "Hey, you know, can you actually work with me, build evals, build synthetic data,

Sweaks: 是的。那些更倾向于前沿部署的解决方案架构师。

Original English

Yeah. The more forward deployed solutions architects.

Speaker A: 我们正在做的就是这些。随着时间推移,会有越来越多的东西从定制化转变为非定制化。但嗯,这就是目前的现状。

Original English

that's what we're doing. And and as and more things will transition from like being custom to to not. But um that's sort of how it is today.

The Paradigm Shift: Data and Generic Agents

Speaker A: 回到最初的问题,我认为我们的一个论点实际上是,如果你能把数据放在正确的地方就行。现在的 AI 模型已经变得非常优秀了。通用的智能体也相当不错,我的意思是,我想我听到你谈论过 AGI(通用人工智能)已经到来了。它们具有相当好的推理能力。实际上,我认为许多传统的软件都会以这种新范式被某种程度上重写,这种新范式就是:只要把数据准备好。然后他们在上面加上一些智能体。奇迹就会出现。

Original English

Going back to original question, I think one of the thesis we have is actually the ones you can get the data in the right place. The AI models are becoming pretty good. The generic agents are fairly I mean I think I heard you talking about AGI is already here. They have pretty good reasoning capabilities. Actually, I think many of the traditional software will be sort of rewritten uh with this new paradigm, which is just get the data to be there. And then they slap some agent on top. Magic will come out.

Sweaks: 是的。

Original English

Yeah.

Speaker A: 嗯,但是如果没有正确的数据,你真的无法做到这一点。实际上这也是我们进入安全领域的策略,以及我们进入客户数据平台领域的策略。

Original English

Um but without the right data, you can't really do that. And it's actually our approach going to security and our approach to going to the uh customer data platform space.

Sweaks: 是的。

Original English

Yeah.

Speaker A: 就是,比如我们在 Data and AI Summit(数据与人工智能峰会)上发布了两款产品。一款是针对安全团队的,另一款是针对营销团队的。而这些领域都有大量现成的技术存在。我想我们的策略就是,“嘿,一旦你把数据接入进来,在上面加上智能体后,一切都会变得容易得多。”

Original English

Is uh like we we launched two products at Data and AI Summit. One targeting sort of security teams and the other one targeting marketing teams. And those all are have a lot of existing technologies out there. And our I think our approach is just, "Hey, once you get the data in, everything is a lot easier with agents on top."

Closing Remarks and Databricks' Growth

Sweaks: 是的。是的。好吧,而且你们各位真的是非常棒的嘉宾。我太喜欢这次讨论了。我不仅喜欢我们能够深入探讨技术层面的内容,还能探讨文化和战略。我希望这不是我们最后一次交谈。就像,我想说,恭喜你们迄今为止取得的所有成功。

Original English

Yeah. Yeah. Well, and you guys have been fantastic guests. I just love this discussion. I just love the ability to dive in on the tech side, but also culture and strategy. I hope this isn't the last time we chat. Like I mean, congrats on all the success so far.

Speaker A: 谢谢你。也恭喜你取得的成功。[清嗓子]

Original English

Thank you. Congrats on your success also. [clears throat]

Sweaks: 是的。

Original English

Yeah.

Sweaks: 是的。是的,我的意思是,大卫(David)实际上正在支持我的活动,也就是一个,所以我在办……是的,我在办一个会议。实际上我……我已经作为参与者参加 Data AI Summit 很长时间了。我注意到这有点像……这要追溯到 2022 年。当时大概是 90% 的数据内容,然后是 10% 的 AI 内容。我就在想,“好吧,像我们需要一个……我们需要一个基本上是 90% AI 内容的社区活动。而这并不是每个人都在做的。”

Original English

Yeah. Yeah, I mean uh David's actually supporting my uh event, which is a it's a So I run Yeah, I run conference. And it is actually I was I I've been attendee of Data AI Summit for a long time. And I noticed that it was like kind of This is back in 2022. It was like 90% data and then 10% AI. And I was just like, "Well, okay, like we need a we need the community thing that is like just 90% AI. Which like not everybody is.

Speaker A: 是的,是的,这行得通。

Original English

Yeah, yeah, that works.

Sweaks: 所以是的,没错,Databricks 将会……将会参加这个会议,你知道的,我……我只是觉得非常惊讶,看到你们构建出了我所见过的最有趣的云平台,除了你们知道的那些三大巨头之外,并且看到你们发展得如此深远,真是令人惊叹。就像最有洞察力的一点,就像,你知道,我不是一个风险投资人(VC),但我经常在电视上扮演这种角色。就像本·霍洛维茨(Ben Horowitz),当他和你们交谈,就这家公司未来走向给你们建议时,他说不到 1000 亿美元就别卖,或者类似这种版本的故事,对吧?

Original English

So yeah, yeah, so Databricks will be will be at the conference and you know, I I it's just amazing to see you guys build out the most like interesting like cloud that I have ever seen outside of like the you know, the the the the big three and like it's amazing how far you've grown. Like one of the one of the most insightful like you know, I don't I'm not a VC but I play one on TV. Like Ben Horowitz like when he was talking to you guys advising you on just like where is this company going? He was like don't sell until 100 billion or some some version of that story, right?

Speaker A: 他的原话大概是,这家公司应该价值一万亿美元。你们以 100 亿卖掉它简直是在贱卖。

Original English

It was like the the company should be worth a trillion dollars. You're underselling it for 10 billion.

Sweaks: 而且他并不会对每个人都这么说,你知道的,就像……

Original English

And like he doesn't do that for everyone, you know, like

Speaker A: [笑声]

Original English

[laughter]

Sweaks: 出于某些原因,就像你知道的,我……我认为他看到了愿景,但也看到了你们所拥有的无限潜力。

Original English

for some some reason like you know, I I think he saw the vision but also the infinite runway that you have.

Speaker A: 我们很幸运能有本(Ben)的支持。是的,他是一位大支持者。

Original English

We're lucky to have Ben. Yeah, he's a big supporter.

Sweaks: 是的,太不可思议了。好的,那么,非常感谢你们。

Original English

Yeah, amazing. Okay, well, thank you so much.

Speaker A: 好的,非常感谢你,Sweaks。

Original English

All right, thank you so much, Sweaks.

Sweaks: [音乐]

Original English

[music]

📌 文中提及的人物和组织

公司/组织: Databricks

关键字: agentic-workflow data-integration model-customization generic-agents