生产级 AI 行动指南:企业级智能体的部署与落地 AI Engineer 2026-06-18

嘉宾介绍与演讲背景

桑迪: 好的,非常感谢大家来参加我的这场分享。谢谢大家!我是桑迪(Sandy),目前担任 Databricks 的数据与人工智能(Data and AI)技术负责人。在加入 Databricks 之前,我在 Amazon Web Services (AWS) 工作了五年,担任数据与人工智能领域的首席架构师。在过去的几年里,我专注于利用分布式系统和技术来构建和扩展数据与 AI 平台。特别是最近这几年,我一直在与各种客户合作,共同探索如何应用这项全新的人工智能技术。当我说“新 AI”时,其实 AI 技术已经存在了很长时间,但在过去的几年里,我们所有人开始呈指数级地进行实验。通过与 B2B 软件行业以及金融服务等监管行业的不同客户合作,我在构建 Demo(原型演示)以及如何将这些 Demo 推向生产环境的过程中积累了大量的经验教训。因此,在这场分享中,我想分享一个我在一线实战中总结出来的 playbook(行动指南)和框架。当你们思考如何将自己的 AI 系统投入生产环境时,可以直接拿去应用。

Original English

Sandy: All right. Um, thank you for joining my session. Thank you, man. Uh, I'm Sandy. Uh, I'm a technical lead uh, for data and AI at Databricks. Um, prior to working in Databricks, I worked in Amazon Web Services uh, for 5 years as a principal architect for data and AI. Uh, in the past few years, I worked extensively uh, building and scaling data and AI platforms using distributed systems and technology. And in the past couple of years, specifically, I've been working with customers trying to figure out what we do with this new AI technology. Uh, when I say new AI, AI has been here for a long time, but we all started experimenting quite exponentially uh, in the past couple of years, right? And I have learned a great deal of lessons from building demos and how to take those demos to production working with different customers in B2B software industries and then regulating industries like financial services. So, in this session, I want to share a playbook, a framework that I put together from lessons that I have learned working in the trenches uh, that you can take and apply when you think about how to put your AI systems into production.

生产落地三大鸿沟

桑迪: 我觉得这场分享被安排在下午非常好,因为你们现在可以用这个框架,把今天一天在不同分论坛听到的各种零散知识串联起来,看看它们在框架的各个要素中分别处于什么位置。两年前我刚开始做这件事时,我注意到在与客户的每一次交谈中,都存在着一个相同的模式。每个人都想用 AI 做点什么,管理层施加了巨大的压力,要求“必须做点什么”,必须先做出一个 Demo。而每一次交谈几乎都是从“我们应该选择哪种模型”开始的。这不能怪任何人,因为当时的市场环境就是这样,所有人都在讨论模型,模型对我们来说是一项全新的技术。每次谈话的开头都是:我们是用 GPT,还是用 Claude?组织内部会因此产生巨大的争论。然后,你选好了一个模型,构建了一些功能,并以此作为该应用要实现的功能子集。接着,你会在一个受控的环境中去构建它,使用可以预测的数据集、受限的场景。在 Demo 演示时,它看起来效果棒极了,管理层非常高兴,爽快地签字批准,然后把它部署到了生产环境中。然而,仅仅过了几个星期,人们就开始质疑:这个 AI 到底在干什么?为什么它的回答和我们做 Demo 演示时完全不一样?这不仅导致投资回报率(ROI)无法实现,还浪费了大量的时间、金钱和精力去构建一个永远无法扩展到生产环境的 Demo。在这些会议中,我总结出了三个核心洞察,它们与我们讨论将 AI 推向生产环境时的方方面面都紧密相连。第一个是可观测性赤字(observability gap):当我们使用 AI 并将其投入生产时,如果我们无法看到它实际在做什么,无法追踪它做出的每一个决策,那么它在生产环境中就是毫无用处的。第二个是评估赤字(evaluation gap):在过去的很多讨论中,我们并没有真正思考过我们究竟在衡量哪一个关键指标。是的,我们会谈论准确性、延迟和 groundedness(接地性/基于事实性),但我们并没有明确定义到底哪一个具体指标对业务最重要,以及如何构建一个系统来持续衡量它是否在改进。第三个是治理赤字(governance gap):我们没有真正思考过当 AI 在生产环境中出现故障时该怎么办。谁来承担责任?如果凌晨三点发生故障,我该去找谁?谁该对喂给 AI 的数据资产负责?如果 AI 对客户胡言乱语,会发生什么?因为没有明确的问责制,所以没有任何治理可言。这三个洞察促使我构建了一个关于如何将 AI 推向生产环境的框架,目前已经在多个客户组织中落地实施,相信对你们也会有所启发。

Original English

Sandy: And I think this session is nicely placed in the afternoon because what you can do now is in this framework, you can fit the different knowledge that you've gathered attending these different sessions throughout the day and see where they fit in each of these elements in the framework. So, when I started 2 years ago, this is the pattern I noticed in every customer conversation, right? So, everyone wanted to do something with AI. Uh, there was immense pressure from the top to do something, to build a demo. And every conversation started with let's choose the model, right? And it was nobody's fault because the market was like that. We were talking about models, the models were new technology for us, right? And every conversation started, shall we use GPT? Shall we use Claude? You know, there was huge debate within organizations. Then, you would choose a model, you'll build some features, offsets of over what features to build for that application. Uh, you would build that in a controlled environment, so predictable data sets, you know, limited scenarios, and then it looked great as a demo, and then leadership would get happy, they would sign it off, and they'll put it into a production environment. Then, after a few weeks, people would start asking questions that what the hell is AI doing? Right? Why is it not answering the questions the way we expected it to answer when we were doing the demos? Uh, it would result in not only no realization in return on investment, but also loss of money and effort in building these demos that can never scale to production. Throughout these meetings, I gathered three insights that connect to everything that we are talking about when thinking about taking AI to production. The first one is the observability gap, right? When we use AI and put it into production, if we can't see what it is actually doing, if we can't trace every decision that it's making, it's no use in production. Second is the evaluation gap gap. A lot of these conversations that we were doing, we were not actually thinking about what is what is that one thing that we are measuring. Yes, we talk about accuracy, we talk about latency, we talk about groundedness, but we were not defining what is that exact thing like that matters to the business, and how can we build a system that can continuously measure that, whether it's improving, whether it's not improving, like what what is that system that we need to build. And that was that evaluation gap that I noticed. And the third is the governance gap. Like we were not actually thinking what happens when AI fails in production. Who's accountable? Who do I go to when something happens at 3:00 a.m. in the morning, right? Who needs to own the data assets that feed some AI responses? What happens if AI talks nonsense to a customer, right? What what happens, right? So, there is no accountability, no governance around it. And these three insights led me to build a framework on how I think AI should be taken to production, and this has been implemented across multiple customer organizations, and I think this is something that you can pick up from here.

生产级AI五大核心支柱

桑迪: 这就是我今天想分享的五大支柱。在开始任何项目之前,你都必须绝对优先考虑这些支柱。接着,你再去逐步构建它们,最好是按顺序进行。虽然在实际操作中,我知道这个顺序可能无法完全死板地执行,但这些支柱是你必须了解并在开始构建时深入思考的。第一大支柱是评估(Evaluation):在动任何代码之前,在讨论任何模型、任何功能之前,你必须想清楚:当我们构建好这个系统后,我们如何进行衡量?成功是什么样子的?什么样的系统能够帮助我们持续地衡量成功是否达成?第二大支柱是可观测性与追踪(Tracing and Observability):我们如何追踪 AI 做出的每一个决策。这不仅对 AI 系统的性能至关重要,对合规监管也同样重要。在欧洲或许多大型企业中,尤其是受到严格监管的行业,如果没有完善的追踪和可观测性机制,你甚至根本无法获批将 AI 部署到生产环境中。所以,这是必不可少的。第三大支柱是数据基座(Data Foundation):我从两个维度来理解数据基座。一个是“问题数据”(Question Data),这基本上是 AI 回答用户问题所需要的基础数据,它可能是你的预训练数据、微调数据,或者是通过 API 获取来回答用户所需的实时数据。另一个是“追踪数据”(Tracking Data),这与可观测性中的追踪数据有关。但从数据基座和数据策略的角度来看,它必须在数据基座这一大支柱中进行集中处理。因为当你需要同时在企业内部运行成百上千个智能体(Agent)时,你必须有一套完整的数据策略来处理这些追踪数据。第四大支柱是智能体编排(Orchestration):如果你只运行一个智能体,它的表现通常会很好,你也不需要过多考虑编排。但当你引入五个甚至更多智能体时,系统的复杂度会呈指数级上升。它们之间会有多种协调模式,它们需要以不同的方式相互通信,需要等待彼此的响应,这带来了巨大的复杂度。这就是为什么思考在特定系统中如何编排你的智能体变得如此重要。第五大支柱是治理(Governance):在这个支柱中,你需要思考当出现故障时该怎么办。谁来承担责任?我们如何治理数据?如何保障数据的安全?如何保障我们系统的安全?如何确保没有人恶意向我们的智能体注入指令导致其失控,或者给公司带来名誉损失。在接下来时间里,我将深入探讨这五个支柱,并告诉大家在实际落地时应该如何思考。

Original English

Sandy: These are the five pillars, and these are absolutely what you need to think about even before starting a project, right? Then you start build them gradually, preferably in sequence, but in real life, I know that this sequence don't work, but these are the pillars that you have to know about and you have to think about when start building. First one is evaluation. Before touching any code, before discussing about any models, any features, you have to think about when we build this system, how do we measure? What does success look like, and what is that system that will help us continuously measure what success looks like for us? Second is how do we trace each and every decision that AI makes. It's not only important for the performance of the AI system, it is also important for the regulators. In Europe or in a lot of companies, especially in regulated industry, you cannot even onboard AI into production without having tracing and observability in place. So, this is a must-have. The third is the data foundation, right? Uh, I think of data data foundation in in two ways. One is the question data, so that is basically the data needed for the AI to answer questions that users ask to it. So, it could be your pre-training data, post post-training data, data that you use APIs to hook onto and and get to the answer that the user needs. The other one is the tracking data, related to the tracing data in observability, but when you think from the data foundation and and uh, data strategy perspective, this needs to be handled in this pillar because you need a whole data strategy now with tracing data, especially when you run hundreds of agents in your organization. Fourth is orchestration. One agent would work pretty well. You don't need to think about orchestration. But when you onboard five agents, the the complexity increases exponentially, right? You will have multiple coordination patterns between these agents, they will need to talk to each other in multiple different ways, they will need to each wait for each other's responses, there's a lot of complexity that comes in. And that's where orchestration patterns and thinking about how you will orchestrate your agents in a particular system becomes really important. Fifth is governance. This is where you think about what happens when something fails. Who's accountable? How do we govern data? How do we secure it? How do we secure our systems? How what do we make sure how do we make sure that no one injects into our agent and leads uh, to you know, misbehavior, right? Or loss of reputation. So, in the rest of the session, I will dive a bit deeper into each of these pillars and tell you how how you can think about when you start working with them, right?

第一支柱:深入评估体系

桑迪: 第一大支柱是评估。评估本质上就是你的 AI 系统的技术规范说明书。你要以此来定义成功。正如我所提到的,这不仅仅是泛泛地谈论“准确率”。你必须用数字化指标来明确定义它——什么样的准确率对你们的具体业务场景是足够好的?用数字把它定义出来。明确你们可以接受什么样的误报率(false positives),以及分流率(deflection rate)应该是多少。这里有一个零售银行聊天机器人的例子:当你使用 AI 智能体来部署聊天机器人时,核心目标之一是将简单的查询分流给智能体,从而让人工客服不需要去处理这些问题。因此,你需要去跟踪这些查询、跟踪这些数据,并将这个评估系统建立起来。其次是构建测试案例,即评估数据集。你们可能在评估中听说过“黄金数据集”(golden data sets)。你需要与业务领域专家沟通,了解在真实的前线场景中到底在发生什么。比如,一个优秀的人工客服针对某个具体问题会给客户怎样的回答。把这些信息收集起来。还有那些灰色地带、边界情况(edge cases),比如当人工发现客户提出一个容易令人困惑的问题时,他们会怎么处理。把这些情况都收集到一个数据集中,然后实现 AI 测试的自动化。也就是说,你向 AI 提问,它给出回答,你获取这个回答,并将其与测试集进行对比。将这整个流程自动化,这样当你的 AI 投入生产时,该流水线就可以实时捕获线上响应,并与你构建的测试数据集进行对比,从而在准确率和业务目标方面给出 AI 表现的量化评估。

当我们谈论评估时,在企业中通常会呈现出三个主要的维度,这是你在构建这些评估系统时需要做出的架构决策。第一个维度是确定性评估(deterministic):这是最简单的部分。比如格式检查,验证邮箱格式、电话格式,这些我们在传统软件系统中早已通过正则表达式实现的方式。另外,你也可以使用传统的机器学习模型来进行命名实体识别(NER)、意图分类,来提取名字、姓氏、敏感信息(PII)检测等。这些都是非常简单且成本极低的内容,你应该迅速将它们落地。这套技术我们已经用了许多年。第二个维度是非确定性的语义评估(semantic):这就是 groundedness 概念的用武之地。我们需要在这个维度实现像 LLM 作为裁判(LLM-as-a-judge) 这样的技术。大家都知道 LLM 裁判是什么吧?我看很多人在点头。这是一个非常典型的 LLM 裁判提示词。所谓 LLM 裁判,就是你使用一个独立的 LLM(不同于你的主 LLM),去评判主模型的输出质量。在这种模式下,你会为裁判模型定义具体的评判准则,比如安全性、事实一致性、回答相关性等。同样,这需要依赖你之前构建的黄金数据集,对比预期答案与实际输出。目前在 Databricks 中,在 MLflow 里你可以找到自动化的 LLM 裁判功能,能够自动针对追踪轨迹(traces)运行自定义的评估任务。这是第二个维度。第三个维度是行为评估(behavioral):在这个层级,你需要关注工具调用(tool calls)。我们的智能体是否调用了正确的外部工具?它们是否陷入了死循环?举个例子:用户提问“我的账户余额是多少?”,智能体完美回答了“您的账户余额是 xx 美元”。单看前两个维度的评估(确定性和语义),这个回答没有任何问题,质量也很高。但当你去检查它的“行为追踪”时,你可能会发现,这个智能体为了获取这一个答案,竟然向数据库发送了三次重复的调用请求!这可能是因为在某些环节中发生了调用失败或无谓的重试。在 Demo 演示环境中,三次 API 调用可能不算什么,但在实际生产中,当每天要处理成千上万个用户的并发查询时,这种重复的 API 调用会导致极其昂贵的带宽与计算成本。这就是行为评估必须要去解决的问题。这一层级极其重要,但我发现很多企业和开发团队在讨论评估时往往忽略了它。

Original English

Sandy: The first one is evaluation. Evaluation is basically specification for your AI system. You define success. As I mentioned, it's not like talking about accuracy. You have to define it with numbers, like what accuracy is is is good for your business use case, right? Define it in numbers. Uh, define what kind of you know, false positives you can handle. What should be the deflection? So, this is this is an example from a a retail chatbot, right? A banking chatbot where when you implement a chatbot with an AI agent, one of the main goals is to deflect simple queries um, to the agent so that a human agent don't need to uh, deal with them, right? And so, you need to uh, you need to track those queries and track those numbers and put that system in place. Second is building those test test cases, like the evaluation data set. You've heard about golden data sets in evaluation. Talk with the domain experts and find what is actually happening in real life on the ground. Like what answer would a support human support agent um, give to a customer on a particular question. Collect those information. What happens in gray areas, in edge cases, like what happens when a human sees a customer asking a confusing question, right? Collect those into a data set, and then automate your AI testing, right? So, you put a question to AI, it answers, take that answer, compare against the test set, and automate this whole pipeline so that when you put AI in production, that pipeline can actually take live responses and evaluate against the test data set that you're building, and then give you the result in terms of how AI is performing against those numbers and the goals that you've defined. When we talk about evaluation, there are three main layers that I see appear across organization, and this is an architectural decision that you need to make when you build these evaluation systems, right? The first layer is deterministic. These are the easy stuff, like you know, checking formats, you know, checking email formats, phone formats, the regular expression things that we have already been doing with our coding systems, right? The uh, the the the other other is like you know, you could use a classic ML models for name entity recognition to for intent classification, for understanding what is first name, last name, PII detection, etc. So, the these these are easy stuff, cheap stuff, you should get them out of the way. We have already been doing this for years. The second layer is the non-deterministic semantic stuff, all right? This is where groundedness comes in. This is where we implement technologies like LLMs judges. We all know what LLMs judges are, right? Everyone? Okay, I see a lot of nods. So, um Again, this is a pretty simple version of how a uh how a prompt would look for an LLM as a judge. Um So, with LLM as a judge, you you you use a separate LLM from the primary LLM to judge the response of the primary model. And when you do that, you tell the secondary, the judge model, on how it should judge the primary model's output. So, it could be around safety, groundedness, you know, relevance to the answer, etc. etc., right? Again, that can feed from a lot of uh the evaluation data set that you have created, right? To look at what are the expected answers, and then it can check against that. This is a sample prompt on how these things work, but I'm sure you've attended some of these sessions where you've seen vendors doing this automatically at scale. Uh for example, in Databricks we in MLflow you'll find automatic LLM as judge, uh where you can create these custom LLM as judges that run automatically on traces. That's your second layer. The third layer is behavioral, right? This is where uh you think about a tool calls, like is our agents calling the right tool? Are they getting into loops? So, for example, um you know, the first layer you you can have a user ask a question, "What is my account balance?" And you could go and check that, okay, this there is no deterministic problem with it. The seman- the agent answered right, that, "Okay, your account balance is this many dollars." And that was right, and you can see this is this is right, but when you go into the behavioral checks, you will see that the agent uh actually made three calls to the database to find that answer. Right? And that is because it was doing making duplicate calls for whatever reason. Calls failed, you know, calls did not work, it went and retried and stuff like that. Now, three API calls in demo environment is fine, but in production, when you get thousands of queries from users every day, and there's like duplication in API calls, that's an expensive operation. And that's where you need to think about behavioral evaluation. And this layer is very, very important. I see a lot of organizations, a lot of teams miss them when when talking about this.

第二与第三支柱:可观测性与数据基座

桑迪: 第二个支柱是可观测性。在这一部分,我们主要讨论的是追踪。你需要收集智能体做出的所有决策信息。我这里想用一个真实的场景来解释。这是我之前参与的一个零售银行聊天机器人项目。显然,在实际的生产环境中,追踪数据并不会像我 PPT 里的这张幻灯片一样整洁漂亮,我是为了演示才把它进行了极大的简化和美化。在这张幻灯片中,用户发起提问:“我被收了一笔透支费,能帮我免除吗?”因为客户认为这笔收费是不合理的。由于你已经开启了可观测性并捕获了追踪轨迹,你可以非常清晰地看到 AI 的每一步操作:意图分类完成了,耗时多久,置信度是多少;接着,它去连接客户的账户,调用客户数据库 API 获取账户信息;随后它去检索政策文件,在 RAG 向量数据库中查找关于透支费减免的政策,来验证客户的要求是否符合条件;接着,它进行了一步推理以决定如何回复客户,在经过最后的安全护栏检查后,将答案输出给用户。如果你没有建立这样的可视化追踪系统,当客户对服务结果产生争议时,你根本无从查起 AI 到底是怎么做出决策的。你只能两眼一黑,最后为了平息客户的不满,不得不妥协给客户一些补偿或折扣。这就是为什么可观测性是监管机构的硬性要求,也是生产系统必不可少的部分。同时,这也是你发现前面提到的“重复 API 调用”这种隐藏缺陷的唯一手段。当你启用了这一追踪能力,你不仅可以事后分析,甚至可以实现线上监控。当在生产中监测到某个调用反复失败或出现重复调用时,可以自动触发降级容错机制:比如限制重试次数不超过三次,若超出则将任务直接路由给人工客服处理。

接下来是第三个支柱,在我看来,这是最核心的支柱——数据基座(Data Foundation)。在我的典型项目中,我往往会把 60% 的精力花在这一部分,很多企业也在这里投入了极大精力。因为在早期的系统建设中,没有人能预料到智能体会突然爆发,并开始直接向数据发起查询。过去的数据系统完全是为人类设计的,而人类是非常具有容错和自我修正能力的。如果人类在报告中看到了错误的数据,他会自己去找人核实纠正。但智能体不会!智能体发现数据错误后,只会非常自信地给你输出一个极其荒谬的错误答案。这就是为什么对于今天的企业而言,数据质量和数据策略变得前所未有地重要。我将它分为“问题数据”和“追踪数据”两部分。问题数据用于生成回复,而追踪数据则需要我们规划好如何收集、存储、以及如何将其提供给审计和监管机构,用于线上的评估和监控。在 Databricks 中,我们利用 Delta Lake、Apache Spark、MLflow 以及 Unity Catalog 等开源和商业技术来帮助客户构建数据基座。Delta Lake 在你的原始云存储上提供了类似数据库的事务和管理属性;而 Unity Catalog 则是我们在整个平台的核心,它可以集中应用权限管理,并在目录级别实现数据发现和元数据标注。当你为数据表、字段添加了清晰的描述,并标注了哪些是敏感数据(PII)时,AI 系统就能够非常轻松地获取这些上下文,并安全地进行查询。无论你的 AI 运行在 CrewAI 还是 LangChain 上,无论是部署在 AWS、Azure 还是谷歌云上,你都能够通过这个中心化的一体化数据层(如 Mosaic AI)来支撑数据分析、文本转 SQL(Genie 助手)以及智能体追踪数据收集等多维度的业务需求。

Original English

Sandy: The second layer is observability, right? Uh and in this pillar, what we're talking about is tracing, right? So, you collect [snorts] all the decisions that an agent is making. So, I want to explain this with a scenario here, right? And this is a scenario from an actual project I worked on with a banking uh retail retail banking chatbot. Now, obviously, if you've seen tracing data, it's not as beautiful as this slide, right? So, I've simplified it and made it beautiful for this slide. But what this slide says is basically, a user comes in and says, uh "You know, I have been charged an overdraft fee, can you waive it for me?" Because the user thinks that the customer thinks that that that is not legitimate. So, the agent does an intent classification, and you all you know about this because you've enabled observability, you're capturing traces, and you're actually seeing what the agent is doing, right? What AI is doing. Intent classification, it is done, this it took this many seconds, this was confidence score. Then it goes and connects to the customer's account, maybe in a database, a customer database, call calls an API, connects to the customer database, gets the account details. It retrieves policy documents. It checks from a rag vector database, um what is uh what is the policy around overdraft, right? Is what the customer claiming is legitimate? So, it checks for policy documents. Then it goes and does a reasoning on what should be uh you know, responded to the customer, and then it does some final guardrail checks, and responds to the customer. Now, if you did not set up a system that helps you look visualize all of these traces, when the customer comes to you and raises a dispute, you have no way to check what the AI did. Right? You have nowhere to go, and you end up saying that I don't have have any idea. Let's Let's give the customer a discount or something, and then make them happy. So, this is why you need this, and this is why regulators are are are basically mandating, because otherwise there's no production system if you cannot do this kind of stuff. So, this is where um you know, you you you detect this the example that I gave around duplicate API calls. This is where you start detecting this stuff. So, when you when you enable these traces, you can actually go and see duplicate calls, and then take relevant actions based on that. Not only that, you can actually do that in online monitoring. So, when it's happening in production, in on you can set up online monitoring, and at that point, if it is doing duplicate calls, you can apply fallback strategies. Or even if it is doing a call that is failing, you can actually go and apply a strategy where it will say, "Okay, go and retry for three times, not more than three times. If it if it is more than three times, then report somewhere, or pass it to a human to take some action." The third pillar is the most important pillar, in my opinion, is the data data data foundation, right? Uh in my typical project projects, I spend 60% of my time, uh and I see I see a lot of organizations spending a lot of time here, because no one expected agents to come suddenly in the market and start querying data. Data was always built for humans, and humans are always forgiving. You find the wrong data in a report, you just go and ask someone to correct it. Agents don't forgive you, right? Agents will go, find it wrong, they'll give you the wrong answer confidently. Right? And you wouldn't know what's happening. And this is why data quality, setting the right data strategy, has become so important for enterprises now. I divide it into two sections. One is the question data, as I was explaining, like data needed for actually serving the AI's uh outcome. And the other one is the tracking data. This is the observability data, the tracing data I was talking about earlier. You need a proper plan on how you collect this tracing data, and how you serve it to auditors, to regulators, to do online monitoring, to run LLM as judges on the tracing, and everything else, right? So, there it needs a proper strategy on how you structure the schema and everything on the tracing data. Um On Databricks, um we create a robust data foundation for our customers using uh some of the technologies that we provide. If you don't know Databricks, Databricks has been built on some open-source technologies like Apache Spark, MLflow, and Delta Lake. Uh we provide a bunch of capabilities on top of it. So, the blue layer at the bottom is basically your cloud storage. Databricks works on the three major clouds, Google, AWS, Azure. Okay. I thought it was for me. So, so once you once you store raw data on your cloud storage, uh the data is then um um we we we bring in a a layer called the Delta Lake layer, which uh which basically brings database-like properties on top of your raw data. So, you have got images, text files, video files, or whatever. We we help you create this um you know, uh table-like structure on top of it using manifest files, right? And and we help you to uh incrementally load data, do all of those um data management tasks in a structured way. On top of that, we bring in Unity Catalog, which is a data catalog. Uh with Unity Catalog, you can centrally apply permissions on top of the data. You can um you can uh share the data using uh Delta Sharing, but also uh what happens with Unity Catalog is uh you you can enable discovery and um you know, um ownership, metadata tagging capabilities at the catalog level. What that means is, when you apply table description, column description, uh tag columns uh like PII columns with metadata, it becomes really easy for AI to then get that context when it queries these tables on top of Unity Catalog. So, everything is governed at one layer through Unity Catalog, and on on top of that we bring in different uh applications. So, whether it's AI through Mosaic AI, so to build LLM, tune LLM, or even build AI applications, we bring in uh data warehousing capabilities, BI capabilities, and um uh and some of the other text-to-SQL capabilities. We have got Genie that uh helps you write natural language to do SQL querying, etc. And one application of that in the observability and tracking tracking data, as I was showing, is is this. So, basically, think about when I was talking about the tracking data strategy. Organizations, especially enterprises, will not be running AI in just one framework. They'll be using different frameworks, CrewAI, LangChain, etc. etc. They'll be using different cloud platforms. And once they do that, you need a centralized layer of collecting that tracing data, so that you can serve sev- several use cases on the right hand side. So, whether it's for operational dashboarding, for first line support, uh a lot of these first line um first line of defense teams need health monitoring uh sort of dashboards, right? These teams can also write SQL using Databricks Genie to do text-to-SQL. But they can also build Databricks apps using coding agents uh to create common workspaces or custom uh UIs that customers might need for different uh different use cases. And then we've got Agent Bricks and MLflow that serves you uh LLM out of the box LLM as judges, and uh proactively monitor a The idea is, no matter where your AI runs, you can create this kind of strategy bringing in data in one common place and serving uh different teams from one shared location.

第四与第五支柱:编排模式与系统治理

桑迪: 第四个支柱是多智能体编排模式(Multi-Agent Orchestration Patterns)。正如我之前所说,单智能体表现良好,但当系统扩展到多智能体时,复杂度会急剧增加。这时候你就需要根据业务场景来选择合适的编排模式。我们常讨论的第一个模式是编排器-工作者模式(Orchestrator-Worker pattern)。在这种模式中,会有一个中心化的编排器充当控制中心,根据各个子智能体的专业技能将任务分发给它们。所有的请求都由编排器集中控制。如果系统出现故障,你只需去调阅编排器的日志,就可以非常明确地查出哪个环节出了问题。这种集中化的模式非常易于控制。与之相对的是编排编舞模式(Choreography pattern):在此模式下,每个智能体都是完全对等和自治的,没有中心化的编排器。它们通过一个统一的“消息总线(Message Bus)”进行异步通信,各自订阅并监听自己感兴趣的事件。举个例子,在贷款申请流程中,处理客户信息的智能体和负责审批的智能体在收到消息总线上的事件后可以并发执行,减少了向中心编排器反复请求和响应的时间。这种模式的显著优势是降低了系统的整体延迟。第三个是人机协同模式(Human-in-the-loop):当智能体做出的决策置信度低于设定的阈值时,自动挂起当前工作流,并引入人工客服进行干预。关于这些多智能体的状态管理、容错策略(如 Saga 模式、补偿机制、熔断器设计等),我已经录制过一个深度的专题视频并上传到了 YouTube,大家可以去深入学习。

第五个支柱是治理。请注意,这里我不是指传统的数据治理,那是一切的前提。从 AI 系统的角度出发,我们需要关注的治理主要包括以下几点:首先是合规审计追踪(Audit Trails),我们必须完整记录系统中的每一次请求、每一次用户连接和 AI 执行的每一个动作。其次是敏感信息屏蔽,在系统前置阶段必须使用 NER 等技术过滤隐私数据。在我们的实际案例中,仅在测试阶段,这个治理层就成功拦截了 47 次潜在的个人隐私数据(PII)泄露。接着是提示词版本管理(Prompt Versioning),在企业级应用中,提示词必须被“视作代码”,经历严格的版本迭代和变更审批,绝对不能简单地修改一下就直接提交到 Git 仓库,必须记录每次变更的原因和解决的具体缺陷。最后是模型变更管理(Model Change Management)。大模型厂商会高频地对模型进行升级,但在公共基准测试(Benchmarks)中表现好并不等同于在你的特定业务数据上表现优秀。因此,必须有一套固定的评估数据集(Evaluation Data Sets)来对新模型进行准入测试,以此来降低对单一模型厂商的依赖风险,确保系统具备随时切换底座模型的能力。Databricks 正是将这五大核心支柱融入到了我们的 Agent Bricks 产品中,为企业提供开箱即用的生产级 AI 应用支撑。

Original English

Sandy: The fourth pillar is multi-agent orchestration patterns. As I said, one agent is good, multiple agents increases complexity. That's where you start thinking about, okay, what pattern is good for my use case. The first one I describe here is the orchestrator worker pattern. Where you have one orchestrator which orchestrates all the work, which controls all the work from a centralized plane, and then distributes this work to different agents based on their specialized skills. And then every request goes through the orchestrator, so you have got central control. If something goes wrong, you can go to the orchestrator logs and look into them and see what has happened. Right? So, that's the orchestration data uh pattern. There is this choreography pattern where each agent is independent, they're autonomous, they don't depend on an orchestrator. All of them talk to a message bus and they listen to the events that they are interested in. Right? So, think about agents that are independent of each other, right? They can run parallelly. So, they are not sequential, like one agent is not dependent on another. So, they run parallelly, they listen to the message bus for the for the for the events that they are interested in. Maybe it's a trigger for, let's say, a mortgage application, and it says, uh you know, uh the mortgage application agent one of the agent looks customer details, right? The other agent looks at approval details and everything else, right? They can work in parallel, and the advantage it brings you is the latency is reduced because they are not dependent on an orchestrator and sending messages back and forth. Right? So, this is the choreography pattern. And the third one is human in the loop, which is where when an agent crosses a threshold or serves below threshold a confidence threshold, then a human is called in the workflow to look into the pattern uh so, look into the looking into what the agent has done and then take action based on that. I have done a deep dive video on multi-agent orchestration pattern uh for the online track of this conference. Uh it's already on YouTube, so you can look into it. I talk about the real implications of when you think about multi-agent patterns. One is state management, the other is fault tolerance, like what happens when things fail, like how do you manage them? I talk about different patterns. And then talk about how how you think about scaling them in large scale on enterprises. Pillar five is governance, right? Now, here I'm not talking about data governance at all. That's given, we need that, right? From AI perspective, what what are we thinking about? Regulatory, right? Audit trails, have we got the trail of every action, every user connection, every request, everything that happens in the system? Are we capturing everything? Are we doing pre-validation of personal information? Are we using name entity recognition? The the easy stuff, the rejects and all of those things, right? In our example, the work that I was doing with the customer that I mentioned, we already detected 47 PII breaches during the testing phase by applying this layer. So, that's that's really important. Fourth is um prompt versioning. You have to treat prompt versioning as change management in enterprise grade solution. It cannot be just change to a prompt and commit to get. It has to be is it has to go through proper change management processes as you do with code. So, basically treating prompt as code. Third is model change management. So, as models change, the model providers upgrade these models, you have to have a system to understand whether that upgraded model will be good for your use case, for your data. Right? Model providers you will put evaluation benchmarks on three uh benchmark uh boards, but those are not really useful when you put them in your context, in your enterprise. So, that's where these evaluation data sets come in handy, where you try these different models on this evaluation data set and try to understand which one performs better. And that management needs to be done because from a risk perspective, you cannot really rely on one single model. You have to have the flexibility to switch to different models and also test them on your own data. That management needs to be done. Uh in Databricks, uh we have taken all of these these points, these pillars that I've been talking about into Agent Bricks. We are building Agent Bricks to make uh all of these operations out of the box for you, uh so that it's easy to implement production grade AI applications on uh on in in your enterprises.

案例研究:零售银行聊天机器人

桑迪: 接下来,我希望能通过一个实际的案例研究,让大家对这些系统是如何落地的有一个更直观的感受。在 18 个月前,我与一家零售银行客户进行合作,当时他们正尝试构建一个零售银行客服聊天机器人。他们面临的业务痛点是:其线上聊天机器人每个月会收到大约 2 万次客户咨询,通过业务梳理,他们发现其中 60% 都是一些非常常规和简单的查询——比如查询账户余额、询问透支规则等。这些都是可以通过系统自动化直接解答的问题。因此,他们希望能利用 AI 智能体来自动化这部分业务,减少对人工客服的依赖。为了验证可行性,他们此前已经花费了 6 个月时间以及大约 8.5 万美元来进行概念验证(POC),但最终以失败告终。当我们受邀介入该项目并进行复盘时,我们发现了之前提到的一系列问题:没有人知道生产系统里 AI 为什么会频繁出错;没有任何量化指标来评估它的成功与失败;甚至在发生业务争议时,找不到相关的责任归属。针对这一情况,我们重新确立了项目目标:实现 AI 智能体能够完全接管 60% 的简单咨询,并建立起全套的可追踪和评估闭环。与普通项目最大的不同在于,在我们的 8 周交付计划中,我们直到第七周才开始进行具体的模型选择。在此之前,我们把所有的精力都放在了前期的基础建设上。

在第一和第二周,我们全力以赴构建了评估层。我们从人工客服的历史对话记录中抽样收集了 200 个典型案例,整理并精细化标注出一个标准的黄金数据集。我们定义了明确的业务考核指标:比如目标分流率需要达到 60%,准确率需要达到 85% 以上,同时设定了严格的响应延迟目标。在此基础之上,我们开发了一套自动化的评估流水线。这套系统在运行中会自动捕获用户的提问以及 AI 智能体给出的回答,将其与评估数据集进行对比评分。如果评估得分低于设定的阈值,该对话会被自动挂起并推送给人工客服进行复核。如果发现了错误,开发团队会针对性地修正提示词或调整工具调用逻辑,然后把这个全新的测试案例补充进评估数据集中。这意味着,我们的评估数据集不再是一个静态的文件,而是一个会随着业务运行不断演进和扩充的生命系统。它会从最初的 200 条记录成长为包含成千上万条记录的庞大测试库,测试库越大,你的 AI 系统的鲁棒性就越强。

在第三周到第六周,我们重点解决了数据基础架构的问题。我们构建了安全的数据库 API 连接,启用了分布式追踪来完整记录每一次 API 调用,这让我们在测试阶段就成功捕捉并修正了由于重试机制设计不合理导致的重复数据库调用问题,消除了潜在的性能隐患。

直到第七周,我们才将目光投向大模型本身。因为我们手中已经握有了一套高质、自动化的评估流水线,我们可以在极短的时间内将市面上主流的几种大模型跑一遍测试集,并根据输出的准确率和延迟数据直接得出结论。模型选择这个原本需要反复开会拉扯数周的决策,我们仅用了几天时间就通过数据支撑做出了最优抉择。

在项目最终部署上线的六周后,系统的各项指标表现非常优秀。然而,真正的考验发生在上线后的第四周:银行临时调整了一项利率政策,并通过手机 App 向所有客户推送了通知。随后大批客户涌入聊天机器人咨询这一新政策的细节,但由于后台向量数据库的嵌入向量(Embeddings)更新延迟,AI 智能体无法获取最新政策,只能拿着旧文档进行错误回答,引来了大量客户的差评。幸运的是,因为我们预先建立了可观测性追踪和负反馈监控机制,系统在负反馈指标触发阈值的第一时间就发出了告警。开发团队通过追踪日志仅用了几分钟就精确定位到是“由于 Embeddings 未同步导致引用了过期的政策文件”这一技术根因,并迅速修复了数据管道。这一切的快速响应和自愈能力,完全得益于我们在项目初期就坚定不移地构建了这一套“可见、可衡量、可问责”的系统支撑。

Original English

Sandy: So, I wanted to quickly touch upon a case study, just to give you a uh a flavor of how these things go, right? So, when I was working with this client um they were a retail banking they were building a retail banking chatbot you know, one and a half 18 months ago. Uh their their problem the the problem they wanted to solve is they had got around 20,000 odd calls per month from customers on their chatbot. They wanted to deflect they they they saw that there were like 60% of them were simple queries, what is my account balance, you know, what do I do with my overdraft and all of those stuff, like that can be answered simply. So, they wanted to the reliance on human agents for those answers. So, they identified those queries and they wanted to automate them. Right? They spent around 85K in 6 months doing a POC which did not succeed. When we got involved, we found those insights that I was talking like no one knew why things were failing when it was in production when when when when it was in production. No one could actually measure why why it's not succeeding and no one could actually understand who is accountable for what when things go wrong. Right? So, the goal we set for them is AI agent handles 60% of user queries, right? Which were simple user queries and then a way to identify and track them. The key difference in this project that we when we did is that we selected the model in week seven, like in a eight weeks POC. Right? And this is how it turned out. For the week one and two, we built the evaluation layer. We collected 200 cases on their actual human agents answering to their customers on simple queries and understand how they are responding to them. We created that database. Then we defined the success metrics. What does success look like to you? So, out of let's say 100 queries, you need 60 queries or the 60% of the queries that are simple queries to be uh to be handled by the agent, right? They needed some sort of accuracy. So, 85% it was around 85% accuracy target. They needed latency, all of the operational targets that you need. They were there. Then we created this automated evaluation pipeline for them. And what what I mean by that is an automated system where you can capture a user's uh a user's question and the AI agent's response. You take that, compare that against your evaluation data set. You rate that, and if the rating is below certain threshold, you get it checked by a human. And you if if if something goes wrong, you make sure that you find the solution. So, it could be a change to the prompt, it could be change to a tool calling system or something else. Once you have done that, you add that test case in the test data set. So, that when it happens next time, the test cases cases catch them. So, the the the the the summary of that story is that your evaluation data set is a living system. You start with 200, maybe there is no correct number here, but once you start, as you start building in production, this is a living system. This will keep growing. And the and the bigger it grows, the better your system will be. In the second week, we talked we thought about the foundational layer, right? So, the the question data, we thought we we thought that, okay, if you have to call the database, have you got the API connections right? Have you got a system in place that can trace the API connections? Are those secure, right? We were not talking about MCP at that time. Right? It was just direct API calls to database to run queries. Have you got the distributed storage? Have you got the Have you got the Are you collecting traces? And this is where when when we started testing after building these systems, we could catch those duplicate API calls, right? We could catch why customer satisfaction was dropping and stuff like that. And then comes in week seven to eight, we started talking about models. Now that we had the evaluation data set, we could run different models on that data set to see the responses, compare them against the expected responses, and calculate a number on on on the accuracy, right? That helped us to decide which model to use. Now, that decision didn't take long, right? We in as I explained in the introduction, like we spent weeks debating on which model to use, but when you took the other approach, you can actually uh do that in a very quick way. So, once once that's done, we we stitched everything that I was talking about around observability, evaluation, the layers of evaluation. Once we had that system that can make AI visible, measurable, and accountable, that's when we started launching it to production. And that's when uh so, this is this is the result uh six weeks post launch, uh we we calculated the operational metrics, of course, you know, the accuracy, the deflection rate, the response time, uh the customer CSAT. But what's important here is in few weeks time, when uh there was a problem with uh so, one of the one of the things that happened was that the bank changed some uh interest rate related policies. So, when that when they changed the policy, they they actually sent emails to customers or notifications in the application in the in the in the in the mobile banking app about the policy change. But when the customers came and queried on the chatbot for further questions, they couldn't get the right answers and they were like putting thumbs down on the answers. So they were getting this feedback, right? So feedback decreased. The problem with this kind of system, if you did not have this measurement system, is that you couldn't actually know what's happening. But because we had the measurement system in place, these the drop in C set was detected, right? Because we were getting negative feedback from customers. We could actually look into the tracing decisions and see that the agent was looking at a policy policy document that was outdated. So the new policy document was not updated in the vector database. The embeddings did not come through. Because it did not come through, it was giving it stale answers. And that's when we went and fixed that. But it it was all possible because we built that those systems that led us to to detect this.

第八支柱:事故响应机制与核心要点

桑迪: 在结束今天的分享之前,我照例为大家准备了一些可以直接带走落地的业务工具模版。在幻灯片的最后有一个二维码,扫描它就可以下载包括我的“生产事故响应指南(Production Incident Playbook)”在内的多个文档。这是大多数开发团队在推进 AI 项目时经常遗漏的内容。这个响应指南明确定义了当生产环境中的 AI 出现不可控行为或故障时,团队应该如何一步步应对。首先是通过评估看板及时发现异常,随后通过我们的追踪系统进行根因诊断。而在控制阶段(Containment),你可以利用提示词的版本管理,迅速将出错的提示词版本回滚到上一个稳定版本,并将该受影响的业务流程强制分流给人工客服处理。或者,如我在多智能体视频里详细介绍的,在系统架构中引入熔断器模式(Circuit Breaker)、补偿事务模式(Saga)等容错设计。在修复问题后,还要将这个导致系统出错的边缘案例作为全新的测试用例补充进我们的评估测试集中,使评估流水线不断进化。同时,该响应指南建议将这套机制与企业现有的 IT 服务管理系统(ITSM)进行深度集成,确保告警能够第一时间触达对应的运维和开发人员,从而保障下游业务系统的稳定性。

对于大家明天回到工作岗位后具体该怎么做,我给出以下三点建议:第一,如果脑海中已经有了一个明确的 AI 项目,请立即着手定义它的成功指标。注意,不要从纯技术指标(如 F1-Score 等)出发,而是要从实际的业务成效出发。第二,找业务专家梳理并整理出一些好回答的真实案例,建立起你们第一版的评估黄金数据集。第三,哪怕只用简单的 Python 脚本,也要把自动评估和测试的代码流水线写出来,实现测试反馈的闭环。

最后,我在实际交付中还总结出了三个极易被忽视的教训,供大家参考:一是评估数据集必须有明确的业务归属和分类管理,你需要清晰界定哪些测试用例对应安全分类、哪些对应登录流程、哪些对应业务咨询,以便在测试失败时快速对应到具体业务模块。二是必须建立提示词的治理规范,任何提示词的改动不仅要在 Git 中进行版本提交,更要在提交日志中详细写明变更原因、对应修复了评估集中的哪一个测试用例,将提示词的开发完全纳入标准的软件工程变更流程中。三是高阶行为评估的成本控制,当你的评估数据集增长到几百上千条记录时,如果每一次代码微调都全量运行非确定性的语义评估和工具调用链追踪,其调用大模型产生的成本是难以承受的。因此,建议在持续集成(CI)阶段只运行精简的子测试集进行准入测试,只有在最终合并到主分支(Main Branch)并进行发布时,才触发全量测试,以此来优化研发成本。扫描屏幕上的这个二维码,就能获取这些评估模版以及基于开源技术搭建追踪系统的技术指南。另一个二维码是我的 LinkedIn 个人主页,我每周都会在上面的免费通讯中分享自己在客户前线实战中总结的技术洞察。再次感谢大家的聆听,谢谢大家!

Original English

Sandy: Before you go, I I generally in these sessions I share different artifacts that you can take away. I have a QR code at the end for you to download and you will find multiple artifacts. One of the important artifacts that I want to talk about is the production incident playbook. This is something that a lot of us tend to miss when we work in AI projects. And this playbook is basically a definition of what needs to happen when things fail in production. First, you detect using your eval dashboard. Then you diagnose using your tracing as I explained. Then you contain. So basically, you you are versioning your prompts. Is there a Is there is a If there is a problem with the prompt, you you take that prompt out, right? And start the changes, deflect it to a human. Or in my multi-agent orchestration video, I've talked about multiple fault tolerance failure recovery patterns around saga pattern, compensation pattern, and circuit breaker pattern that you can look into the video. I've explained them in details on how you can handle them. And then you use the test case library to fix. So you look into LLM's judge reports, you look into your evaluation data set reports. Then you fix your problem. Once you fix your problem, you put those test cases in your data set, right? And create that eval suite that is a living system that will keep growing. And and and you you keep improving your AI system based on that, right? But this playbook needs to be in place. When it runs in production, you will need to integrate it with your ITSM system so that it alerts the right person at the right time. You know, a lot of these organizations would have existing ITSM systems, right? So which which is used for alerting and you know, making sure that the downstream systems don't get affected, etc. So once once you have this in place, you can go and stitch it together to other systems. So what can you do tomorrow, right? Start with If If you have a project in mind, start with defining success. Success not from the technical sense, from the business sense. What it means for the business, right? Come up with a few examples of what good answers look like. And and create a data data set of that. And then build that pipeline using simple Python code. See if you can automate that so that when you run AI and get some response, it can become it can go and compare You can go and compare the answer against that data set and then that can be delivered to to the customer. Now, these are three lessons that I have learned while doing these things with my you you know, my easily miss. The test case library, as I explained, is a growing system. It will grow over time. And because it grows over time, you need some sort of governance around it. You you need a owner, right? You need to You need to figure out which test cases relate to what kind of problem. So that whenever you go back to it, you can you can relate your answers to those sort of problems, right? If it is a security If it is login, so you can say that the agent did not ask for login credentials when the customer asked the answer. And all those kind of issues can be put under a security category within that data set. So to categorize the rows in your data set so that you can pick up what changed and compare it with them. The second is prompt versioning. Now, when you start versioning prompts using Git, you know, we all know when you put Git message commit messages tend to be simple commit messages. But you have to put governance around what kind of commit messages you are putting in when you're changing these prompts because you need to understand when a prompt was changed, for exact what reason it was changed, right? What was the failure that caused this prompt to be changed? What kind of failure would it address and what would it correct, right? In the next version. That needs to be documented. Otherwise, it becomes difficult because when you go back and look into prompt versioning and look at different versions and you cannot trace why why those changes were made, then it becomes difficult to track what's happening. The third, the layer three evals, right? So the behavioral evals that I was talking about around tool calls and stuff like that, they can be really expensive as you grow your eval data set as well. So when you have a wrong tool call, for example, and you want to correct that system, when you correct it and run it against the eval data set, you have to basically run it against let's say if you have got 300, 400, 500 rows in the data set, you have to run it against them. And you do all the testing again and again and again and again and again, that can cost you a lot of money. So you have to put some governance around that. So for example, when in your continuous integration pipeline, when you do the prompt change, you can actually put some checks around just just selecting a small subset of the eval data set to do the testing. And you only do the full test when you merge to the main branch. So you can put these kind of decisions in place so that you can reduce cost around around you know, expensive eval decision. If you scan this QR code, it'll take you to a Google Drive link where I have put some examples on some of these how these templates look like, what evaluation checklist should look like. I've given you some guide on set setting up tracing with open source technologies so that you can quickly set up some tracing and start testing in the test environment before you decide on what kind of tools you want to use. Thank you very much for listening to me. This QR code will take you to my LinkedIn profile. So I share I have a newsletter where I share this kind of topics every week. So if you're interested, you can join. It's free. I basically share what I learn in in the field working with customers, right? So it might be useful for you. Thank you very much. >> [applause]

📌 文中提及的人物和组织

关键字: production-ai observability llm-as-a-judge multi-agent-orchestration ai-governance