企业智能体项目的破局之道:从人类速度到机器速度
在大型企业(如电信、公用事业、医疗)的运作中,传统的规模化经营依赖于控制、流程、可重复性及多层治理。虽然这种架构在过去创造了辉煌的成功,但其核心痛点在于所有流程均以**“人类速度”**运行。
目前,企业正面临一场深刻的范式转换,进入了一个以**“机器速度”**运行的世界。在我们的研究中,仅有 12% 的公司达到了所谓“AI成就者”(AI Achiever)的标准;绝大多数企业仍陷于昂贵的原型开发(Piloting)阶段,未能获得实质的投资回报。更严重的是,这些企业正落后于那些能够以机器速度计算和交付的敏捷竞争者。
Original English Source
So, we work in the world of enormous enterprises. So, telecoms, utilities, serving entire nations, government, healthcare that you just heard a little bit about, consumer products in your home right now. And when you operate at that scale, actions have consequences. So, bad deployment for example can take down critical national infrastructure. And so, over time these organizations have built structures for this reality. Control, process, repeatability, governance, layers and layers of it. And this has worked really well, right? Like for years these companies have seen massive successes and growth. But always at human pace. Human pace. That is what shifted. We're entering a world at a machine speed and transforming everything we know. Our work, our clients, and ultimately the societies these enterprises underpinning. Our research show that 12% of companies reach what we called AI achiever. It means that most of the company are still stuck piloting and spending millions and perhaps not getting too much in their return. 80%. The tragedy is not just wasted spend. It's about falling behind in a world that accelerating beyond what they can compute.
企业转型的核心桎梏:被拖慢的工程与治理
即使CEO们已意识到AI的重要性,企业运作速度依然未能提升。这并非因为AI编写代码的能力不足,而是源于企业自身的底层脚手架(Enterprise Scaffolding)——这一为人类速度设计的人类操作系统,在处理数据访问、安全审核及部署流程时,其平衡点过度倾向于企业合规与繁琐的利益相关者会议,而非工程投入。
交付一个Agentic(智能体)解决方案往往极其缓慢,即便应用本身构建仅需两周,投入生产却可能历经十二个月,原因在于基础设施、安全、AI网关、数据治理及应用团队需要多方对齐。这导致了严重的工程技术债务,即长期缺乏对CI/CD等工程自动化的投入。要打破这一困局,唯一的路径是将每一个人类流程转化为可执行的、适应性强的代码,而非增加会议或审批链。
Original English Source
Many of you here may ship here on Fridays and then rollbacks on Saturdays. A decision you take an afternoon could easily take an enterprise 6 months or more. Octopus, Klarna, Shein, they think that's insane. And then they go on and redefine the games themselves. Others studied the games, crafted the playbooks, and then ran the workshop. But they went home. We stayed. We shipped through the reality. And that is our mode. When you stay, you learn things that the slide decks don't warn you about. So, for example, it's not just about data availability or API availability that impacts AI success. It's the entire enterprise scaffold itself. The very thing that has made these companies so successful, which is increasingly becoming the drag, the thing that is holding them back from capturing AI value at scale.
Um 18 months ago, you were probably still needed to explain why AI mattered, why speed mattered. But that battle's gone. The C-levels are convinced. You know, the CEOs are terrified of being left behind now. But yet, the enterprise speed has not really shifted. It's not because AI cannot write good code. It is not because our engineers can't solve the context problem. I think it's something a lot more deeper. It is the actual enterprise scaffolding itself. A human operating system that designed for human and running at a human speed. The automation behind every delivery and that Jess mentioned about. Think about data access, security reviews, you know, deployment process. Most of the enterprises never needed to invest like a tech company. Corporate process balanced with minimal engineering investment with the complementary of stakeholder meetings. And that is how enterprise runs today. Fit for enterprise, fit for human.
We had the pleasure of delivering agentic solution in a large corp and integrating their centralized AI gateway. You know, we've been given the essentially testing configuration templated to us. Every single configuration change required a manual review before you can actually hit the test to run. And we eventually have to automate it at and the whole application we built took about 2 weeks. It then took another 12 months to get that into production. And because their infrastructure team, their security team, their AI gateway team, their data governance team, their application teams, they all needed to align. Um the best way I can describe this is think about Google search. Before you see the results come out, they're going to be three teams review the results first. Need to be a legal sign-off for the results. And then they say, wait for 2 weeks because we're quarter end, it's change freeze. That is how AI in enterprise delivery today. So, how do you go faster? I guess you hand an AI coding agent to your developers and the next thing that you find is a massive bottleneck at the code review and the deployment stage, right? And this is going to get worse because these coding agents are turning everyone into a builder, PMs, designers, domain experts. And so, the the supply of deployable code is exploding. Some of you might have seen the GitHub stats. In 2025, they reported 1 billion commits. So far this year we're averaging 275 million per week, which means we're on track for 14 billion by the end of the year. And this is super exciting, but approval infrastructure, deployment infrastructure hasn't changed because these processes were ultimately designed for human speed. The real tech debt here goes beyond the legacy code that exists within applications. And it's the years of underinvestment in the engineering automation, CI/CD, et cetera that allows companies to move faster while maintaining control. So, what's the pathway? Every single human process needs to become adaptable, executable code. Not another meeting, not a sign-off chain, code. And the good news here is that AI can help you build this faster and cheaper, right? Like these aren't new capabilities, but it does represent a fundamental mindset shift for a lot of the organizations that we work with.
重新定义投资与交付:像风投(VC)一样思考
传统的企业财务模型倾向于“确定性”,要求预先明确ROI、范围及交付成本。然而,在AI时代,执行成本趋近于零,真正的价值在于解锁此前经济上不可行的全新功能。
因此,财务部门必须像风险投资人(VC)一样思考。VC不追求单一项目的固定回报,而是通过组合投资(Portfolio)来寻求跨越确定性、依靠幂律(Power Law)增长的价值。项目成功的关键不是预先证明其ROI,而是询问:“不做这个的成本和代价是什么?”以及“我们在组合投资中是否投入了足够的赌注,以捕捉那些能改变企业游戏规则的机遇?”
在交付层面,数据科学团队应打破传统的瀑布式 Milestone 规划,转向假设驱动的交付(Hypothesis-driven delivery)。Agentic交付的核心在于通过小型循环(构建-评估-迭代)构建统计学置信度(Statistical Confidence),并招聘、培训那些能够翻译这些数据以获取利益相关者信任的复合型人才。
Original English Source
Okay, next up. Who here has had to start a project with a business case to unlock internal funding? Cool. Welcome to our world. Um now, business cases aren't wrong per se. They, you know, they raise a lot of the right questions, they create oversight, they ensure that someone has thought about ROI. All of these are really good things. But they assume that three things are knowable up front. The scope and the solution, the expected value, and the cost and time to deliver. Now, with AI, this is often backwards, right? You learn the solution and the business case by doing the work. And more importantly, when the execution cost for prototyping, experimentation, building, et cetera drops down to near zero, this is no longer just about efficiency. This is now unlocking capabilities, entire new categories of things that weren't possible previously. It means you can now attempt things that were previously economically impossible for the organization. And that means things like new products, new services, customer experiences, they are just waiting to be reinvented. Um and we see this born out in the stats. So, in terms of the AI achievers that we mentioned earlier, we see them achieving about 50% higher revenue growth than their peers. And that's not from cost cutting, that's from doing entirely new things. If you think about as well recent AI product successes, often they've been emergent. So, you think about Cursor's user base of live coders, they didn't exist when they started building the product or when they released it. Cloud code wasn't something that was planned out on a product roadmap months and months in advance. And on the enterprise side, you have examples like Walmart for example, who built out a social media trend scanner and generative designer that's now allowing them to compete in entirely new ways with the likes of Shein and Temu. Um or JP Morgan who started building something as an internal productivity tool that they've now been able to productize and generate entirely new revenue stream. Now, currently the enterprise finance is wired for certainty, which means that generally your project starts life justifying itself in terms of committed benefits and predictable cost phasing. And that framing can kill projects before you even begin because it's asking a question based on can we justify this specific thing based on predictability, rather than asking about what now becomes possible. And so, the right question to be asking is, what is the impact of not doing this? What is the cost of not doing this? So, your CFO needs to think like a VC. At least when it comes to agentic transformation. Um a VC doesn't bet on one project and demand like 3 years fixed guaranteed payback because they know the certainty from a business case is a fantasy. Rather, they back a portfolio, they're knowing most of the bets may not pay off, but they will basically looking for those ones that compound. They knows that where the true value really lies is in that beyond certainty and that power law exponential growth. Enterprise investment works the same way for AI. The question is not can we justified this project, but are we placing enough bets across the portfolio that we going to hit the right ones that's going to change everything for us. So if your finance function cannot think like that, that is where your transformation should start and because everything else is downstream.
Cool. Next one. Um do we have any data scientist or machine learning engineers here? Brilliant. I think hopefully you guys going to like this part. You guys been doing something different everyone else, I hope. Um hypothesize and and experiment, statistical confidence. Um most of enterprises probably treated you guys like the modern IT crowd. Brilliant, quietly right and just very kindly ignored. Um kept you guys in the basement while the upstairs doing the real work through the Jira board at PI planning. However, I think this is your opportunity. Agentic delivery is your world, not theirs. Models are non-deterministic. Agent behavior is emergent. And you should not scope it like just a feature build like software traditionally. And you cannot milestone it like a fixed program. And yet that's what entire enterprises are trying to do. When you are in that delivery trench, what we see is that the more enormous effort we have to spend to really not about building the things, but to really try to bridge that gap between and how any system actually work versus what our stakeholders actually expect. You know, the um the never-ending utopian design up front and the constant conversation you have to talk about guaranteed performance and those kind of endless status updates to for those decisions that never gets made. And those are the things that is currently consumed the energy when we come to delivery this the agentic systems. The IT crowd is never the problem, right? The organization just need to learn your language. And it's the only language we think is matter for agentic delivery. Um the team needs to upskill themselves to learn hypothesis driven delivery. You need to reshape your program around one goal, which is building that statistical confidence. Small loops of build, evaluate, iterate, fast evidence. And then and and actually your delivery team needs to look quite different, too. People who are comfortable with ambiguity who can articulate what they have learned, not just what they have delivered. And most importantly, they can translate those statistical number into stakeholder confidence. And those are different kind of skill set. You need hire for it, you need train for it, and you probably need to value that as well.
构建信任与防御壁垒:反馈回路(Feedback Loop)即护城河
在Agentic交付中,完成具体功能并非最核心的目标,**在输出中建立信任(Trust)**才是。企业需要通过“渐进式自主”(Progressive Autonomy)来管理信任账户:从不能影响结果的影子模式(Shadow Mode),到仅提供建议的咨询模式,再到在低风险场景下触发行动的受控自主模式。这种模式要求每一步都由结果的证据支撑,而非项目计划的完成度。
最后,企业必须明确自身的“护城河”(Moat)。在一个AI递归编写AI的世界里,所有功能均可被轻易克隆。真正的护城河不在于昨天的知识储备(Transactional Memory),而在于**“活的记忆”(Living Memory)**——即通过客户在边缘场景、情绪意图及特定规模下的实际行为所产生的反馈回路(Feedback Loop)。交付不是终点,而是竞速的起点。每一项功能都应旨在生成反馈信号或基于这些信号交付价值。如果做不到这一点,企业只是在制造易被复制的商品。
核心总结: 像风投一样下注,为机器速度升级架构,从第一天开始就围绕反馈回路为信任而设计。
Original English Source
All right, next one. Now, as a society, we are collectively learning to trust AI. I mean, no one cross-checks their Google results anymore, right? And we in this room are probably quite comfortable with using AI tools. Um probably all of us don't review every single code output that we generate. Um now, AI is on that same trajectory, right? But it's not there yet, and it's our job as AI engineers to bridge that gap. And for large enterprises, that trust gap is not small. Um and what we have learned is that the completion of individual features is not necessarily the most valuable thing that you ship. Um the trust in the outputs and the AI that you build over time is the more valuable thing. And and when I say trust, I mean in the broader sense, you know, in the content, in the accuracy, responsible use, privacy, all of those things that collectively allow end users to trust an AI system. You can think about agentic delivery in some ways as a deposit or a withdrawal into a trust account with your stakeholders, your leadership, with your end customers. And what survives over time isn't necessarily a specific feature. It's that trust that you've built as things evolve and things change. And so the question we ask is, how do you build trust at speed, deliberately, and with evidence? So we spend a lot of time talking to companies about progressive autonomy. Many companies still see agents as basically the same as traditional automation where you complete some tasks, you run them, you deploy them, and they run, and that's that. But agents aren't just built and then turned on, right? Like you can't foresee every single response or behavior up front and test for it. Um their behavior is emergent. And this is particularly relevant in the context of autonomous processes, which is something we do with quite a lot of our clients. Um and so the eval suite, we've heard a lot about it over the last few days, that's super important, um and you have to evaluate the right things, but we also talk about how you actually deploy into production and increase autonomy over time. So you follow this exposure ladder. So you start with a shadow mode where an agent might run alongside human processes, but it can't actually affect outcomes. Um you compare the human decisions that are made to what the agent is saying, and you use that as a signal to iterate um and build up confidence in the specific behaviors that you're you're trying to achieve. Um you keep iterating, you then move up to more of an advisory mode where the agent runs live, um but it only recommends. So the humans are still playing an active role in the workflow. They can approve or reject the outcome. And again, this provides you with another signal that you can incorporate and you iterate again. Then you shift up to controlled autonomy where the agent is able to run and trigger actions, but in a narrow, low-risk uh scenarios. Um and it has clear limits, kill switches. And over time, you can extend that up to um to to wider autonomy based on achieving the right level of confidence in the target behaviors that you're trying to drive. And the key here is that each step is gated by evidence in outcomes. It's not based on completion of activities in a project plan or pass-fail testing. It's entirely about the confidence, the trust in those outcomes. So, engineer for trust, not just for completion.
Right, last one. Um we mentioned our moat earlier. Now, what is yours? In a recursive world where AI codes AI, anything you ship can be cloned a minute it goes viral. So ask this, what is unique only to you? I think your existing enterprise knowledge, the CRM, the ERP, the SOPs, they got you to the table. We call it the your transaction your transactional memory. However, every competitor has one version of that. It's a floor, not a fortress. The real moat is in the moment when your customer touches your product. Edge cases, corrections, emotional intent, and actual behavior at your specific scale in your specific context. Those signals belong to you. And we call it your living memory. The day the day you ship is not the finish line. Far from it. It is when the race actually begins. How quickly you can compound and iterate. How fast you can turn a signal into value. It is a race against yourself to engineering your own competitive edge recursively and constantly. The pathway requires a fundamental shift in your engineering vision. Every feature you ship should either generate that feedback signal or deliver on what the signal has already taught you. Um because if it does neither, you're building something anyone can copy. Feedback is not the option. Feedback is the only moat.
So, five enterprise tensions, speed, value, delivery, trust, moat. How do you then apply what we've learned here to make sure that your next agentic project is a success? So first, we say start now. Deliver differently, measure in terms of confidence, uh shape the project around hypotheses rather than requirements or specific features, and run delivery in small loops of experimentation, iteration, and evaluation. Second, make finance a transformation partner, not just a gatekeeper. Um create a portfolio of different AI bets across the organization rather than justifying each project in isolation. And see value beyond the certainty of cost out. And third, um make the governance speed your CTO's top engineering problem. The ultimate technical debt you want to rectify. And the not last one, um that's for CEOs. Your moat is not in what you hold from yesterday. It is what it is in what you are learning and compounding every day. The technology will accelerating. Those who will survive won't be the ones that has to be the earliest adopter, but they will be the ones that learn to learn. They'll be ones that living and building the the memories through the feedback loops. They will be the ones to cultivate the trust for their people and their customer. And you cannot buy that. And you can you cannot copy that. You can only start building that by now. And by never treating the journey as finished. So, prescription is simple. Bet like a VC, upgrade for machine speed, and engineer for trust with the feedback loop from day one. Thank you. Thank you.