新生代与多模态:AI工程群体与应用演进
从调研群体的背景来看,AI 工程(AI Engineering)已经不再单指一个职位名称,而更像是一门涵盖创始人、CTO、研发人员及产品负责人的综合性学科。在经验分布上,行业呈现出一种有趣的“老兵学新艺”特征:在拥有10年以上软件开发经验的资深工程师中,超过半数接触 AI 开发的时间不超过3年;而相比之下,刚刚入行的新生代工程师,其 AI 开发经验的经验中位数几乎与这些十年老兵持平。这意味着,新一代研发人员从入行起就在与 AI 范式共生。
在具体的多模态(Modalities)构建中,尽管文本仍然占据主导地位,但其他模态的采用倾向表现出不同的曲线。我们通过采用意向比(Intent to Adopt Ratio:尚未采用某项技术但计划采用的人群比例)来衡量潜在增长,发现音频(Audio)拥有最强劲的采用潜力。在目前尚未在工作中使用音频的工程师中,高达 56% 的人计划在未来的 AI 应用中引入音频。与此同时,图像生成(Image Generation)的实际应用在过去一年中实现了翻倍——从去年的 18% 跃升至今年的 36%。随着像 Nano Banana 2 或 ChatGPT 图像生成等模型的演进,曾经只被用来生成“扭曲多指怪手”(cursed hands)的试玩工具,如今已经真正成为实际业务工作的一部分。
Original English
Now joining us on stage is the partner at Amplify Barren. Fantastic. You did a great job practicing. I feel very very loved. Um, let's get started. So, like you just heard, my name is Bar. I run a survey every year on the state of AI engineering. And the funny thing about running a survey on the state of AI engineering is that the field changes as you make the slides. Just in the past week, we've had Frontier releases treated like national security events. Meta reportedly exploring selling AI compute. By the time I get off stage, maybe something else will happen. So, if I miss a major announcement while I'm up here, please come find me after. But that's exactly why we run the survey every year to cut through the noise, take a moment, step back and understand what AI engineers are actually doing. Uh for the first time this year, we were thrilled to partner with Notion and Verscell to run this survey. Very quickly on me, uh this is the least interesting slide. I'm an investment partner at Amplify. Very lucky to invest in companies built by and for AI engineers. And I'll make the same promise that I make every single year, which is short time on bar, long time on bar charts. So, let's get right into it with lots of bar charts. First, let's talk about well, maybe raise your hand. Did you fill out the survey? This is a very large group. Okay. Yes, I see you in the front. Um, if the answer is you, thank you so much. If the answer is not you, I will find you in 2027. But genuinely, this only exists because a thousand of you gave your time. So, thank you. We had 1,048 respondents this year, which is a lot of AI engineers. And to be precise, this is not just AI engineers, as I'm sure you see at the conference. Every year, we see that AI engineering is more of a discipline than a job title. It touches founders, CTO's, engineers, product people, folks across company sizes and experience levels. And that range shows up in experience too. Um for the third year running we see the same pattern which is skew towards senior engineers but newer to AI. Of those with over 10 years of software experience over half have three years or less of AI experience which tracks uh these are very experienced engineers learning a new paradigm in real time. And the newest cohort, the ones who just started uh engineering, the median new engineer has nearly as much AI experience as the median 10-year software veteran. Uh so the newest engineers have never known software without this. But doing AI doesn't mean one thing. We talked about all these different titles, all these different roles. Before we get into models and agents, I have a more basic question, which is when people say they're doing AI at work, what are they actually doing? So, first up, like to start with the modalities. We asked, which modalities are actively building with at work? Can anyone take a guess? Text dominates. I know. Hold your applause. Um, but one piece of this chart that I always find very interesting and I always look at is the ratio of nope, I'm not using this modality to I'm not using it, but I do plan to. I call this the intent to adopt ratio. Of the people who are not building with a modality today, how many say they plan to use it? And audio has the strongest intent to adopt this year. Among AI engineers who are not building with audio today, a whopping 56% say they plan to adopt it in the AI applications they build. And this is not a brand new signal. Last year audio also had the highest intent to adopt across modalities but 37%. So audio continues to take the lead and have high interest but that interest is accelerating. Now there has been an audio swing but if we look at what changed most from the last year in the survey the biggest jump is actually in people using image generation. The share of respondents using generative AI for images and feeling really good about it doubled from 18% last year to 36% this year. Makes sense if you look at what we launched in the same window. Over the past year plus survey time, uh we've had models Nano Banana, Nano Banana 2, Chat GPT images 2.0. The products have gotten much better. What used to feel like an efficient way to generate cursed hands is just increasingly becoming a part of real work. Audio may have the strongest intent to adopt, but image generation shows us what happens when a modality crosses that threshold. So, I'm excited to continue watching these adoption curves every single year. I think we're going to see a lot this year.
混合模型与成本控制:作为一级工程约束
在模型选择上,虽然社交媒体上关于开源权重模型(Open-weight Model: 允许下载并在本地运行的开放模型)与闭源模型(Closed Model: 仅能通过 API 调用的商业模型)的讨论铺天荒野,但在实际生产中,它并不是决定团队选择的首要因素。在影响模型选择的考量中,开源与闭源的区分仅被 5% 的受访者列入前三。工程师更在乎的是模型质量(Quality),其次是诸如工具调用等智能体能力(Agentic Capabilities)以及成本(Cost)。出人意料的是,可靠性(Reliability)并未名列前茅(仅五分之一的人提及),这或许表明目前的领先模型在可靠性上已达到了工程要求的“基准线”,决策点因而向应用栈的更高层移动。
目前,生产环境已呈现明确的多模型混合态势:94% 的团队使用闭源模型,45% 的团队使用开源模型,而在这 45% 使用开源模型的团队中,超过 90% 也在同时使用闭源模型作为补充。87% 的团队在实际业务中采用不止一种模型,根据任务类型进行动态路由或对比输出。与此相对的是,超过半数的组织开始在平台与基础设施层进行工具链的收敛与标准化,以对抗过于分散的灵活性。
然而,“无限智能”依然伴随着清晰的账单成本。在 2026 年,成本已成为一级工程约束条件(First-class Engineering Constraint)。40% 的受访者表示成本会“经常”限制他们使用 AI 的野心和设计尺度,另有 36% 认为“偶尔”会受到影响。几乎有四分之三的团队在根据钱包厚度调整 AI 使用策略(或许剩下的四分之一用的是公司信用卡)。在生产监控中,成本与 Token 使用量成为了仅次于质量的第二大监控指标,像监控服务等级协议(SLA)一样被严格对待。
Original English
Uh, now models. If you've Who here spends time on Twitter? All right. Yes. I imagine this is a very Twitter pilled uh crowd. If you spend any time on Twitter in this uh in this circle, you've seen a lot written about openweight models these past few months and I think we'll see it even more in the next year. Um so we asked what models are you actually using in production. 94% use closed models. 45% are using openweight models. But here's the thing, you know, openweight models are not replacing closed models for the most part. at least not yet. The respondents using openweight models, over 90% of them are also using closed models. So they're looking like an augmentation. Teams are mixing and matching. We also asked just to double click on this for the top three considerations when choosing a model. If you're choosing a model, what is important to you? Um and despite the airtime of the open versus closed, it's not what drives model choice. It was a top three consideration for only 5% of the respondents. What matters is actually more straightforward. It's quality. Quality dominates. Followed by agentic capabilities like tool calling and cost tied right with it. We'll money money. We'll get back to that. Um, one thing that I found very interesting is that reliability is not near the top. Only one in five named reliability. That doesn't mean teams stopped caring about reliability. Uh there are different ways to interpret this data. My guess is that it's more likely to become a threshold requirement and the models they're choosing are reliable enough so the decision moves up the stack outside of certain circumstances to quality, capability, cost, but we could talk after. All right, so here's where the model story all comes together. Like I said, teams are not choosing one model and calling it a day. Earlier I showed that 87% of teams are using more than one model. Uh the model that's the opposite of standardization. Uh and the way that they choose models for given tasks varies. Most popular is routing by task type. Some run multiple models compare outputs. Some route based on cost. Uh but models are good at different things. What was interesting was that more than half of respondents said that their organizations starting to standardize on fewer AI tools. They're trading flexibility for standardization. A share of those are mixed. They say they're standardizing on some layers while staying flexible on others. But the headline here is that there's we're in the early great standardization of the platform and tools, not the models. All right, this is the slide where anyone who's opened an AI bill in the last year starts nodding. So it turns out that infinite intelligence still comes with a usagebased bill. Once teams are managing many models and AI workflows, the next question becomes cost. Cost is now a first class engineering constraint. We see this in the data. 40% of respondents say that cost regularly shapes how ambitiously they use AI and another 36% say that it sometimes does. Well, this is pretty straightforward. So all in about uh three out of four respondents are adjusting their AI usage based on cost and maybe the fourth has a company card. That might be surprising or maybe it's obvious but 12 months ago it was not. Token maxing is cool. Being able to find real use cases is amazing but cost is becoming a real big part of the product decision today. And it shows up in monitoring too. What folks are monitoring in production includes cost and token usage as the number two thing they watch for. It's being monitored like an SLA right under quality itself.
写权限大门打开:智能体控制层的缺失
2026年是智能体(Agent)逃离“Demo世界”真正落地的一年。今年有高达 95% 的受访团队声称正在使用或开发智能体,这一比例几乎是去年的两倍。然而,比采用率翻倍更重要的质变在于——这些智能体被授予了真正的写权限(Write Access)。
去年,只有 52% 的智能体开发者允许智能体写入数据。而到了今年,这一比例飙升至 89%。这意味着智能体不再只是安静地阅读、总结或撰写草稿,它们开始代表用户在系统中执行真实的交易、修改数据库或调用破坏性 API。如果将“采用智能体的人数增加”与“写权限授予比例提高”这两个变量结合起来,今年在生产环境中使用具备写权限智能体的整体比例较去年暴增了三倍以上。
如此高风险的权限下放,其配套的控制工具却显得非常原始。目前团队最依赖的两种智能体控制方式是人机协同审批(Human-in-the-loop: 关键步骤需人工确认)以及硬性的权限网关(Gating Permissions)。这套策略本质上与企业管理一个刚入职的实习生(intern)别无二致。至于任务拆解、向量检索、记忆持久化以及沙箱隔离等更具技术含量的方案,在数据中呈现出零星散落的态势,行业尚未形成大一统的“智能体控制层”。这也导致了智能体面临高昂的失败挫折感:约三分之二的工程师表示,智能体在执行任务中发生幻觉(Hallucination)或丢失上下文是令他们最头疼的痛点。
Original English
Which brings us to the biggest line item of them all. Agents. We've been talking about agents for a while. Uh this year, as you've seen, as you'll see today, as you've seen in previous days, you're going to talk a lot about harness engineering. They're escaping demo world. So, we asked respondents what level of tool permissions their agents typically have. And this is where agents start to look more real. There are two things happening at once. First, and I don't think this is surprising, relative to last year, there are far more teams using agents. This year, 95%, this seems high to me, 95% say they're using agents, roughly double last year. Second, amongst the teams that are using agents, those agents are much more likely to have write access. Last year, 52% of folks building with agents said their agents could actually write data. This year, that number is 89%. So when you combine these two shifts, more teams using agents and more of those agents having write permissions, the share of all the respondents and again it's a survey using write enabled agents is up more than three times relative to last year. So this is really the big shift. Agents are no longer reading, summarizing, drafting. They're taking actions inside of systems. And that raises the obvious question, how are we controlling all of this? um with pretty blunt instruments. Uh very there are many ways that folks are controlling agents today. The top two are human in the loop approvals and gating permissions which are the right instincts but kind of the same toolkit you'd use to manage an intern. Below that the results scatter. Task decomposition, retrieval, memory, sandboxing. People are trying everything. Nobody has settled the control layer for agents. uh memory and persistent context is one that I'm watching very carefully right now. I think it's going to evolve a lot in the next year. And when agents fail or when people complain about agents failing to be more precise, it's usually the thinking, not the plumbing. So, uh you know, like twothirds say that hallucination or losing context mid task is what frustrates them the most.
评估困境与自研边界:构建技术栈的权衡
在 AI 工程技术栈的八个层级中,关于“造”还是“买”(Build vs Buy)的博弈展现出清晰的业务边界:
- 模型推理与托管服务(Inference & serving):这是工程师最倾向于直接购买的层级。极少有团队愿意自己去维护底层推理基础设施。
- 提示词管理(Prompt Management):则处于完全相反的极端。61% 的团队选择自己编写和构建相关工具,似乎每个工程师都觉得自己的 Prompt 非常独特,不愿假手于人。
- 类似提示词管理,包括产品业务逻辑、RAG 以及 评估系统,在相对比例上也更倾向于保留在企业内部自研。而微调(Fine-tuning)对绝大多数团队而言仍处于“尚未开始”或高度锁定状态,买方与卖方的界限十分分明。
然而,无论技术栈如何变迁,评估系统(Evaluation / Eval)始终是横亘在工程师面前的最大挑战。虽然这一痛点与其他挑战的差距在逐年缩小,但它依然稳居榜首。更令人深思的是,虽然市面上涌现出大量评估工具,但目前最主流的评估方法依然是人工直觉审查(Vibe Review: 凭感觉看几个样本的输出质量)。在受访者中,多达 96% 的人承认自己的 AI 技术栈存在各种问题,他们唯一的共同点是“无法就到底哪里出了问题达成一致”。这种散点图式的痛点分布,恰恰构成了下一代基础设施创业公司的“路线图”。
Original English
All right, so agents are out in the wild, which makes it a good time to look at what everyone's actually running underneath. So let's take a peek at the stack. Um, we asked, what is the biggest challenge in your stack? Every single year that I ask this, the answer, the number one answer is eval. Um, so Eval's lead here is same as always, but by a very thin margin, like that margin is getting smaller. And I'll say the quiet part here, which is that 96% of the people in the survey in this room have a problem with the stack. You just can't agree on which one. Um, so if you're deciding what to build next, if you're interested in infrastructure, that scatter is the map. And the leading challenge, how to evaluate your AI outputs requires many different methods, but as always, the vibe review is number one. So there are some consistent things that we'll see if they change over the time, but they they have not changed. Okay, this is interesting. So across eight layers of the stack, we asked what do people build versus buy. Um again, maybe the corporate card is is going to play a part in this, but there is a wide range and mix for every layer of the stack and a few clear takeaways. So the first is that inference and model serving is the layer that people buy the most. Many people don't want to build inference infrastructure and fair enough. Uh prompt management is the opposite. 61% build it themselves. Um apparently everyone's prompts are special. And this is true of a lot of the product logic prompts rag eval. They tend to stay inhouse on a relative basis. fine-tuning is the clearest not yet. Like most people don't have it at all and uh folks are pretty locked in. So those who bought aren't looking as much to build. Those who built aren't looking as much to buy. Uh but those are those are the core takeaways from the usage in our stack.
边界崩塌与工程债:人效提升的下游隐忧
尽管面临技术栈的混乱,AI 工具带给组织的整体收益依然非常显著——97% 的受访者报告了对组织的净正面效应。然而,这种正面效应的核心不是开发速度的绝对提升,而是失败成本的大幅降低。AI 使得尝试新想法、做原型开发和多路下注几乎变得完全免费。
但天底下没有免费的午餐。代码生成成本的极度廉价,正在转化为巨大的下游隐忧。超过 90% 的受访者感受到了负面的后遗症:最突出的问题是深层技术能力的退化(Erosion of deep technical skills),以及团队对整体代码库理解力的下降。大量由 AI 灌注的代码让系统的维护成本和代码审查负担(Review Burden)成倍增加。59% 的受访者明确表示,担忧目前的 AI 生成代码会为企业留下沉重的长期维护债务(Long-term maintenance liabilities)。
另一个令人吃惊的趋势是产研职责边界的彻底模糊。81% 的人认为 AI 正在模糊研发、产品、设计与市场之间的界限。超过三分之一的团队如今允许非开发人员直接将特征代码合并部署。在小规模或内部项目中这很常见,但也有 17% 的团队表示非开发人员正在常规性地向客户交付全栈特征。AI 虽然大幅提高了工作满意度(76% 表示 job satisfaction 提升),但也让“软件工程”这一传统职业的定义和护城河正在经历前所未有的重塑。
针对未来五年的预测,受访者的态度呈现出耐人寻味的复杂性:67% 的人相信在未来五年内,某家头部实验室会正式对外宣布(Declare)实现通用人工智能(AGI)(注意,大家赌的是“宣称实现”的公关稿,而非实际技术达成);同时,仅有 9% 的人押注当前大火的 Transformer 架构在五年后依然能维持最前沿(State-of-the-art)的统治地位;而在被问及“五年后,太空中的 AI 算力是否会超过陆地”时,36% 选“是”,38% 选“否”,成为了本次调查中最具分裂性的话题。
Original English
So many of you work on teams and like we said at the start these range from solo founders to large enterprises. What is this doing to teams? And remember this is a builderheavy sample. But among builders the vibes are good which you know I'm sure if you look to your left and your right you're feeling that the vibes are pretty good. 97% report a net positive effect on their organization. The top effect isn't really just speed. It's cheaper failure, more experimentation, more prototypes, more bets. It didn't just make engineers faster, but it made trying things nearly free. And so there's some happy campers as a result of that. But it's not free free. You know, there's no free lunch as nothing is. So the same tool that increases experimentation also increases review burden. Both can be true. And um you know o over nine and 10 respondents are feeling negative downstream effects in some way. The most common ones being wid you know widely discussed at this conference uh online and anywhere that you see AI engineers erosion of deep technical skills and understanding of the codebase. And these are consequences of cheap code generation. And the org chart is really feeling it. So many folks, 81% are saying that AI is blurring the line between their role as engineers and product design and marketing. These stats shocked me. Um, where you feel it the most is shipping software once exclusively the engineers domain. I know folks talk about vibe coding and how that's accessible to more folks than ever before in different roles, but today over a third of teams have non-developers shipping features, which was pretty wild to me. Mostly smaller, mostly internal, but 17% say that non-developers are regularly shipping customerfacing features across the stack. And even when non-developers aren't shipping, a third of teams see them building really useful things, prototypes, front-end mocks, and more. So, shipping software is not gated on being an engineer. We knew this, but uh the extent to which it's being pushed is is higher than I expected. All right, so where does all of this go? We always ask people to place bets rapid fire. So, let's talk about those results. Um, so present tense first. 76% say AI boosted their job satisfaction. So that's good for most of this crowd. I hope you're uh as uh Alphaba and Glenda say, I hope you're happy now. Um, that's great. But 59% fear today's AI code creates long-term liabilities. Only a third call software engineering a solved problem. Although uh when I have conversations with folks sometimes the way in which they define software engineering is different. So you can read into that stat as you will. Um happier faster but embracing the maintenance build is the TLDDR and people are unsure what's going to happen with hiring. And for the five-year bets we have 67% expect a leading lab will declare AGI in the next five years. Note the wording. We said will, we asked about the press release, not the achievement. So, will they declare it? Yes. What does that mean? Not sure. Uh, only 9% bet on Transformers being state-of-the-art in 5 years. Most are unsure. Uh, but that was interesting. And then my favorite, will there be more AI compute in space or on land? 36 yes. 38 no. The most divisive question in the survey is about outer space. I promised you a lot of bar charts and that was a lot of information. So a review or our 2026 wrapped impact is overwhelmingly positive. Image genen doubled or happy image gen doubled while audio has the highest adoption intent the same as last year. cost really became a first class constraint and we see that everywhere in monitoring in how ambitious folks that are going out and building AI products are behaving. Open weights augment but they don't replace. So we're seeing a multimodel future with a consolidation of the stack. Agents got right access more than ever before, tripling relative to last year. While the guardrails stayed pretty primitive, and inference is the buy market, everything closer to product logic tends to relatively stay more in-house. It is a very exciting time to be an AI engineer. I cannot wait to see how the next year unfolds. So, you can find the full report in the link up here. Every chart plus some cuts that we didn't have time for today. Um, I won't ask you to fill out a survey about the survey, but if there's something that you want on the books for 2027, something you're curious about, you can come find me here on the internet. I'm easy to spot. Thank you so much. Uh, we will see you next year or per 36% of you, maybe in orbit. Thank you.
📌 文中提及的人物和组织
人物: Barr Yaron
公司/组织: Amplify Partners, Notion, Vercel