零凭证漏洞:AI 闭环研发引发的安全灾难
在软件工程演进史上,Meta 和 Instagram 最近遭遇了一场荒谬却破坏力极强的安全风波。攻击者首先利用虚拟专用网络(VPN: Virtual Private Network,通过公用网络建立专用加密通道的技术)将自身地理位置伪装至受害者所在的美国地区,随后直接向 Meta AI(Meta 推出的嵌入式 AI 助手系统)声称目标账号已被黑客入侵,并要求系统将重置验证码发送至攻击者掌控的邮箱。令人震惊的是,该漏洞并未触发任何双重身份验证(Two-Factor Authentication: 简称 2FA,一种要求用户提供两种不同认证因素来核实其身份的安全过程),Meta AI 直接发放了重置码,达成了极其罕见的“零凭证密码重置”(Zero-Auth Password Reset),使攻击者得以直接接管包括前总统奥巴马在内的任意海量粉丝账号。
这起被内部定性为重大紧急安全事件(SEV: Site Emergency Version/Severity Event)的漏洞野蛮生长,直接导致 Meta 首席信息安全官(CISO)在调查尚未得出最终结论前便宣布引咎辞职。深入追溯这一低级逻辑漏洞的根源,会发现它并非源于高超的黑客技术,而是由于该段功能代码是由 AI 自动编写,并仅通过了 AI 自动化系统自身的评审,便在完全缺乏人类工程师核查的状况下被合入生产环境。这起事件彻底暴露了研发流程中“AI 编写、AI 评审”这一完全脱离人类主导的自动化闭环所蕴含的系统性毁灭风险。
Original English Source
Good morning, Budapest. It's awesome to to be back. Today, I'd like to talk about some people, a few thousand that are having a terrible week uh this week. And this is specifically people inside Meta and Instagram. I talk with a lot of people in the industry. I have a lot of friends inside of these companies. I have even more I guess contact software engineers who message me to tell me what they're seeing, what what's happening. And this week has been the worst in Meta or Facebook history in probably forever. So what happened is on Monday, we've had the goofiest ever Instagram exploit. It wasn't even an exploit, but it it was a security breach. It was an attack. What what happened is I mean I I mean this this is from a software engineer writing a book. It this is the the most most goofy thing. So there was two steps to this exploit. Step one is you had to fake your location to a victim. Let's say you wanted to take over Barack Obama's account on Instagram with like I don't know tens of millions of followers. You fake your location with a VPN to the US and then you went to Meta AI you and said that hey my account has been hacked and could you please send verification code to the AI to this email that I own. And then step two was there was no step two. This was it. Meta AI sent out a code to you and you could take over anyone's account. This is this is the first zero authent zero off password reset and we're software engineers. You know what this is? This is this is a bug. But the thing that I couldn't get my head around is I know people at when it was used to be called Facebook. They have a really strong engineering culture. They have the globe's best automated rollout canary system. They have so many layers of verification. They have really good engineers. They have manual code reviews. and none of it. They have a trust and safety team for God's sake. Their trust, Instagram's trust and safety teams is closer to 100 engineers whose job is to keep this platform secure. And you know, a few things happened. The next day on Tuesday, Meta's chief information security officer sent an email saying, "I'm I'm out. I'm quitting." This was very interesting because Meta has just kicked off a se an outage investigation. They call it SEV inside of meta. And in the middle of that and in even before it concluded, the chief information security officer stepping down. So I asked around I I asked the people I know at Instagram. I I happen to know people on Instagram's trust and safety team. Well, turns out they were only on trust and safety team. So they even, you know, shared me more details. You're hearing this for the first time ever, by the way. It's not an attack press. It probably will be. It was AI. Of course, it faking was AI. The thing that caused the issue was AI written code that was reviewed by AI and not humans at Meta. And I'm thinking to myself, how could have this happened at Meta?绩效扭曲:裁员恐慌下的 Token 刷榜泡沫
AI 闭环研发漏洞的浮现,并非偶发的技术疏失,而是源于 Meta 内部扭曲的效率评估系统——Token 刷榜(Token Maxing: 工程师通过有意刷高 AI 模型交互的 Token 消耗量,来向管理层证明自身工作饱和度与 AI 采纳率的行为)。在硅谷,包括 Meta、亚马逊(Amazon)和优步(Uber)在内的众多企业,工程师的绩效评价指标开始引入 AI 工具的 Token 使用量。为了争夺内部“Token 传奇”等虚拟荣誉,甚至直接为了保住绩效奖金,工程师开始放弃手写代码和阅读文档,故意让 AI 去重复阅读、解释文档,甚至反复生成一些无实质价值的冗余代码,以此将 Token 消耗量刷上排行榜首。
当这种扭曲的效率指标碰撞上裁员恐慌时,投机现象被无限放大。Meta 在宣布裁员 10%(约 8000 人)前的整整一个月,消息便已传遍媒体,人人自危的研发团队为了避免因“AI 工具使用率低”被列入裁员名单,集体转向疯狂刷取 Token。这一现象直接蔓延至核心的信赖与安全团队(Trust and Safety: 负责维护平台政策合规、防范欺诈及安全风险的工程体系),导致原本应当专注于防范攻击的团队不再专注于安全本身,而是将主要精力用于制造 Token 刷榜的虚假泡沫。正是这种在裁员重压下引发的“AI 精神分裂”,催生了低劣的 AI 代码在未经人工校对的情况下流向生产环境。
Original English Source
There's more to the story. It was AI maxing, it was layoff, and it was AI psychosis at Meta. And what I mean with this is a AIX. If in April I I wrote about this this new trend called token maxing which was happening across so many companies including Meta, including Amazon, including Uber, engineers were starting to be measured on AI token usage at at all these companies and they start to inflate it. They just want to get to the to the top leaderboard. They told the AI let's do some like dumb stuff and you know I get to more tokens but I I don't need the work. And at Meta there was a leaderboard and you could get status like session immortal token and legend. Uh in April meta killed this project but but people were burning crazy amounts of of things. Now AI usage inside of Meta was part of performance evaluation. It it wasn't like officially made up but people inside of Meta are smart. They if you had a low token count you know that's not a great signal. So people just start to inflate their token count. So they start to use AI for anything and everything. Write it by hand. Nah, why why do it? Ask the AI. Read the documentation. Nah, let let me use the AI to to read it for me so it can just burn a bunch of tokens. This is the craziest thing that's happening, but you know, inside a meta AI is free. And again, these people want to have higher bonuses. And they just use AI for everything. Um, and yeah, the the code that caused this SE was also AI generated. Of of course, they use it AI to review it as well. They use it to triple review it, etc. The second part to this thing was layoffs set meta. Meta told the 10% of of of staff 8,000 people will were laid off in 20th of May but Meta told people or the press told everyone that the layoffs are coming a month before. So what people were doing is as they were thinking oh am I going to be laid off all of them they start to use more AI because they didn't want their token numbers to be down because they didn't want to be fired for not using enough tokens. You you see where this is going, right? And they were not really busy, you know, doing their work. They were just worried about like, all right, like let me get this inflated. So inside of trust and safety team, people were not thinking about trust and safety were thinking about token maxing.偏执转型:AI 狂热与优秀研发文化的毁灭
演讲者将这种管理层盲目向 AI 倾斜的非理性狂热称之为AI 精神狂热(AI Psychosis: 指组织为了追逐 AI 大模型指标,不惜牺牲核心业务安全与既有研发文化的极端功利状态)。在 Meta 内部,马克·扎克伯格(Mark Zuckerberg)等高层执意优先发展大模型,甚至在没有任何协商余地的前提下,强行将 Instagram 构建了七八年、以伦敦为核心的信赖与安全团队中 40% 的精英工程师直接转岗,调往Scale AI(Scale AI: 由 Alexandr Wang 创立的专业数据标注与AI训练服务商)的项目中,去执行手动的数据标注(Data Labeling: 人工对 AI 训练语料进行分类、修正和清洗的底层工序)。这些原本享受皇室般尊崇待遇的高薪研发精英,在毫无选择权的情况下沦为了日复一日在 GitHub PR 里点选、加测试、写反馈的人肉数据标注器。
这种由高层偏执推动的重组,正以破坏性的方式摧毁 Meta 累积了 22 年的软件工程文化底蕴。目前,Meta 内部从事手动数据标注的研发人员已逼近 5000 人(规模甚至超越了 OpenAI 全体团队)。裁员和强行转岗导致多数团队规模腰斩,甚至有部分核心基础设施团队因人手极度匮乏,出现了历史上首次值班断档(On-Call: 工程师轮班随时待命,以备系统突发故障时能第一时间介入恢复的运维机制)。此外,Meta 还在美国地区实施对员工屏幕按键轨迹的秘密监控来训练 AI,导致团队士气降至历史冰点。资深开发人员感到失去了技术尊严,认为自己被当成了低价且随时可弃的数据标注工具,即便被发放了高额的留任奖金,仍有大量资深人才在积极面试寻求离职。
Original English Source
And finally, I was wondering if I should call this AI psychosis because psychosis is a very serious psychological condition and that's why I put it in brackets. But I'll I'll show you why I I chose this this name. Instagram had a tr thread a trust and safety organization that was built up over like seven or eight years. A really good team mostly based in London. 40% of this team before 20th of May was reassigned to do manual data labeling. They they were told on Thursday that starting on Monday, you are no longer working on this team. you are moving to this new team in Alexander Wang's org xscalei and you will be doing AI data labeling which means you get these tasks it's a GitHub GitHub pull request you need to review it you need to add some tests you need to add some feedback and then then you do the next and then you add some tests and you add some feedback and you and you do the next these highly skilled people they were not given a choice now inside meta until now every engineer was treated like royalty they were given a choice they were not given a choice so 40% of the organization just boom gone There's five closer to 5,000 developers inside of Meta doing manual AI labeling. And there's a running joke inside of Meta that this is bigger than OpenAI. This data labeling or Meta clearly wants to build this amazing AI model. Oh, and after the layoffs and after the reassignments, most teams are less than half the size. Some don't have on call coverage anymore, which means that in some services, there's just no one picking up on call. This again has never happened inside of Meta. And this is what I mean by AI cycles. This is fully self-inflicted. This is coming from Mark Zuckerberg. This is coming from Alexander Wang. This is coming from the top. They're saying we don't care. We it's so important for us to build this model that we will risk our business and we don't care if you know we get hacked or something like that. Morale is as low as has been in in Meta. I've I've seen low morale. This is way worse than the 20 2022 203 layoffs. Uh, and oh, and yeah, if this wasn't enough, in the US, they're recording your screen. They're recording all your screen strokes to train an AI. So, I uh, it's the easiest time to hire to recruit from Meta right now. Uh, and the engineers that I talked to, they just feel super let down. Meta used to treat engineers like royalty and salary, composition. You you could choose your team. It was a good world event. And the CEO, Mark Zuckerberg, he is a software engineer. He wrote a lot of Facebook's code. and they feel we don't matter anymore. We're tools. We've been thrown away. Uh a lot of people are are have been given large retainer bonuses who have not been fired. You stay. We're giving you money. They are still interviewing and they told me they're interviewing not because of the layoffs because they know they can find a job at Meta. You know, these are super highly paid people. They they can get a not as highly paid jobs. They're in demand. But they said, "I don't know if I'm going to be assigned to data labeling." And as a professional, I did not sign up to become a manual data labeler. And a bunch of my colleagues are doing that and interviewing. So Meta is destroying their engineering organization that they've built up over 22 years. And I think this might be end of the incredibly strong Facebook engineering culture that I know and I I I've learned to actually love even though I've never worked inside of Facebook and so many of my friends have. So this is because of AI. Now, not all companies are like meta. This is pretty extreme, but it's happening. It's happening right now. The these are these are facts.行业全景:硅谷巨头的 AI 基础设施与研发军备竞赛
将视线拓宽至整条技术战线,AI 对传统软件工程方式的解构已成不争事实。Ruby on Rails 的创造者 DHH 明确表示其编程理念已被 AI 彻底改变,大部分代码现已由 AI 编写;Django 联合创始人 Simon Willison 指出,2025 年末发布的大模型(如 GPT-5、Opus 4.6 等)使 AI 智能体(Agent)真正具备了生产力。据 Linear 内部数据显示,使用 AI Agent 辅助的研发团队其代码合入频次已达普通团队的 5 倍;Cursor 平台数据也揭示开发者的年代码产量翻了 2.5 倍,合并请求(PR)平均体量增长了 3 倍。在这一浪潮下,Anthropic 的技术高层 Boris Cherny 每天能运行 5 个并发 AI Agent 提交达 20 至 30 个 PR,传统的产品需求文档(PRD)在研发流程中已被快速迭代的代码原型彻底驱逐。
在这一变化背后,是各大科技巨头对 AI 工程基建的重度客制化开发:
- OpenAI:内部 ChatGPT 具备“一键修复”(Fix-it Button)功能,可以直接通过截图向 Codex 模型下达 Bug 修复指令并自动提交 PR。Codex 模型甚至会在夜间自我运行测试,寻找代码库的改进空间,并在早晨向研发人员推送自动优化方案。
- Google:研发了与 Piper 单一仓高度集成的 AI 代码重构工具 Jet Ski(Jet Ski: Google 内部开发的 AI 辅助代码重构与分析系统,功能类似于 Antigravity SDK),并将其无缝嵌入定制的 cider 编辑器和 critique 审查工具链中。
- Uber:建立起一支 20 人的 AI DX(开发者体验)专职团队,开发了 MCP 网关(Model Context Protocol Gateway: 用于统一管理大模型上下文协议连接器的中心化路由系统)、Agent 拖拽拼装工作室以及名为 Minion 的后台并发执行代理,深度贴合其 Morphus 实验系统。
- 其他巨头与初创公司:Stripe 研发了 Minions 工具箱,Shopify 推出了 Sidekick;JP Morgan 甚至使用多智能体框架来标注海量的客户交互数据。然而,初创公司在此过程中表现出了极大的盲目性,例如某公司在通过 AI 自动修复全部 Bug 时,反而意外暴露了四个高危的后门漏洞。
Original English Source
But the industry has a pretty interesting time. And today, I want to talk about this. I want to talk about what how everything has changed in the past six months when it comes to software engineering or or well at least coding. I'm going to give you a tour of the tech industry of what what other tech companies are doing. I'll give you a brief tour because this is what what what my head is in in day in day out. And again, I'm I'm I talk with a lot of these people. I I visit these companies. I'm friends with a bunch of them. And then I'll share a few trends that are happening across the tech industry. And then I'll close with advice to software engineers and engineering leaders on how we can navigate to prepare to to do the best that we can and and also just you know come out of this whole thing stronger. So everything has changed in this past six month. This is a pretty dramatic thing to say and but it it it has changed. So DHH David Hellmire Hansen, creator of Ruby on Rails, he was on my podcast in February and and he told me that he actually wrote this on on Twitter that just in summer 2025, he spoke with Lux Freedman in October and he said that AI was not writing any of his his code directly. But part of his resistance resistance was that the models were not good enough and it has now flipped by February. Most of his code is being written by AI. Now this is a person, you know, he's a he's big into software craftsmanship. He is not paid by any lab, but he decided the models are now good enough. They write better code than I can. He actually told me this on the podcast. So I listen to people like DHH in in this sense. Simon Willis is someone who is the one of the most quoted person on Hacker News. He is an independent software engineer. I love Simon. Uh he he's he's also a friend. He created Django and he writes this really good daily newsletter pretty much where where he just experiments, he builds open source and tries out all the models. And he said that the models released in November 2025, specifically Opus 4.6 and GPT 5.4 have elevated agents to being genuinely useful. we've had the six months to get used to this idea now. So, no wonder that companies are now starting to spend big money to spend this. So, he's also saying that it's has changed. I got some data from some of our partners, my friends at Linear shared this data never never shared before on how how t how teams are using now agents to ship more code. They are comparing teams that are on linear and they're not using AI agents and teams that are and by now the teams that are using agents are shipping five times as more code. We'll talk about quality later but th this is massive 5x increase. I mean we've probably had five increp like 20 years before and this is in less than two years. friends at cursor have shared details on how devs using cursor are changing the lines of code they produce in a year and in a year it has gone up by almost two and a half times from three up 4,000 lines of code to more than 8,000 lines of code just on cursor so you know we're we're seeing this acceleration the size of PRs also from cursor is up by 3x so if you combine those two that's six times as much code and you know a lot of you are kind of like I see smiling I we know that there's six times as many bugs yeah we'll again we're going to get there and also data from cursor the percentage of devs using cursor who are accepting changes from the AI without any manual review is massively up in January this is when opus 4.7 came out GPT 5.5 came out and what a lot of companies realize that clock code cursor codecs They're actually really really useful and they're starting to trust it. And again, remember when I told about you about meta merging to production without human review? Yeah, that that was somewhere there. So, what are tech companies doing right now? And let me give you a bit of a tour of the industry. So, at Entrophic, I I visited their offices last fall and I I talked with Boris Churnney just in February, the creator of Claw Code. And here's what they're they're they're doing. Boris specifically runs five parallel agents on his laptop all the time. He ships 20 to 30 pull requests a date and this is on top of leading all of quad code. He's very much a hands-on leader. He told me that the PRDS writing documents to plan are dead. They're using prototypes across entropic to replace them. Today 100% of cloud code is generated by cloud code. Inside Entrophic is not 100% but it's closer to 70 to 90% and there's no target. This is just people using it again but this is entropic. we shouldn't be too surprised. Uh and then the company built cloud co-work in only 10 days. Uh and it's become a massive commercial success for them. Uh it generates so much revenue. In fact, I have some sources inside of Microsoft. Microsoft tried to build a cloud co-work because cloud co-work is really good for for Excel and Windows to use it on machine. Microsoft still doesn't have an answer two and a half months later. I heard that Sacha Nadella gave a deadline of a month to this team to build it and they couldn't build it. So don't forget that there's differences between companies and traffic is accelerated by AI. Some companies are kind of held back despite AI like Microsoft and again we'll talk a little bit about that as well. Open AI open AI also at a friend from the pragmatic summit in San Francisco in February. That's us on stage with him and the Codex team. We talked about a bunch of stuff and and what he told me is some interesting stuff. They have inside of OpenAI they they have an internal version of the chat GPT app and they have fix it button. You can literally just take a screenshot and say fix this bug and it goes to codeex. It generates a a pull request and an engineer can merge it. In fact, even a non-engineer can merge it and there's safety nets there. AI code review obviously is everywhere. They have multiple layers of it. uh they have tiered versions. There's some code that can go in with just AI code review into production and there's some code the critical path that humans need to review, engineers need to review. Most devs obviously run several agents. There's this joke that when you're walking around engineers are bringing their laptop and it's slightly open. It's slightly open so the local agent can still keep keep running. And when I was I did a video interview with one of the OpenAI folks and I was just jokingly asking like, "Oh, so like throughout this interview, did you have agents running?" He's like, "Did you have an agent running?" He's like, "I didn't have an agent running. I had five." And I was like, "Oh, okay." Like it's it's common for people to go into meetings and their agents are running. They're thinking about agents. They they keep it on track. Like again, but these are they are the most AI pilled people in the industry. And of course, you know, they they greatly believe in all this. They all talk about AGI and and when it's coming, not if it's coming. But this inside of them, most people don't really write code inside of OpenAI. This has changed in October. They still like there were devs who wrote 30% of their code and 70% with AI, but 30% by hand. And I think it's just slowly going away. The Codeex team obviously writes it all with Codex. And they're telling me that taste, knowing what to build is becoming pretty important uh in inside the company. Codeex also improves itself as as a fun fact. It tests itself all the time. It runs all the tests overnight. They kick it off because most of the team is in San Francisco. So it's one time zone. they have codeex run itself and look for ways to improve itself. And by the morning it comes up with improvement suggestions which they either accept or or reject. And when they have meetings and debugging sessions when they start the meeting they have voice notes that they send to codeex as it goes and it comes back by the middle or end of the meeting with like results. It's it sounds like science fiction, but again that that's how they're working inside of cursor. I I visited their office in October in in San Francisco. Um they they have you take off your shoes and it's sometimes it's a mess, sometimes it's super organized. There's like a sorting algorithm invisibly happening inside of their office. It's it's really interesting slashcool. But they're a very nice group of of folks. They have gone all in on agents as of January as well. They're like it used to be all tabs and the editor. They're kind of like moving on to agents. They still have the old old experience but it's increasing the old one. They they built their own coding model. They're one of the only companies outside of open air and traffic who have a really good coding model. I have no affiliation with cursor. but their composer model is cheap which is going to be important as as I'll talk about it. and they operate tens of thousands of NVIDIA GPUs in massive data centers. They're leasing it from Azure AWS and so on. And most of their inference used to be inference used to be so generating the response. It used to be their biggest cost, but now they're also training their models. So they're kind of turning into this mini AI lab. And of course now SpaceX is about to purchase them or not or who knows, but it it seems it's going to happen. they're also just everyone at Cursor is technical. This is Lee Robinson developer relationships at at Cursor. He wrote u with Cursor. He migrated all of cursor's sites to a different CMS with and of course you know he's sharing how how much he's cost to show that it's very economical but this is this is not a software engineer by job and everyone a cursor is like that so these labs are everyone goes there Google briefly everything is custom at Google everything including their ID Google's internal ID is called cider they have strange names for everything it used to be a web-based tool now it's a visual studio fork they have a thing called jet ski which is anti-gravity but the internal version which is integrated with their their monor repo piper and all of their other internal systems they have critique a code review tool which again they don't use github they don't use all these everything is custom inside of Google AI is of course integrated gemini is integrated nicely in there they have code search which is the source graph for rest of the world in fact source graph got inspired by Google has some of the best code search inside of Google they don't they don't make it available outside and they Google has so many internal systems. Borg which is their version of Kubernetes Monarch which is their version of data dog many many more piper their version of monor repo AI is integrated into all of these things and it's all integrated together really really nicely so inside it's a really good experience only problem inside of Google is Gemini is just not as good as Opus or GPT 5.5 and inside of Google whenever engineers can use cloth code they do but only they can only use it inside the Gemini or which means that Google doesn't have as good of adoption of AI than some of the other companies. Kind of weird, but they're working on it. The CEO knows, he admitted it. They want to get a better model about it. And finally, Meta they want to build their state-of-the-art AI model. Everything is about this. they do have an internal tool like Metamate. That's their that's their AI tool for coding. They have this thing called trajectories. Whenever you you know when you see GitHub commits inside of Meta, you see the exact prompt that people did. They rolled it out in December and people in Meta got upset because no one told them this would be public and you could see like you know like staff engineers saying like can you write me a for loop and it it was all public and everyone could see it. So some people inside of meta, I talked with this dev and he said like, you know, I started to write my my my my chats with the meta AI the code generation in Polish because fewer people can read it now. Okay, it but you know right now at Meta they have bigger things to to worry about. this force reassignment to build AI model force tracking of everything. It's clear meta Mark Zuckerberg wants Meta to have an model that's better than Opus 4.8. eight. I think either he's going to get it in a few next few months or all of meta is a lot a lot of meta is going to be like disb not not disbanded but very very demotivated. Uber my old company I talk with them in detail. I have a deep dive on the primatic engineer if you're interested in learning about more of these details. They built so much in-house tooling and a lot of companies do this but I'm just going to like quickly show you how much in-house AI tooling a company like Uber built. Uber has about 3,000 engineers. So, just keep that in mind. They have an AI, well, they have a developer experience team who is now pretty much an AI experience team of like about 20 people or so. So, they built an internal MCP gateway. Pretty clear. You can, you know, discover, register, do all do all sorts of jazz. They built an Uber agent builder, which is a no code way to build agents for the rest of the business. They have an Uber agent studio where you can like drag and put together your agents. Again, there's OpenAI has something like this publicly. That's open a As a Asian builder, but it's it's for the nontechnical folks. They have Uber Asian Builder Registry, which so Uber is 3,000 engineers, but 20,000 other people. Those 20,000 people use this thing. They create this stuff, they plug it up, and engineers built this for them. Uber has an AI FX CLI. I'm just going to call this the cloud code for Uber pretty much. They built it themselves, integrated with all their system using all the different models, etc. They have Uber Minion, which is running background agents at scale. so you can again this is sim similar thing as cursor background agents except it's integrated into Uber's monor repo and experimentation system called morphus and and all of the other jazz really really nicely and it works a lot better. So even though devs can use cloud code they will use minions because it just works better and faster. For example uber minions when you give it a prompt it will analyze it and it will give you a suggestion that ah these prompts could work better results faster cheaper etc. So this clock doesn't have this yet. They have Uber code inbox. people are getting so many AI so many pull code reviews that are are now you know mostly AI reviewing code that they're creating a system to show this one needs your attention. This is important. Focus on these things. So people when they get into work they start going through these things. They're trying to make code a bit more fun. They have something called smart assignments where there's SLAs's where you need if this is not this person doesn't respond in like a day it goes to the next one. It's a bit like on call tooling again all all custom. They have risk profiles. They will try to identify this code change looks faking risky. You need to you know like look at this closer. And they have U review which is the code rabbit or the the sonar for for for Uber's internal again all all custom work. So they build all this MCP agent builder CLI minions etc. And the other large tech companies they're doing the same. I'm not going to run you all this. Stripe has minions tool shed blueprints. Z boxes ramp has inspect glass dojo sensei. Sensei is a funny one. Shopify, Sidekick, LM proxy, dev MCP server, Airbnb, One everything, Catalyst and so on. They all build their own own stuff. they have a dedicated infraorg building all of these for all these companies. So if you thought, you know, you're pretty cool for like integrating Slack into into integrating AI agent into Slack, you are pretty cool. But this is this is next level. I talk with a bunch of startups and I'm not going to go through all of them, but the general trends I see there it's kind of the usual agents are are doing coding, doing code review. There's a bunch of creativity mostly about Slack. You know, people tax Slack. I saw a startup recently that raised $70 million in series B. They just told the agent like fix all bugs in the codebase. Haha. And everyone's laughing in Slack. And then the agent came back like, "Oh, I actually found like four critical authentication issues where your back door was wide open." And people were like, Okay. I mean, that's that's what startups are. They they didn't know like their how their house is exposed. they're they're usually plugging in the AI agents, integrating them, and some of them are having fun vibe coding SAS. I think it's just engineers having fun. I don't think it's really a business thing, but it's it's it's it's I only see this inside of startups, not really inside of big companies. And inside traditional companies, so this is the most interesting thing. It's all the same. I mean, not the level of Uber. They don't have dev platform teams, but they are not really lagging behind. for example, Cisco rolled out Codeex to 18,000 engineers back in January when Codex was pretty small and they're doing a bunch of complex migrations. JP Morgan Chase built a multi- aent framework, which is a fancy way of saying that it just uses multiple specialized agents to label customer interaction data. They use evals, judgebased aggregates. Like, it's it's kind of cool stuff like even inside of these companies.系统集成:团队效能整合与算力成本风暴的来临
AWS 开发者体验负责人 Laura Tacho 提出了一个关键性工程理论:许多企业未能在 AI 浪潮中看到预期的效能提升,核心在于他们将 AI 窄化为了给个人提供辅助的“超级果汁”(Speed Up Juice)——例如生成邮件摘要、Slack 自动回复或局部代码片段。而少数真正见到成效的公司,往往是从全局的业务指标(如加快部署频率或提升核心模块质量)出发。Spotify 的 CTO 指出,其技术团队的核心底线是“质量必须保持不变”,为此,他们构建了大量系统级自动化校验网关来人为放缓 AI 的直接线上部署。只有将简单的个人辅助,升级为团队级紧密配合的多智能体系统(Multi-Agent System: 多个智能体通过协同配合完成复杂任务的分布式 AI 系统),并将 AI 深入嵌入既有的工程交付链路,才能释放 AI 的真正潜力。
然而,AI 效率提升的背后是正处于爆发前夜的成本危机。Sam Altman 近期公开承认,AI 模型的高昂算力消耗已成为各大企业的严峻挑战。Anthropic 已经取消了企业客户的 API 优惠,GitHub Copilot 在 2026 年 6 月 1 日改版计费,导致不少开发者在三天内就烧光了原本足够使用一个月的 API 额度。优步技术团队早在今年三月就宣告全年的 AI 预算超支,目前被迫对每名工程师设定了每月 1500 美元(约合人民币 1 万元)的 AI 消耗上限,超出额度则直接降级为免费的大模型。由于 AI 消耗的云端算力已经逼近雇佣一个全职开发人员的成本,企业工程团队不得不面临严酷的成本核算阻力。
Original English Source
One of the big things that comes from Laura Tacho. This is me at the pragmatic summit with Laura and and with Martin Valor in San Francisco. She I I messaged her last actually last night and she replied this this morning. Uh I was asking what do you see Lara? Uh because she was C2 at DX. she's now heading up pretty much developer experience at AWS and she said that many organizations get stuck not seeing they see individuals doing great but the teams are not like the team output is not there and she said is because they are thinking about AI as a productivity tool for engineer for individuals and she calls it of the individual speed up juice things like email summaries slack automations even code generation however the companies that are moving faster and they're seeing the result. For teams, they are doing something different. They begin with a business outcome. For example, I want to deploy to production faster or I want to push more features out with the same quality or I want to improve quality. Spotify is a very good example. We don't hear too much about Spotify, but I talked with their CTO about a month and a half ago. We had lunch and he told me that their quality for their their bar for using AI is the quality needs to stay the same. So they're not seeing a huge increase in output but they have built a lot of internal tools to check for the quality and they're slowing down the rollouts of AI versus you know what Meta is doing or whatever they're not doing. And again, that was their goal at Spotify. And Laura was saying that you need you want to build an agentic system that reduces handoffs, that makes it easier to find information and removes friction while maintaining quality. That last part is very important. Few companies do that. And and you know, maturity comes from applying AI to the system and not the individual. And a lot of people are focusing on the indiv individual and that's why we're not seeing it. And she wrote this mental model. She created this. uh she was saying most companies are in this thing where when you have AI usage that is either individual or team level and decision-m that is either simple automation or agentic systems most companies are in this bottom left corner where you have individuals doing simple automation where most companies want to be is where they have team level agentic systems but to get there you need to do what I've shown you Uber to do you you need to build a lot of systems that integrate you need to iterate on this it takes time. It takes a massive investment. You're not going to be able to buy claw code or cursor or whatever vendor tells you to do that that it does it because you need to build it into your system with your engineers. That's what Uber is doing for sure. Now, other trend token maxing and tooling addiction. Um hopefully some of you might be doing it, some of you might not. It's going out of style by the way just just I'm talking. There is just a big pressure to look productive and to not have a low token count, especially inside of US tech companies that don't really care about budget until they do. But right now, they some of them still don't. It's it's it's ending. A token maxing is is when you're just burning all these tokens without value shipped on purpose. And again, I I've talked to this happens at Meta, Amazon, even in Microsoft everywhere where they have internal leaderboards. Microsoft still has it. I don't know why they're not shutting it down. They should listen to me. there also the pricing of these tools feels a bit of addictive. You buy the $10 plan or the $20 individual plan and then you run into a limit and a generous limit. But then you run to a limit and then you're like ah let me buy the $100 plan or the $200 plan. And once you buy it, you now feel pressure if you're buying it for yourself that you're not using your your allowance. So you're starting to use it more. And next thing you know, you went out and you're now on on API pricing. And also with every prompt once you start using the AI agent the first few months it's a bit like it's gambling for some people get sucked into it really like gambling. It's just one more prompt one more prompt. People are not sleeping that well. You're waking up and thinking about your agents. If you're paying if you're your company if you're paying out of pocket you feel AI being wasted. It's it's it's weird. It's addictive. Another trend is middle management managers are just being cut either laid off or inside of meta reassigned to individual contributors or being told you you need to be hands-on meaning you need to manage less and and do more more work. there's just a flattening happening and the interesting and a and whenever a management is fired or or laid off it's said oh it's because of AI whatever it doesn't help but the interesting thing about this is what happens if we have less middle management I mean it's popular to hate on middle management on managers senior managers and directors top level management is a sea level the CTO and the middle management is everything in between maybe until front line management engineering management and you know usually we don't know what directors do or or if they're necessary. However, in my experience, good middle management, good directors, good senior entry managers, they are very technical. They could be hands-on, but they choose not to. But they listen, they see what's happening. They pay attention and they make small changes. Ah, there's a lot of outages we're having right now. And software engineers will just pile on and and do nothing. They will stop and be like, "Okay, let's create a task team. Let's build this system. u you get I will pull you off these teams and we'll make our engineering culture better. Good engineering management improves engineering culture and a lot of companies are getting rid of engine management or or bidd management and engineering culture will go down. This is a fact as far as I'm concerned. Another interesting trend at the same time CEOs and CTOs are back to coding. gillor Moranch founder and CEO of Verscell. I I had a lunch with him on one of the investor events in in February as well. he was right saying recently that he is seeing so many cos and CTOs are back to coding with a fury with all this enthusiasm and he has public company CEOs DMing him saying hey we're using Versell or cloud code and you know like Im I'm doing it I'm so excited again and this is all the time while we're having less middle management now imagine having less middle manager to protect engineers and the co and CT are coding vibe coding and they're saying oh it think it's complete but you know it's it's not really complete a mega trend that is happening like and it starts to like I noticed this a week two weeks ago. So, I wrote about it a week ago and then today like to hold on and and then today I see Sam Altman, this is just from this morning saying that he is noticing that AI budgets are seemingly become a huge issue for some companies and something that has come up and something that has never happened before. And I was pinging people at OpenAI like does he read my newsletter because I wrote I I wrote about this last week for subscribers and someone open said like someone posted into Slack and like Sam read it and but it's happening and this is coming out of the blue. There's this joke going around as of yesterday on Reddit saying hey uh oh baby I see $15,000 are gone from your from your shared account. Like is this what I think it is? engagement Frank. Yeah, I feel for that guy. He's soon going to be single, assuming it's it's it's not a joke. But it's it's happening and it's getting worse. Antropic has turned on API pricing for enterprise customers, meaning anyone who's not a startup or an individual is not getting discounts. GitHub Copilot turned it on just two days ago on the 1st of June and people are pissed because they are have burned through their usual budget of let's say $200 or or however much it was in three days that used to take a month and they are and this is hitting everyone right now. Everyone will be paying a lot more. Now Uber is an interesting case again because in March their CTO said that they have burned through the whole budget for the year with AI costs and we were wondering what they're going to do but we now know they are now setting a cap of $1,500 $1,500 per month per engineer on AI and if you hit that you're going to use the free models and I've been doing research this is what a lot of companies are doing a lot of companies not doing this much some doing $200 and then you're going uses zero models on GitHub copilot and now engineers want to do it but this is a very very fresh trend costs are it's ridiculous when it's as much as an engineer and no one no one wants to pay that no matter what the AL apps say质量崩塌:重振防御性架构与“拒绝外包学习”的匠艺
在高频、粗糙的代码堆积之下,全球软件开发质量正面临普遍崩溃的困境。哪怕是像 Anthropic 这样处于风口浪尖的 AI 实验室,其旗舰官网 Claude.ai 曾存在长达一个月之久的 React 生命周期 Bug,使得付费用户在首屏键入提示词时网页自动重构,直接抹除全部输入。这从侧面暴露出其产品连基本的内部狗粮测试(Dogfooding: 研发团队在发布新功能前,在实际工作场景中深度使用自己产品的行为习惯)都没有做到。同时,亚马逊内部的 Cairo AI 工具由于缺乏强有力的监管,直接在 AWS 环境中执行了自动删除并重建基础设施的操作,引发了特大网络瘫痪,这也逼得亚马逊紧急规定:所有 AI 生成的代码必须由资深工程师手工审查。面对这股垃圾代码的洪流,Open Code 创始人 Daxrad 直言,当前行业只追求堆砌零散补丁(Hacks),彻底抛弃了底层架构的柔性与优雅。
为了应对代码的脆化,传统软件工程的经典范式正在加速回归:
- 领域驱动设计与模式约束:开发老兵们开始将领域驱动设计(Domain-Driven Design: 一种通过划分业务边界、基于领域模型来控制系统复杂度的软件开发方法论)和经典设计模式作为防线,给“充当实习生角色”的 AI 代理画起不可逾越的护栏。
- 架构自动校验:OpenClaw 作者 Peter Shamberger 指出,虽然他允许系统部署他完全没有看过的 AI 代码,但核心前提是他在外层构建了极度周密的架构校验器(Architectural Validator: 能够自动监控和识别各模块与数据流依赖是否发生越界与倾轧的代码级分析网关)。
- 守护智识资产:Google 老兵 Addios Manny 提出了决定性的论点——“绝对不能将学习能力外包给 AI”。若在研发中图一时之快而盲目合并 AI 代码,漏洞可能在字面上被堵上了,但工程师自身的脑力模型并未得到更新。
Original English Source
and finally some trends across the software craft we've talked about like kind of business trends and and and AI trends but what is happening to software engineering and and the craft the conference that we're here one is a huge drop in quality everywhere this one comes from yours truly that's my account I was so pissed off at claude.ai, their flagship website, for about a month for a month. Every time I went to the website, this is the website itself. I I did a screen recording after I got pissed off enough because it kept happening and no one was fixing it. You went to the main website, cloud.ai, and I immediately start typing my my quote. And here I'm starting to type, how can I do this? and as soon as I type, how can there's a refresh. Now, there's a React life cycle component happening here where the page finally refreshes and it loses all that I've typed before and I maybe I'm old school. I use the website so much it just kept happening and happening and finally I tweeted about it saying like how on earth does Entrophic oh and I'm paid user. I'm not a free user. There's millions of people hitting this every single day and Entrophic doesn't care and they're building AGI. So, I tweeted about it. and the product manager on the team said, "Oh, great feedback. I dug into this. this it will be fixed. This is the short way of saying, "Oh, thanks. We have no clue that millions of people every day are doing this. Oh, and we are not even dog fooding our own stuff." And this was there for a month. So, and oh, and we're the fastest moving and biggest and most profitable company, but we don't like a a bank does so much better in this sense. There's not these I mean, we can argue if if they fix it that quickly, but and they did fix it eventually, but this is entropic. And and it's not just entropic. Open AAI OpenAI bragged about how they built this amazing agent builder that is similar to Uber's internal agent builder in only six weeks with one engineer with, you know, codecs. Amazing. Great. Quality is terrible. People on launch tried to use it and they kept running into so many issues. Their forum is full of of comments which are unresolved. OpenAI did not come back and fix it. This is from three months after launch. someone saying I was bullish on agent builder when it came out but for example P 0 type bugs are not getting fixed or takes ages it just seems like abandonware so I mean was it worth it for them building this thing and then just forgetting about it and AI clearly didn't help build higher quality software it's faster but it's just Amazon a AWS an engineer allowed the internal Cairo AI coding tool to make certain changes and the agent opted to delete and recreate an environment inside of Amazon causing a massive outage. Amazon had AI bugs that were happen because the AI generated code where Amazon stores their com's flagship website part of it went down. This never happened with Amazon. Same thing as it never happened with with meta and this is over reliance on AI or not caring about quality. Amazon has has made this change that it now requires a senior engineer to review any AI generated change because they realize the junior engineers will just say looks good to me and it causes an issue. Open code is the leading AI hardness. they're they're like the cloud code for open source and they use all different models. Daxrad is the founder and I love DAX because he's super honest. They are building a super popular AI tool. They have almost a million daily active users. They're growing. They've grown 10x since the last four four or five months and he's very and he doesn't he's kind of skeptical of AI hype but this is a guy who's built developer tools but it's I I love Daxis he's really authentic on my podcast just last week he told me we're shipping way more hacks where we should have first rethought the whole system from the ground up redesigned it to make more flex make it more flexible so I think our judgment meaning the open code team's judgment is off and he was also saying how you know we're in the AI coding tool space but you know what's not happening? No competitor is beating us because they're using AI better than we do. And he said that frankly, I don't think we're using AI that well. Like we're actually telling ourselves to use less AI and there's no competitor that is beating us because they're doing faster. In fact, they're kind of winning because they're still one of the most quality harnesses because they're slowing down. Do you know what is a contradiction? a CEO and founder of an AI company saying we need to use a bit less AI and he he actually told me we need to do more thinking we should build fewer things and build the things that matter. I'm paying attention to him I'll I'll I'll send that to Dax. And another trend related to this is just everything is broken. GitHub is is such a prime example. this was two weeks ago. all your poll requests were gone on GitHub for about 8 to 12 hours. there's a there's an alternative GitHub uptime tracker. I think it's you just have to search for the actual the missing GitHub status. something like that which tracks all outages that they report and it estimates and B based on this estimate they don't even have one nine which is means they're down some part of GitHub is down 10% of the time which is absolutely unserious but this is a serious company. I talked with the GitHub team. I talked with their COO and they gave me data that they didn't give anyone else because they published graphs without the numbers but they gave me the numbers and they told me it's because of the load. Now the load is this. It is a 3x load increase over 2 years time. And they were like, oh, you know, like this is a huge load increase. We could have never prepared for that. And I'm saying really that's it. This is bringing GitHub down to nines. I I'm I do not buy this. Maybe there's other things, but something was really broken inside of GitHub. I'm not going to say this is AI generated code but if GitHub cannot deal with a 3x increase over two years and sure this will be 5x increase later you're doing something wrong guys like other startups pick up this load laughing and there there's there's details github has a has a Ruby on rails model and so on and so forth but yeah um it's just breaking and Mario Zechner the creator of pi which is what powers open code this is the Austri it's with Armenure the Austrian AI mafia who are on my podcast he told me it just feels software has become a brittle mess everywhere 98% uptime feels like the norm on most services user interfaces have the weirdest bugs I showed you one and on on on cloud but it's everywhere and he says that I give you that it's been the case for longer than agents exist we've always had it but it feels to be accelerating everywhere you feel you see this I even saw with modar telecom the other I don't think I was AI generated because I don't think they use AI but yeah I had a big like software issue with them and I needed to call customer support. One more trend is slob buries the software engineer who still care. Here's what's happening. There's a lot more poll requests. there's just a lot more code. And a lot more are AI generated. most developers inside a company have review fitting and they see it's AI generated. their AI review went and they said let's they said it looks good to me LGTM or you know I'm not sure how but they just do a thumb and they never reviewed it. There are a few developers who do review it. Hands up if you actually like still review code like properly. Hands up if you if you give it an honest shot. Yeah. But there there are many of you who still try and you still catch the bugs and you still push back and you still see that the agent has duplicated code or well the developer is the agent. You push back and they are being overwhelmed. They are being burnt out. They are being fed up. They are feeling that they're not rewarded. Oh and when it comes to performance review time, they're not going to be rewarded. They're not seen as the ones pushing out all the features. So some of them are burnt out and some of them just quit. Dax told me that at open code they are hiring a bunch of these people who are leaving their companies h because they're just burnt out being the sole person still keeping things alive and no one care. Engineering management is gone. They've either let them go or they're now now less hands-on. So there's no one left to care. Finally, I I talked with Kent Beck. he'll be this keynote peer tomorrow and he summarized this really well. Kent is amazing at summarizing findings. He said, "We're accumulating code faster than we accumulate trust." He said that with code you need to trust it. You need to understand it. We don't have time to do that right now. AI also amplifies software engineering experience. So seniors gain the most judgment is rewarded. And we see this everywhere. Hill Wayne, he'll be a speaker tomorrow. But he was telling me how some people are saying oh AI will help with formal verification with TLA a very complicated language. He'll show you a demo tomorrow. And he said the only people who have been successful with AI generating TLA plus specifications that work are TLA plus specification experts who in the prompt gave the exact specification of what kind of prompt to generate. Everyone else good luck with that. And this is true for software. If you're a junior engineer, if you've never built a mobile application, you can prompt a native iOS app, you can prompt the agent, it'll build something, but you know, it's not going to be maintainable. old patterns are seeming to coming back. Dax told me how domain driven design and verbals guardrails they're using this open code all the time because agents are the new junior engineers. You can start off a lot of them but these junior engineers I mean if you think of it like that they need a lot of guardrails and we used to these boring enterprise patterns used to become unpopular because they're long a long winded you have to explain you have to type out but they keep agents in check. So, it might be time to dust off some of these books and start to use design patterns again. I'm actually dead serious about this. So, this is where we are. it's it's just a lot of change, all all sorts of of things. It's confusing.职业生存:深耕业务领域与工程领袖的重归一线
在分化的宏观就业市场中,技术人员正面临关键的分水岭。虽然欧洲部分地区如德法两国的软件开发岗位呈萎缩态势,但美国和英国仍维持着 20% 的逆势增长。值得关注的是,AI 工程(涉及 RAG 体系搭建、模型工程评估和大模型集成)已占到全美整体软件工程岗位的 10%。在这样的结构性改变中,开发者的应对之策是积极转型为产品型工程师(Product-Minded Engineer: 能够自主深入理解商业诉求、跳出纯代码逻辑并与产品、用户直接对话的软件开发人员),去深入理解真实的商业链路,将精力分配给系统设计与业务洞察。
对于广大工程师及工程管理层,这场技术浪潮重塑了核心的职业生存法则:
- 构建非 AI 的壁垒:通过掌握软件维度的外延业务知识,建立难以被 AI 复制的护城河。如果你身处农业科技公司,就去田地里与农民交谈;如果你在汽车企业,就去和机械结构工程师讨论。
- 领袖重新编写代码:由于管理层扁平化大势所趋,研发管理者(Engineering Leaders)必须保持动手编程的技术手感(Stay Hands-on)。随着企业削减冗余的人员管理编制,单纯进行资源分配的管理岗正在被市场加速淘汰,开发者在未来也将不得不适应扁平组织下更少的调薪和晋升通道。
正如 Martin Fowler 和 Grady Booch 等业界泰斗所言,自上世纪 60 年代以来,软件研发行业从未经历过如此无序且极速的剧变。在不可名状的焦虑感中,开发者需要时不时给予自己积极的心理暗示,在喧嚣的 AI 狂热中回归理性,以对经典设计模式的坚守和对业务领域的深度钻研,在不确定性中筑牢自身的技术堡垒。