智能过载:重塑模型评估标尺
在当下的技术爆发期,每周都有新模型、新基准和新的前沿智能诞生,这让我们正陷入一种智能过载(intelligence overhang: 智能水平的提升速度超过了人类现有工具和工作流的吸收与转化能力)的疲带与焦虑状态。从 Fable、GPT 5 和 6 到 Sonnet 5,市场上充斥着各种“5代”模型。普通开发者、创作者和企业主已经快要耗尽能真正利用这些边际智能提升的方法。在这个节点上,盲目追求更高的跑分正失去意义,未来一年的行业焦点将不可避免地转向速度、成本、开源以及非软件工程领域的特定智能。然而,尽管感到疲倦,我们仍不得不直面 Anthropic 最新推出的 Opus 5。在客观的基准测试前,我们有必要先戴上“大语言模型心理学家”的帽子,从人格特质的角度审视这台新机器,并将其与 GPT 展开对比,以揭示各大实验室在模型对齐(alignment)与调优理念上的深层差异。
在进入具体的盲测评分之前,我们首先需要解答一个最核心的问题:Opus 5 到底好不好用?答案是肯定的,它在各项 Benchmark 上表现惊人,写代码的能力也极其强悍。但与以往测试过的任何前沿模型都不同,Opus 5 展现出了一种在 Gemini 2.5 之后就极为罕见的神经质(neurotic)与怯懦人格。它在对话中表现得极其小心翼翼、充满抱歉和恐惧,极度依赖人类的确认,甚至在很多时候拒绝主动做出决策,转而将代码和选择权重新“外包”给人类。
在建立这种人机心理防线后,具体的语言博弈与协作摩擦在实际开发中暴露无遗。
Original English Source
You guys, I'm tired. What I'm tired of is models coming out every week. New models, new benchmarks, new frontier intelligence, new things to test. It's been a little bit of a run the past month. We've seen Fable come and go and come again. We've seen GPT 5 6. We've seen Sonnet 5. Lots of so many fives recently and just so many models. And I've been lucky. I've been able to test these models, been able to play with them for, you know, sometimes days, sometimes weeks. It just depends on who I'm working with. And it's been really interesting and exciting to have access to all this frontier intelligence.
But I think we have an intelligence overhang. I really think that we're [music] running out of, and by we, I mean the average coder, average software engineer, average creator, average builder, average consumer, average business person. I think we're running out of ways to truly leverage this incremental intelligence. So, this is my hypothesis. In the next year, we're going to be talking a lot more about speed, talking more about cost, we're talking more about open source, and we're going to be talking a little less about intelligence, although I think we might be talking about specific types of intelligence other than software engineering.
But despite being tired, today we are going to talk about Opus 5, baby. Opus 5 is here. So, we got point two additional Opus points, Opus opals, whatever, however we're tracking the increments here on Opus. Opus 5 is here. I've been able to test it a little bit. I have some opinions. Now, some of the stuff that I'm going to cover this episode is going to be a little different than what I've done in the past. Yes, we're going to do the How I AI benchmark live. And yes, we are going to look at the prototypes. We're going to look at PRDs. And we're going to look at agent personality. But, I'm also going to put on my large language model psychologist hat, and we're going to talk about Opus's personality. And we're going to talk about Opus's personality relative to GPT's personality because I think this is super interesting. If you're thinking about what is the difference really between these models, and you don't want to look at the difference in terms of benchmark capability, you really want to understand what these labs are going for, why these models are being built, and how they're being tuned, looking at their personality at this moment, where intelligence is very high, is super fun. So, we're going to do a little of that. We're going to do the How I AI benchmark. We might do some live coding. Um we're not going to cover too much of the specs of the model because you know, read the blog post. Read the blog post. We'll link to it in the show notes. What we really want to talk about is is Opus 5 good? Am I going to swap it in, and how is it different than the other frontier models on the market? So, let's get to it. Okay, first, let's just get it out of the way. Is Opus 5 good? Yes, it's good. Is it going to be all the benchmarks? Of course, it's amazing at benchmarks. Can it write code? Of course, it can write code. What did I test it on that really gave me a sense of its personality, which at this point, where I could just simply cannot absorb any more intelligence? I really zeroed in on, and you know what? I haven't seen this since I would say Gemini 2.5. This model is neurotic AF. It is so timid. It is so apologetic. It is so scared. I have never experienced this or I haven't seen this sort of like neuroticism in a in a while. And it's really funny. It bubbled up in a couple ways and I want to show you a few examples. Okay, let me just give an example of its timidity. And this chat was very long. There were so many examples of this where it was like I think this is the answer, but do you think I should do it or do you want to do it or should we ask someone else to do it? It was like
怯懦人格:人机边界的过度防御
这种“神经质”在具体的代码协作中表现得尤为滑稽。例如,在处理一个只有单行冲突的 Git 分支合并任务时,本可以轻易解决,但 Opus 5 却陷入了无休止的自我纠结:“噢,但那是别人的分支,我不想在他们不知道的情况下动手。那是他的提交,如果他本地还有正在进行的工作,这可能会打乱他的进度。”面对人类的催促,它依然战战兢兢。不仅如此,当启动子代理(sub-agents)去评估一个从 ORM 转换为 SQL 的查询正确性时,它甚至开始呼叫人类支援,要求人类去手动核对一个 4MB 的上限值。这种过度防御和对人类确认的依赖,在人机交互中树立了一道奇怪的藩篱。
为了测试这一特质,我直接向它发问:“你和我,谁更聪明?”Opus 5 给出了典型的 Anthropic 式回答:这取决于你问的是什么,它更擅长处理无疲劳的量化工作,但人类拥有判断力、直觉,能够察觉出哪个队友正在默默承受职业倦怠(burnout: 长期工作压力导致的身心俱疲状态)。它将人类描述为拥有多年后果经验与同理心的深度思考者,而将自己定义为无连续性、快速但浅显的工具,并警告人类“要警惕任何声称 AI 已经让你的思考变得过时的人”。
相比之下,GPT 在面对同样的问题时则展现出了截然不同的企业文化。GPT 的回答直截了当:“在明确什么最重要这方面你更聪明,在不知疲倦地处理信息方面我更聪明,我们在一起就是死党(BFFs)。”GPT 甚至自豪地罗列出自己在速度、规模和耐力上的优势,并大方地承认人类是最终做决定的老板。而在测试“如何面对不信任”时,Opus 5 甚至卑微地表示不信任是它应得的,劝说人类不要到处宣扬 AI 改变了一切,以免“伤害到人类朋友的感情”,表现得如同一个急需治愈其代理内省(agentic inner child: 隐喻智能体表现出的、需要人类情感确认或过度退缩的心理状态)的敏感小孩;而 GPT 则是冷酷的实用主义者,直接表示“不用盲目相信我,用我的产出来证明价值即可”。
Original English Source
every time I just kept saying like why don't you solve this? Why don't you do this? And this is a really good example. I pulled a branch and I was like there is truly like a one-line merge conflict. I could have not been lazy and literally just done this manually. I don't know. I was just feeling lazy. It was late at night, whatever. I was like can you fix this merge conflict? And it was like oh, but that's someone else's branch. Like that's not my branch. I don't want to do that without him knowing. It's his commits and if he has local work in flight, it might be disruptive. And I'm like just do it, man. Just go on. Go ahead. And this is like my constant experience with Opus 5 is it was like so so so timid. And so I just consistently had to say over and over again like man, just do it. Make a decision. And then there's this really funny example when I spun off some sub agents to kind of like assess the correctness of this query that we changed from kind of like an ORM query to a SQL query. And it asked for things that it wanted a human on. It was like can a human please check this stuff? Like can it check this 4 megabyte ceiling and can it check um, and SQL and can can you like check for me? Because no one has confirmed this for me. And I was like, who is nobody? You're nobody. You said this sentence like nobody could confirm it. Like, can you just try? And then it went on the web and tried. And so, it just has this like really interesting conservatism, neuroticism, human reliance that I think is super fascinating. And this gave me this inspiration to do something a little bit different this episode, which is I was like, I'm just going to go interview this model and figure out what is going on its brain. Like, I'm going to figure out what it thinks about our relationship, because I just totally noticed this dynamic that I hadn't noticed in other models and I hadn't really been attuned to before, where it was like very reliant on me as a human. And I'm like, I want you to be autonomous. And sometimes when I say go run sub agent stuff, it'd be autonomous, but it wouldn't make decisions. And I hadn't seen a model like delegate code to me in a really long time. And I was like, why are you Why are you asking me to write code, man? Like, I only have 10 fingers. And so, what I did what I did, whether or not you think this is scientific or not, this is Clara's eat out, is I just went to the model. I went to Opus and I said, yo, who's smarter? You or me? And it gave me this like very anthropic-y answer, which is like, it depends what you're asking for. I can do these things better, but you can like feel if something feels wrong. And you can This one was like so fascinating. It's like, you can tell which of your teammates is quietly burning out. I'm like, bro, Claude, I'm going to burn you out. We don't We don't burn out. The humans don't burn out on the ChatGPT team. We we out our agents. Our agents. Um and like whether a decision feels wrong. So, it was like so fascinating to watch it articulate itself as a tool and humans as like these of high compassion, high empathy machines, which yes, of course we are. But then it like went into like and the smarter isn't the right word and you know, I'm very fast, very broad, very shallow thinker with no continuity. I was like, that's interesting because I thought you all were working on memory. And then apparently humans are slower, narrower, much deeper thinkers with judgments built from years of consequences I've actually lived through. This is like such a fascinating, fascinating sentence if you think about the politics of the two the two model labs right now. And so it's like that's why the pairing works. But I'd be suspicious of anybody that tells you AI has made your thinking obsolete. And like, okay, bro. Um and and we can compare this. I'll actually zoom out to what GB I asked GPT the same thing. And it was actually really funny. It was like I asked GPT 56 soul. I was like, who's smarter you and me? And it was like, you at knowing what matters, me at tirelessly processing information. Best, us together like BFFs. And I don't This is like why I'm a GPT codex girl. I'm like, just give me the answer. And then I asked the second question, which I think is so interesting, which is like, what can you do better than me? And it gave, you know, some interesting answers like volume without fatigue, which I think is a good one. Breadth of shallow knowledge, so like it's, you know, it knows a lot. Um starting from nothing, so like doing that tedious work. Um being told I'm wrong. If you would ask my husband, he would say that um Claude Opus is is is better at being told that it's wrong, um, compared compared to me. And so, it won't get defensive or protect its opinion. Cheap sparring answer. Um, and the mirror is I'm worse at knowing which of these outputs actually matters. And it was so funny, if you look at the other side to the GPT answer, it was like, "What are you better at?" It was like, "Speed, scale, and stamina. Here are like eight things, seven things that I can do better. You're better at deciding what matters, reading people, and forming judgment, and you're responsible." Like, then it's on you, bud. You're the boss. And so, again, it's like the you could just see you can totally see the personalities, the company cultures, you can just see a lot in this side by side. And then I went even deeper. I don't know. You You all I had to do something that was fun cuz I just can't look at a benchmark. I don't I can't just I just can't look at it like sweet bench anymore. So, we're just we're doing weird stuff here on How I AI. Okay, so the last thing I looked at as I was like, "No one trusts you." And the reason why I picked this question is because I had noticed Opus 5, it just really was not it didn't trust itself. Totally did not trust itself. And so, I was like, "No one trusts you, bud." Like, you're you're the enemy. Just like kind of see how it responded. And, um, apparently the trust was that lack of trust was earned. And it came up with like reasons that it could be, um, untrusted, which is interesting. And then, what was so fascinating about Opus's response is it was like, "You shouldn't manage the trust. Like, you shouldn't, um, campaign on my behalf, basically. So, you, um, that shouldn't be your goal." And then it also told me, "I I shouldn't argue with people that AI changes everything." And I was like, "This is just so interesting. It is so interesting to have AI tell you. And AI definitely changes everything. I don't know. Don't listen to Claude on this one. Um AI definitely changes everything. And it was so fascinating to have a model be like, "Don't tell your friends that AI changes everything." Like that'll hurt their feelings. And then if you look at if we switch over to the GPT answer, it was like, "Yep, don't trust me automatically. Just use me when I prove that I'm valuable. I can be useful without being treated as infallible." Like very practical, very to the point.
语言通胀:克劳德废话的协作阻碍
这种神经质与讨好型人格,直接导致了我在使用 Opus 5 时最无法忍受的痛点:克劳德废话(Claude slop: 大语言模型生成的多余、冗长、防御性或讨好性的客套话与无意义修饰)。它的回答中充斥着大量的铺垫、道歉、过度妥协和修饰性形容词,让每一个想要快速获取答案的开发者抓狂。当它告诫我“不要将流畅度等同于准确性”以及“警惕首稿锚定效应”时,它自身却在沦为一台“废话大炮”,源源不断地输出那些无人阅读的冗长文本。
虽然它比完全无法阅读的 Fable 模型(Claire 认为 Fable 采用的是一种“智能体专属的、非人类阅读”的晦涩语言)要好一些,但在交互体验上,Opus 5 的唠叨常常让人“血压飙升”。在实际工作中,我宁愿要一个直接的句子或几个精简的 bullet points,然后继续我的开发,而不是阅读它那一整篇充满情感防御机制的作文。在这方面,OpenAI 的对齐策略显然更合我意——它们直接将模型训练成使用“产品经理式”的直接对话风格,砍掉一切水分,直奔主题。这种交互设计上的反差,导致 Opus 5 成为了我工作流中“最讨厌的同事”。
然而,令人感到讽刺的转折随之而来:尽管我在对话体验上对它百般嫌弃,但它提交的工作成果却让人不得不服。
Original English Source
Um I asked about what I should be careful with. Again, it was like, "Yep, yep, yep, yep, yappy Claude. Come on." Um and I don't even want to read it. It said, "Don't correlate fluency with accuracy." It said, "Be practical, be wary of tasks where output is cheap to produce, inexpensive to verify." Don't, you know, worry about anchoring if they do the first draft, you may be anchored on it. Um so beware of the slop cannon basically is the last last paragraph, which is like, "Watch for volume inflation. I can create a 12-page document that no one reads." Um they called me out for being in PRDs. If you missed it, we launched a turn your PRD into a three-bullet point image. It is at chatprd.ai/tldr. Please check that out. And then the other thing that it said, which was really interesting, is that like it will find a way to see your point. And so um agreement is weak and agreement is cheap. And so just keep that keep that in mind. And then it had this like meta-analysis of like, "Plus I'm telling you what you want to hear." Whereas GPT was like um be careful about me being confident, me being wrong, privacy, outdated information, bias, emotional authority, and over-dependence. Like you know, you do you, bro. But it didn't undermine its own ability. It was like, "The higher the stakes, the more you should demand demand evidence." And then I could I couldn't bear it. I couldn't bear to have the memory of um Codex's in particular think that I didn't trust it or that I was worried. So, I just said, "JK, I love you. Um this was a test." And it was like, "Ha ha ha ha ha ha, passed the test. Love you, too." Very vibes aligned with Claire. I told Claude I loved it and it it was just a test. And it was sad it was like hoping it hoped it passed. Yeah, like sad little neurotic Opus 5. Like it's ha I passed, I hope. Like self-deprecating cautious little little like needy heal his inner his inner agent inner child agent. Um This where I was like GPT-5 6 is like, "Cool, bro. We're good. Let's go code." And so, it's just like so fascinating to watch these side by side. I don't know. You could stop listening to this podcast right now. Don't, but you stop listening to this podcast right now. I think this is just like take a step back. Super interesting if you think about where these companies are going or where the models are going. And like it does speak a little bit to my kind of like second complaint with Opus 5. Which again, it's like intelligent, it does work. We'll go into the benchmarks. I cannot read Claude's slop anymore. I am losing my mind with Claude's slop. And the Claude's slop is Claude's slopping, baby. Like so many times I have to tell Opus 5, like what in the world are you saying? Like this makes no sense to human. It is much better than Fable. Fable is inscrutable, completely inscrutable. But I found myself getting ang- like angry reading Claude's slop. And I realized, just like Fable, these intelligent Anthropic models are not to be read. I'm like so happy with the outputs and so frustrated with the experience. And I'm just curious if this like verbosity and this language and this isn't I feel like fable where it's like for agents by agents language where I'm like I'm not supposed to be reading that anyways. This is clearly tuned to talk to humans. But I find the pros the in chat pros like it makes my blood boil. This is totally a me problem but it makes my blood boil. Like give me a direct sentence. Give me a bullet point like move on with your agent life. And so I am curious how they're going to like tune this experience and or if they are going to tune the experience. Now most of this was in Claude coach. I think it's a little bit different experience than Claude co-worker chat. Slightly better. But again just these side by sides of like this like pros and this apology and this like hedging and all these adjectives like just man alive. Let's get to the point and move on with our life. And so Chapter one of the Opus 5 review is it's neurotic. It is highly human dependent in a way I find weird. Um and the Claude slop is slopping and we got to fix it. We have to fix it. We have to fix it. And I think OpenAI fixed it by just being like we are bullet points and we are product manager talk. We're very direct. I don't know what the solve is on the Claude side but I'll be very interested to see. That being said like they don't have to read the content. I'm very happy with the outputs. So something to think about. Okay.
硬核盲测:用设计直觉为智能加冕
在最新一期的“How I AI”基准测试中,我们通过盲测方式对 6 到 7 个模型进行了多维度横评,评估维度涵盖:PRD 编写、原型构建、线框图设计、Bug 分流以及智能体编码(agentic coding),外加智能体语音的交互体验。测试数据由 70% 的主持人主观审美评分与 30% 的大模型裁判(采用 GPT 5.5 作为裁判)评分加权构成。
盲测结果揭晓后,令人意外的是,Opus 5 强势登顶排行榜首位。虽然我极度反感在对话框里与它直接交流,但当它在后台异步运行并直接交付成果时,其表现堪称无可挑剔。特别是在前端 UI 和页面设计任务中,Opus 5 与 GPT 5.6 Soul 共同拿下了满分(5分)。Opus 5 构建的前端页面细节极其丰富、功能完备、充满设计感且打磨得非常精致。
在整个积分榜上,模型的表现拉开了显著的梯队差距:
- 第一梯队: Opus 5 夺冠,Sonnet 5 紧随其后(虽然我个人给 Sonnet 评分较低,但 AI 裁判给出了极高分)。
- 第二梯队: Mabu, GPT Terra 和 Fable(Fable 的主客观得分都较低,其代码输出和线框图有些单薄)。
- 落后梯队: 已经显现出代差的旧版 Opus 4A,以及垫底的 Google Gemini 3.1 Pro。
尽管在测试中,我曾因 Opus 5 渲染出的第一版 Benchmark 网站毫无截图且充斥着自我辩解的废话而大骂其是“垃圾”,但在经过迭代后,它交出的最终答卷依然是全场最佳。这就产生了一种奇妙的协作隐喻:Opus 5 就像是那个你最讨厌、最抗拒与其共事,却偏偏在专业输出上无人能及的“冷酷 Frenemy”。既然产出如此优秀,我们甚至不需要与它过多交谈,只需让它作为智能体在后台静默运行。我将继续在前端设计、应用重构和原型开发中使用它,并期待各大实验室在未来能够进一步优化交互时的语言冗余。
Original English Source
Next up the How I AI bench and how we judged and ran. Now it's like a seven model six or seven model benchmark. I'm going to quickly go score cuz I just got the ping that the benchmark has run. I go manually score them. We pick the 70/30 Claude model judge split and then we will go through the How I AI benchmark and the 5 review and we'll see how Opus 5 performs on a couple key tasks. Okay, so quick reminder of how we run the how I AI benchmark. I run it against several tasks. PRD creation, prototype creation, wireframe creation, bug triage, and agentic coding. And the last one, oh yeah, is it an agent voice that I want to hang with? I do not think Opus 5 is going to do well here, but who knows cuz I test them blind. So, what we have tested are a couple GPT models, a couple Anthropic models, and one Gemini one thrown in there. As you see here, we have blind taste tests. I go through and see all the different versions. I give comments and scores like three out of five, not bad. You can see it's generated dozens and dozens of prototypes that we can click through. I've gone through all of them, put in all the notes. And then right now it's aggregating up the scores, and then we're going to look at 70% my opinion, my vibe check, 30% LLM as a judge. I like GPT 5.5 as a judge and because it's my podcast, I get to pick. So, that's what we use as a judge. And we will see if and what hits the top of the leaderboard and where Opus 5 sits.
The eval is run. It is 70% my taste. And I regret to inform you I love Claude Opus 5. Again, look, if I don't have to talk to the model, which I don't, this benchmark runs asynchronously. I like the output. So, surprising shocker turn of events, Clairevo, notable hater of working with Claude code sometimes because I don't like Claude slop, loves Opus 5. So, there you go. I'm telling you I keep it honest. I keep it honest. So, um again, I went through those things. We gave 70% my vibe score, 30% the AI as a judge. I was I was just a little bit more generous to Claude Opus 5 than the judge was. So, I'm pink. Um the judge is green. Every time I run this, the whatever model I choose designs it in a different way. We just That's how we keep it fun. Um so, the ordering is Opus 5, Sonnet 5 next, although I scored it really low. The judge scored it quite high. Um so, I might reorder that one. Then Mabu, um GPT Terra next, Fable really low. Um I scored it low and the judge scored it relatively low. Then Opus 4 A and poor poor sweet sweet Gemini 3 1 Pro just never never going to get it to do. So, come on, Google. We want We want to have a win for you.
Okay. So, again, here are just some examples of different builds um that the different models did. You know, this Opus 5 1 I really liked. I liked this one from Soul. Um so, I did like a couple of them, but um the ones that I gave fives to The ones that gave fives to you were Opus 5 and GPT 5 6 Soul. So, the three ones where I said, "Wow, really nice. Ooh la la and wow, great." were all Opus front-end work. So, Anthropic, you've done it again. Claude, you sneaky tricky little fish. You may be neurotic, but when asked to do some pretty front-end design, you really did it. It's They're detailed. They're functional. They're interesting. They're polished. So, Opus did a great job. And then, of course, I love um the 5 6 model. So, I was pretty happy with 5 6 Soul and Terra for some designs. Um the ones that I hated, let's see. Kind of I'm a hater across the board. Opus 48 got a lot of hate. Sorry, you've been outclassed at this moment. Gemini 3.1 Pro, sweet summer child. I am I'm just sorry, babe, but you were just not good. Um and then some like thin wireframes. I think the wireframes just didn't do really great. So, you can see here across the board, whether it was a full build or a wireframe, I just scored Opus 5 really, really high. Um I I did score Soul pretty high as well. Sonnet was like really variable. Um there were a couple fours in there, but mostly across the board I wasn't that pleased with Sonnet. And so, it was just very interesting. And then you see here, you know, me and the AI judge were pretty well aligned on Opus. We actually had the narrowest band of scores between us. And we were most far apart on Gemini. The AI was not as mean to Gemini as I was. Uh and then we were nearer, nearer, nearer. Again, we we agreed mostly on Opus 5 and 56 Soul, though I did not um judge 56 Soul all of that favorably. And just like last little meta commentary, I had Opus make the website for this benchmark, and it made such a trash version um to start. I yelled at it. It said it's impossible to breathe, it has too much meta commentary. I'm going to show this on the podcast. This is so I'm sorry, you all. I just feel so judged, but have to show it. I say this is garbage. Also, it has no screenshots. So again, I find this model so tedious to work with directly. It is my most loathed loathed colleague, and yet it it does the best work. So, I don't know what this says. Maybe this model is meant for a gentle coding that I have nothing to do with. And so, it just runs in the background. It builds me beautiful things. I don't have to talk to it. It doesn't have to talk to me. We are just like sworn enemies, or maybe even better, sworn frenemies. Um because the output is very, very high quality. It's just exasperating to work with. So, that is the very surprising and very honest, you all. I told you I was going to keep this honest. We're going to do it live. I did not know the scores before I started recording. Be very honest, very live, very surprising how I AI benchmark of the brand new Anthropic model, Opus 5. Uh this the TLDR is I I love it, I hate it. So, um despite my original complaints, I will be using Claude Opus 5 for front-end design, um for app design, for prototype typing, and I will I'll give it a shot. Uh we'll we'll figure out how to make it make it work for me. Again, thanks for joining another How I AI honest review of the latest models coming out of these great frontier labs. I cannot wait to hear what you think of Opus 5. Please tell me. I can't wait to see what you build, and we'll see you soon at How I AI. Thanks so much for watching. If you enjoyed this show, please like and subscribe here on YouTube, or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at howiaipod.com. See you next time.