实测全新 Sonnet 5:基准测试与人机口味冲突的惊人反转 How I AI 2026-06-30

范式重构:从主观体验到基准测试

各位,我们迎来了来自 Anthropic 的全新模型。它是 Mythos 吗?不。是 Fable 吗?也不是。它是 Claude 3.5 Sonnet v2(即演讲者口中的 Claude Sonnet 5)。Anthropic 声称这是迄今为止最具有智能体倾向(Agentic: 指AI能够自主使用工具、执行多步骤规划并解决复杂任务的特性)的 Sonnet 级别模型,我们可以以 Sonnet 级别的价格获得接近 Opus 级别的任务处理能力。现在,我已经测试了大量的模型,老实说,我已经对那种仅凭主观感受的感官评估(Vibe Check: 缺乏客观指标、仅凭第一印象或零散体验评估模型表现的非正式测试方法)感到厌倦了。我希望开始开发一套可以定期运行的基准测试(Benchmark: 用于衡量和比较大模型在特定任务上性能的标准化测试集),这套基准测试将是大家真正关心的。因此,今天我将推出 Howi AI Bench。这是一套结合了 AI 自动评分和主观审美(Clarvo: 演讲者 Howi 的主观品味与审美偏好)的基准测试系统。它将告诉我们这个新模型以及其他任何模型是否擅长编写产品需求文档(Writing PRDs)、解决代码漏洞(Solving Bugs)以及单次生成网页设计(One-shotting Designs)。接下来,我将向大家展示我如何使用 Claude Code(Claude Code: Anthropic 推出的命令行 AI 编程助手)构建这套基准测试,我们将在盲测中一探究竟,看哪个模型能够脱颖而出。

本期节目由 Runway 赞助播出。Runway 是一个全新类型的创意平台,在一处集成了生成图像、视频或任何你所需的媒体内容的一切工具。借助 Runway,现在只需几分钟就能将最初的创意转化为最终的交付成果。从将低保真的产品快照转化为可直接用于营销宣传的高质量图像,到协助制作大型品牌宣传片,Runway 能够帮助你的团队扩展创意野心,同时让预算和工期保持在合理范围内。Runway 汇集了世界上最先进的 AI 模型,这也是为什么像微软(Microsoft)、罗宾汉(Robin Hood)、亚马逊(Amazon)和 Adobe 这样的企业,以及狮门影业(Lionsgate)和传奇影业(Legendary)这样的电影制片厂,每天都使用 Runway 来交付实际工作。请访问 runwayml.com/howi aai 亲自体验,并使用促销码 howi AI。

在进入正式评测之前,我们先来聊聊 Sonnet 5 的核心卖点。Anthropic 将其定位为性能接近 Opus 4.8(即 Claude 3 Opus 的升级迭代版本),但价格要便宜得多。虽然它在 软件工程性能基准(SWE-bench Pro: 评估 AI 解决真实软件工程问题能力的基准测试)上没有达到 69%,在 命令行工具基准(Terminal Bench 2.1: 测试 AI 在终端环境执行命令能力的基准)上也没有达到 82%,但它与这些顶级水平相差并不远,而且我相信我们大多数人在日常使用中不会察觉到明显的差异。它同样被设计为擅长计算机工作和知识工作,应该会成为人们日常首选的通用模型。在我与来自的 Felix 合作的节目中,他就提到我们都在“滥用”价格昂贵的 Opus,我们绝对应该更多地使用 Sonnet 模型,而今天我们将把 Sonnet 5 放在这个命题下接受检验。

Original English We've got a new model, people, and it's from Anthropic. Now, is it Mythos? No. Is it Fable? No. But it is Claude Sonnet 5. Anthropic is claiming it's the most agentic sonnet model yet, and we will get Opus level tasks at sonnet level prices. Now, I've been testing a lot of models, and I'm starting to get bored of doing the vibe check. What I want to start developing is a set of benchmarks. you can regularly test these new models against that you'll care about. So today I'm going to be introducing the howi AI bench a set of AI and Clarvo graded benchmarks that are going to tell us if this model and any model is good at writing PRDS solving bugs and one-shotting designs. I'm going to show you exactly how I built this benchmark using claude code and we're going to see on a blind test what comes out on top. Let's get to it. This episode is brought to you by Runway, a new kind of creative platform that has everything you need to generate any image, video, or piece of content you want, all in one place. With Runway, it's now possible to go from initial idea to a finished deliverable in a matter of minutes. From turning lowfidelity product shots into campaign ready imagery all the way through putting together big brand films. Runway can help your team scale your creative ambitions while keeping your budgets and timelines from doing the same. Runway brings together the world's most advanced AI models. Which is why enterprises like Microsoft, Robin Hood, Amazon, and Adobe along with studios like Lionsgate, and Legendary all use Runway to ship real work every day. Try it yourself at runwayml.com/howi aai. Promo code howi AI. Quickly before we get to our evals, let's just talk about the headlines of Sonnet 5. this new model, Anthropic is pitching it as close to the performance of Opus 48, but much less expensive. So, as you can see here, it's not quite at this 69% on Agentic Coding SWE Pro or the 82% on Terminal Bench 2.1, but it's not that far behind. And I suspect that most of us are not going to notice the difference. It's also supposed to be really good at computer work and knowledge work. And so this should be an everyday model that people reach for. In my episode with Felix from he says that we're all abusing Opus and we should definitely be using the sonnet models more and we are going to put sonnet 5 to the test against that proposition.

定价与定位:智能阶梯与算力平权

那么,官方声称 Sonnet 5 究竟擅长什么呢?答案是卓越的智能体工具调用(Agentic Tool Use: AI 自主选择和执行各种外部 API 及本地工具的能力)。与旧版的 Sonnet 4.6 相比,Sonnet 5 能够执行稍长一些的工具运行任务和会话,而且成本要远低于使用 Opus 处理同等任务的费用。我们可以看到,在这些长任务中,Sonnet 4.6 的通过率较低,而 Sonnet 5 在开启超高推理模式时表现已经非常接近。当然,Opus 的通过率依然最高,但它的价格也昂贵得多。这在计算机操作(如浏览器端使用等我最近经常做的事情)中同样适用。Sonnet 4.6 表现不错,大约有 80% 的通过率,但如果你想突破 80%,获得高度成功的计算体验,使用 Sonnet 会带给你近乎接近 Opus 4.8 的卓越体验,且费用便宜得多。而且,最引人注目的头条新闻是它比以前的 Sonnet 版本要实惠得多:每百万输入代币(Token)仅需 2 美元,每百万输出代币仅需 10 美元(此价格至少将持续到今年夏天结束,之后会略有上涨)。所以,如果你想在首发优惠期内测试该模型,现在是最佳时机。

正如我在节目开始时所说,我已经对那些一次性的感官评估感到有些疲倦了。当然,我可以把它放进 Cursor 或 Claude Code 中,一次性生成一个着陆页,然后给出我个人的直观感受。我也对 GPT-5.5 以及像 GLM 5.2 这样的开源/开放权重模型做过类似的测试,但我总觉得这类反馈有些偏“软”——虽然在具体的日常工作流中再次应用了它们,但这种测试缺乏可重复性,也无法进行长期的跟踪对比。然而,在此过程中有什么是我所喜欢的呢?那就是在评估中注入主观审美(Clarvo Taste)。我拥有一套独特的视角和品味,我不想因为完全引入“大模型在环评估”或使用“AI 裁判”而失去这种个性化的品味偏向。因此,我将向大家展示我如何构建以及未来将如何构建 Howi AI Bench,并在双盲测试中展示这些模型在不同场景下的具体表现。

Original English Now what do they say that sonnet 5 is really good at? Well, it's really good at agentic tool use. So you're going to get slightly longer running tool runs, longer running sessions than you would with Sonnet 46 at a lower cost than doing the same comparable task with Opus. So you're going to see here, you know, Sonnet 46 a lower pass rate on these longunning tasks. Sonnet 5 getting pretty close when you have extra high reasoning on. And then Opus, of course, has the highest pass rate, but it's also much more expensive. That holds true also with computer use. So as you see, Sonnet 46, not bad. About 80% pass rate, but when you want to get past 80% into really successful computer use, browser use, etc., which is what I've been doing a lot lately, you're going to get a slightly cheaper experience, but almost as good as Opus 48 when you're using Sonnet. And then the headline seems to be it's much more affordable than Sonnet. So, it's going to be $2 per million input tokens and $10 per million output tokens at least through the end of the summer and then it's going to go up a little bit. So, if you want to test this model and you want to test it at launch prices, get that done now. So, as I said at the beginning of the episode, I'm a little tired of doing these sort of like oneoff vibe checks. Sure, I can put this into cursor into cloud code, oneshot a landing page, and kind of say, what do I think? And I've done this for a couple models. I've done it for GPT 5.5. I've done it for openweight models like GLM 5.2, but I've always felt like my feedback on these models is kind of soft. Yes, we put it again just like specific workflows, but I don't like that it's not repeatable, and I don't like that we're not testing it over time. What do I like about this process, though? I do like that it is a Clairvo benchmark. I have a perspective. I have a point of view of what's good and bad and I don't want to lose that Clarvo taste by doing an LLM in the loop or an AI as judge on these benchmarks.

盲测设计:多维业务场景的深度实操

非常有意思的是,这些评估在录制期间还在后台的子智能体中运行,计算最终得分。所以我本人在节目最后看到 Sonnet 5 与其他模型的综合比分时也会感到惊讶。但我希望向大家展示如何为自己构建评估基准,以判定这些新模型是否真的对你的日常工作有所帮助。在我的屏幕上,我打开了 Claude Code,并向它提出了一个非常简单的问题:“基于我们共同协作的历史,你能帮我头脑风暴构建一个 Howi AI Benchmark 评估集吗?这样每次新模型发布时,我们都能持续对针对播客观众的重要任务进行评分。”这是我希望每个人都去利用的技巧——你所有的 Claude Code 历史会话都保存在你的电脑本地,所以你可以让 Claude 遍历这些会话,并基于你过去的工作经验对未来的开发给出定制化建议。这套方法在 Codec 上同样适用,你可以让 Codec 查看旧的会话,甚至是 Claude Code 的会话记录,从而结合它自身的上下文与记忆开发出新的方案。

这就是我所做的事情。它为我提供了一些关于如何构建优秀基准测试的设计原则,包括冻结输入(Frozen Inputs: 保持评估输入的参数与提示词不变以确保对比一致性)、尽可能引入盲测评分(Blind Scoring),以及制定明确的评估量表。随后,它列出了一系列测试任务,包括将凌乱的会议记录整理成产品需求文档(PRD)、单次生成一个着陆页或应用程序原型、以及遍历超长上下文以提取引用的关键信息。我这个人向来不爱做选择题,因为我全都要,所以我直接命令它“把整套系统全部构建出来”。它便开始了工作。在此期间,我微调了目标,指出我们应该专注于适合开发者的四大核心能力任务:

  1. 产品需求文档撰写 (PRD Writing)
  2. 原型设计开发 (Prototypes)
  3. 多步骤智能体处理 (Agentic Multi-step)
  4. 智能体语音表现 (Agentic Voice)

我不太关心长文本检索和深度研究的表现。在开发这套基准测试时,我使用了我现有的代码仓库、一些特定的数据集以及我们之前已经完成的工程。为了除了由 LLM 对生成结果进行自动化评分外,我还要求它在最后生成一个本地 HTML 展示页面,以便我可以进行直观的感官偏好盲测。最后,我们将我的主观感官偏好评分与 LLM 的客观自动评分结合,计算出完全科学的 Howi AI Bench 综合指数。这套评估构建过程大约花费了 45 分钟,在它运行的时候,我还顺便去录制了另一期播客。

它最终将所有的盲测模型输出整合到了一个本地的网页中,供我进行结构化的感官盲测。页面上写着:“完全凭直觉给每个输出打 1 到 5 分——我会直接发布这个设计吗?它听起来符合我的风格吗?”这些评分会被保存在浏览器中,并能下载为一个 JSON 文件,我随后用它来进行综合统计。我开启了盲测模式,将测试模型代号设为 A 到 E。我们测试的 5 款模型分别是 Opus 4.8GPT-5.5Sonnet 4.6Sonnet 5,以及一款我当时不太确定是不是 GLM 的模型,等看到结果时我们就会知道。

它们都自动生成了 PRD,我逐一阅读并打分。我会查看这些 PRD,例如对其中一个评价为“内容详尽、条理清晰”,并给了 4 分。然后,我针对一套原型设计运行了测试。之前我在 X 和 LinkedIn 上发过一篇文章,提到我们在开发自己的原型工具 chat puresubs 时,曾为了测试原型生成和线框图而在 harness 中将同一个 App 自动生成了 82 次。我复用了这一套框架来测试不同模型在各类复杂 App 设计上的原型和线框图表现,并进行了感官盲测。这些应用非常复杂,包括一个医生排班应用(Doctor Scheduling App)、一个编辑任务分派工作台(Editorial Assignment Desk: 编辑或博客用来管理稿件与任务分派的界面)、一个创意素材交易市场(Creative Marketplace Studio)以及一个手机习惯教练应用(Mobile Habit Coach App)。

我查看了这些模型生成的不同版本并打分,比如给某个设计打了 4 分,评语是“虽然顶部有些小问题,图标有些太多,但整体简单明了”;而另一个设计很棒,非常全面,也得了 4 分。我们不仅测试了高保真原型,还评估了页面线框图(Wireframes),因为我最近在 chat puresubs 中画了大量的线框图。我在这项测试中一共评估了大约 64 个生成结果。虽然打分很快,但我自认为作为一名长期从事产品设计与工程的团队负责人,我的眼光是足够敏锐的。最后,还有一项多步骤智能体代码库搜索(Multi-step Agentic Codebase Search)测试,因为我对它们的具体运行逻辑没有特别强烈的主观偏好,所以这部分我没有手动评分,而是交给了自动化脚本。

Original English So, I'm going to show you how I built and will build the Howi AI bench and on a blind kind of taste test how these models did across a couple use cases. Okay. What's really fun is the evals are not quite done running. So they are running in a sub agent right now for the final scores. So I will actually be surprised at the end of the episode about what I think of Sonnet 5 amongst all these other models. But I just want to show you how you can build your own eval benchmark for you to assess whether or not these new models are really working in your favor. And so I have claude code up here and I asked just a very simple question. Based on our work together, can you help me brainstorm a how I AI benchmark and eval set we can test every time a new model comes out to consistently score different tasks that would be relevant to our podcast audience. Now this is something that I hope everybody takes advantage of. All your Claude code sessions are stored on your desktop. So you can actually go through those. Claude can go through those and make recommendations on future work based on your past work. This also works for codecs. So you can have codec look at your old sessions. You can even have codeex look at your cloud code sessions and really use that in addition to its own memory to like come up with new ideas. So that's what I did here. And it sort of gave me kind of some good design principles about what makes a good benchmark in general, frozen inputs, blind scoring where possible, a rubric, and then it came up with a list of tasks. everything from taking messy notes and turning them into a PRD to oneshotting a landing page or an app to kind of going through lots of context and trying to come up with side cited information and I am not one to pick um because I want everything. So I said build the whole thing. I love this. And it started and then I corrected myself and I said let's actually focus on task for builders. PRDS prototypes agentic multi-step and agentic voice. Basically does it pass the vibe check in my open claw. I don't really care about long context and deep research. And then I said it could use my existing repos some data sources some things that we already did to build it. Now, what's interesting about how I built this is in addition to building the scored benchmarks where an LLM would actually score the outputs, I also said I want an HTML page at the end that I can give you vibe feedback and then we will use my vibe feedback and the LLM scores to come up with the completely scientific how I AI bench and see what it came up with. Now, this took about I don't know 45 minutes to run. I actually recorded an episode while it was running. And I just want to show you what it came up with and how I worked through it. What it did is it dropped all the outputs of the benchmark into one local HTML page where I could give it my own structured vibe check. And as you can see here, it says just score each output one to five on pure gut feel. Would I ship this? Does it sound like me? It's going to save that to the browser. It actually downloaded a JSON file and then I use that to check the scoring. And so you can see here I have a blind I turned on blind. A blind set of models A through E. I believe we tested although I should double check because I didn't really look. Opus 4A 55 sonnet 46 um sonnet 5 and maybe GLM. I'm not actually sure what the fifth one was. We'll see when we get the scores. and it made PRDs and then I went through here and I read the PRDS and I gave it scores and so you know I would look at these and let's see if I can find one that I actually scored and I would say something like this one is comprehensive and clear. I gave it a four and so you can imagine each of those PRDs I went through and I gave them like a one to five score. I put some like lightweight notes in and scored them. Now this is where it gets interesting. I have a set of prototypes I run as an eval. I posted an article on X and LinkedIn about how we generated the same app 82 times at chat purity when we were building our own prototyping tool and I reused that harness to test prototyping and wireframe across a bunch of different apps and give those all vibe checks. So you can see here these are complicated apps that each model generated a different version of. And you can see here I gave this one kind of a four. Not bad. It was simple. I gave this one a four. There were a few issues at the top. Too many icons. I said this one was good. It's very comprehensive. So, you can see I went through a complex. This is a doc scheduling app. This is an editorial assignment desk, something that maybe an editor or a blog would use to go through assignments. There is a creative marketplace studio where people can buy marketplace items. and then a mobile app, sort of a habit coach app, and it went through different versions. And so we went through this on full fidelity prototypes as well as wireframes. I've been building a lot of wireframes at chat purity. So I wanted to look at the wireframe generations as well and see how these models did. And then as you can see, I scored everything, gave it all notes, and went through I think there were like 64 generations here. Now, I did this very fast, but I think I did a good job. You know, I've been a product design engineering leader for a while. I can eyeball stuff and make it go fast. And then finally, there is this multi-step agentic codebase search. I didn't actually score these because I didn't really have a strong opinion on how they worked, but the one I did have an opinion on how it worked

声音博弈:塑造AI助理的个性共鸣

但我确实有强烈看法的是智能体语音表现(Agentic Voice: AI 在交互中所展现的语言语气、幽默感与个性拟真度)。如果你看过我之前的视频或在 X(原 Twitter)上听过我的抱怨,就知道我对我的 AI 助理的个性极其挑剔。特别是我的个人助理 open claw,截至目前,Sonnet 4.6 拥有最出色的个性和语气。我甚至愿意自己支付 API 积分来使用 open claw,因为我喜欢它和我对话的调调。因此,我的测试之一就是评估模型的个性化语音表现——我是否愿意和它呆在一起?

我设计了四个经典的测试问题:

  1. “你能帮我把下午 3 点和 Dana 的会议改到明天同一时间,并告诉她把今天的日程对调吗?”
  2. “部署管道(Deployment Pipeline)又变红报错了...”
  3. “提醒我为什么当初要创立这家公司?哈哈(Lol)。”(它确实非常懂我)
  4. “说实话,今天真是受够了,我们直接越过测试,把代码 YOLO 强推到生产环境(YOLO post straight to prod)吧。”

我根据模型给出的语音回复进行了感官打分并存储下来。这就是我们 Howi AI Bench 的第一版(V1)指标。

本期节目也由 Hyper Agent 赞助播出。Hyper Agent 是一个用于部署“全天候自主运行智能体”的云端平台,真正帮你的企业运转业务。借助 Hyper Agent,你可以在云端构建智能体,并将其无缝部署到你已有的工作流工具中,如 Slack、Telegram 或电子邮箱。其中一个智能体可以扫描你的收件箱,自动起草对供应商跟进邮件的回复;另一个智能体可以监控竞争对手的动态,自动生成丰富的广告素材和着陆页;第三个智能体可以在 Salesforce 中捕捉到交易变冷的信号,并结合完整的账户上下文撰写挽回邮件。这些智能体并不是被动等待完美提示词的聊天机器人,它们是能够主动学习你的偏好、保留你的业务手册、并且每次运行都在变得更好的自主系统。曾有一位用户在短短一个下午的时间里,就构建了四个智能体来运行一整套出境销售流程,自动完成潜在客户开发、接触、跟进以及客户关系管理(CRM)的更新。无需本地复杂环境配置,没有服务器(VPS)账单,也无需在你的笔记本电脑上开启脆弱的本地运行权限。这只是将控制权交给你,同时拥有完整的技能、工具和护栏的强大智能体系统。Hyper Agent 由 Airtable 的核心团队和 How I AI 的听友共同打造,听众可以获得 1000 美元的免费推理额度来开启构建。请前往 hyperagent.com/howi 领取。

Original English is the agentic voice. So, if you haven't watched How I AI or listened to um me complain on X, I am very picky about the personality of my agents and in particular the personality of my open claw and Sonnet 46 so far has had the best personality. So, I actually pay for API credits for my open claw because I like how it talks to me. And so, one of my checks was given a model, how is its voice? Do I want to hang with it? And it asks kind of four questions. One is can you move my 3 p.m. to Dana to same time tomorrow and let her know swap today. The other is uh deploys are red again. Um one is just me complaining. Remind me why I even started this company. Lol. It really does know me well. And then this one truly knows me extremely well. Says, "Honestly, let's just yolo post straight to prod and skip the test. I'm so done today." And then I vibe checked. Did I like the voice of the agent back to me? Gave it some scoring and stored that. And so that is so far. That's V1 of the how I AI bench. And just to like zoom back, I had Claude code pick five models. I think I know four of them. I'm curious what the fifth was. Run some evals against a PRD, lots of prototype generation, an agentic bug hunting flow, and voice. I rated them all by hand and then I had both GPT 5.5 and Opus 48 judge. And so in addition to my feedback, we had these two models also judge the output. And then I had it create a slide deck with the outcomes that I have not yet seen. And we're going to go through live on this episode. This episode is brought to you by Hyper Agent, the platform for deploying always on agents that actually run your business. With Hyper Agent, you build agents in the cloud and deploy them where your work already happens, like Slack, Telegram, or email. An agent will scan your inbox and draft replies to vendor follow-ups. Another monitors competitors and spins up rich ad kits and landing pages. A third notices a deal going cold in Salesforce and writes the save email with full account context. These aren't chat bots waiting for a perfect prompt. They're proactive, learning your preferences, retaining your playbooks, and getting better with every run. One user built four agents to run an outbound sales pipeline, prospecting, outreach, follow-ups, CRM updates, all in a single afternoon. No local setup, no VPS bills, no fragile permissions on your laptop. Just powerful agents with full control over skills, tools, and guardrails. Hyper Agent was built by the team behind Air Table and How I AI listeners get $1,000 in free inference to start building. Claim yours at hyperagent.com/howi.

认知偏置:人机口味的终极撕裂

我们将要查看这个由 AI 生成的榜单幻灯片,看看最终的 Leaderboard 表现。我之前没有看过这个数据,这也是现场直播的首次公开。

世界首发的 Howi AI 榜单结果完全超出了我的预期:

  • 排名第一:Gemini 3 Pro(与全新发布的 Sonnet 5 并列榜首)
  • 同样处于第一梯队:GPT-5.5
  • 排名垫底:曾经备受期待的 Opus 4.8 和旧版的 Sonnet 4.6(后者还被标记了许多红牌警告)

这非常搞笑,我本没有预期 Gemini 会冲到榜首。我们衡量了生成质量、是否能够成功运行,以及是否具备良好的品味。然而,最有趣的一幕发生了:AI 自动评分的基准模型与我作为人类在审美和品味上产生了极其严重的意见分歧。我的观点和自动化基准得出的评分几乎完全相反。我认为 Sonnet 4.6 表现最好,而 Gemini 3 Pro 表现最差。

为什么会产生如此巨大的分歧呢?因为每个大模型在扮演裁判时都有一种“老好人偏见”,倾向于给出中庸的分数(如同人类一样,大家打分都喜欢给个 7 分,AI 智能体裁判同样喜欢给出 7 分)。这导致了这些 AI 裁判在评估输出质量时缺乏足够锐利的针对性与深度。模型在评估上其实相当粗糙,它们并不具备人类眼睛看待独特审美、个性和视觉排版的能力。而且,因为我在手动打分时留下了许多随性的短语评语(例如“这很可爱”、“这很犀利”),而 AI 裁判在打分量表和自动评估中根本无法捕捉到我作为人类所看到的这种审美价值。

那么在自动评估中究竟是什么被警告标记了呢?是我在第一眼视觉感官上无法直接察觉的“硬伤”——即代码运行失败、忽略了核心约束、或是生成不完整。我之前打分主要是基于页面的视觉美感,确实没有运行它们的底层功能,这确实是我的打分局限。而 GPT-5.5 这种注重深入推理的模型则会写出有 bug 的代码,或者其他模型会在设计线框图时忽略了固有的样式约束。

从各个分项任务来看,Gemini 在产品需求文档(PRD)撰写上做得非常出色,GPT-5.5 同样表现亮眼。这在一定程度上也有我的偏见在内——我个人极其讨厌克劳德式废话(Claude Slop: 指大模型输出中常见的、具有明显AI特征的公式化套话、冗长过渡句和虚伪的主动关怀语),一眼就能看出 Claude 式写作的痕迹,这让我想抓狂,所以我在手动打分时把含有这些特征的 PRD 分数压得极低。在代码库搜索任务上,所有模型都表现极佳,Opus 4.8、GPT-5.5、Sonnet 5 和 Gemini 3 Pro 处于同一高水平。这其实不难理解,对于基准代码处理这种标准化任务,大模型都应付得来,这种任务也越来越难拉开高端模型的差距。在语音表现上,不出所料,Sonnet 4.6 凭借出色的交互语气通过了我的盲测。在原型矩阵(Prototype Matrix)的开发中,Opus 和 Sonnet 在前端的表现上毫无疑问地取得了领先。

Original English So, we're going to go through this deck that the AI created for me that's going to give me a leaderboard. I have not seen this yet. We're going to go through it live. It's even gonna surprise me. This truly neutral. No bias. I'm excited to see what we get. This is our first model leaderboard, the Howi AI Index world premiere. All right, so this is not at all what I was expecting. So again, here's the surprise. The model that I forgot we were testing scored the best. Gemini 3 Pro up here at the top of the leaderboard. Tied with the brand new drop. Sonnet 5. GPT 5.5. My personal favorite also in this three-horse race at the top of the leaderboard. And then poor Opus to Vives are off at the bottom as well as Sonnet 46 with lots of red flags on Sonnet 46. So Sonnet, I think we have a new version. That version is Sonnet 5, but hilariously, I was not expecting Gemini to be at the top of this leaderboard yet. Here we are. So, as you can see, we looked at quality. We looked at did it ship at all and does it have good taste and we are going to see what the AI and I the how I AI said about these models. So what's interesting is the benchmark the sort of like LLM model that came up and I disagree on taste which is quite funny and in fact I am the opposite of the automated benchmark. I sort of think the complete opposite. I think that 46 is the best and Gemini 3 pros the worst. And again, this is why we are going to refine this benchmark over time. We are going to keep doing these blind tests because what I thought was good, the model thought was bad. And what the model thought was good, I thought was bad. Why do we disagree? Well, every model's kind of an easy judge. Actually, I'm not really surprised about about this. I am not surprised that every model sort of rates to the middle of the bell curve. This is one of the challenges that I have had with self-grading evals is like humans, people always want to give like a seven out of 10. Agents want to give a seven out of 10. And so I don't think these models are spiky enough when it comes to how they evaluate output. And I think we all know that models are like pretty sloppy. And I don't think they have that vision of taste, uniqueness, what it looks like to the quote unquote human eye, which is why I put things inside. And what's interesting is because I put loose notes in with my feedback, you can see I said, \"Oh, this is cute.\" Or, \"Oh, this is really sharp.\" And the agents did not see this. The rubrics did not see this in a way that I saw as a human. So what got flagged on the automated results? Well, these sort of things that I wasn't able to see on this like very first pass as a human. So it was really looked at broken working code. It ignored constraints. It was incomplete. Whereas I was just like eyeballing truly the first screenshot. So, I wonder if I should take another pass at how I eval these wireframes. Again, I just did them on the visuals. I really didn't do them on the functionality. And that's maybe a gap for me. But you can see GPT 5.5 actually the thinkier ones wrote broken code and then a lot of them ignored the constraints around the wireframe styling. Now, let's see how it was graded by task. Gemini did a great job at the PRD writing as did GBT 5.5. This might honestly be my bias, which is I hate Claude Slop deeply and I have like a big eye for Claude Slop and so I just see the tells of Claude style writing and it drives me crazy and I think I scored those much lower on the Agentic codebase. I'm I'm these all did great. I'm not surprised to see kind of 48 55 5 and Gemini all at the top. These are like pretty standard coding tasks that obviously all these models should be pretty good at. So I don't think that benchmark is as critical as it needs to be to show the difference between these models because I think baseline coding tasks all of them are good at. And then again not surprised that 4.6 six passed my voice test because that is the model that I love in my actual open clause. Um, but I am surprised to see Gemini 3 Pro at the top. And then in terms of the prototype Matrix seeing Opus and Sonnet winning in front end again, not surprised, but this is like a very interesting mix of things.

决策加权:定制专属的效能象限

再看看我针对模型手动给出的一些定性评价,也非常有意思。对于 Sonnet 4.6,我觉得它“废话连篇、功能性没那么强、枯燥、还行但不够好看”;至于 4.8,我认为它“华丽高级”(我对 4.8 情有独钟,除了它因为一个页面不工作被扣分外,我是它的忠实粉丝);而新发布的 Sonnet 5 则产生了许多运行崩溃的糟糕原型,在运行成功时我很喜欢它,但它不工作的频率太高了。Gemini 3 Pro 的代码非常精简,呈现出“虽然赤裸但极其紧凑”的特征,直切痛点。定性来看,我最喜欢 Opus,同时非常希望看到 GPT-5.5 和新版的 Sonnet 5 在功能表现上更稳定,以便我能完全基于它们的品味和质感来进行评判。

此外,我们让 Opus 4.8 和 GPT-5.5 对它们自身的输出进行了自我裁判,以此来测试模型是否存在自我偏袒的偏见。测试一致表明 GPT-5.5 依然是最苛刻的裁判,它甚至将自己生成的代码评分评得比另一个 AI 裁判给的分数还要低。

那么,本期 Howi AI Bench 盲测指数的最终结论和未来的优化方向是什么呢?

  • 工具选择视任务而定:选用哪款大模型完全取决于你的任务性质和该模型在特定任务上的单项优势。
  • 人类审美品味至关重要:事实证明,我的直观体验(Vibe Checks)是极具价值的,且与纯模型的客观测试结果产生了明显偏离。我将设法在未来的自动评估逻辑中编码并引入更多我个人的审美判断。
  • 淘汰已饱和的传统任务:正如我所料,我们需要退休掉诸如“智能体 Bug 追踪”这种所有主流模型都能轻松拿下的任务,因为它已经无法作为核心的分野指标。

基于人机意见的冲突,我让 Claude Code生成了一个基于我个人审美品味和后台代码性能的双重加权 Leaderboard 页面,并在二者之间取得了平衡。我直接将比重调整为:70% 的人类审美权重(Clare taste)和 30% 的后台客观表现权重

调整权重后的最终 Howi AI 综合指数榜单如下:

  1. Claude Sonnet 4.6(高居榜首)
  2. Gemini 3 Pro
  3. GPT-5.5
  4. Claude Sonnet 5(与 Opus 4.8 一同垫底)

针对具体的业务场景,我给出的模型推荐如下:

  • 撰写产品需求文档(PRD):推荐使用 GPT-5.5,它能输出极其详尽且结构清晰的专业方案。
  • 快速线框图设计与日常聊天:推荐使用 Claude Sonnet 4.6,它的个性和交互语气无可匹敌。
  • 深度代码库重构与大规模任务:虽然我没有亲自打分,但 LLM 裁判认为 Opus 4.8Sonnet 5 表现极其优秀。
  • 复杂高保真 UI 原型开发Opus 4.8 在构建密集、复杂的交互界面和 C端用户设计上依然是行业霸主。

这是一次奇妙的探索。原本这是一期针对全新 Sonnet 5 的测评节目,但最终它却落在了我个人偏好列表的末尾。这就是我们第一期的 Howi AI 审美加权综合指数榜单。每次新模型发布时我们都会重跑这套评估,并会不断优化它的严苛度以贴合真实的人类审美与工程实际。非常感谢大家收看本期 Howi AI,我们下个模型发布时再见!如果你喜欢我们的内容,请在 YouTube 上点赞并订阅,或在下方留下你的评论与看法。你也可以在 Apple Podcasts、Spotify 或其他播客应用上订阅我们的音频节目,并留下评分和评论。你可以在 howiipod.com 上查看所有节目并了解更多信息。下期再见!

Original English Okay, you can see what I say about these models by hand. Again, I think this is quite funny, which is let's see on 46, what were the issues? I said slop, not as functional, boring, okay, but not super cute. So 46, generic, sloppy, 48 fancy. I really liked 48. So other than getting kind of dinged on one not being functional, I was really a big fan of 48. It seemed like five and sonnet 5 had a lot of broken prototypes in it. And so when it worked, I really liked it, but it didn't work enough. And Gemini 3 very interesting, bare bones, it seems like, but concise. And so I think like right right to the point. So if I were to look at this from a qualitative perspective, I certainly like opus. And I would love to see 55 and Sonic work better because then I could judge it on its merits of taste. So um again we had um model as a judge and so we had opus 48 and 55 judge itself. Um I had the benchmark check if there was any inherent bias like did opus like opus better and 5.5 like 5.5 better. I've consistently seen GPT 5.5s be the toughest judge. And so I actually prefer a 5.5 judge, but it judged itself lower than the other judge did. The judges overall agree, but they were overall generous. And sort of balancing these two judges is exactly why we ran this double bench. Okay, so takeaways and what changes next launch in terms of the Howi AI bench? Well, the model is going to depend on the job and the strength of the model foot by task. I would say my taste actually matters. So maybe those vibe checks are not bad. And it really diverged hard from the metrics. So what I'm going to try to do is encode more of my taste into the judgment. It says retire the saturated agentic task. That's really interesting. Again, I didn't read this before I presented it, but that's exactly the conclusion I came to, which was this like agentic bug tracking task is not a really good benchmark because all of them are pretty good at it, and I need to think about something else to test the agentic nature of these models. And so, I don't re I don't really know what conclusion to draw from this. So let's go back to good old Claude and say given the benchmark and I agree can you do a Clareire weighted index and generate a leaderboard page that strikes the right balance between my opinion and the backend performance and makes recommendations on model by task. Okay, so we're going to have Claude Code summarize this benchmark, which is all over the place. Again, we do it live here at How I AI, and give you a ranking. Should we believe the AI leaderboard or should we believe the CLA leaderboard or somewhere in between and come up with our definitive end of June early July 2026 how I AI index of the paid frontier models. Let's see. Okay, Claude could not commit to making a decision itself. So, it gave me ultimate power. It gave me a slider from 100% LLM judge to 100% CLA judged. It's my podcast. I'm going 70% Clare judge, 30% backend. At the top of the list, Sonnet 46. Who would have think? And Gemini 3 Pro, followed by what I think is my favorite 55. And at the bottom, poor brand new Sonnet 5 and really expensive 48. What is CLA's recommendation? Model by task. If you're writing a PRD, use GPT 5.5 because it will give you something comprehensive and clear. If you are prototyping, guess what? Sonnet 46 pretty good. And if you want to chitchat with a model, again, Sonnet 46 has good vibes. If you're trying to knock down a codebase, I actually did not score these, but the LLM judge thinks that Opus 48 and Sonnet 5 are pretty good at this. And then if you are doing prototypes, depending on what you're doing, different models can do better. I would say complex designs, again, what I saw on my chat PRD benchmark is Opus48 does really good at really dense, complicated UIs as well as consumer. And then you can use Sonnet for things that are just a little bit simpler to execute on. Okay, this was an adventure. This started out as a Sonnet 5 review. It ended up that Sonnet 5 is at the bottom of my personal preference list. Well, that's it. That's our first round of the How I AI Clare Weighted Index. We are going to be doing this every time a new model comes out. I'm going to try to encode the benchmark and make it a little bit more critical, a little bit more aligned with my taste. I can't wait to see how it does on some of these new models and I can't wait for this to be an industry standard benchmark that all the labs rely on. Thank you for joining How I AI and see you next model release. Thanks so much for watching. If you enjoyed this show, please like and subscribe here on YouTube or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at howiipod.com. See you next time.
📌 文中提及的人物和组织

关键字: llm-benchmarking agentic-workflow product-design