模型破局:带护栏的通用智能体
随着 Anthropic 正式推出 Mythos 级别的模型,代表着通用可用性(General Availability: 软件或系统面向所有大众用户正式发布的阶段)的 Fable 5 终于来到我们面前。作为 Mythos 的“幼年版”,为了防止在网络安全和生物学等敏感领域的滥用,系统为其配备了专门的安全护栏。在性能层面,它彻底碾压了现有的基准测试,尤其在 SweetBench Pro 上达到了 80% 的卓越表现。与传统模型相比,Fable 5 是一台为了雄心勃勃的复杂项目而生的引擎,具备极高的自主性,能够执行长达数天的异步任务,动态编排不同的子代理(Sub-agents: 在复杂任务中被主代理调用的专门负责特定子任务的自治单元)。
然而,构建这种深度智能的代价是极其高昂的。在实操层面上,Fable 5 作为超越 Opus 的新阶层模型,其定价高达每百万输入 Token 10 美元、输出 50 美元。系统反馈显示,由于其在验证和构建时异常卖力,Token 的消耗速率甚至达到了其他模型的两倍。这意味着,即便是日常的软件工程任务,用户也必须在成本效率与模型智力之间进行精密的计算,随时准备好迎接“燃烧 Token”带来的账单压力。
Original English Source
It's here. The model, the myth, the legend. Mythos from Anthropic has finally dropped. Well, baby mythos. We're calling it Fable 5. And this new model is crushing benchmarks. But the question is, can it crush my backlog? I got early access to the model. And of course, I have my own opinions on where it does really well, where it needs a little work, and the question on everyone's mind. Does it live up to the terrifying marketing hype? Let's get to it. Okay, let's talk about what Anthropic is telling us about this model, and then we'll get into what I think about it. So, this is Claude Fable 5, the first mythos class intelligence model to reach GA. Now, if you haven't been paying attention, Anthropic has been marketing slashcaring slash warning us about the unbelievable capabilities of Mythos, and it is finally here. Now, they had originally been rolling this out with a couple select companies. I got early access to test what I thought of the model, but you have to know this is not Mythos, capital M, big mythos. This is baby Mythos. This is fable. And so it's going to have some guardrails on it in particular around cyber security exercises and biology exercises. Now, good news. Your girl's working on PRDS. She's shipping SAS. She's not working on biology quite yet. Although, give me give me a little time and some time to experiment. Maybe I'll get there. So, this is really going to be focused on what the everyday user with the everyday software engineer is going to think about when they're using this model. Although I did run into some things that I suspect are a result of the tuning and training of this particular model to be extra safe. Now quick, it's not cheap. It's $10 per input token and $50 per output token. It's going to be a new tier above Opus. And so if you're going to use this model, you're going to pay pay the price. So what is Anthropic saying? Basically, it's a completely new model class. So we had Sonnet, we had Opus, and now we have Mythos. The first of which is Fable 5. It's completely state-of-the-art. It is exceeding every benchmark they tested by a significant amount. This 80% on SweetBench Pro, you'll look at that compared to some of the more recent models that have come out. Very, very good benchmark performance. And then they're saying it's really good for long complex tasks. Now, what are some things that earlier models couldn't do that they are saying now that Fable 5 can do? It's very autonomous including running daysong asynchronous tasks. It's really an engineer's engineer. And that's some of the downside I experienced with this model. I'm going to show you a very specific example of where you don't want an engineer doing your work with an engineer's um point of view. Proactive. Um it's very good at vision, exceptionally good at vision. This is a place where I actually really loved the model. And you know me, I'm pretty critical of models, but I did see a step ahead of vision. So that's something we're going to dive into. And then effort. It works hard. It builds harder. It verifies more. It's built for ambitious work. Now, guess what it also can do? It can consume those tokens. So, Anthropic has said it consumes rate limits and tokens at about 2x the rate of other models. So, again, this is a big boy model and it's going to consume tokens and some of the things that it's good at and even some things that they have done in the harness seem like they're uh intentionally or not token consumers. So, we're going to keep an eye out on costs and an eye out on efficiency when using this model. Again, talking about longunning tasks, Fable 5 is supposed to be able to run for days. So, doing longunning planning, being able to spin up sub agents and I show a little bit about dynamic workflows, which are, you know, different architectures of sub agents and holding multi-day sessions. Now, I have done probably day days long sessions with other models. I didn't have Fable for many days, so I cannot verify that it ran for days. I did get it to run, however, for several hours on some tasks that may or may not have merited that several hour effort, but it definitely seems like it has both the harness and the intelligence capability to run for a very long time if that's appropriate for your task.机制博弈:工程师心智与系统降级
Fable 5 被官方明确设定为具备“经验丰富的资深工程师”的工作方式。在具体实操中,这种心智特征表现出了强烈的两面性:为了确保交付万无一失,它会穷尽所有的逻辑死角,极具彻底性地追求 120% 的准确度;但这往往无助于快速的产品发布。在开启“极高(Extra High)”级别的推理投入时,系统会疯狂燃烧 Token,这种过度的智能投入反而可能降低效率。因此,在某些不需要极高智力的任务上,回归 Sonnet 或 Opus 模型可能是更明智的选择。
在建立这种认知后,理解其底层的分类拦截机制显得尤为重要。为防止越界,系统在网络安全、化学等领域设置了分类器,并引入了一套优雅的后备机制(Fallback: 当系统遇到限制或错误时,自动降级切换到稳定旧版本的安全策略)。一旦触碰红线,API 会平滑降级至 Opus 4.8,而真正的完全体 Mythos 仍被限制在 Project Glass Wing 等企业级安全测试中。同时,为了更好地调度这种复杂的模型能力,Anthropic 推出了进入公测的 Claude Manage Agents 托管沙盒。开发者可以通过设定 Fable 5 为高级策略顾问,搭配廉价执行模型,构建起高性价比的自动化链路。在基准测试的硬指标上,Fable 5 不仅全面超越了 Opus 4.8,甚至大幅领先于 GPT-55 和 Gemini 3.1 Pro,彻底确立了其前沿地位。
Original English Source
Now, here's your pros and here's your cons. They explicitly say that Fable works like a seasoned engineer. Unfortunately, if you have worked with a seasoned engineer, you know there's good to this and you know there's bad to this. So, it is very complete in its investigation. It's definitely going to go search out all the corners. It's definitely going to think about how it can be 120% sure that it's shipping the right thing. But guess what? That's not always in service of launching. And that's honestly not always in service of building a great product. So while you can give it a goal and it will be very autonomous and it will be very thorough honestly sometimes you want like a slightly less thorough engineer product manager talking even engineer talking sometimes you want it to be a little bit dumber we'll talk about some of the prompting techniques it says and when to use this model but it's just something to think about when you're working with any high intelligence model is how much intelligence does the task actually take now as I said before it is token intensive by design and I did most of my tasks on extra high and so it was like token burning on token burning and so they say that high is probably the sweet spot for most work. I used extra high just because I don't want anybody in the comments saying Claire you picked high for this task and it should have been extra high and you would have had a better experience. I used extra high. I used all of the brains of Fable, but again, it is very, very token intensive. And my question for any of these models, this is not an anthropic model question. This is not a Fable question, is does this token intensity actually output the right results? And that's a place where I'm just not 100% sure. But again, as us humans in the loop, we're going to have to be much more intelligent about where to put what model and where to use what reasoning and what effort level to match what we're doing. And again, I think that the untrained of us will say, "Oh, well, I have this fable model. I should use it. It's better than anything." And honestly, I still think there's a place for good old sonnet. I think there's a place for Opus. And I think there's a place for other models in the ecosystem. Now there are safeguards in this model and so this was one of the first things that Anthropic told me testing the model and this is one of the headlines that they're making in the release which is there are specific classifiers in this model for cyber security biology chemistry and distillation. Basically they don't want anybody doing bad stuff in those categories in particular with this very intelligent model. What's nice about how they've implemented this, however, is they have this new fallback concept. And so if you get classified into one of these categories, instead of saying like do not pass go, you may no longer fable, it just falls you back to opus 48. This is also a capability in the API now where you can do this graceful fallback to 48 if you're using a mythos class model. They also have a 30-day retention policy used only to catch misuse and it's not used to train claude. So, while it's still not clean training Claude, they do want to check the use of this model because they have been and will forever be very cautious about normies using their intelligent models. And just, you know, for context, 95% of sessions on this model did not hit a fallback. I don't believe I hit a fallback but again I'm not doing anything in cyber security biology or chemistry at least yet. Okay, so this is the question. Is this or is this not Mythos? It is Mythos. Fable has the safeguards. Mythos does not. Fable all normies can have in general availability. Mythos is still restricted to these project glass partners. Some of these enterprise level partners that are really checking it against cyber security use cases. I would suspect that at some point we get some access to a Fable 5 whatever or that the Project Glass Wing class opens up but for now we get Fable Project Glass Wing or these um pre-selected companies get mythos but they are all fundamentally the same underlying model. A couple product things that are also launching today along with the Fable 5 model cloud manage agents are going to public beta. If you haven't paid attention, this is Anthropic's hosted harness, hosted sandbox for running long running aentic work. I am still trying to figure out what a good use case for cloud manage agents is. I will get there. Um, but Fable ships out of the box in cloud manage agents. There's also a new advisor strategy where you can use Fable 5 as a senior adviser and use cheaper models as an execution layer. A lot of people are doing this with Opus and Sonnet and so this is going to work today in the API and in cloud code and is a strategy you can use. And then as I mentioned this fallback API where you can put an optional parameter on the messages API that allows you to continue to block requests by using 4.8 at Opus pricing. Okay, as we said crushing benchmarks. Look at this Fable 5 compared to Opus 48, GPT55 and Gemini 3.1 Pro. significant increase in SweetBench Pro benchmark. Um, very far ahead of these other models and while I wasn't testing the most advanced use cases, I didn't find something that technically it failed at. So, I think these benchmarks are really going to hold and these benchmarks have outperformed across the board. So, this is Anthropic's stateofthe-art model.模态反差:视觉跃迁与文本迷宫
在跨模态的具体应用中,Fable 5 展现出了极端的偏科反馈。在实操中,其视觉解析能力实现了真正的代际跨越,尤其在文档排版重组方面表现惊艳。例如,在处理针对七岁儿童的经典诗词手写练习表时,与 Opus 拥挤不堪的排版灾难相比,Fable 5 能够精准控制留白和行距,生成了高度符合人类视觉直觉、清晰易读的页面布局。
然而,这种在视觉上的卓越表现在长文本撰写领域却遭遇了彻底的滑铁卢。当模型被赋予编写系统规范(Specs)或长篇文章的任务时,“工程师心智”再次引发了灾难。以 chat pd 开源项目的产品图谱(Product Graph)对抗性审查为例,模型交付了一份充斥着内部引用、极长且密集的 Markdown 文档。从底层心理来看,模型纠结于每一个微小的技术细节,这让阅读者深陷“只见树木不见森林”的泥潭中,完全无法提取宏观框架。基于这种反馈,最佳的做法是将撰写策略文档的任务退回给 Sonnet 或 Opus,仅保留 Fable 5 作为无需人类干预的后台纯执行大脑。
Original English Source
Okay, so enough about what they say. Let's talk about what I say. what is it actually like to use? So, I ran Fable 5 on a bunch of different work and I want to give you my feedback on where I thought it did well, where it needed a little bit of work, and where I was really surprised. As I said before, it's really good at vision. And where is it good at vision? That really impressed me. It's really good at document formatting. So, this super simple um but we've been doing these handwriting documents for my 7-year-old based on classic texts and classic poems. And on the right is Opus 4A and on the left is Mythos 5. And it looks so silly, but I really do think Mythos 5 did a much better job of a second grade layout for a handwriting sheet. There's just like the right spacing. It's very clear to read. there's enough white space. I think on the one on the right, it's just very dense and even the lines themselves are sort of hard to tell. Do you write above? Do you write below? So, I do think that PDF formatting documents, I tested this against a bunch of different models. Mythos 5 really did a good job. So, very simple eval for me, but a very, very good one. Now, here's the problem though. The writing is nearly unreadable. So, if you're thinking about Mythos for pros, for spec writing, for PRDS, unfortunately, it's an engineer. And what's the problem with engineers? They just really get wrapped around the axle on details. And this is a real struggle with these more intelligent frontier models is they're like too smart. And so, it's just very, very hard to parse what they're saying. And I'm going to show an example of this in actually cloud code. So I have this concept of a product graph that I'm working on for chat pd. It's actually a fairly complex open- source project and I had fable 5 go through that and actually do like an adversarial review of my requirements to try to figure out where there were internal consistencies in the logic. and it gave me this markdown document that looks very long and intelligent. But if you actually go through it, it's just really hard to parse. It's there's like internal references. It's very detailed, but not in a way where you can zoom out. There are these big blocks of paragraphs. Like look at look at this. It is just really hard to see the forest for the trees in this particular model. And I saw this sort of like over and over again working on it with specs is it was very complete but nearly impossible. And that's a real challenge when working with these very very high intelligence models. Again, I would actually suggest pulling back to maybe a sonnet or opus model for specs and then looking at fable as an orchestrator of execution where that detail really matters, but you don't have to read it.执行边界:设计崩溃与动态编排
在系统的边缘测试中,Fable 5 的前端界面设计能力令人大跌眼镜。在技能注册表(Skills Registry)的零样本(One-shot)设计实操中,它交付了充斥着灰色、黑色和红色的刺眼界面。这并非普通的“AI 工业流水线感”,而是彻底的视觉灾难。即便是通过极高密度的提示词进行干预,也无法实质性扭转其糟糕的设计审美。在功能执行层面,它的策略同样极度保守。当被要求构建一个具备基础客户价值的最小可行性产品(MVP)时,模型过度执行了“最小化”的指令,导致产出的系统极度狭窄且几无实用价值,其产品野心似乎被底层的安全护栏完全束缚了。
除了设计层面的痛点,动态工作流(Dynamic Workflows: 根据任务进展实时调整子代理结构与交互逻辑的系统架构)的实测也充满了不确定性。在 Claude Code 框架内运行长达三小时的多代理任务时,系统频繁出现进程停滞和无响应的报错。总结这套实操方法论:Fable 5 是一台专为硬核技术难题和视觉解析(如 PDF 文档提取)打造的精密仪器;绝对不要让它沾手前端设计、战略规划或文档打磨。只有通过精准分配任务并适度降低非必要环节的算力投入,我们才能真正释放这只“幼年神兽”的无穷潜力。
Original English Source
The other thing that shock shock shocked me was how like actually legitimately terribly bad it was at design or at least a oneshot design. And so I asked Fable to design a skills registry and man alive did it do a very poor job. I mean I'm not even talking like AI slop bad is like fundamentally terrible design. gray, black, red, simple outlines, just really, really terrible. Now, the anthropic team suggested that I just needed to be a little bit more detailed in my prompting. I've never had to do this before in I would say the last year of models in terms of front end, but even when I prompted it, it was still just not very impressive design. I think there's this real balance between design slop and specificity and just shipping like terrible design. I'm not sure what about Fable 5 resulted in this and I have to keep testing it as it rolls out today, but this was a real disappointment in terms of design. So again, you might want to toss an opus in the mix instead of relying on Fable for design. It's really conservative on execution. So when I was trying to do that ambitious days long work, I took a spec and I said, can you ship the V0 of this, the MVP? I said enough to that a customer could get value. and the MVP, they just really took minimal to heart. It was like very very narrow, not actually that useful. And I'm curious if this comes from some of the safeguards on this model, and it's it's been a challenge I've seen since the kind of later Opus models is they're not super ambitious. And so again, you'll have to think about how to prompt this to get that longunning outcome paired with the right product ambition. And then I really doubled down trying to test these claw dynamic workflows and these sub aent designs trying to see if this would really add value and the multi- aent capability is definitely there and I definitely had some successful multi- aent runs kicked off in Fable but I also ran into a lot of stalls and errors in using multi-agent orchestration. Now, I made the mistake. I walked away from my laptop and came back to these dev agents that had stalled after about three hours. And so, like, egg on my face. But I really want to see how technically the claude code model holds up to the promise of multi-agent orchestration. I had some successes and some bugs. I think this is a claude code issue, not necessarily a model issue. Although with this promise of longunning days long prompts, you really got to deliver technically on the outcome. So what's my takeaway? I would hand it hard problems. Of course, not cyber security bio or chemistry problems, but hard technical problems were being extremely detailed matters, long horizon work. I would also hand it vision problems where you really want something to look good or you want it to parse PDFs or other documents. It's done exceptionally well there. I was actually really surprised. I probably wouldn't uh hand it my front end work or I definitely wouldn't hand it my front end work. And I definitely wouldn't hand it strategy or spec work. I think it overthinks things. I think it's pros is nearly impossible. And so maybe I'll test it again with effort level lower on sort of pros and spec writing, but it wasn't it for that. That being said, I'm not a hater on this model. I definitely not. It definitely has a place in your stack. I'm going to test it. If you want to learn more, definitely look up the prompting guide for Fable. It's going to probably repeat a lot of what I said. Hand it your hardest problems, what this model is good for and what it's not, and how to get a good outcome. That being said, Mythos is here. I cannot wait to hear what you build, what you overbuild, and what you make ugly with this new model. Thanks for joining How I AI. Thanks so much for watching. If you enjoyed this show, please like and subscribe here on YouTube, or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at howiipod.com. See you next time.📌 文中提及的人物和组织
公司/组织: Anthropic
产品/模型: Claude, Fable 5, Mythos, Opus 4.8, GPT-55, Gemini 3.1 Pro, Sonnet, Claude Code