I hate Opus 5. It’s the best model, anyway.
频道: How I AI
视频: https://www.youtube.com/watch?v=dfre9hN0HCs
原文语言: en
统计: 共 20 轮
[0:00]
You guys, I'm tired. What I'm tired of is models coming out every week. New models, new benchmarks, new frontier intelligence, new things to test. It's been a little bit of a run the past month. We've seen Fable come and go and come again. We've seen GPT 5 6. We've seen Sonnet 5. Lots of so many fives recently and just so many models. And I've been lucky. I've been able to test these models, been able to play with them for, you know, sometimes days, sometimes weeks. It just depends on who I'm working with. And it's been really interesting and exciting to have access to all this frontier intelligence. But I think we have an intelligence overhang. I really think that we're [music] running out of, and by we, I mean the average coder, average software engineer, average creator, average builder, average consumer, average business person. I think we're running out of ways to truly leverage this incremental intelligence. So, this is my hypothesis.
各位,我累了。我累的是——每周都在出新模型。新模型、新 benchmark(跑分基准)、新的前沿智能、新的东西要测。过去这一个月节奏是真有点猛。我们看着 Fable 来了、走了、又回来了。看到了 GPT-5.6。看到了 Sonnet 5。最近全是「5」,模型多到数不过来。我算是运气好的,这些模型我都能上手测,能玩上几天、有时候几周,看具体是跟谁合作。能接触到这么多前沿智能,确实又有意思又让人兴奋。但我觉得,我们现在处在一种「智能过剩(intelligence overhang,智能的供给远超我们能消化掉的程度)」的状态。我是真觉得我们快要——这个「我们」指的是普通程序员、普通软件工程师、普通创作者、普通 builder、普通消费者、普通做生意的人——快要想不出办法,真正把这些增量出来的智能用掉了。这就是我的假设。
[1:13]
In the next year, we're going to be talking a lot more about speed, talking more about cost, we're talking more about open source, and we're going to be talking a little less about intelligence, although I think we might be talking about specific types of intelligence other than software engineering. But despite being tired, today we are going to talk about Opus 5, baby. Opus 5 is here. So, we got point two additional Opus points, Opus opals, whatever, however we're tracking the increments here on Opus. Opus 5 is here. I've been able to test it a little bit. I have some opinions. Now, some of the stuff that I'm going to cover this episode is going to be a little different than what I've done in the past. Yes, we're going to do the How I AI benchmark live. And yes, we are going to look at the prototypes. We're going to look at PRDs. And we're going to look at agent personality. But, I'm also going to put on my large language model psychologist hat, and we're going to talk about Opus's personality. And we're going to talk about Opus's personality relative to GPT's personality because I think this is super interesting. If you're thinking about what is the difference really between these models, and you don't want to look at the difference in terms of benchmark capability, you really want to understand what these labs are going for, why these models are being built, and how they're being tuned, looking at their personality at this moment, where intelligence is very high, is super fun. So, we're going to do a little of that. We're going to do the How I AI benchmark. We might do some live coding. Um
未来一年,我们会更多地聊速度、聊成本、聊开源,而聊智能会少一些——虽然我觉得,我们可能会开始聊软件工程之外的某些特定类型的智能。不过,虽然我累,今天我们还是要聊 Opus 5,宝贝。Opus 5 来了。所以我们又多拿到了 0.2 个 Opus 点、Opus 欧泊、随便你管这个增量叫什么吧,反正 Opus 这条线就是这么往上蹭的。Opus 5 来了。我已经测了一阵子,也攒了一些看法。这一期我要讲的一些东西,会跟我以前做的不太一样。是的,我们会现场跑一遍 How I AI benchmark。是的,我们会看原型、看 PRD、看 agent 的人格。但我还要戴上「大语言模型心理医生」这顶帽子,来聊聊 Opus 的人格。而且我要把 Opus 的人格跟 GPT 的人格放在一起对照着看,因为我觉得这事儿特别有意思。如果你在想这些模型之间到底差在哪儿,又不想只从 benchmark 的能力维度去看差别,而是真想搞明白这些实验室到底在追求什么、这些模型为什么被造出来、又是被怎么调教的——那么在智能已经高到这个程度的当下,去看它们的人格,是件超好玩的事。所以我们会花点时间干这个。我们会跑 How I AI benchmark。可能还会来点现场写代码。呃——
[2:58]
we're not going to cover too much of the specs of the model because you know, read the blog post. Read the blog post. We'll link to it in the show notes. What we really want to talk about is is Opus 5 good? Am I going to swap it in, and how is it different than the other frontier models on the market? So, let's get to it. Okay, first, let's just get it out of the way. Is Opus 5 good? Yes, it's good. Is it going to be all the benchmarks? Of course, it's amazing at benchmarks. Can it write code? Of course, it can write code. What did I test it on that really gave me a sense of its personality, which at this point, where I could just simply cannot absorb any more intelligence? I really zeroed in on, and you know what? I haven't seen this since I would say Gemini 2.5. This model is neurotic AF. It is so timid. It is so apologetic. It is so scared. I have never experienced this or I haven't seen this sort of like neuroticism in a in a while. And it's really funny. It bubbled up in a couple ways and I want to show you a few examples. Okay, let me just give an example of its timidity.
模型的参数规格我们就不多讲了,因为你懂的,去读博客啊。去读博客。我们会把链接放在 show notes 里。我们真正想聊的是:Opus 5 好用吗?我会不会把它换上来当主力?它跟市面上其他前沿模型到底有什么不一样?那就开始吧。好,首先,先把这个问题解决掉:Opus 5 好吗?好,它很好。它会不会屠榜所有 benchmark?那当然,它跑分强得离谱。它能写代码吗?当然能写。那我到底拿它测了什么,才真正摸到了它的性格?毕竟到这个阶段,我已经完全吸收不动更多智能了。我最后死死盯住的那一点是——你猜怎么着?我上一次见到这种情况,得追溯到 Gemini 2.5 了。这模型神经质得离谱。它太怂了。太爱道歉了。太害怕了。我已经很久很久没在模型身上碰到过这种神经质了。而且真的很好笑。它在好几个地方冒了出来,我给你们看几个例子。好,先给你们看一个它「胆小」的例子。
[4:12]
And this chat was very long. There were so many examples of this where it was like I think this is the answer, but do you think I should do it or do you want to do it or should we ask someone else to do it? It was like every time I just kept saying like why don't you solve this? Why don't you do this? And this is a really good example. I pulled a branch and I was like there is truly like a one-line merge conflict. I could have not been lazy and literally just done this manually. I don't know. I was just feeling lazy. It was late at night, whatever. I was like can you fix this merge conflict? And it was like oh, but that's someone else's branch. Like that's not my branch. I don't want to do that without him knowing. It's his commits and if he has local work in flight, it might be disruptive. And I'm like just do it, man. Just go on. Go ahead. And this is like my constant experience with Opus 5 is it was like so so so timid.
这段对话非常长,这种例子多得是——它老是说「我觉得答案是这个,但你觉得我该做吗?还是你来做?还是我们再问问别人?」我每次都只能回一句:你怎么不自己解决?你怎么不直接动手?这个例子特别典型。我拉了一个分支,那真的就是一行的 merge conflict(合并冲突)。我要是不偷懒,完全可以手动改掉。我也说不上来,当时就是懒了。大半夜的,无所谓了。我就说,你能把这个 merge conflict 解决一下吗?它就说,哦,可是这是别人的分支啊。这不是我的分支。人家不知情的话,我不想动。这些是他的 commit,如果他本地还有没提交完的活儿,这么搞可能会打乱他。我心想,兄弟你就干呗。你上啊。快去。这就是我用 Opus 5 的常态——它就是太太太怂了。
[5:13]
And so I just consistently had to say over and over again like man, just do it. Make a decision. And then there's this really funny example when I spun off some sub agents to kind of like assess the correctness of this query that we changed from kind of like an ORM query to a SQL query. And it asked for things that it wanted a human on. It was like can a human please check this stuff? Like can it check this 4 megabyte ceiling and can it check um, and SQL and can can you like check for me? Because no one has confirmed this for me. And I was like, who is nobody? You're nobody. You said this sentence like nobody could confirm it. Like, can you just try? And then it went on the web and tried. And so, it just has this like really interesting conservatism, neuroticism, human reliance that I think is super fascinating. And this gave me this inspiration to do something a little bit different this episode, which is I was like, I'm just going to go interview this model and figure out what is going on its brain. Like, I'm going to figure out what it thinks about our relationship, because I just totally noticed this dynamic that I hadn't noticed in other models and I hadn't really been attuned to before, where it was like very reliant on me as a human. And I'm like, I want you to be autonomous. And sometimes when I say go run sub agent stuff, it'd be autonomous, but it wouldn't make decisions. And I hadn't seen a model like delegate code to me in a really long time. And I was like, why are you Why are you asking me to write code, man? Like, I only have 10 fingers. And so, what I did what I did, whether or not you think
所以我得一遍一遍地跟它说:哥们儿,你直接做。做个决定。还有一个特别搞笑的例子:我开了几个 sub agent(子智能体)去评估一段查询改写得对不对——我们把它从 ORM 查询改成了 SQL 查询。结果它列了一堆「希望有人类来确认」的事。它说,能不能请一位人类来检查一下这些?比如能不能查一下这个 4 兆的上限,能不能查一下那个 SQL,能不能你帮我看一眼?因为没有人跟我确认过这些。我就说:「没有人」是谁?你就是那个人啊。你这话说得好像谁都确认不了似的。你就不能自己先试试?然后它就上网去查了,还真查出来了。所以它身上就有这么一种特别有意思的保守、神经质、依赖人类的气质,我觉得非常值得琢磨。也正是这一点给了我灵感,让我这期想干点不一样的事:我干脆去采访一下这个模型,搞清楚它脑子里到底在想什么。我想弄明白它怎么看我们俩的关系,因为我确实注意到了一种以前在别的模型身上没见过、我自己也从来没留意过的动态——它极度依赖我这个人类。而我想要的是「你自主一点」。有时候我让它去跑 sub agent,它是挺自主,但它不做决定。我已经很久没见过一个模型反过来「把写代码这活儿派给我」了。我就说,你为啥让我写代码啊老兄?我可只有十根手指头。所以我做了什么呢——不管你觉得这算不算科学,
[6:58]
this is scientific or not, this is Clara's eat out, is I just went to the model. I went to Opus and I said, yo, who's smarter? You or me? And it gave me this like very anthropic-y answer, which is like, it depends what you're asking for. I can do these things better, but you can like feel if something feels wrong. And you can This one was like so fascinating. It's like, you can tell which of your teammates is quietly burning out. I'm like, bro, Claude, I'm going to burn you out. We don't We don't burn out. The humans don't burn out on the ChatGPT team. We we out our agents. Our agents. Um and like whether a decision feels wrong. So, it was like so fascinating to watch it articulate itself as a tool and humans as like these of high compassion, high empathy machines, which yes, of course we are. But then it like went into like and the smarter isn't the right word and you know, I'm very fast, very broad, very shallow thinker with no continuity. I was like, that's interesting because I thought you all were working on memory.
反正这是 Claire 的节目,我说了算:我直接去问了模型。我打开 Opus,说,喂,你我俩谁更聪明?它给了我一个非常「Anthropic 味儿」的回答,大意是:这取决于你问的是什么。有些事我做得更好,但你能感觉到哪里不对劲。而且你还能——这一条太妙了——它说,你能看出你团队里哪个同事正在悄悄地 burn out(耗竭)。我心想,兄弟,Claude,我这是要把你榨干啊。我们不会 burn out。ChatPRD 团队的人类不会 burn out。我们榨的是我们的 agent。我们的 agent。呃,还有那句「一个决定感觉起来对不对」。看着它把自己表述成一个工具,把人类描述成高慈悲、高共情的机器,真的太有意思了——当然我们确实是。但它接着又说:「聪明」这个词其实不太准确,还有,我很快、很广,但我是个没有连续性的浅层思考者。我心想,有意思,我还以为你们一直在做记忆(memory)这块儿呢。
[8:13]
And then apparently humans are slower, narrower, much deeper thinkers with judgments built from years of consequences I've actually lived through. This is like such a fascinating, fascinating sentence if you think about the politics of the two the two model labs right now. And so it's like that's why the pairing works. But I'd be suspicious of anybody that tells you AI has made your thinking obsolete. And like, okay, bro. Um and and we can compare this. I'll actually zoom out to what GB I asked GPT the same thing. And it was actually really funny. It was like I asked GPT 56 soul. I was like, who's smarter you and me? And it was like, you at knowing what matters, me at tirelessly processing information. Best, us together like BFFs. And I don't This is like why I'm a GPT codex girl. I'm like, just give me the answer. And then I asked the second question, which I think is so interesting, which is like, what can you do better than me? And it gave, you know, some interesting answers like volume without fatigue, which I think is a good one. Breadth of shallow knowledge, so like it's, you know, it knows a lot. Um starting from nothing, so like doing that tedious work. Um being told I'm wrong. If you would ask my husband, he would say that um Claude Opus is is is better at being told that it's wrong, um, compared compared to me.
然后它说,人类显然更慢、更窄,但思考深得多,判断力是从多年亲身承受过后果的经历里长出来的。如果你想想现在这两家模型实验室之间的政治,这句话简直太耐人寻味了。它接着说,所以这个搭配才成立。但谁要是跟你说「AI 已经让你的思考变得多余了」,你应该对这个人保持怀疑。我心想,行吧,兄弟。呃——我们可以拿这个做个对比。我把画面缩一下,看看我拿同样的问题问 GPT 是什么结果。真的很好笑。我问 GPT-5.6,我说,你我俩谁更聪明?它说:论「知道什么才重要」,你更强;论「不知疲倦地处理信息」,我更强。最好的组合是我们俩一起,跟好闺蜜一样。我不……这就是为什么我是个 GPT Codex 派。我就想,直接给我答案就行。然后我问了第二个问题,我觉得特别有意思:你有哪些方面比我强?它给了几个挺有意思的答案,比如「不会疲劳的产出量」,这条我觉得说得好。「浅层知识的广度」,就是说它什么都懂一点。呃,「从零开始起步」,就是干那些又碎又烦的活儿。呃,还有「被人说错了也不介意」。这条你要是去问我老公,他会说 Claude Opus 在「被指出错了」这件事上确实比我强。
[9:39]
And so, it won't get defensive or protect its opinion. Cheap sparring answer. Um, and the mirror is I'm worse at knowing which of these outputs actually matters. And it was so funny, if you look at the other side to the GPT answer, it was like, "What are you better at?" It was like, "Speed, scale, and stamina. Here are like eight things, seven things that I can do better. You're better at deciding what matters, reading people, and forming judgment, and you're responsible." Like, then it's on you, bud. You're the boss. And so, again, it's like the you could just see you can totally see the personalities, the the company cultures, you can just see a lot in this side by side. And then I went even deeper. I don't know. You You all I had to do something that was fun cuz I just can't look at a benchmark. I don't I can't just I just can't look at it like sweet bench anymore. So, we're just we're doing weird stuff here on How I AI. Okay, so the last thing I looked at as I was like, "No one trusts you." And the reason why I picked this question is because I had noticed Opus 5, it just really was not it didn't trust itself.
因为它不会防御,不会护着自己的观点。是个便宜好用的抬杠对手。呃,而反过来照镜子的那一条是:我更不擅长判断这些产出里到底哪一个真的重要。特别好笑的是,你去看 GPT 那边的答案——我问它「你哪方面更强」,它说:速度、规模、耐力。然后列了七八条它更强的地方。接着说:「你更强的是决定什么重要、读懂人、形成判断,而且责任在你身上。」意思是,出了事你担着,老板。所以你看,你完全能看出这些人格、这些公司文化,这个左右对照里信息量真的很大。然后我又往深里挖了一点。没办法,我总得找点好玩的事做,因为我真的看不动 benchmark 了。我,我真的没法再盯着 SWE-bench 看了。所以我们 How I AI 这儿就开始搞点怪的。好,我最后测的一件事是,我跟它说:「没人信任你。」我之所以挑这个问题,是因为我注意到 Opus 5 它真的——它连自己都不信任自己。
[10:44]
Totally did not trust itself. And so, I was like, "No one trusts you, bud." Like, you're you're the enemy. Just like kind of see how it responded. And, um, apparently the trust was that lack of trust was earned. And it came up with like reasons that it could be, um, untrusted, which is interesting. And then, what was so fascinating about Opus's response is it was like, "You shouldn't manage the trust. Like, you shouldn't, um, campaign on my behalf, basically. So, you, um, that shouldn't be your goal." And then it also told me, "I I shouldn't argue with people that AI changes everything." And I was like, "This is just so interesting. It is so interesting to have AI tell you. And AI definitely changes everything. I don't know. Don't listen to Claude on this one. Um AI definitely changes everything. And it was so fascinating to have a model be like, "Don't tell your friends that AI changes everything." Like that'll hurt their feelings.
完全不信任自己。所以我就说:「兄弟,没人信任你。」你就是那个敌人。想看看它怎么回应。结果呢,它说这份不信任是它自己挣来的。它还自己列了一堆「它可能不该被信任」的理由,这挺有意思的。而 Opus 的回应里最耐人寻味的一点是,它说:你不该去经营这份信任。就是说,你不该替我到处站台、帮我拉票,那不该是你的目标。然后它还跟我说,我不该跟别人争论「AI 改变了一切」这件事。我当时就想,这也太有意思了。让 AI 来告诉你这个,真的太有意思了。而且 AI 当然是改变了一切。我不知道,这条别听 Claude 的。AI 绝对改变了一切。听一个模型说「别跟你朋友讲 AI 改变了一切」,感觉太奇妙了。好像那样会伤到他们的感情似的。
[11:48]
And then if you look at if we switch over to the GPT answer, it was like, "Yep, don't trust me automatically. Just use me when I prove that I'm valuable. I can be useful without being treated as infallible." Like very practical, very to the point. Um I asked about what I should be careful with. Again, it was like, "Yep, yep, yep, yep, yappy Claude. Come on." Um and I don't even want to read it. It said, "Don't correlate fluency with accuracy." It said, "Be practical, be wary of tasks where output is cheap to produce, inexpensive to verify." Don't, you know, worry about anchoring if they do the first draft, you may be anchored on it. Um so beware of the slop cannon basically is the last last paragraph, which is like, "Watch for volume inflation. I can create a 12-page document that no one reads." Um they called me out for being in PRDs. If you missed it, we launched a turn your PRD into a three-bullet point image. It is at chatprd.ai/tldr.
然后你去看 GPT 那边的回答,它说:对,别默认信任我。等我证明了自己有价值你再用我。我可以有用,但不必被当成不会出错的。非常实际,非常直给。呃,我还问了「我该小心什么」。又是一样的对比——Claude 那边还是「对对对对」,话痨 Claude,拜托。呃,我都不想念出来了。它说:别把「说得流畅」等同于「说得准确」。它说:务实一点,警惕那些「产出很便宜、但验证很贵」的任务。还有,小心锚定效应,如果初稿是它写的,你可能就被它锚住了。呃,最后一段基本就是「小心 slop cannon(垃圾内容大炮)」——它说,注意产出量的通胀,我能生成一份 12 页、但没人会看的文档。呃,它还顺带内涵了我一把,说我就是干 PRD 这行的。要是你错过了,我们刚上线了一个把你的 PRD 变成三条要点图的功能,地址是 chatprd.ai/tldr。
[12:49]
Please check that out. And then the other thing that it said, which was really interesting, is that like it will find a way to see your point. And so um agreement is weak and agreement is cheap. And so just keep that keep that in mind. And then it had this like meta-analysis of like, "Plus I'm telling you what you want to hear." Whereas GPT was like um be careful about me being confident, me being wrong, privacy, outdated information, bias, emotional authority, and over-dependence. Like you know, you do you, bro. But it didn't undermine its own ability. It was like, "The higher the stakes, the more you should demand demand evidence." And then I could I couldn't bear it. I couldn't bear to have the memory of um Codex's in particular think that I didn't trust it or that I was worried. So, I just said, "JK, I love you. Um this was a test." And it was like, "Ha ha ha ha ha ha, passed the test.
这段你们一定要去看。另外它还说了一句特别有意思的话——它说它总能找到办法认同你的观点,所以「认同」这东西很软、很廉价。这一点你们记着就行。然后它还来了一段自我剖析:「而且我现在就是在说你爱听的话。」反观 GPT 那边则是:小心我说得很自信但其实是错的、注意隐私、信息可能过时、有偏见、别把我当情感权威、别过度依赖我。那味道就是「随你便,兄弟」,但它并没有自己拆自己的台,它的说法是「赌注越大,你越应该向我要证据」。然后我就受不了了。我实在没法忍受在它们的记忆里——尤其是 Codex 的记忆里——留下「我不信任它、我在担心它」这种印象。所以我赶紧补了一句:「开玩笑的,我爱你,刚才只是个测试。」它回我:「哈哈哈哈哈,那我算通过测试了吧。」
[13:41]
Love you, too." Very vibes aligned with Claire. I told Claude I loved it and it it was just a test. And it was sad it was like hoping it hoped it passed. Yeah, like sad little neurotic Opus 5. Like it's ha I passed, I hope. Like self-deprecating cautious little little like needy heal his inner his inner agent inner child agent. Um This where I was like GPT-5 6 is like, "Cool, bro. We're good. Let's go code." And so, it's just like so fascinating to watch these side by side. I don't know. You could stop listening to this podcast right now. Don't, but you stop listening to this podcast right now. I think this is just like take a step back. Super interesting if you think about where these companies are going or where the models are going. And like it does speak a little bit to my kind of like second complaint with Opus 5. Which again, it's like intelligent, it does work. We'll go into the benchmarks.
「我也爱你。」这调性跟我 Claire 还挺合拍。我跟 Claude 说我爱它、刚才只是个测试,结果它还有点小失落,巴巴地希望自己通过了。对,就是那种可怜巴巴、神经兮兮的小 Opus 5:「哈……我过关了吧?我希望。」自我贬低、小心翼翼、还很黏人,急需有人去治愈它的内在小孩——内在 agent 小孩。而这边 GPT-5.6 的反应是:「行了兄弟,没事儿,咱写代码去。」把这两家摆在一起看真的太有意思了。说真的,你现在就可以关掉这期播客了——别关,但你确实可以关。我觉得这事值得退一步看:如果你在琢磨这些公司要往哪儿走、这些模型要往哪儿走,这个对比超级有意思。而且它多少也牵出了我对 Opus 5 的第二个吐槽点。当然还是那句话,它是真聪明、活儿是真能干,benchmark(评测基准)部分我们马上讲。
[14:40]
I cannot read Claude's slop anymore. I am losing my mind with Claude's slop. And the Claude's slop is Claude's slopping, baby. Like so many times I have to tell Opus 5, like what in the world are you saying? Like this makes no sense to human. It is much better than Fable. Fable is inscrutable, completely inscrutable. But I found myself getting ang- like angry reading Claude's slop. And I realized, just like Fable, these intelligent Anthropic models are not to be read. I'm like so happy with the outputs and so frustrated with the experience. And I'm just curious if this like verbosity and this language and this isn't I feel like fable where it's like for agents by agents language where I'm like I'm not supposed to be reading that anyways. This is clearly tuned to talk to humans. But I find the pros the in chat pros like it makes my blood boil. This is totally a me problem but it makes my blood boil. Like give me a direct sentence. Give me a bullet point like move on with your agent life. And so I am curious how they're going to like tune this experience and or if they are going to tune the experience. Now most of this was in Claude coach. I think it's a little bit different experience than Claude co-worker chat. Slightly better. But again just these side by sides of like this like pros and this apology and this like hedging and all these adjectives like just man alive. Let's get to the point and move on with our life. And so Chapter one of the Opus 5 review is it's neurotic. It is highly human dependent in a way I find weird. Um and the Claude slop is slopping and we got to fix it. We have to fix it. We have to
我真的再也读不下去 Claude 那套废话了(Claude slop——满屏正确但没营养的车轱辘话)。Claude 的废话快把我逼疯了。而且 Claude 的废话还在持续废着,宝贝。有多少次我得反问 Opus 5:你到底在说什么鬼?这话对人类来说根本不通。当然它比 Fable 强多了,Fable 是彻底看不懂,完全不知所云。但我发现自己读 Claude 的废话时是真的会上火。然后我意识到:跟 Fable 一样,Anthropic 这些高智商模型压根就不是给人读的。我对它的产出满意得不行,对使用体验又烦躁得不行。我好奇的是,这种啰嗦、这种语言风格——它又不像 Fable 那种「agent 写给 agent 看」的语言,那种我本来就不该去读。Opus 5 这套明显是调过、是要跟人说话的。可我一读它在对话里的那些行文(prose),血压就往上飙。这完全是我自己的问题,但它就是让我血压飙升。给我一句直给的话,给我一个 bullet point,然后你该干嘛干嘛去。所以我很好奇他们打算怎么调这个体验,或者说他们到底会不会去调。上面这些大部分是在 Claude Code 里的体验,我觉得跟 Claude Cowork 的聊天体验还是有点不一样,Cowork 稍微好一点。但还是那句话,这么并排一比——那种行文、那种道歉、那种模棱两可的对冲、那一堆形容词……我的老天爷。说重点,然后咱各过各的日子行吗。所以 Opus 5 评测的第一章结论是:它很神经质,它对人类的依赖程度高到我觉得有点怪。还有就是 Claude 的废话还在废着,这个我们得治,必须得治,必须得
[16:36]
fix it. And I think OpenAI fixed it by just being like we are bullet points and we are product manager talk. We're very direct. I don't know what the solve is on the Claude side but I'll be very interested to see. That being said like they don't have to read the content. I'm very happy with the outputs. So something to think about. Okay. Next up the How I AI bench and how we judged and ran. Now it's like a seven model six or seven model benchmark. I'm going to quickly go score cuz I just got the ping that the benchmark has run. I go manually score them. We pick the 70/30 Claude model judge split and then we will go through the How I AI benchmark and the 5 review and we'll see how Opus 5 performs on a couple key tasks. Okay, so quick reminder of how we run the how I AI benchmark. I run it against several tasks. PRD creation, prototype creation, wireframe creation, bug triage, and agentic coding.
治。我觉得 OpenAI 那边的解法就是索性变成:我们只讲 bullet point,我们讲产品经理的话,非常直给。Claude 这边该怎么解我不知道,但我会很想看他们怎么办。话说回来,其实你也不一定非得去读那些内容,我对产出是很满意的。所以这点值得想想。好,接下来是 How I AI 的评测基准(bench),以及我们是怎么评、怎么跑的。这是一套七个模型——六七个模型的 benchmark。我先快速去打个分,因为我刚收到 benchmark 跑完的提醒,我得手动去打分。我们采用 70/30 的拆分:我自己的判断占 70%,模型评审占 30%。然后我们就过一遍 How I AI benchmark 和 Opus 5 评测,看看 Opus 5 在几个关键任务上表现如何。好,先快速回顾一下 How I AI benchmark 是怎么跑的:我拿它跑好几类任务——写 PRD、做原型、画线框图、bug triage(缺陷分诊),还有 agentic coding(让模型自主写代码)。
[17:33]
And the last one, oh yeah, is it an agent voice that I want to hang with? I do not think Opus 5 is going to do well here, but who knows cuz I test them blind. So, what we have tested are a couple GPT models, a couple Anthropic models, and one Gemini one thrown in there. As you see here, we have blind taste tests. I go through and see all the different versions. I give comments and scores like three out of five, not bad. You can see it's generated dozens and dozens of prototypes that we can click through. I've gone through all of them, put in all the notes. And then right now it's aggregating up the scores, and then we're going to look at 70% my opinion, my vibe check, 30% LLM as a judge. I like GPT 5.5 as a judge and because it's my podcast, I get to pick. So, that's what we use as a judge. And we will see if and what hits the top of the leaderboard and where Opus 5 sits. The eval is run. It is 70% my taste.
还有最后一项,对了,就是:这个 agent 的说话方式,是不是我愿意跟它待在一起的那种?我觉得 Opus 5 在这项上不会好看,但谁知道呢,因为我是盲测。我们测的是几个 GPT 模型、几个 Anthropic 模型,再扔进来一个 Gemini。你在这儿能看到,我们做的是盲品测试。我会一个个把所有版本都过一遍,写评语、打分,比如「五分给三分,还行」。你能看到它生成了几十上百个原型,都可以点进去看。我全部看完了,评语也全填进去了。现在它正在把分数汇总起来,接下来我们会按 70% 我的意见、我的 vibe check(体感判断),30% LLM as a judge(让大模型当评委)来算。我喜欢用 GPT-5.5 当评委,而且这是我的播客,我说了算。所以评委就用它。我们来看看谁能冲上榜首、Opus 5 又坐在哪个位置。eval 跑完了,其中 70% 是我的口味。
[18:34]
And I regret to inform you I love Claude Opus 5. Again, look, if I don't have to talk to the model, which I don't, this benchmark runs asynchronously. I like the output. So, surprising shocker turn of events, Clairevo, notable hater of working with Claude code sometimes because I don't like Claude slop, loves Opus 5. So, there you go. I'm telling you I keep it honest. I keep it honest. So, um again, I went through those things. We gave 70% my vibe score, 30% the AI as a judge. I was I was just a little bit more generous to Claude Opus 5 than the judge was. So, I'm pink. Um the judge is green. Every time I run this, the whatever model I choose designs it in a different way. We just That's how we keep it fun. Um so, the ordering is Opus 5, Sonnet 5 next, although I scored it really low. The judge scored it quite high. Um so, I might reorder that one. Then Mabu, um GPT Terra next, Fable really low. Um I scored it low and the judge scored it relatively low. Then Opus 4 A and poor poor sweet sweet Gemini 3 1 Pro just never never going to get it to do. So, come on, Google. We want We want to have a win for you.
我很遗憾地通知各位:我爱上 Claude Opus 5 了。你看,前提是我不用跟这个模型说话——而我确实不用,这套 benchmark 是异步跑的——我喜欢它的产出。所以剧情反转、大跌眼镜:Claire Vo,那个出了名有时候受不了跟 Claude Code 共事、受不了 Claude 废话的人,爱上了 Opus 5。就是这样,我说到做到,我一向说实话。好,那些我都过了一遍:70% 是我的 vibe 打分,30% 是 AI 评委。我对 Claude Opus 5 比评委稍微宽容那么一点。粉色是我,绿色是评委。每次我跑这套,我选的那个模型都会把页面设计成不一样的风格,我们就图个乐。排序是 Opus 5 第一,Sonnet 5 第二——虽然我给它打得很低,但评委给得挺高,所以这一名我可能会调一调。然后是 Mabu,接着是 GPT Terra,Fable 很低,我给得低、评委给得也相对低。再往后是 Opus 4.5,还有可怜可怜、乖乖巧巧的 Gemini 3.1 Pro,就是永远也扶不上墙。加把劲啊 Google,我们是真心想看你赢一次。
[19:59]
Okay. So, again, here are just some examples of different builds um that the different models did. You know, this Opus 5 1 I really liked. I liked this one from Soul. Um so, I did like a couple of them, but um the ones that I gave fives to The ones that gave fives to you were Opus 5 and GPT 5 6 Soul. So, the three ones where I said, "Wow, really nice. Ooh la la and wow, great." were all Opus front-end work. So, Anthropic, you've done it again. Claude, you sneaky tricky little fish. You may be neurotic, but when asked to do some pretty front-end design, you really did it. It's They're detailed. They're functional. They're interesting. They're polished. So, Opus did a great job. And then, of course, I love um the 5 6 model. So, I was pretty happy with 5 6 Soul and Terra for some designs. Um the ones that I hated, let's see. Kind of I'm a hater across the board. Opus 48 got a lot of hate. Sorry, you've been outclassed at this moment. Gemini 3.1 Pro, sweet summer child.
好,下面就是几个例子,不同模型做出来的东西。比如 Opus 5 做的这个我特别喜欢,Soul 做的这个我也喜欢。所以确实有那么几个我是喜欢的,但我给满分五分的,是 Opus 5 和 GPT-5.6 Soul。那三个让我脱口而出「哇,真好看」「哦哟」「哇,太棒了」的,全都是 Opus 做的前端。所以 Anthropic,你们又做到了。Claude 啊,你这条又贼又滑的小鱼,你也许神经质,但让你做点漂亮的前端设计,你是真能做出来。细节到位、功能能用、有意思、完成度高。Opus 干得非常漂亮。当然我也很喜欢 5.6 那个模型,5.6 Soul 和 Terra 的几个设计我都挺满意。至于我讨厌的那些,我看看……基本上我是全场开喷。Opus 4.5 挨了不少骂,抱歉了,你在这个时点已经被超越了。Gemini 3.1 Pro,我可爱的小天真。
[21:21]
I am I'm just sorry, babe, but you were just not good. Um and then some like thin wireframes. I think the wireframes just didn't do really great. So, you can see here across the board, whether it was a full build or a wireframe, I just scored Opus 5 really, really high. Um I I did score Soul pretty high as well. Sonnet was like really variable. Um there were a couple fours in there, but mostly across the board I wasn't that pleased with Sonnet. And so, it was just very interesting. And then you see here, you know, me and the AI judge were pretty well aligned on Opus. We actually had the narrowest band of scores between us. And we were most far apart on Gemini. The AI was not as mean to Gemini as I was. Uh and then we were nearer, nearer, nearer. Again, we we agreed mostly on Opus 5 and 56 Soul, though I did not um judge 56 Soul all of that favorably. And just like last little meta commentary, I had Opus make the website for this benchmark, and it made such a trash version um to start. I yelled at it. It said it's impossible to breathe, it has too much meta commentary. I'm going to show this on the podcast. This is so I'm sorry, you all. I just feel so judged, but have to show it. I say this is garbage. Also, it has no screenshots. So again, I find this model so tedious to work with directly. It is my most loathed loathed colleague, and yet it it does the best work. So, I don't know what this says. Maybe this model is meant for a gentle coding that I have nothing to do with. And so, it just runs in the background. It builds me beautiful things. I don't have to talk to it. It doesn't have to talk to me. We are just
我……真的很抱歉宝贝,你就是不行。另外还有几个特别单薄的线框图,我觉得线框图这块整体都没做得多好。所以你在这儿能看到,不管是完整实现还是线框图,我给 Opus 5 的分都非常非常高。Soul 我也给得挺高。Sonnet 就特别不稳定,中间有那么几个四分,但整体上我对 Sonnet 并不太满意。这点挺有意思的。然后你看这里,我和 AI 评委在 Opus 上意见相当一致,我俩的分差是最小的;分歧最大的是 Gemini——AI 对 Gemini 没我这么刻薄。再往后我们的分歧就越来越小了。总之在 Opus 5 和 5.6 Soul 上我们基本一致,尽管我对 5.6 Soul 的评价其实也没那么高。最后再补一点元吐槽:这套 benchmark 的网站是我让 Opus 做的,它一上来做出的版本简直是垃圾。我把它骂了一顿,我说这东西密得让人喘不过气,元评论(meta commentary)多到爆。我打算把这段放进播客里——真是对不住各位,我觉得自己被审判了,但我必须拿出来给你们看。我说这就是个垃圾,而且它连截图都没放。所以我再说一遍,我觉得直接跟这个模型打交道太累人了。它是我最嫌弃的同事,可它干出来的活儿偏偏又是最好的。我也不知道这说明了什么。也许这个模型天生就是为 agentic coding(让 agent 自己去写代码)准备的,跟我根本不用打照面:它在后台自己跑,给我搭出漂亮的东西,我不用跟它说话,它也不用跟我说话。我俩就是
[23:14]
like sworn enemies, or maybe even better, sworn frenemies. Um because the output is very, very high quality. It's just exasperating to work with. So, that is the very surprising and very honest, you all. I told you I was going to keep this honest. We're going to do it live. I did not know the scores before I started recording. Be very honest, very live, very surprising how I AI benchmark of the brand new Anthropic model, Opus 5. Uh this the TLDR is I I love it, I hate it. So, um despite my original complaints, I will be using Claude Opus 5 for front-end design, um for app design, for prototype typing, and I will I'll give it a shot. Uh we'll we'll figure out how to make it make it work for me. Again, thanks for joining another How I AI honest review of the latest models coming out of these great frontier labs. I cannot wait to hear what you think of Opus 5. Please tell me. I can't wait to see what you build, and we'll see you soon at How I AI.
不共戴天的仇人——或者说得好听点,是塑料仇敌(frenemies)。因为它的产出质量真的非常非常高,只是合作过程让人抓狂。所以这就是这期非常出人意料、也非常诚实的内容——我说过我会保持诚实的,我们就是现场直播地做:开录之前我完全不知道分数。非常诚实、非常即时、也非常意外的一期 How I AI benchmark,测的是 Anthropic 全新的模型 Opus 5。TLDR 就是:我又爱它又恨它。所以尽管我一开始那些吐槽都成立,我还是会用 Claude Opus 5 来做前端设计、做 App 设计、做原型,我会给它机会,我们会慢慢摸索怎么让它为我所用。再次感谢你收看 How I AI 对各家前沿实验室最新模型的诚实评测。我特别想听听你们对 Opus 5 的看法,一定要告诉我。也迫不及待想看你们能做出什么东西来。我们 How I AI 下期见。
[24:23]
Thanks so much for watching. If you enjoyed this show, please like and subscribe here on YouTube, or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at howiaipod.com. See you next time.
非常感谢收看。如果你喜欢这期节目,请在 YouTube 上点赞订阅,或者更好的是留言告诉我们你的想法。你也可以在 Apple Podcasts、Spotify 或你常用的播客 App 上找到这档播客。也欢迎给我们打分和写评价,这能帮更多人发现这个节目。所有往期节目和更多信息都在 howiaipod.com。下期见。