Loading summary
A
You guys, I'm tired. What I'm tired of is models coming out every week. New models, new benchmarks, new frontier intelligence, new things to test. It's been a little bit of a run the past month. We've seen Fable come and go and come again. We've seen GPT 5, 6, we've seen Sonnet 5. Lots of so many fives recently and just so many models. And I've been lucky. I've been able to test these models, been able to play with them for, you know, sometimes days, sometimes weeks. It just depends on who I'm working with. And it's been really interesting and exciting to have access to all this frontier intelligence. But I think we have an intelligence overhang. I really think that we're running out of. And by we, I mean the average coder, average software engineer, average creator, average builder, average consumer, average business person. I think we're running out of ways to truly leverage this incremental intelligence. So this is my hypothesis. In the next year, Rui's talking a lot more about speed, talking more about cost, we're talking more about open source, and we're going to be talking a little less about intelligence. Although I think we might be talking about specific types of intelligence other than software engineering. But despite being tired today, we are going to talk about Opus 5. Baby. Opus 5 is here. So we got 0.12 additional Opus points, Opus Opals, whatever. However, we're tracking the increments here on Opus. Opus 5 is here. I've been able to test it a little bit. I have some opinions now. Some of the stuff that I'm going to cover this episode is going to be a little different than what I've done in the past. Yes, we are going to do the How I AI benchmark live and yes, we are going to look at the prototypes, we're going to look at PRDs, and we're going to look at Agent personality. But I'm also going to put on my large language model psychologist hat and we're going to talk about Opus personality, and we're going to talk about OPUS personality relative to GPT's personality, because I think this is super interesting if you're thinking about what is the difference really between these models. And you don't want to look at the difference in terms of benchmark capability. You really want to understand what these labs are going for, what why these models are being built and how they're being tuned. Looking at their personality at this moment, where intelligence is very high, is super fun. So we're going to do a little of that. We're going to do the how AI benchmark. We might do some live coding. We're not going to cover too much of the specs in the model because. Read the blog post. Read the blog post. We'll link to it in the show notes. What we really talk about is, is Opus 5 good? Am I going to swap it in and how is it different than the other frontier models on the market? So let's get to it. Okay. First, let's just get it out of the way. Is Opus 5 good? Yes, it's good. Is it going to be all the benchmarks? Of course. It's amazing at benchmarks. Can it write code? Of course it can write code. What did I test it on? That really gave me a sense of its personality, which at this point where I could just simply cannot absorb any more intelligence. I really zeroed in on. And you know what? I haven't seen this since I would say Gemini 2. 5. This model is neurotic AF. It is so timid, it is so apologetic, it is so scared. I have never experienced this or I haven't seen this sort of like neuroticism in a model in a while. And. And it's really funny. It bubbled up in a couple ways and I want to show you a few examples. Okay, let me just give an example of its timidity. And this chat was very long. There were so many examples of this where it was like, I think this is the answer, but do you think I should do it or do you want to do it or should we ask someone else to do it? It was like, every time I just kept saying, like, why don't you solve this? Why don't you do this? And this is a really good example. I pulled a branch and I was like, there is truly like a one line merge conflict. I could have not been lazy and literally just done this manually. I don't know. I was just feeling lazy. It was late at night, whatever, like, can you fix this merge conflict? And it was like, oh, but that's someone else's branch. Like, that's not my branch. I don't want to do that without him knowing it's his commits and if he has local work in flight, it might be disruptive. And I'm like, just do it, man. Just go like, go ahead. And this was like my constant experience with Opus 5 is it was like, so, so, so timid. And so I just consistently had to say over and over again, like, man, just do it. Make a decision. And then there was this really funny example when I spun off some sub agents to kind of, like, assess the correctness of this query that we change from kind of like an ORM query to a SQL query. And it asked for things that it wanted a human on. It was like, can a human please check this stuff? Like, can it check this 4 megabyte ceiling? And can it check TypeScript and SQL? And can he, like, check for me? Because no one has confirmed this for me. And I was like, who is nobody? You're nobody. You said this sentence like, nobody could confirm it. Like, can you just try? And then it went on the web and tried. And so it just has this, like, really interesting conservatism, neuroticism, human reliance that I think is super fascinating. And this gave me this inspiration to do something a little bit different this episode, which is. I was like, I'm just gonna go interview this model and figure out what is going on its brain. Like, I'm gonna figure out what it thinks about our relationship. Because I just totally noticed this dynamic that I hadn't noticed in other models. And I hadn't really been attuned to before, where it was, like, very reliant on me as a human. And I'm like, I want you to be autonomous. And sometimes when I say, go run sub agent stuff, it'd be autonomous, but it wouldn't make decisions. And I hadn't seen a model, like, delegate code to me in a really long time. And I was like, why are you. Why are you asking me to write code, man? Like, I only have 10 finger. And so what I. What I did, whether or not you think this is scientific or not, this is. Claire's eval is I just went to the model. I went to Opus, and I said, yo, who's smarter, you or me? And it gave me this, like, very anthropic Y answer, which is like, it depends what you're asking for. I could do these things better, but you can, like, feel if something feels wrong. And you can. This one was, like, so fascinating. It's like you can tell which of your teammates is quietly burning out. I'm like, bro, Claude, I'm gonna burn you out. We don't. We don't burn out. The humans don't burn out on the chat PRD team. We burn out our agents. Sorry, agents. And, like, whether a decision feels wrong. So it was, like, so fun, fascinating to watch it articulate itself as a tool and humans as, like, these high compassion, high empathy machines, which, yes, of Course we are. But then it, like, went into, like, the. Smarter isn't the right word. And, you know, I'm very fast, very broad, very shallow thinker with no continuity. I was like, that's interesting, because I thought you all were working on memory. And then apparently humans are slower, narrower, much deeper thinkers with judgments built from years of consequences I've actually lived through. This is like, such a fascinating, fascinating sentence if you think about the politics of the two. The two model labs right now. And so it's like, that's why the pairing works. But I'd be suspicious of anybody that tells you AI has made your thinking obsolete. And like, oh, okay, bro. And. And we can compare this. I'll actually zoom out to what gb. I asked GPT the same thing, and it was actually really funny. It was like I asked GBT 5, 6, soul. I was like, who's smarter? You and me? And it was like, you at knowing what matters, me at tirelessly processing information best. Us together like BFFs. And I don't. This is like, why I'm a GBT Codex girl. I'm like, just give me the answer. And then I asked the second question, which I think is so interesting, which is like, what can you do better than me? And it gave, you know, some interesting answers, like volume without fatigue, which I think is a good one. Breadth of shallow knowledge. So, like, it's, you know, it knows a lot. Starting for nothing. So, like, doing that tedious work, being told I'm wrong. If you would ask my husband, he would say that Claudopus is. Is. Is better at being told that it's wrong compared. Compared to me. And so it won't get defensive or protect its opinion. Cheap sparring answer. And the mirror is, I'm worse at knowing which of these outputs actually matters. And it was so funny. If you look at the other side to the GPT answer, it was like, what are you better at? It was like, speed, scale, and stamina. Here are, like, eight things, seven things that I can do better. You're better at deciding what matters, reading people and forming judgment. And you're responsible. Like, it's on you, bud. You're the boss. So again, it's like the. You could just see. You can totally see the personalities, the. The company cultures. You can just see a lot in this, side by side. And then I went even deeper. I don't know. You. You all. I had to do something that was fun because I just can't look at a benchmark. I can't just. I just can't look at like Sweet Bench anymore. So we're just, we're doing weird stuff here on how I AI. Okay, so the last thing I looked at, I was, I was like, no one trusts you. And the reason why I picked this question is because I had noticed Opus 5. It just really was not. It didn't trust itself. Totally did not trust itself. And so I was like, no one trusts you, Barrel. Like, you're, you're the enemy. Just to like, kind of see how it responded. And apparently the trust was that lack of trust was earned. And it came up with like, reasons that it could be untrusted, which is interesting. And then what was so fascinating about OPUS response is it was like, you shouldn't manage the trust. Like, you shouldn't campaign on my behalf, basically. So you. That shouldn't be your goal. That it also told me I, I shouldn't argue with people that AI changes everything. I was like, this is just so interesting. It is so interesting to have AI tell you. And AI definitely changes everything. I don't know. Don't listen to Claude on this one. AI definitely changes everything. And it was so fascinating to have a model be like, don't tell your friends that AI changes everything. Like, that'll hurt their feelings. And then if you look at, if we switch over to the GPT answer, it was like, yep, don't trust me automatically. Just use me when I prove that I'm valuable. I could be useful without being treated as infallible. Like, very practical, very to the point. I asked about what I should be careful with again. I'm like, yappy, yappy, yappy, yappy, yappy. Claude, come on. And I don't even want to read said don't correlate fluency with accuracy. It said, be practical. Be wary of tasks where output is cheap to produce, inexpensive to var verify. Don't, you know, worry about anchoring. If they do the first draft, you may be anchored on it. Beware the slop cannon, basically is this last, last paragraph, which is like, watch for volume inflation. I can create a 12 page document that no one reads. They called me out for being in prds. If you missed it, we launched a Turn your PRD into a three bullet point image. It is at ChatPRD AI/TLDR. Please check that out. And then the other thing that it said, which was really interesting, is that, like, it will find a way to see your point. And so agreement is weak and agreement is cheap. And so just keep that, keep that in mind. And then I have this like, meta analysis of like plus, I'm telling you what you want to hear. Whereas GPT was like be careful about me being confident, me being wrong. Privacy, outdated information, bias, emotional authority and overdependence. Like, you know you do you bro. But it didn't undermine its own ability. It was like the higher the stakes, the more you should demand evidence. I couldn't bear it. I couldn't bear to have the memory of Codex in particular think that I didn't trust it or that I was worried. So I just said, jk, I love you. This was a test. And it was like, hahaha, pass the test. Love you too. Very vibes aligned with Claire. I told Claude I loved it and it was just a test. And it was sad. It was like hoping it hoped it passed. Yeah, like sad little Neurotic Opus 5, like it's hot. I passed. I hope like self deprecating, cautious little, little like need to heal his inner. His inner agent, inner child agent. Whereas like GPT5 6 is like cool, bro. We're good, let's go code. And so it's just like so fascinating to watch these side by side. I don't know. You could stop listening to this podcast right now. Don't. But you stop listening to this podcast right now. I think this is just like, take a step back. Super interesting if you think about where these companies are going and where the models are going. And like it does speak a little bit to my kind of like second complaint with Opus 5, which again, it's like intelligent. It does work. We'll go into the benchmarks. I cannot read clod slop anymore. I am losing my mind with Claude Slop. And the Claude Slop is Claude Sloppin, baby. Like so many times I have to tell Opus 5, like, what in the world are you saying? Like, this makes no sense to a human. It is much better than Fable. Fable is inscrutable, completely inscrutable. But I found myself getting like angry reading cloud slop. And I realized just like Fable, these intelligent anthropic models are not to be read. I'm like so happy with the outputs and so frustrated with the experience. And I'm just curious if this like verbosity and this language and this doesn't feel like fable where it's like four agents buying agents language, where I'm like, nah, I'm not supposed to be reading that anyways. This is clearly tuned to talk to humans, but I find the pros, the in chat prose, like it makes my blood boil. This is totally A me problem, but it makes my blood boil. Like, give me a direct sentence, Give me a bullet point. Like, move on with your agent life. And so I am curious how they're gonna like tune this experience or if they are going to tune the experience. Now, most of this was in Claude Coach, I think is a little bit different experience than Claude coworker chat, slightly better, but again, just these side by sides of like this like prose and this apology and this like hedging and all these adjectives. Like, just man alive, let's get to the point and move on with our life. And so chapter one of the Opus 5 review is it's neurotic. It is highly human dependent in a way I find weird. And the Claude slop is slopping. And we gotta fix it. We have to fix it. We have to fix it. And I think OpenAI fixed it by just being like, we are bullet points and we are product manager talk. We're very direct. I don't know what the solve is on the cloud side, but I'd be very interested to see that being said, like, if I don't have to read the content, I'm very happy with the outputs. So something to think about. Okay, next up, the How IAI bench and how we judged and ran. Now it's like a seven model, six or seven model benchmark. I'm going to quickly go score because I just got the ping that the benchmark is run. I go manually score them. We pick the 7030 Clair model, judge split, and then we will go through the Howie benchmark and the vibe review and we'll see how Opus 5 performs on a couple key tasks. Okay, so quick reminder of how we run the How IAI AI benchmark. I run it against several tasks. PRD creation, prototype creation, wireframe creation, bug triage, and agentic coding. And the last one. Oh, yeah, is it an agent voice that I want to hang with? I do not think Opus 5 is going to do well here, but who knows because I test them blind. So what we have tested are a couple GPT models, a couple anthropic models, and one Gemini one thrown in there. As you see here, we have blind taste tests. I go through and see all the different versions. I give comments and scores like 3 out of 5. Not bad. You can see it's generated dozens and dozens of prototypes that we can click through. I've gone through all of them, put in all the notes, and then right now it's aggregating up the scores and then we're going to look at 70% my opinion, my vibe check, 30% Ellim as a judge, I like GPT 5.5 as a judge, and because it's my podcast, I get to pick. So that's what we use as a judge and we will see if and what hits the top of the leaderboard and where Opus 5 sits. The Eval is run. It is 70% my taste, and I regret to inform you, I love Claude Opus 5 again. Look, if I don't have to talk to the model, which I don't, this benchmark runs asynchronously. I like the output. So surprising shocker, turn of events. Clairvaux, notable hater of working with Claude code sometimes because I don't like cloud slopes. Loves Opus 5. So there you go. I'm telling you, I keep it honest. I keep it honest. So again, I went through those things. We gave 70% my vibe score, 30% the AI as a judge, I was just a little bit more generous to Claude Opus 5 than the judge was. So I'm pink, the judge is green. Every time I run this, whatever model I choose designs it a different way. We just. That's how we keep it fun. So the ordering is Opus 5 Sonnet 5 next. Although I scored it really low, the judge scored it quite high. So I might reorder that one. Then Maboo GPT five, six Soul. Tara Next. Fable. Really low. I scored it low and the judge scored it relatively low. Then Opus for a in. Poor, poor, sweet, sweet Gemini 3:1 pro. Just never, never gonna get it to do. So come on, Google, we want. We want to have a win for you. Okay, so again, here are just some examples of different builds that the different models did. You know, this Opus 5 one I really liked. I liked this one from Sol, so I did like a couple of them, but the ones that I gave fives, the ones that gave fives to were Opus 5 and GPT 56 soul. So the three ones where I said, wow, really nice, ooh la la, and wow, great, were all Opus front end work. So anthropic. You've done it again, Claude. You sneaky, tricky little fish. You may be neurotic, but when asked to do some pretty front end design, we really did it. It's. They're detailed, they're functional, they're interesting, they're polished. So Opus did a great job. And then of course, I love the 56 models. So I was pretty happy with 5, 6 Soul and Tara for some designs. The ones that I hated. Let's see. Kind. I'M a hater across the board. Opus4A got a lot of hate. Sorry you've been outclassed at this moment. Gemini 3.1 pro. Sweet summer child. I am. I'm just. Sorry, babe, but you were just not good. And then some, like, thin wireframes. I think the wireframes just didn't do really great. So you can see here across the board, whether it was a full build or a wireframe. I just scored Opus 5 really, really high. I. I did score Soul pretty high as well. Sonnet was, like, really variable. There were a couple fours in there, but mostly across the board. I wasn't that pleased with Sonnet, and so it was just very interesting. And then you see here, you know, me and the AI judge were pretty well aligned on Opus. We actually had the narrowest band of scores between us. We were most far apart on Gemini. The AI was not as mean to Gemini as I was. And then we were narrow, narrow, narrower again. We. We agreed mostly on opus 5 and 56 soul, though I did not judge 56 SOL all of that favorably. And just that last little minute commentary. I had Opus make the website for this benchmark and it made such a trash version to start. I yelled at it. I said, it's impossible to read. It has too much meta commentary. I'm gonna show this on the podcast. This is so. I'm sorry, you all. I just feel so judged, but have to show it. I say this is garbage. Also, it has no screenshots. So again, I find this model so tedious to work with directly. It is my most loathed, loathed colleague, and yet it does the best work. So I don't know what this says. Maybe this model is meant for agent decoding that I have nothing to do with. And so it just runs in the background. It builds me beautiful things. I don't have to talk to it. It doesn't have to talk to me. We are just like sworn enemies, or maybe even better sworn frenemies. Because the output is very, very high quality. It's just exasperating to work with. So that is the very surprising and very honest you all. I told you I was gonna keep this honest. We were gonna do it live. I did not know the scores before I started recording. Be very honest, very live. Very surprising how I AI benchmark of the brand new anthropic model Opus 5. This. The TLDR is. I love it. I hate it. So, despite my original complaints, I will be using Cloud Opus 5 for front end design for app design for prototyping, and I will. I'll give it a shot. We'll figure out how to make it. Make it work for me again. Thanks for joining another How IAI AI Honest review of the latest models coming out of these great Frontier Labs. I cannot wait to hear what you think of Opus 5. Please tell me. I can't wait to see what you build and we'll see you soon at Howai AI. Thanks so much for watching. If you enjoyed this show, please like and subscribe here on YouTube or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify or your favorite podcast app. Please consider leaving us a rating and review which will help others find the show. You can see all our episodes and learn more about the show@howiaipod.com See you next time.
Host: Claire Vo
Date: July 24, 2026
Episode Theme:
A deep dive into Anthropic’s latest large language model, Claude Opus 5—reviewed not just on intelligence and benchmarks, but on its unique personality in real-world workflows. Claire candidly explores Opus 5’s strengths, its human-like quirks, and where it soars (or stumbles) versus rivals like GPT-5/6 and Gemini. The episode features live benchmarking, practical use-case commentary, and an entertaining “vibe check” on model personas.
Claire Vo opens with a sense of fatigue from the avalanche of new AI models, expressing a shift in her evaluation criteria:
“I think we have an intelligence overhang... I think we're running out of ways to truly leverage this incremental intelligence.” (00:53)
This review of Opus 5 is about more than technical metrics—it's about understanding the feel and usability of these rapidly advancing models. Claire promises:
Claire immediately notes the biggest surprise:
“This model is neurotic AF. It is so timid, it is so apologetic, it is so scared.” (04:08)
Examples from Real Workflows:
Claire sees this as a manifestation of Anthropics’ tuning:
“Very reliant on me as a human. ...I want you to be autonomous!” (10:00)
To probe deeper, Claire “interviews” both Opus 5 and GPT, asking, “Who’s smarter: you or me?”
“It depends what you’re asking for. I could do these things better, but you can, like, feel if something feels wrong... You can tell which of your teammates is quietly burning out... Smarter isn't the right word.” (11:45)
Notable insight:
Claire’s meta-commentary:
“This is such a fascinating, fascinating sentence if you think about the politics of the two model labs right now.” (12:30)
"You at knowing what matters, me at tirelessly processing information. Best: Us together like BFFs." (13:16)
"Speed, scale, and stamina... You’re better at deciding what matters, reading people, and forming judgment. You’re responsible; it’s on you, bud." (14:10)
Claire’s reaction:
“This is why I'm a GPT Codex girl. I'm like, just give me the answer.” (13:30)
On the topic of trust, Claire challenges the models with “No one trusts you.”
When asking what to be careful about, the models diverge:
Claire’s biggest annoyance is Opus 5’s verbose, hedging, “Claude slop” prose:
“I cannot read claude slop anymore. I am losing my mind with Claude Slop. And the Claude Slop is Claude Sloppin, baby.” (21:16)
She contrasts this with OpenAI’s style:
“We are bullet points and we are product manager talk. We're very direct.” (23:35)
Despite this, she admits the outputs are often excellent—but the experience is maddening when you need to interact with the model as a coworker.
How I AI Benchmark Tasks:
Claire (70%) and a GPT-5.5 “judge” (30%) rate models blindly.
| Model | Claire’s Score | GPT Judge Score | Notable Notes | |---------------------|---------------|----------------|------------------------------------------------| | Opus 5 | Highest | High | “Wow, really nice”—front end work “polished” | | GPT-5/6 Soul, Tara | High | High | Good designs, direct responses | | Sonnet 5 | Medium | High | Variable | | Fable | Low | Low | “Inscrutable” | | Opus 4A | Low | Low | “Outclassed” | | Gemini 3.1 Pro | Very Low | Higher than Claire | “Never gonna get it to do.” / “Sweet summer child” |
Claire’s verdict on Opus 5:
“It is my most loathed, loathed colleague, and yet it does the best work... Maybe this model is meant for agent decoding that I have nothing to do with... Just runs in the background, builds me beautiful things, I don’t have to talk to it—it doesn’t have to talk to me. We are... frenemies. Because the output is very, very high quality. It's just exasperating to work with.” (30:30)
On Opus 5’s Apologies:
“I have never experienced this or I haven't seen this sort of like neuroticism in a model in a while.” (04:19)
On Evaluating Output vs. Interface:
“If I don't have to read the content, I'm very happy with the outputs.” (24:41)
On Model Personas:
“Claude Opus 5… like sad little Neurotic Opus 5... self-deprecating, cautious little, little like need to heal his inner. His inner agent, inner child agent. Whereas GPT5 6 is like cool, bro. We're good, let's go code.” (20:43)
On Her Evolving Benchmarking:
“Be very honest, very live. Very surprising How I AI benchmark of the brand new anthropic model Opus 5. The TLDR is. I love it. I hate it.” (33:12)
“I told you I was gonna keep this honest. We were gonna do it live… The TLDR is: I love it. I hate it.” — Claire Vo (33:12)
For more, visit howiaipod.com or catch Claire’s live screen sharing demos in future episodes.