Loading summary
John
I need some soundboard. Here we go. Yes. Today on tvpn, we're talking about model Mayhem. Everyone's launching new models. Slow summer, but not for the AI race. You got Xai unveiling Grok 4.5, the first model built specifically for coding and AI agents developing collaboration with Cursor. Talked about it a little bit yesterday, but we have some more benchmarks, some more discussion on the timeline about where this model fits in on the Pareto frontier. Also why it might be outperforming so well on Cursor Bench. Lots of debates there. Metta announced Muse Spark, a new agentic coding model with Mark Zuckerberg returning to X for the first time in basically a decade. Three years ago he posted one joke post about launching threads, but he has not been an active user. But the AI vortex sucked him in and he's got a post.
Ben
Oh, I think he's an active user, John.
John
You think so?
Ben
He's just not an active poster.
John
He's just not an active.
Ben
He's not an alert contributor.
John
You're calling him a lurker?
Ben
I'm calling him a lurker.
John
Calling him a lurk.
Ben
I'm calling him a lurker. I think he's absolutely glued.
John
You think so?
Ben
I think so.
John
You really think so?
Ben
I think so.
John
I feel like, I don't know, so busy, so much other stuff going on. I feel like he. I feel like most people, the busiest
Ben
people I know are not active on X. Yeah, but they, they are on X a lot sometimes.
John
But there are. There's a different class of person. You can just quiz them, screenshots come to them via Slack or via text message because they have a team that's monitoring the timeline and then is delivered. This is the important, this is the important stuff.
Ben
They're calling him Mark Lerkerberg.
John
But the other big news, OpenAI just released GPT 5.6. Let's go. A new general purpose model with expanded coding and agent capabilities alongside GPT Live, which we talked about yesterday. A new real time interactive voice experience. Reactions are great to 5.6. Bunch of interesting details here. You had. People have been identifying that while there is a frontier and there are just a few companies that are actually on the frontier. The frontier is spiky and they have different flavors to them and different reasons to pull different tools off the shelf. People are drawing analogies between Fable 5 being some, you know, recluse genius and 5.6 being a, you know, collaborative co worker that you love chatting with or something like that.
Ben
I Said I don't know how else to describe it, but Fable 5 is like Kendrick on Good Kid Mad City and 5.6 Soul is like Chief Keef on Finally Rich.
John
Now it makes sense to me. Thank you.
Ben
So I just wanted to put it into 2010 Hip hop terminology really, really clear there.
John
Thanks for clearing that up.
Ben
I mean, the funny thing is that will be very explicit for like 100 people in the whole world. This one's for you.
John
The most interesting benchmark to me has always been Ark AGI V3. We've interviewed the team over there many times and had a lot of fun understanding what goes into that that benchmark. And 5.6 Seoul scored a massive 7.78%, which is tiny considering that the whole point of arc AGI is that a human should be able to get 100% on it and basically any human. So it is a true test of AGI in this sense of, you know, can you give this test to just actually anyone not, you know, the crazy math projects, the crazy hard programming projects, the hacking, all of that stuff is very economically valuable, of course. But there's a more interesting question where when there's less of a spiky frontier and there's just this question of what is something that anybody can do that AI can't? Because we've been searching for those and the ARC AGI team has done a fantastic job building out these puzzles that AI has historically struggled with Arc AGI one, the model sort of climbed. Two became a little bit more complicated and now three, we're starting to see glimpses of progress. Although 7.76% isn't 99%, we're nowhere near saturation, but it's still a huge jump. Opus 4.8 had 1.5%, so GPT 5.6 SOL is showing more generalization, more spatial reasoning, more puzzle solving abilities. So fun, fun stuff. The blog post is also very, very fun because it includes games. I'm a big fan of the GPT 5.6 launch games. I got immediately sucked into the to the Sailing mini game, which is very high fidelity but also delightful to actually play.
Ben
Should we play it?
John
Yes, we should definitely play it. Yeah. Salt Wind, you guys play it. I want production team to see what they can get. I think my time was 25 seconds.
Ben
And is this hosted on a site?
John
I think this is. I mean this is hosted on the OpenAI blog, but I think the idea is that you could vibe code this in the latest GPT 5.6 in the app in in ChatGPT and then deploy it and have someone. Are you trimming the sails appropriately because it looks like you're losing speed, you're losing wind. It's not working. I'm gonna smoke you. I got 25 seconds. Wow. Amateur hour over here. Look at this boost. Yeah. Yeah. Well, the whole game, which you probably missed, is that there is a little bar there where you have to trim the sails to be in the sweet spot of the wind while you're turning. So as you turn, see the bar. There's a recommendation for where you put the sails. You gotta keep that in. See, it's moving over. You got to. You got to press the down ass. Yeah, exactly. To keep trimming those sails while you steer the ship. This stuff is very, very fun.
Ben
One interesting data point from the live stream, which was just an hour ago they said already Sol has been transforming our research program. As one example, GPT 5.6 Seoul autonomously post trained 5.6 Luna.
John
Yeah, that's fair.
Ben
A lot of people are having fun with that. Dylan Field says a lot of people want to compare Fable vs 5.6 SOL. This is a mistake. They're apples and oranges. Despite all the research achievements, we are still very, very early in exploring the tech tree for model training.
John
Cool. Sorry, I'm just getting set up again. Oh yes, I do think that. Didn't Dylan Abrascado write something about this? What was the essay he wrote about interactive memes and this idea of like generative AI enabling these vibe coded mini games. Like we've been seeing a bunch of them with like the copy bear simulator, the coconut simulator, where it's something that's just a joke, that's funny for like a few people. But. And normally you would instantiate that in a, in a tweet. Or maybe if you were getting really crazy you'd do a Photoshop edit of a meme. But now you can go and create a full mini game, something that runs in the browser and soon something that runs in Unreal Engine and can actually be distributed on the Steam store. We are already seeing that with the data center simulators and all these funny simulator games that are going on Steam. All the advances in the coding model certainly speeds up the ability to actually deliver polished software. I'm particularly excited for like.
Ben
Yeah, Dylan's title was the future of entertainment is interactive. Yes, yes, but, but yeah, that's part of what I honestly love about AI is there's a lot of things you can make now that never would have made sense to make because they would have taken you four days and it was good for like a small laugh. Now you can do it in four minutes and it's just fun.
John
Yeah, I think there's going to be, there's. If you have some sort of like small custom, some sort of custom functionality in your business, it feels like there's
Ben
this the David Senra simulator.
John
Why is this David Senra Late nights
Ben
in a Miami abandoned apartment complex rooms in 2015 just recording podcasts and reading.
John
This is very creepy. Like horror backrooms, liminal space game.
Ben
Stanley Tang, co founder and CPO over at DoorDash says I have an insane magic trick that so far none of the models can figure out, including Mythos. It's a bullet bulletproof trick that I've shown to 100 plus people including magicians that couldn't figure it out. It's not anywhere on the Internet. Only way to know it is through first principles reasoning. Told everyone I'll believe in AGI when it can crack this trick. Well GPT 5.6 just did how I want him to. I want him to actually open like well now. Okay, like give us now that now that a model cracked because I feel
John
like a lot of magic tricks are like sleight of hand. So is he uploading a video or. Or something like.
Ben
Well yeah. So John Palmer says I have a hilarious joke that so far none of the models think is funny. It's a bulletproof joke that I've told to 100 plus people including comedians and no one laughed. It's not anywhere on the Internet. Only way to know it's funny is a first principle sense of humor. Told everyone I'll believe in AGI when it tells me a joke. The joke is funny. Well, 5.6 just did.
John
Huge, huge news. Huge news. Huge GPT 5.6 is a Porsche Fables like warp drive. I had a different experience. If the fable is an F1 car 5.6 SOL at ultra as a Tesla Model X Plaid. Does it find things that Fable misses during plannings and coding? Yes, most of the time. But for the hardest problems, does Fable routinely find things that 5.6 doesn't also yes, some of the time is 5.6 way faster and affordable? Yes. With an unlimited token budget. What am I currently using 95 plus percent of the time? GPT 5.6 from SIKI Chen. So interesting. Take that the the parade of Frontier is alive and well and everyone's duking it out for their slice of the the AI opportunity. Very interesting seeing how the how the market share is shifting while During a time of acceleration, you have multiple companies that are growing revenues, even accelerating revenues while market share is declining because the overall market's growing so fast that, that if you're only growing at 300% and someone else is growing 400%, you're losing market share. But you have like one of the greatest businesses by modern metrics. Very, very interesting dynamics in AI.
Ben
It's also funny because yesterday with Ben Thompson you were like some slow summer and then in the span of 24 hours you get rock 4.5muse 1.1.
John
Yeah, I mean this isn't as dramatic as the AI talent wars.
Ben
It's not as dramatic as rippling deal.
John
Yeah, yeah. This is, this is new technology and there's only so much to, there's only so much of a take to be given around these things. Although AI 2040 launched today, the sequel to AI 2027, that's something that's more of a thought provoking piece that you can debate and interrogate and talk through. I'm sure we'll go through some of it because they pose a couple interesting ideas of where AI might go and where they want it to go and how they want the industry to develop. Sort of advocating for a slowdown generally. But it's an, it's an interesting way they puzzle piece all the different geopolitical chips on the table. Of course people are joking about the lead is widening because the anthropic and OpenAI version numbers over time GPT6 is predicted and it is the model numbering. We were talking about this this morning that the numbers, they sort of don't mean anything anymore. Do the model numbers mean anything in particular? It used to be the model number was the pre train and then the version number was the post train. But then that sort of got flipped around and now it's just like are you, do you feel like you're competing at a four class or a five class? So I wouldn't be surprised if we saw like Muse Spark not release Muspark 2 but Musespark 6 or 5 and jump straight. I mean Samsung wound up doing this where they jumped to the year like sort of like the car manufacturers where you know there's a 5 series BMW but then there's also just the 2027 because that's the actual model year that's relevant.
Ben
2027, 5 series.
John
Yeah. Which is sort of odd. And we're sort of like duking it out between those. Do you have.
Unidentified Guest
Yeah, I mean I think post reasoning models you just have like a different way to scale the models besides just pre training.
John
Yeah.
Unidentified Guest
So it's hard to bake that all into one number that, like, is evocative of both those two ways.
John
Yeah. So the number is becoming closer to the year in the second decade of the 21st century, basically. It's just like, is this on the frontier in 2026? You'll probably see a 6 by the end of the year in front of the models that are leading in the year 2026. Something like that. I'm very interested with Google strategy because the rumor is that 3.5 Pro will be coming out this next week, I believe. But it's very odd going into the Gemini app right now and seeing that there's 3.5 flash, but then you have to go back to 3.1 Pro. I think 3.1 Pro is the most advanced model, but they default you to 3.1 flashlight and I would expect them to jump just forward to four. But I think that they're going to do 3.5 Pro. But it's been a little bit of a slower cycle there as silly. I mean, obviously all these numbers don't really mean anything. They're marketing terms. But I still think they do actually stick in people's mind and so there should be some strategy around them. Mark Zuckerberg is on a press tour. He's talking to the legacy media for the first time in a long time. Andrew Bosworth, the CTO of Meta, also did an interview with the head of the Atlantic, dug into some of the launches around the glasses, and then also had a whole discussion in that podcast around the goals of the keystroke logging thing. It was. It was interesting. I mean, it was framed as like, you know, like a tough interview around surveillance in the workplace. And it certainly. The headlines were very scary. I don't know where I sit on it because I've. I kind of always assume that everything you do at work is logged in the sense that, like, if you're on a work computer and every webpage you visit is going through the network and monitored for traffic and security purposes, and all the code you write and all the emails you write and all the documents are stored in a shared document, it doesn't seem that crazy to go to keystrokes because everything is already so monitored. But he was framing it as more of an experiment, something that they weren't sure was going to pan out, something that they allowed everyone in, everyone at Meta. So there were certain sections of the workforce that were by default opted out. So anyone who is working on Confidential or sensitive information was opted out of that program by default. He said he himself, Andrew Bosworth was opted out of that program because he has a bunch of legal holds, because they're getting sued all the time, so they can't be recording everything I guess that he's doing because then that would be admissible in court. And so all of a sudden the lawyer who's suing him would say, okay, great. In the email you said, you know, we don't want to do this, but
Ben
before you do that, let's see your writing process.
John
Exactly, yeah, let's see what sentence you typed and then deleted. Like what word did you use before? Minimal impact. Did you say medium impact or whatever? So he was opted out. And apparently I think all of the meta employees who were part of that program were able to turn it off indefinitely. Like you could toggle it on and off. And the idea was that they wanted to collect information on how work plays out over a 12 to 18 month period. And they couldn't get that from any sort of data labeler because they needed to have very high skilled workers actually chopping wood on projects for a long, long time to see how projects go from start to finish. So basically like, how do you compact the longest possible rollout, not just a, like a single chain of code, but an actual series of meetings and decisions and trade offs and everything that goes into making a decision in a white collar workplace? Like how do you actually reason through all of that? It's hard to distill that from just oh well, the code got written this way so that's the right way to write the code. The code might have gotten written that way because a lawyer said, hey, oh, we have to do this. And then the marketer said, oh well, you know, we have an activation with this person so we need to integrate it this way. And then the business people came in and said, oh well, like the margins will be better if we write it this way. And so it's not entirely first principles software engineering all the time when you're actually building real products. So interesting to see him sort of step into the, you know, a tough interview and sort of lay out his side of the story. But Mark Zuckerberg is in Bloomberg today pledging aggressive pricing with Meadows first pay to use AI, which is a funny framing for just an API for a model, but that's the way Bloomberg put it. In a crowded market for AI tools, Mark Zuckerberg wants to win on price. Meta Platforms unveiled a version of its most advanced artificial intelligence model, Muse Spark 1.1 that includes a new paid tier for developers, marking the first time Meta has charged businesses for access to its models and providing a new revenue stream. It'll be among the most affordable options on the market, Zuckerberg said in an interview ahead of the release. Quote, since this is not an open source model, this is I think the first time that we're doing a real serious API and the pricing is going to be very aggressive and attractive. Makes sense. I mean they own the data centers, they're very efficient at building data centers. They should be able to serve a model efficiently. The new model standout improvement is in its agentic capabilities. The Meta Chief Executive Executive Officer said agents are a big theme of AI this year, with the label applied to systems that can complete multi step tasks on behalf of the user. Zuckerberg described Musespark 1.1 as having, quote, state of the art or very close to it, agentic reasoning and tool use. The model is also greatly improved when it comes to coding and Meta employees are using it internally to build products and features for various apps.
Ben
Yeah, my big question is how, how quickly do they move all of their internal workloads onto their own models? They're buying, they're getting access to models through Google, anthropic and OpenAI. Yeah, I think that a lot of companies will look to Meta's own actions as a way to basically validate whether or not they should be using this model themselves. Right. Because it was just within the last month that Google had said like, hey, we don't have capacity, we don't have enough capacity for all of Meta's demand for our models.
John
Yeah.
Ben
And so, yeah, they can't get enough AI elsewhere, at least from some providers. And so how much of their workloads will they be able to run themselves is a big question.
John
Yeah, Meta was one of the first companies to sort of reportedly be token maxing and have a leaderboard and all of that. If you have your own model and your own data centers, the incentive to token max is much, much higher because you're just paying the electricity on the cards that you're already appreciating. So you should sort of lean a little bit back into that. Not that you want to be fully token maxing, but you do want your employees using the tools that you've built as efficiently and as effectively as possible. It's just way cheaper to explore when you're not paying margin on another closed source model and you're not paying anything else and you're actually improving the model. So it makes a lot of sense for them to roll this out broadly. The interesting take that Ben Thompson had, which we didn't get to yesterday because we ended up spending the whole interview talking about Xbox. But the interesting dynamic is that when you are willing to sell API access, you're willing to sell compute directly and then you're also using your own tool internally. It creates this economic incentive internally that you have an incentive to always go with the most profitable, the most, the most economically efficient outcome that can be very good for business, very good for the investments that they made. The trick is that you can wind up in a little bit of a situation where your business team or your enterprise sales team goes and sells all your compute capacity or all your chips and then internally your team is frustrated that they're not making enough progress. So there's a little bit of a dance there. But in general it's, it's a forcing function on the internal use of their tools to say, hey wait, why is someone willing to pay five times as much than what we're willing with the value that we're creating here? We spent a billion dollars on energy consuming our own LLM and someone showed up and said wait, we'd pay you 5 billion for that same compute power to run a different model and do a different task. It's like, why is their model not economically valuable internally? That would be the question. The flip side is that they do have low cost so they should be able to say oh yeah, we actually did, yeah, we inferenced Muspark 1.1 internally and we improved the AD model and boom, we made a bunch of money.
Ben
And these are the same trade offs and decisions that every lab is having to make is how much compute do we allocate towards research, towards internal use, towards to the API, subscriptions, to free plans, et cetera.
John
Yeah, there was that funny semi analysis deep dive into anthropics forecast and in there, I mean some staggering numbers. Really, really optimistic. But the flip side was who is it Ed Zitron was taking shots at
Ben
the fact that they had ebtit.
John
EBTIT earnings before training, training interest, no trading inference and, and everything. No earnings before training, interest and taxes. And what was odd about it was that Ed Zitter was saying it's like the new community adjusted EBITDA and it is always odd when a new non GAAP metric pops up. In this case I think it makes a lot of sense because training runs do fit a depreciation profile. It's a little bit different. I don't know why you wouldn't just put it in in depreciation though. Like just figure out how to account for training runs through a depreciation schedule and then maybe it's like a non GAAP depreciation metric, but it's still in there instead of trying to get everyone up to speed on a different like sounding phrase entirely.
Ben
Yeah, I was looking back at Ben Thompson's earnings transcript or a script that he wrote for Mark Zuckerberg. He has a good segment on why AI matters. Ben writes Forgive the long preamble, but this is necessary context for me to properly explain why AI is so important to Meta and why I'm making the right choice to invest so heavily in both talent and infrastructure. And he goes on and on and on. But he says what I've come to realize as I've embraced our status as an entertainment provider and ad purveyor is that our nature as a digital business notwithstanding, we are remarkably well placed to thrive in an AI era. Remember we what we learned about humans. They are obsessed with other humans and they want to connect with them. That obsession and desire are only going to increase as we interact more and more with AI. AI is going to make our properties more essential, not less. Moreover, AI is a productivity tool, but productivity is not the end all, be all of the human experience. I've talked over the last year about building super intelligence that helps you get things done, but that's a business story. What we can do uniquely is give people the experiences they want, from connection to entertainment to shopping when they are off the clock. The fact that we are investing in AI but not selling solutions to businesses is actually one of our business biggest advantages. So of course this is just a sort of fan fiction for an earnings transcript. Meta is in fact selling to businesses now, but who knows over time how big will the API business be relative to how much value they can unlock across their broader business with all of their infrastructure?
John
Leave us 5 stars on Apple Podcasts and Spotify. Sign up for a newsletter tvpn.com and we will see you tomorrow at 11am sharp. Goodbye.
Hosts: John Coogan & Ben (co-host, Jordi Hays not present)
Date: July 9, 2026
Episode Focus:
The rapidly escalating AI model race: new product launches from OpenAI (GPT 5.6 "Soul"), Meta (Muse Spark 1.1), and Xai (Grok 4.5), with live benchmarking, competitive dynamics, and the implications of agentic models, API pricing, and model versioning.
This episode dives deep into the surge of new frontier AI models during the "slow summer" that isn’t so slow for the AI space, highlighting:
(00:00-01:35)
Notable Quote:
“Slow summer, but not for the AI race.” — John (00:03)
(01:39-06:17)
Notable Moment:
John explains the mechanics of the GPT 5.6 Sailing mini-game and challenges the production team to beat his time. (04:43-05:46)
(06:17-07:38)
(07:51-09:03)
(10:08-12:23)
Notable Quote:
“It used to be the model number was the pre train and then the version number was the post train. But then that sort of got flipped around and now it’s just like are you, do you feel like you’re competing at a four class or a five class?” — John (10:31)
(12:24-21:34)
(21:34-24:19)
Conversational, fast-moving, highly technical but laced with humor, inside jokes, and references to hip-hop, gaming, and Silicon Valley startup culture. Hosts pivot rapidly between product specifics, competitive strategy, and the business/economics of massive-scale AI.
This Diet TBPN episode captures the current AI gold rush, with OpenAI, Meta, and Xai launching major new models in rapid succession. The hosts break down the real advances (agentic coding, generalization, internal benchmarks), business strategies (pricing, API plays, internal dogfooding), and cultural impact (interactive memes, “vibe code” games, and the humor/hype around AGI progress). The fun, irreverent tone and spicy analogies make it clear: the AI “frontier” is more crowded—and weirder—than ever, and the next year’s model numbers might mean less than ever before.