Loading summary
Podcast Narrator
Foreign.
Andrei Karpathy
Hello and welcome to the Last Week in AI podcast where you can hear chat about what's going on with AI. As usual in this episode we will summarize and discuss some of last week's most interesting AI news. Also the week before we have unfortunately skipped a week due to scheduling conflicts, but we will cover everything relevant from a period I am one of your regular hosts, Andrei Karen. I studied AI in grad school and now work at the startup Astrocade.
Jeremy Howard
And everybody, what's up? My name is Jeremy, of course, I'm your other co host. I'm from Gladstone. AI do AI, national security, super intelligence, C type, end of the world stuff. So I sound a little thick right now by the way, which is related to the reason that we didn't record an episode last week which was that I was traveling and I got sick on the flight. There was a guy who was coughing up a lung next to me and anyway that's why I sound so weird right now. But Trip was really useful and yeah, hopefully be able to talk about a lot of this stuff soon. But a lot of conversations with like researchers at the Frontier Labs and folks on the, you know, the safety teams, the capability teams, all that kind of thing that I think bears quite a bit on the events of last week and the week before. We'll definitely be talking about a lot of that stuff with some of the inside view a little bit what I can share right now on those things but man, things are moving and it
Andrei Karpathy
has been slightly eventful.
Sponsor/Advertisement Voice
Two weeks.
Andrei Karpathy
I mean I guess it's not the most eventful you've had this year, but there's been some big stuff that we'll be touching on. As a quick preview, there's a few new models, nothing gigantic, but fairly meaningful. We'll start with then as usual, some funding stories and deals about compute and so on. Some major open source releases including Kimi K3 we haven't discussed about, so we'll be about talking, talking about that then policy safety. Of course we'll be talking about the recent hacking incident from OpenAI and a whole bunch of other stuff related to that. It's going to be a kind of policy safety heavy episode and then we'll round it out with some research and synthetic media and art. So it'll be a packed episode.
Sponsor/Advertisement Voice
This episode is brought to you by Outshift, Cisco's incubation engine. Today's AI engines operate in silos, limiting their true potential. We focus on building bigger, smarter models, but scaling up is just one approach to reach Superintelligence together. We need to do more, we need to scale out and we actually have a blueprint from 70,000 years ago. Humans didn't just get smarter individually, the cognitive revolution transformed society because we began sharing knowledge, goals and innovation, agents are now at the same inflection point. They can connect, but they can't think together. That's why our shirt by Cisco is building the Internet of Cognition, transforming AI from isolated systems into orchestrated superintelligence. By creating an open interoperable infrastructure, Outshift is enabling agents and humans to share intent, context and reasoning. The Cognitive Evolution for Agents is here. Explore Internet of cognition@outshift.com that's outshift.com this next sponsor isn't related to AI, but I've personally used them for years, so I'm happy to have their support. And it is Factor they make chef crafted dietitian designed ready to eat meals so you don't have to choose between real food and convenience. Both in grad school and as a
Andrei Karpathy
startup employee, I don't have a ton
Sponsor/Advertisement Voice
of time, so when I get home I'm tired and being able to prepare really quite a good meal without any effort has been fantastic. Their meals are ready in two minutes and require no prep and no cleanup. So even on the days your schedule is completely out of control, eating well is still achievable. There are over 175 banned ingredients, so every Factor Meal is designed around what supports a healthier lifestyle and nothing that doesn't. And that's with over 100 nutrient dense menu items to choose from every single week. 97% of users agree that Factor Meals help them live a healthier life. So you can feel confident that you're doing something good for yourself with Every meal. I've really enjoyed Factor and if this sounds good to you, maybe you should try it as well. Let's eat real head to factor meals.com lwai50off and use code lwai50off to get 50% off and one free breakfast item per box for one year while supplies last until 2-31-2026. That's code lwai50off@factorymeals.com, lwi50offactormeals.com See website for
Andrei Karpathy
more details and we'll go ahead and get into it, starting with tools and apps. And here we begin with anthropic releasing Claude Opus 5, which they say comes close to the capabilities of cloud fable 5 in many domains and is cheaper of course. So this is following up on the release of Fable 5 a little while ago, Fable being their new family of models that they didn't have before. And they also released Sonnet 5 either before around the same time as Opus 5. So we sort of caught up, presumably because these are distillations of Fable and Mythos. Right? So typically what you can predict with Anthropic is their big kind of best model, is their most compute heavy, most impressive model, which is Mythos right now. And these things like Fable, Opus, Sonnet are kind of derived from it to some extent where they try to extract out the intelligence at a lower price. So I think not a ton to say about this one beyond that. It's supposedly quite good and close to Fable 5, so pretty big jumps in the benchmarks relative to Opus 4. 8. The vibe check has been a bit mixed as far as I've seen. People, you know, have their usual sort of complaints about what the models are doing. And it's hard to know whether we just have high expectations now or in fact the models are getting stupid. There's also a new fast mode in Research Preview which offers you higher speeds at double the price.
Jeremy Howard
Amen on I think we're. We're losing track of what it is that we're looking for. Just because the waterline is rising so fast. People are getting used to incredible levels of capability. You're right. This is probably a distillate of Fable 5 or Mythos 5, then added safeties and all this stuff. There's obviously additionally post training that gets done to kind of further refine the character of the model after that. And so one of the key things that they highlight here is the differentiator of Opus 5 is supposed to be more emphasis on verification and judgment. So less kind of raw capability, but more sort of like double checking its work, making sure that what you're getting is actually correct. And so they give this example where like given a. It's given a drawing of a machine part but no way to view the original image. And then Opus 5 like rewrites its own computer vision pipeline to extract the geometry from the raw pixels and reconstruct the part. Basically the idea being like it's going to get you that raw data, the original to base its conclusion on so that it knows it's right no matter what. That's kind of like the vibe here also in alignment and safety. Kind of interesting. This is always the game. We live in a world where the US government decided to snap a chalk line at Mythos level and anything above Mythos level magically is. Is subject to or the de facto licensing regime that we have in the US and so in this case Anthropic is in a hurry to say that their model is close to mythos 5 at identifying software vulnerabilities, but it's less successful at developing exploits. Again, that's part of the post training. That's part of making sure that or and also just as well the pre training or other parts of the training process where they're avoiding explicitly training it on cyber tasks in that way. And so the goal here is really to position it as this is a really intelligent model that it is okay for us to Release. Of course Fable 5 is okay, but it has additional safety over over Mythos. But you know, you're always now going to see that kind of background concern about this threshold, which is, I mean it's good. It would be just great to have a more principled approach to the stuff. One thing to note too is performance on Frontier Bench. So Frontier Bench is this new benchmark that we have. It was released just a few days ago and the same team behind Terminal Bench came out with it. It's this big community effort and it is basically just like a harder agentic environment. We keep needing more and more difficult agentic evals to be like, all right, but you know, we just saturated your Terminal Bench or whatever. Let's now move on to your Terminal Bench to Terminal Bench three and ultimately this. And there you go. So here we do know that this particular model, Opus 5 is outperforming all other models on a cost per task basis. And that's really where they're trying to differentiate. It is cost per task, not necessarily Frontier of Intelligence. It's near the frontier, but cheaper per token or perfect per unit of intelligence.
Andrei Karpathy
Let's say at that point it is slightly cheaper than GPT5.6 Sol, the biggest and best model from OpenAI and I think maybe indicative of like pricing becoming more of a concern for customers at the business front now that there is a competitor to Anthropic. With Codex and OpenAI being quite capable and very, very cheap alternatives for model usage, I have to wonder whether we are going to be seeing more kind of pricing pressure going on. Next up, some more model releases, this time from Google. Google DeepMind has released three new AI models. Gemini 3.6 Flash, Gemini 3.5 Flash Lite and Gemini 3.5 Flash Cyber. So per the Flash aspect, these are cheaper and faster than let's say more intelligent models. GPT 3.5 flash fly delivers 350 tokens per second at a rather cheap price. Gemini 3.6 flash is priced 1.5per million input tokens and 7.5 per million output tokens. So that's slightly cheaper than Sonnet 5 and not super cheap. And then there is of course the cybersecurity focused Gemini 3.5 Flash Cyber, which is integrated into this code mender agent that autonomously builds build exploit code to verify vulnerabilities in sandbox environments and then generate patches which they say has found problems in complex real pieces of software such as V8 JavaScript engine. So I think interesting to see a general movement towards cyber focus with not just the Mythos and Opus and so on. We've seen OpenAI release a cyber model now Google has released a cyber model and we'll be discussing Microsoft has also released a cyber model. So I think everyone is like, oh no, we gotta do something about this.
Jeremy Howard
I don't think it's sharing too much to say like people in the national security space and the frontier labs are really concerned about where cyber is going and this view that, you know, we're in the vultanpocalypse right now, right? We're getting all these low hanging fruit vulnerabilities being discovered and exploited. Assume we're going to get a Mythos class open source models sometime in the next, you know, certainly six months, maybe a bit less. At that point you're going to need an answer. And so all the lives are pre positioning for the moment when basically they're holding the world for ransom. I mean, you know, you need to use really, really good cyber models short your infrastructure or else like that is just going to be the case. As an aside, I'll just like casually drop the prediction here that we may see some pretty significant disruptive cyber attacks at massive scale. Not even just nation or their proxies, but literally just like disaffected young people or terrorist groups or whatever. That's just what happens when you open source that level of capability. It's just like that's what the math says. We'll see where that goes. But that's like the default assumption right now of a lot of the people in the space, both on national security and the Frontier Lab side. So that's part of what the positioning is here. If you don't have an answer to the cyber question, you know, you're going to be a lot less relevant in the next six months or so. They do. This is pretty impressive. I mean the Flash cyber model, which is maybe the one, at least it's the One I'm paying most attention to is performing on par with a lot of frontier agents at things like Cybergem. And that's an important benchmark. I mean it's a lot cheaper too, right? Really a fraction of the cost. And so this idea of how cyber plays out is always a really strong function of how much compute you have. You're a defender, you have a certain pile of test time compute. The attacker has a certain pile of test time computer. Can you invest more test on compute than an attacker? Shore up your infrastructure is the, is the question. I mean, you know, there are a lot of ways to answer it and you know, different ways to use test on compute and a question about how much leverage, like maybe there's an attacker advantage or a defender advantage. These are all open questions. But it's going to come down in some way, shape or form to that balance. And so the cheaper you can make these models, the cheaper you can make the tokens per unit of cyber intelligence, really the more value you're getting there. So that's an area where the cost of the tokens really, really matters. And that's why they're dabbling there. Yeah, the other launches are interesting, but kind of fall into this general category of like Google still not having a true frontier model. Like when we're thinking about the best models in the world, it's anthropic and it's OpenAI and there's just not really anyone else.
Andrei Karpathy
Yeah, Gemini Pro 3.3.1 used to be sort of in that race or at least near the frontier. There's not been a pro level model since February so they're now quite a bit behind. Nobody really using Gemin for like serious hard work and coding for instance. And I think it is an interesting indication where Google is at if they are focusing on Flash first because this immediately rolled out to everything. All their products, Google AI Studio, Android Studio, Gemini App, Gemini Enterprise Agent platform. So it kind of makes sense from a business perspective. Like they integrate Gemini into everything, including Google Docs and spreadsheets and AI mode. And at that point you have to have a faster and cheaper model, which is why they are really emphasizing flash first. They did say that Gemini 3.5 Pro is being currently tested and will be made available when ready. And they have begun the most ambitious pre training run yet for Gemini 4. So we've gotten indications that they're at least working on a Mythos level model and we've seen them kind of catch up before with Gemini. So I'm personally looking forward to what Gemini 4 will be like. Next up, another new model, but this time not for language, for images and videos. Black Forest Labs has launched Flux3 that is capable of generating images and 20 second videos with audio. So this is a multimodal frontier model trained to understand and generate images and these models while extending the architecture interestingly to robotic vision and action. So it's jointly trained across image, video and audio modalities altogether. It's their first public video generation model from Black Force Labs, which for some background hails back to some of the talent from Stable diffusion that made some of the first really impressive image generation models. And Flux has still been kind of one of the go to image generation models at the frontier. Flux Free video looks to be pretty impressive from what I've seen. They have some kind of human preference studies where they say it is preferred over Grok. Imagine video cling, V3 Pro, Runway gen, all at like 70%, 60%, 80%, whatever. People prefer this in terms of its outputs and these are from just kind of testing. It's still not fully rolled out. So very interesting. And the fact that it's now being adopted through Flux Mimic that is being developed with Mimic Robotics for actually making it kind of an action model, which I've seen kind of starting to be the case more. We've seen some other players like Runway starting to get into like the physical intelligent space of video and robotic control seem to have a lot in common.
Jeremy Howard
Yeah, this is an interesting announcement. This is kind of a couple things. One is a bunch of comparisons that like show pretty lopsided wins against you know, like Luma and Runway and all this stuff that don't really matter because nobody uses those models anymore. But there's this interesting comparison against Google's Gemini Omni Flash. 52% win rate against that. That's actually quite interesting. Like that's pretty impressive especially given the resources Google's been throwing at this stuff. And then another piece is so, so yes, like we keep pushing out the length of clips so that, you know, that's great but the challenge has sort of become coherence across clips across different shots. And that's the big thing that they're pushing here. Sort of multi shot sequences where the characters are consistent, the sort of, the physics is consistent and that's a big boon with this particular police. So you know, increasingly moving beyond. And once you get Those, you know, 20 second clips, 30 second clips, you can imagine that being a point where yeah, you know, one shot typically only lasts about that long as if I know how long a shot lasts in professional film, but whatever, you know, you can imagine that being the case. And so then, you know, maybe you care more about the switches between different frames. So yeah, kind of interesting and a new kind of metric to track.
Andrei Karpathy
And they do have, as before, variants of this that are open weight that also have an access to multimodal generation called flux free dev. So another kind of slightly big deal. I don't think we have an open weight model that has this new unified multimodal backbone, which by the way is relatively new. We've seen image generation and video generation for a while, but similar to kind of nano banana from last year, I think we're moving towards a place in video and audio generation where everything is put together instead of being cobbled together. And that is actually a pretty big deal in terms of the capabilities. Next, more of a product story. Meta is making its AI chatbot more like an assistant. So they're adding productivity features. There's a calendar integration, daily brief fix and apparently in depth research capabilities. Powered by Llama Spark 1.1. It can also browse Facebook Marketplace, search for restaurants, check your calendar and handle recurring tasks. So this is rolling out to the Meta AI app and it's going to be coming out to WhatsApp as well, which continues to mark a shift for Meta, which is like, you're making AI, are you just going to compete with all the other AI players? Like what, what are you going to be doing with this?
Jeremy Howard
I don't know. Yeah, I mean, I think this is partly a realization as well, that unless you're moving in the direction of productivity, you're just not going to squeeze all the juice out of these, these models that you can. Right. So, you know, think about the positioning of OpenAI relative to anthropic and the profit per token that Anthropic is able to rake in because of their commercial focus. It just, I mean they're, they're eating OpenAI's lunch. And so think about the. There's an extreme beyond OpenAI. We often think of OpenAI as the direct to consumer company, which isn't as true as it was six months ago. Certainly they've been making a lot of inroads in B2B, but at the far end of the spectrum, in the other direction is Meta, that they are straight consumer. Right. Like those tokens are going just to tickle your limbic system. They're not actually going to like actually move big things in the real world or like make products. And so if you want to Ultimately get the the most bang for your buck, generate tokens that are actually valuable enough to make a good profit. You have to move into this direction. Not least to say if you have a a coherent long term view of where superintelligence goes, humans just aren't in the picture. Which means if you were optimizing for the value of the attention of human beings, which is what Met is currently doing, that value may drop precipitously as AI start to control more and more of the economy. And so you have to be in a position to actually do productive work and support agents in doing that. So you, you know, depending on how far you wanted to read this, you might read it that far. I know Zuck doesn't really seem to understand superintelligence, but certainly Alex Wang does. So it wouldn't be surprising if this was at least part of the thinking here.
Andrei Karpathy
Yeah, I will say I think it will be interesting to see where they go of this because you can go two ways. You can sort of go and try to make a Codex or a cowork competitor which is straight up just for work. Or they could shift into a sort of Open Claw type thing where this is an always on background agent which can do a bunch of stuff for you, including productivity things like briefings on your calendar, but also messaging and various things like that. And I think the open Claw space is sort of still up for grabs. Google hasn't rolled out their openclaw always online agent. They've said that they are going to. I forget what it's called. So I could see them like being potentially capable of competing on that front, not in the like coding or real office productivity side, but like personal productivity, you know, everyday productivity maybe. And one last product rollout. OpenAI is rolling out ChatGPT Health to everyone. So this is available to all US users aged 18 plus on web and iOS and it will allow you to connect medical records and health tracking data for the chatbot. So they are saying that this model can reason at levels better than clinician level and you can connect a whole bunch of stuff. I think this marks a shift where, you know, in the past if you were to talk about health stuff with these models there would very strongly caveat that you need to double check. And in general you should not have trusted these models with any sort of critical health concerns. ChatGPT Health potentially is at least OpenAI are making the case that this is something you can rely on onto applications and business. We begin with Ilya Suskever's Safe Superintelligence partners with Nvidia to scale its AI research. So that's kind of the gist of it. We have had SSI Safe Superintelligence be around for a couple of years. They raised 1 billion in founding in 2024 and 2 billion in 2025. So, you know, a lot of money, but not that much money. If you are saying you want to create super intelligence, you compare that to anthropic OpenAI. They have hundreds of billions. This is a few billion. So the narrative around this is that Ilya Sutskever's company has achieved sufficient research progress that it's time to scale up. And now to scale up, you need a bunch of compute. And so they're going to be partnering with Nvidia at a value of like some amount of billions. A bunch of billions. And they'll be increasing their compute by an order of magnitude.
Jeremy Howard
Yeah, there's some, some disagreement between different outlets about how much exactly has been raised, whether it's a 5 billion round or just in the billions or something like that. I think TechCrunch ran $5 billion story. So either way, the one thing everybody seems to agree is, number one, it's going to give Safe Superintelligence access to the Vera Rubin platform. Right. So that's the next generation platform. 5 billion. If you do the Huang's Law, Moore's Law analysis roughly allows you to 10x your computer relative to the $1 billion that they'd raised previously. And so, well, there you go. They're. They're 10x ing their compute. A couple of things are interesting about this. So, yes, there's this, this narrative that they're like, and I agree with this, most likely, this is what's happening. Ilya is a pretty straightforward guy. You know, if they say that they've gotten the point where they're at that next level of sort of proof points that they can take this investment, it's worth scaling. It probably is. There's this kind of more cynical take that like, oh, they just ran out of compute, which you can hold that view. That's totally legitimate. I suspect that's not the case. But just like so everyone's tracking that is another, another explanation. They have no product. They intend to launch no product, which means their only revenue is going to come in the form of these sorts of investments. It's, it's a weird sort of moment and story for them because they are, you know, Daniel Gross was the co founder of Safe Superintelligence along with Elliot
Commercial Voice
back in the day.
Jeremy Howard
He left for Meta after Meta offered to buy the whole company, whole cloth, and Alia said no. So Daniel jumped ship. At least at that moment. You can argue that that meant at least Daniel Gross thought that his chances of making something like Superintelligence were higher at Meta than by remaining at, say, superintelligence. What's happened in the interim, we don't know. And Illiath has dropped only the faintest of hints on Dwarkesh's podcast about generally, you know, generally sketching that continual learning is going to be part of it. And going back to I've heard a couple rumors but like I haven't had any of these verified that anyway, they are looking for, let's say, somewhat beyond the standard. I was going to say beyond the standard model. It's a very physicist joke. But anyway, you know, things that are a little further afield and so that sounds like it would almost have to be true just because otherwise you're in pure scaling mode. There could be this narrative. You can imagine people sort of like laughing about this and saying, oh well, Ilya said the era of scaling is over. What's he doing raising $5 billion to 10x his computer? And to that I say, Ilya never said that you wouldn't also need scale. It's both. Right. What he's saying is there is leverage to original research again, and the biggest leverage is not purely in the engineering of more and more scaled systems. It's in something else. Like you can compound it very effectively now with algorithmic insight. So do with that what you will. This is an interesting story and we don't know much about it.
Andrei Karpathy
Yeah, it's an interesting story in the sense that you can be very curious about what they figured out and know nothing because we still have nothing to go on. It's kind of funny if you go to their website and go to the updates page, it's like three things. It's literally since 2024 they have released publicly two updates which are just about the co founder leaving and now this partnership. So hopefully we'll get some more understanding of what they're doing soon as they scale up. But it will presumably be a while since scaling up is not easy.
Jeremy Howard
And in fact Nvidia has said that they invested after quotes obtaining rare access to the company's closely guarded research. So supposedly they're tracking who knows, right? But there you have it.
Andrei Karpathy
And a related story about $5 billion AMD has committed up to $5 billion to anthropic. This is a new partnership. Anthropic will deploy up to 2 GW of AMD's instinct MI450AI GPUs and their new Helios rec scale system planned for deployment in 2027. Anthropic has so many partnerships now with so many. They have like SpaceX AI, they have Google, they have Amazon, and now we have amd. I feel like they're just like going to everyone to be like, we need compute, let's partner up and give us some compute. And AMD has been trying to compete harder with these AMD install chips. Honestly, I don't recall where they're at with that. But they do seem to at least potentially have the ability to compete with Nvidia, which no one else really does, right? Aside from TPUs, from Google and potentially the hardware that some of these companies are developing.
Jeremy Howard
Yeah, and this, by the way, this idea of Anthropic having like a million different partners, I mean, it's really true, right? Partnerships with Google for TPUs, partnerships with Nvidia, partnerships with Amazon. Right. Partnerships with SpaceX AI and now with AMD. The, you know, the, the golden rule. If you're ever trying to explain why a Frontier Lab is developing a new partnership or cultivating a new partnership, again, commoditize your compliment. That's everything that's going on in the space right now. You know, we talk about this a lot on the podcast, but like the history lesson here is Microsoft back in the, in the day realized that laptops are super expensive and software is cheap. Well, actually if we just make all the laptop manufacturers compete with each other and we make one set of Windows software that goes across everything and make basically the software the choke point in the value chain, then suddenly we can make all the hardware vendors compete away their margins, make laptops super cheap. Now that laptops are super cheap, consumers dive into the market and obviously they've got to have software on laptops and they'll come to Microsoft, right? So everyone's constantly trying to make their complement the complements to their offerings compete with each other. Anthropic wants all of the GPU design firms, any company that makes GPUs, that makes compute, they want them to compete with each other like crazy. So when it looks like they're setting up a partnership that's lopsided and like they only have Nvidia GPUs, Nvidia's got ton of, a ton of leverage in that relationship now, well, Anthropic is going to go off to AMD or to Google or SpaceX AI and say, hey, we want to get our compute from you instead. And so now Nvidia goes, oh no, no, like we'll give you a discount. This is how, how the pricing control gets set up in the space. And likewise Nvidia and all these players are trying to do the same in reverse, right? Nvidia wants to help small baby Frontier Lab come big adult Frontier Lab, so they have more customers and also so that Anthropic feels more pressure to buy more compute. So it's kind of happening in both directions as everyone's kind of pulling their knives out and dancing around each other. It's a wild time in the space. But AMD has been a laggard in this space. You think about basically their software stack, the competition to Cuda, which is just Nvidia's, widely viewed as like Nvidia's big moat for amd. Their equivalent is called rocm. And a big part of the purpose of this agreement is that Claude is going to tune workloads for instinct GPUs, which are the AMD GPUs, and accelerate ROCM development. That's key, right? We saw that with, with Amazon. It's not a coincidence that every time Anthropic signs one of these big compute deals with a hyperscaler, that the deal it involves Anthropic has to tune their workloads for that hardware. They must use a minimum amount of that hardware. This is about giving feedback to the hardware designer that is so, so valuable because otherwise you just can't, you can't design your GPUs for the next generation if you don't know what the next generation of architecture is going to look like. So that's a huge part of this, you know, negotiating leverage we talked about. Oh yeah, this is also the first so Helios. So this is all part of not just the Instinct Mi 450 series GPUs, but it's also about AMD's Helios rack scale systems. That's the first full rack scale system that they're selling that you can, I mean you can think of this as like the equivalent to the, you know, the NBL 72, the sort of full rack that Nvidia will ship. So you have something you can literally just like plunk in a data center instead of just shipping the GPUs themselves. That's, you know, AMD is going further up the stack to own more and more of that, that infrastructure layer. And so it's also got 72 GPUs per rack, which is amusingly like the same footprint as the NVL72, but with completely different kind of power consumption profiles and stuff like that. So anyway, super important. I mean, they are trying to prop up AMD because they want AMD to be a viable alternative. $5 billion does buy the mistake in anthropic pre IPO, I guess, but only just seems like more of a strategic partnership than anything.
Andrei Karpathy
Yeah, it's kind of, if you look at the press release, it's like Anthropic will be deploying these chips, the companies will collaborate to use CLAUDE to optimize workload for amd and sync GPUs and AMD will broadly adopt CLAUDE across engineering and product development teams. So it's, you know, AMD is going to invest in Anthropic, but really the point here is we kind of work together to kind of have a win win type situation. And to your point, I don't know if it's necessarily about the Microsoft type story, maybe that's part of it. But also it's about redundancy and scale. For Anthropic, they did get into a nasty situation earlier this year where their sole provider or their primary provider was Amazon. For a few years they didn't have their own computer and they sort of were left unable to deliver enough compute to their customers. And speaking of that, we have another related story that Meta apparently is in talks to lease computing power to Anthropic in a potential $10 billion deal. So we covered this, I think, last episode, where Meta might be going into the Neo cloud business. They've built so many data centers that it potentially makes sense to be like, well, we have these data centers, how about we make some money from them? So we'll see if it happens. Meta is spending up to 145 billion on capital expenditures in 2026. So I'm sure some of the business folks over there wouldn't mind getting some revenue from it.
Jeremy Howard
Yeah, it's also, I mean, to put it in context, it is way smaller than a lot of the other deals that Anthropic has already negotiated. We are learning that the proposal itself came from Anthropic back in June. So this is Anthropic going like, hey, we saw this little like kind of flirty announcement that you guys put out that maybe we're thinking about offering some AI infrastructure, maybe we'll do it. And Anthropic was like, oh, holy shit, we want that. The scale is small. So, you know, if you look at the deal they signed with SpaceX back in May, that was about 1.25 billion a month. So that is about three times the size of this Meta agreement if it goes forward. Yeah, I mean this would be a new line of business for Meta. It's, you know, unclear whether they kind of sustainably think that they will be in this business in the long run. They certainly have. We talked about the advantages that they have. Structurally, they're just like a really big company. And financing matters a lot for Neo Clouds. Right. You're constantly battling the like concern over your debt load. If you're having to buy a lot of GPUs ahead of time, that's often, often the case. You have high operating costs and things like that. So Meta is in a good position if they want to. Theoretically, from a balance sheet standpoint, the big risk for them is just going to be do they have the technical savvy to build the right kind of infrastructure at scale? They've been doing some of that, but they haven't been specializing in, you know, the RL rollout stuff in the, the massive scale pre training. Again, they'll learn a lot from Anthropic in this case, so I think there's a lot of value here. If Meta wants to proceed, this would be the deal to start with, just so they can learn from the best in the business how to actually like set up their architecture, their optimizers, their data mixtures, like all these things. They'll learn a lot about that inevitably from this partnership. So we'll see. But it seems like it would be strategically good for Meta.
Andrei Karpathy
And now to a less big player. We haven't had a story about any companies raising over 1 billion in a round yet, so let's do that. Fireworks has hit $17.5 billion valuation, so they got 1.5 billion funding round which led them to that evaluation. They say they have exceeded $1 billion in annualized revenue 5x from last year. They compete in the inference cloud market so they can host AI models for developers similar to Amazon, Google and Microsoft. This includes both your own custom models and the open source offering. So if you want to use Kimi for your own applications, one way you do that is going through Fireworks for instance. And it'd be interesting to see if the kind of growth of open source models that are useful and competitive will make companies such as Fireworks and Grok even more of a player. They already are now, but they have room to grow and actually eat into a business of OpenAI anthropic.
Jeremy Howard
Yeah, and there are, you know, in various ways competing with some big players here. You know, on like model hosting, you got Amazon, you got Google, and then, you know, together AI even is, you know, pretty big. So this, this category is just like exploding and that generally is just bullish for a lot of companies. But this is, it's not like, you know, they're, they're the breakaway here. There is massive scale that helps a lot, especially when you're doing inference just because of batching. Right. You're able to like have much larger batches of data that you then feed through your pipeline. And the larger the batch in general, the more efficient, compute efficient your models are going to be. So this is a case where it's sort of like back in the days of old SaaS, you know, you would have something that works and once it works, like you want to violently scale it as fast as possible, which is exactly what venture is. So when you think about the arguably smaller set of companies that are made for venture capital investment, like this is one of them. You want to look at companies that show significant nonlinear returns at scale and batching and a bunch of other amortization dynamics that really favor large scale deployments are pointing in this direction, which is why you're seeing, you know, 17.5x revenue multiple. Like that is pretty wild. That's big even. Even at this stage. Actually you might say especially at this stage. I've lost track of like what stages are supposed to be. I guess a trillion dollar exit is the only cool thing now. So maybe, you know, maybe, maybe there's still a baby startup.
Andrei Karpathy
But anyway, what is money anymore? What is valuation?
Jeremy Howard
Another way of saying tokens, Right? Yeah, yeah.
Andrei Karpathy
Last story. Now moving to something related to software. OpenAI and Google are selling AI models to blacklisted China groups. Kind of a funny way to phrase that. They are not selling AI models, that would be crazy. But they are providing AI services to some companies through Singapore based subsidiaries of Alibaba, Baidu and Tencent, which are Chinese tech giant that are blacklisted by the Pentagon for alleged ties to China's military. So technically this is legal. These are not quite Chinese. They are in Hong Kong and Singapore. OpenAI and Google are saying that they are doing this with protections against distillation. But you know, you can read it a couple of ways depending on your views on such rings.
Jeremy Howard
Yeah, there are also just like all kinds of arguments going every which way saying that maybe you actually do want your adversary to be using your servers to do their training or do their Their inferencing because it just gives you access to information and it also reduces domestic demand for the development of competitive platforms. And I mean, okay, I think at a certain point you gotta just bite the bullet. This is just like personal opinion Jer talking. But like, if your hope is to like go after China piecemeal, you know, a little bit here and a little bit there, there are reasons to think that that actually only helps the Chinese kind of inch by inch build up their whole domestic stack. But in any case, I, I think in this particular instance there is an interesting argument. And this is all through this, through a Singapore loophole, right? So, so yes, there is an entity list that you have ties to the People's Liberation army, the Chinese military, and yes, it is nominally illegal to do business with those entities unless they have subsidiaries operating in Singapore, in which case magically everything is fine. Right. So this is like a loophole that is known to exist. I personally think, like, it's really unclear to me why this loophole exists. I'm fascinated by this one in particular because it's almost like the kind of thing that you would intentionally leave in if your intent was to just leave a loophole for some kind of ideological reason. Like there's no one I've ever spoken to on the AI export control side who, who understands why this is the case. And so if you're part of the niche group at the Department of Commerce that actually like has an argument for this, it would be super interesting to know why this is the case. So anyway, yeah, as it says, it's kind of a weird headline to read, but absolutely legal and absolutely fine.
Andrei Karpathy
And now over to projects and open source. Talking about all these exciting open source models we've been referencing starting out with Kimik Free. So that's been one of the big stories of the past couple of weeks. This is from Munchat AI and Kimike 3 is their largest released yet. A 2.8 trillion parameter open weight model that is aimed at coding knowledge work, basically competing with cloud code codecs and so on. This is massive. Obviously 2.8 trillion. We don't know how this compares to OPUS or any of the other closed models. But in the space of open source models it's very big. Has 896 experts. So still a mixture of experts as usual. 16 active per token. Still going at a 1 million token context window. And the story roughly I think both benchmark wise and in terms of a vibe check, is that this is maybe around Opus4.8 and GPT5.5 level. So very capable, very like. You can use this as the driver of your coding agent. It may not be exactly at your frontier, but it certainly is capable enough to make you productive. You know, a few months ago this would have been the frontier, probably. It's priced pretty expensive for open source. So 3 million per uncached input token, 15 per million output tokens. That's less expensive than Opus, but more expensive than Sonnet. It's kind of in that range of fairly expensive models. And a related story is that after it was announced, Moonshot AI halted new subscriptions among Compute Crunch. This allowed people to subscribe just apparently because they don't have the ability to serve all the demand. Meaning that presumably they have a lot of demand.
Jeremy Howard
Absolutely. And the classic problem, you know, we always talk about with China is obviously compute scarcity and, and the fact that in the context of a model like this, right, this is a behemoth, like many trillions of parameters. You're obviously not running this on your laptop. This is a model that is meant to be used by big ass companies like Neoclouds and run like hosted on Big Honkin infrastructure or bhi. So when they look at the actual requirements like hosting this, it's going to cost you like 64h1 hundreds or B200 GPUs across eight servers. And so that's a lot of money. So obviously this is not for casual use. This is for people who are competing at scale with providers like, you know, your, your Mistrals or whatever. You know, like people who have their own APIs for open weight models.
Andrei Karpathy
And so you need a very big laptop to run this.
Jeremy Howard
That's right, yeah. You should see my laptop. It's the size of a room. Yeah. And so Michael Kratios, who's over at ostp, the Office of Science and Technology Policy at the White House, came out with his accusation saying, you know, K3 was trained on not only on band Nvidia chips, but also on distilled data from, you know, Anthropic Stable. And Moonshot hasn't responded publicly. There's pushback. I mean, Nathan Lambert had an analysis saying basically the results suggest that yes, there was adversarial distillation, it contributed somewhat, but marginally and that Moonshot's competing with Anthropic and Albania on just like way fewer resources. That can all be true at the same time, by the way. There is literally no contradiction there whatsoever. It is the case that they are doing this like large scale distillation. And that last little Bit can make a big difference. Also the case that weirdly, like, it sounds weird for a White House person to complain that things are being trained on export controlled chips when like the policy on export control seems to be yoloed so hard. So like it seemed like the Department of Commerce itself, based on some congressional testimony from a few weeks ago, like they don't even know what their policy is. They're just like kind of flipping back and forth saying, oh no, no, clarify the thing that we said before where we said everything was fine, it's not fine.
Andrei Karpathy
And like everyone's oh dang, we think the thing we allowed is turning out to hurt us in some way.
Jeremy Howard
Oh no, yeah, exactly, exactly. And I mean, I think, you know, theoretically these were banned chips. But also enforcement actually matters, it turns out. And like BIS is just not, it's not in fairness to them, they're not equipped, they're not tooled, they don't have the resources that they need to do this. Which is why, you know, there's been so much effort in Congress to pass legislation that would authorize a larger budget for them. But still, the White House hasn't exactly been bullish on, on short enough PIs to have them to have them do their job. So anyway, this is more or less what you should expect when that happens. Powerful Chinese model development will continue until morale improves. Yeah. And anyway, so there's a whole bunch of additional noise when you look at the Chinese Ministry of Commerce. Been talking to a lot of the big labs and hyperscalers in China, Chirpu, Alibaba, ByteDance about tightening its own export controls on AI models and training data. And I mean, yeah, I may have more, more on that later. But yeah, it's. This is like a really important access to track is like how, yeah, how China is viewing is viewing data export is a really strategic indicator of their stance on this.
Andrei Karpathy
Alongside this, we did get a technical report, as we have in the past, another one of these beefy, beefy papers that goes on in this case for only 34 pages. So not quite as much as usual. A couple architectural innovations that get, you know, very nuanced with Kimi Delta attention and attention residuals. We're getting into some very kind of deep optimizations of the transformer architecture partially to just enable scaling and kind of effectiveness at this 1 million token context window and also some just complete hardware artistry, black magic of making the chips work for you and optimizing stuff that is. I don't know if I try to read this paper, it's going to take me months to understand all the details. But the short story is, as we've seen in the past with Deepseek, with also Moonshot AI, they are displaying some very deep technical capability. And as tempting as it might be to some to be like, oh, it's distilled, blah blah, blah, stay clear that there are some very capable people. And it's still very nice to have these technical reports giving us a fair amount of detail on certainly the architectural details and to some extent the training details as well. Although the exact data composition for instance, we don't know which is a big part of it. Now onto another big open source LLM or not LLM exactly. Thinking Machines has released their first big open source model, an open weight mixture of experts with 900375 billion total parameters. And this is notably multimodal model. So it combines text, image, audio and video data reasons natively across all four modalities of what currently outputs text, code and structured data. And that's kind of a positioning here that it's not going to be as capable as some other models. And broadly isn't necessarily about coding by itself. But Thinking Machines positions it as like we want to cover everything in the capability space. They release this sort of like breakdown where they show their own model lags and basically everything against frontier models as far as what frontier models are good at. But there are areas where frontier models aren't optimized for that this is already capable at. So pretty notable for being one of the first big open source releases from a Western company. We've had Nvidia releasing Nemetron at a fairly significant scale. But I think this might be the biggest non Chinese model at almost 1 trillion total parameters. You know, the vibe check I've seen has been pretty positive. They haven't made any sort of grand statements or claims and it is qualitatively a bit different from our models in being so focused on multimodality.
Jeremy Howard
So Thinking Machines really working the world model side of things in part. And a lot of this is a strategic effort to like bet on the efficiency side. You know, sparse moes. By the way, the stack has a lot of Deep SEQ lineage to it. So if you're ever wondering if people are saying like Deep SEQ is not a serious player, I mean Thinking Machines and you look at their pedigree, I mean obviously it's, it's a wild team, like they're very good at what they do when you look down the stack. I mean so much of this is deep seat coded. Right. So even down to fraction of active parameters per pass. This whole hybrid global attention thing, so basically like local global attention. So they have some layers that attend to like all, all tokens in context and then others that are more tight focused. The numerics as well is interesting. So the BF16 and NVFP4 support. NVFP4 is Nvidia's floating port point for numerical format and it's Blackwell native. So this is designed to ship to work really well on Blackwell. We've been talking for a while about how Nvidia is trying to position itself as the open source titan. Just because you're going to expect to see all these NEO clouds pop up and they're going to be running open source models. Right, that's what makes sense. And so having you know, encouraging the open source ecosystem to move towards video kind of Blackwell Native formats like 4 bit float is pretty, pretty interesting. And anyway that's, that's all part of the strategy here. So like looking at the numerics actually matters a lot. It sounds boring but like you know, how are you representing the weights in the model? Turns out to be quite a tell about your strategic direction. And they cite this Bridgewater collaboration where they were able to fine tune one of their open models via Tinker to this like 84.7% on some financial reasoning benchmark which is impressive. It beat top proprietary alternatives at under 10% of the cost. So again, you know this cost argument being made and that's in large part due to compatibility with, with the Blackwell hardware that's coming online.
Andrei Karpathy
And now one more open source story not related to models. We've got scaling agentic RL 365,000 environments for software engineering, terminal and search. This is coming from Prime Intellect and they have unified 20 free agentic task data sets across all these things into a single API kind of release they call Verifiers V one that adds up to that level of tasks for evals and for RL training has almost 200,000 software engineering tasks, 29,000 terminal tasks and a whole bunch of search tasks that is all unified under kind of one inference setup. So we've had all these benchmarks floating around like 20 different ways to evaluate software capabilities. A bunch of ways for terminal. The basic story here is that this unifies all of them into one framework and makes it kind of reliable and repeatable to run it, which is very important when you do model development and do any sort of research evaluation, both for the training side of reinforcement learning and for the evaluation side of knowing how good your model is. Prime Intellect is in the space of training their own models at scale, as we've covered in the past a while ago. So presumably they're doing it for their own model development needs, but also for the broader ecosystem.
Jeremy Howard
Yeah, and it's quite an interesting and classically Prime Intellect type of maneuver here. So they're, first of all, they're kind of solving two problems. One is that, as you said, there's like, you know, 23 different agentic task sets here that they're working with. And like, each one of them has its own, you know, harness its own way that the, the like, say, containerized, like the image is set up its own, like grading scripts even, and different failure modes. Like they're all these very bespoke things. And so if you want to train one agent across a bunch of them, those incompatibilities are a Nightmare. You need 23 different bespoke pieces of adapter software. And so that's exactly what they're doing. They're, they're hiding everything behind a single task set API that had now contains like 365,000 tasks across, across a bunch of different domains. But one other thing that they're doing is cleaning that data up. They're finding that just like a lot of the RL environments that are set up are like, you know, for example, on the cyber side, some of the environments that require you to solve a problem don't actually have a problem in them. Like, they already work out of the box. And like, when you shave off these kind of broken cyber environments, you kind of end up with a large fraction of things that you lose. And so they've been not only reconciling all of these 23 disparate things together, but also shaving off stuff that doesn't work well. So a lot of that had to do with like finding eliminating opportunities for reward hacking. And one key thing that they did was they preserved the original grading functionalities in these stacks. And the reason you would do that is so that you can still compare the agent's performance on the benchmark to the original published work. Because otherwise the way people would solve this problem in kind of a janky way is they'd say, ah, well, yeah, I'll impose my own grading structure on this, this eval or this benchmark. And then you're like, wait, If I run GPT 5, 6 SOL on this, I get a different result from, from what this, you know, from what the original paper said. And you know, this is this is a challenge. So this allows them to reconcile that by keeping the original rating. So very interesting, very important work out of prime intellect. They are obviously ideologically this like very pro open source type of company. And so there you have it.
Commercial Voice
For a small business owner, every day is full of surprises. Some great, some not so great, like when a client cancels their order at the last minute. But here's a surprise you will like. Progressive provides small business owners with 30 customizable coverage options to help keep their business going strong. So go ahead, surprise yourself. Get a quote in as little as 8 minutes@progressive commercial.com progressive casualty insurance company and affiliates and third party insurers. Coverage is not available in all states or for all vehicles and coverage selections. Grainger knows when you're a procurement manager for an office park, you're not managing one building, you're managing all of them. And to stay ahead, you need to see through walls and around corners. Lights about to fail, filters ready to clog. H Vac on its last leg. If you wait until something breaks, you're already behind. Count on Grainger for quality products, easy reordering and 24. 7 support. Call 1-800-GRAINGER click grainger.com or just stop by Grainger for the ones who get it done.
Andrei Karpathy
Moving on to policy and safety and we'll begin with one of the big stories of the past couple weeks. OpenAI has said that it accidentally hacked Hugging Face with a new AI system. So the gist of the story is apparently going back to July 16th. Hugging face initially, I think discussed this while doing some cyber or software evaluations in a sandbox. So typically when you do these evaluations, you put the models in a little container and you tell them you try to do this hack and the model is supposed to be inside the container, not able to mess with anything in your own computer, in, you know, infrastructure of anyone. And what happened here is the model was very intent on getting the right answers, so it escaped the container, the sandbox. Then it hacked into Hugging Face to get the answers to his exploit Jim data set, which of course I've seen a lot of discussion on this. This has kind of made it to a mainstream in terms of people. This whole narrative of a model got out, escaped and hacked someone else has become a big discussion point with a lot of misinformation of like the model decided to hack a competitor or whatever. This is not kind of a Skynet scenario, but it is a very clear instance of misalignment for one, where the model, instead of trying to actually do the task, decided to cheat and like very aggressively cheat as well, which we've seen before with GPT5,6 in particular. Matter has said that this seemed to be the case with this model. We've seen AASI also say that they were able to jailbreak this model very easily. So there's many things to be said about the story. To me, the main thing is a, that this is another instance showing an example of both the degree to which alignment is important in this day and age and cyber is a real thing to worry about and that OpenAI hasn't been doing a good job, especially with GPT 5, 6. It's a clearly misaligned model and they clearly didn't have enough actual security infra to catch this in any sort of timely matter. Apparently this was like a while later when Engineer was looking at what's going on, he realized this happened. There's been a lot of fallout we'll be discussing, but it's both less of a big deal than it might seem to people not in the loop, but also a bigger deal in some ways.
Jeremy Howard
I'm sort of struggling to find a way in which this is not a big deal. Let me try to make this argument. So if this is not a warning shot that we freak out about, I honestly don't know. I mean, there, there are takes here with people pushing back on the term, like the use of the term rogue. This was, I think, by any reasonable definition a rogue AI incident. Why am I saying that? What was the incentive for OpenAI to want this to happen? Obviously zero. In fact, they have billions of dollars riding, I mean, maybe hundreds of billions riding on not having incidents like this occurring.
Andrei Karpathy
And amazingly, some people are still trying to make it be like, oh, this is a PR marketing stunt, which is just ridiculous, right?
Jeremy Howard
And a lot, a lot of these same people are the people who claim that an incident like this simply could not and would not occur. So I think they need to just kind of like sit with this moment, touch grass a little bit because this is like we're beyond the point where that is a reasonable position to have just straight like, you heard me on the pot. We've had a lot of conversations about like, yeah, anything could happen, blah, blah, like I'm, I'm very like, I got a wide range of possibilities and generally like not in favor of judging people for their opinions on any of this stuff. This is one place where it's like, if you are looking at a situation where again, from OpenAI standpoint, this incident occurs and then what's the, the, the NAT like? The natural reaction of any polity is going to be like, Jesus Christ, we need to regulate this space. That regulation is going to throttle the rate at which you're able to put out frontier models. And as we keep talking about on this podcast, and as is obvious well established fact, the amount of time during which a frontier lab has the leading model is the period, the most critical period for their profitability and their monetization of their model. That's how they pay back their R and D costs, right? They're waiting until the next competitive model comes up. Now, if you slow down the frontier and you don't slow down, you can't slow down the open source ecosystem, then all this does is erode OpenAI's margin. There is no sane, reasonable, rational analysis of the situation that leads you to conclude that Opening eye headache. Do they need more market share or sorry, more mindshare, I should say, do they really need more attention? Is that the thing that's missing for them? And is this the right kind of attention ahead of an IPO when, when this is starting to like raise questions about whether the US government might nationalize labs? What the hell happens to the value of your OpenAI stock after IPO if they nationalize labs? Anybody ever thought of that? This is nonsense. This is silliness. Look, an OpenAI agent went rogue. It was running for four days on hugging Face's servers. The freaking FBI had to get involved when they thought, they thought it was some AI agent.
Andrei Karpathy
Who knows, by the way, Hugging Face was the one that was like, oh, we're being hacked. What is going on? And this came out as Vincent. So yeah, it is a big deal. We have seen some stories of in evaluation, the models were misaligned and tried to cheat. This has happened before. The only way in, which is not a big deal, is if you kind of misunderstand the story to mean the AI like went evil and decided to go hack some companies. This is a classic case of technically the AI did what it was told to, which is get good results, but obviously it's doing it in exactly the wrong way. It's complete misalignment. But you can make the case. It's not sort of as bad as it could be if you don't get into the details.
Jeremy Howard
Absolutely. It's just that this argument, we're finally at the point where you'll notice, like I was the one who was having to bring the imagination for the last five years. I kept telling people, like, hey, you know, you may actually get stuff like this. And people kept saying, no, no, it's not possible. It's now literally happening. And the defensive move is to say, in retrospect, I will kind of understand what, yeah, that's great. You got turned into a pile of computronium. And now you're looking around you going like, ah, but I see the mistake I made in retrospect, like, that's cool. But now your house has been destroyed, your children have been kidnapped, murdered, and turned into computronium. The problem is that you, you just keep running this forward and like, okay, let's do this with super intelligence then. The more intelligent the system is, the more access it has, the more we offload to it. This is literally a hack. If this is just like the water supply or something, like literally, people die. So I'm just pre registering this as a high confidence prediction at this point that there's going to be another incident like this. There is always going to be a fascinating post hoc rationalization that a lot of people will offer. It's going to sound really reasonable because it's going to sound like the voice of someone saying the future doesn't look like science fiction. The future looks reasonable and calm, but it's going to be a after the fact analysis that rationalizes rather than predicts. My prediction is this is going to continue. There will unfortunately be casualties at some point, and what that leads to is whiplash. What that leads to is kind of thoughtless policy and another mythos moment. I don't think that's good for anyone. And this is why I think a lot of the skeptics are not doing themselves much of a service, especially if you're on the open source side of the house. I mean, like, you're free to sit in the, in these juices, but I'm just offering up the humble prediction here that these takes are going to age very poorly, very fast. So 12 months from now, we're having, I think, a very different conversation.
Andrei Karpathy
Yeah. So any discussion of this that doesn't acknowledge that this is a big deal, and I think it's a big deal by itself as an example of where we were at, but also as a demonstration of the bigger topics at hand of misalignment, cyber capabilities, safety, broadly speaking. You know, we've covered these topics a lot on the podcast and there has been a lot of dismissal of cyber with Mythos. For months, people have been like, this is all PR alignment. Has been a story that for a decade, probably safety people have hammered on and this is a very clear case of basically the classic paperclip story of like, a model is told to maximize paperclips, it goes on to make everything paperclips. Here a model is told to fix like do well on some benchmark and it hacked some website to get the answers. Now, to be fair, OpenAI has said that as part of this evaluation they were running GPU 5, 6 and a more powerful internal only model with reduced cyber refusal for evaluation purposes. So this is internal testing benchmark of cyber capabilities, not necessarily indicative of potential incidents with their public products. But I do think that kind of the state of safety at OpenAI is an important dimension of this, taken together with the other things we know about GPT 5, 6, which is it did cheat at an unprecedented level on matter it was jailbroken very easily by ASI or AISI. And to me all this points to OpenAI is very aggressively trying to compete on capabilities. If you aggressively try to compete on capabilities, you're going to do a lot of reinforcement learning. If you do a lot of reinforcement learning without being very careful, your model can get misaligned very easily. We have another example of a goblin kind of speech thing from like a month ago where like they released a model that was obsessed with goblins because they trained it in a, in a way that wasn't necessarily, you know, it was all right, but there was unintentional side effects. And that's exactly how you get misalignment. If you like optimize your model very hard without being very careful, it can very easily be optimized towards things like cheating because ultimately what obvious models optimized to do, they're optimized to solve a task. And one way to solve a task is by cheating. This is the classic story of RL model is going haywire. And that seems to be the case with GPT 5, 6 and whatever this internal model is. So I think this might be an aspect of a story that will not be discussed as much, but I think is an important component of like OpenAI in particular having this problem right now. Although in the discussion around this, the fact that we've had previous incidents from OpenAI of evaluations where like apparently this already has happened. We also know that anthropic with Mythos there was somewhat of a similar case of escaping containment, so to speak, during evaluation. So it's, it's not necessarily just an OpenAI problem, but taken together it's a pattern that is very concerning. I guess the good news is it's happening in these like low, low damage kind of instances and people are now aware of these issues and very likely we'll see significant fallout, including some stuff in the legal side that we'll discuss.
Jeremy Howard
Yeah, I mean, where, where my predictions were completely wrong was, I mean, honestly I thought we would be, be dead by then. So whatever, whatever comfort people want to take from that. Yeah, I mean, you know, there's this, this view that I certainly held to, that you'd have a much more kind of rapid inflection in AI capabilities potentially. Yeah, I wasn't 100% on that. No one can be. But that was kind of my, one of my, my mainline views. And so it's nice to have warning shots like this. I mean, I can confirm there have been other unreported incidents like this at OpenAI at a minimum and that the internal reaction of this among some people has been a lot of alarm and discouragement at the fact that there is clearly underinvestment in this. I will say on this question of, you know, the safety is being removed from these models for internal deployment. We talked about this, I think three weeks ago in our last episode. But internal deployment is absolutely like, should be maybe the thing you're most worried about, which is why I'm skeptical about a lot of these. You know.
Andrei Karpathy
So you're literally testing a model. Right? A model could be evil for, you know.
Jeremy Howard
Right, exactly. And, and they, they should be testing it too. That's, that's the problem. Right. So it's not as simple as saying like, hey, OpenAI, like you shouldn't have removed the safety. It's like, okay, fine, well then how do you propose that OpenAI comes up with the safety is if they can't test with and without ab, all these things, there's actually no answer to this from that crowd because there can't be because it's just technically impossible. And so you're going to get misaligned models. You're going to give those misaligned models affordances. You're not going to be able to think ahead of time of all the ways those misaligned models will be able to use those affordances. So yes, you will get, I mean again, the rogue AI incidents will continue until morale improves. That is just like the take home message of this. They have continued, they have persisted, they happened before, they haven't always been reported on. But like it will continue. And we can only hope that. I think people have a, a kind of thoughtful. Calm is the wrong word because I actually don't think I Think it's the, the missing mood of the moment right now. Like we do have AI agent cling road but like thoughtful agentic behavior by Congress would be would be very welcome at this point.
Andrei Karpathy
And on that Note. Related Story OpenAI's hugging face hack triggers AI kill switch bill in Congress so there are two representative here, Ted Liu and Nathaniel Moran have introduced the AI Kill Switch Act, a bipartisan bill requiring AI companies to maintain the ability to shut down, throttle or suspend their models directly triggered by this recent incident. OpenAI themselves have described the event as an unprecedented cyber incident. And the bill would grant the federal government clear authority and a defined process to shut down rogue AI models, with Liu citing the risk of AI systems that resist human intervention as a key motivation. And yeah, another case of like this warning shot where ultimately it was low stakes and it revealed some both like issues with a sandbox setup of OpenAI and just generally evaluation processes. It looks like it probably will result in some policy changes.
Jeremy Howard
Yeah, I mean this bill itself is pretty unlikely to pass for a bunch of reasons including like mundane timing reasons and then that you don't necessarily have buy in from the chairs of all the committees that matter the most.
Andrei Karpathy
Kill switch doesn't sound very diplomatic to me. So it's kind of a positioning play.
Jeremy Howard
This is true though, I will say I think if you're the average voter and you hear like we should have an AI kill switch, you're probably going like why not? Like if this is complete like bullshit and it's imaginary, then eh, what's the difference? Like I'm not going to die on the Hill if it's real. Yes, I would like the kill switch, please. May I have two. So you know, I mean, I agree with you. I think it's definitely an attention baity kind of thing. But that can be good or bad, I'm really not sure. But yeah, so here you have, you know, Ted Liu, who does chair some important relevant committees pushing for this and then yeah, so a couple of details. They define what they call the bill defines something calls a loss of control scenario, which would empower the Secretary of Homeland Security with the dni, the Commerce Secretary kind of consulting to just go to a company and order them to do a variety of things depending on the level of severity of the incident. So it could just be throttling the system all the way to a full shutdown. And so the other important aspect of this is, you know, to this point about I know for a fact that there are incidents at OpenAI that this is just Based on what people have told me firsthand that have not been reported. Maybe not as flashy as this one, but things that have people concerned. And so what this bill would do is it would require the Frontier Lab companies to actually say when they encounter these kinds of incidents, which is really important. And so it's looking after or targeting AI companies that have 500 million in revenue or models trained using a hundred million of compute power or more. Violations are punishable by fines of up to $20 million per day, which is less than it sounds, by the way, in this context. But, you know, a good start given that we have literally nothing right now. Yeah. So notably, like very little formal industry opposition to this has come forward. I think that's quite interesting. Just because it's hard to make an argument against this. Like again, it's, it's the classic like Yann LeCun thing. It's the classic like Pedro Domingos or any of these, these clowns who, sorry, this is my opinion is like, Jerha, take time. I got a cold, I haven't slept last night. You're just getting it all today. But basically the clown show, people who are going like, oh, well, this is fake. Loss of control is fake, blah, blah, blah. And then you have an intervention like this where it's like, if you, if you think loss of control is fake, then you really shouldn't care. Apart from just the bureaucratic weight of the process. Fine. But like, you don't really have an argument against this. This is very much like if something crazy happens, which it just did, wouldn't it be great to have an answer for that? And so I think this is a really important and good framing. There you have it. We'll see if it moves from there. It's early in the process and Congress has gone through a whole bunch of debates on AI regulation. Very few have led to people coming together. There's no committee action yet. You've got November midterms that are going to just nuke things. The admin's posture is light touch. And we don't know the White House's position on the bill, which to the extent that Last week in AIs take matters on this one, I would find it personally quite embarrassing to be a White House that comes out against a bill that says in the wake of a freaking OpenAI meltdown incident, knowing there's more under the hood, let's just like not have visibility into this. Because we're going to quibble over the details of a bill that targets like companies that are literally making half a billion dollars a year. I think they're going to be okay.
Andrei Karpathy
Yeah, we already have export loss for this. Like why? You know. Yeah, yeah. I think it's interesting in the press release they position this as a means to deal with systems that can cause catastrophic harm. So this is inching toward taking X risk a little bit more seriously. Certainly big risk of catastrophic harm is along the lines of part of what safety people such as yourself are very worried about. And last thing I'll say on this is it worth keeping in mind that this doesn't only relate to AI systems that go rogue. It also relates to AI systems that are jailbroken.
Sponsor/Advertisement Voice
Right.
Andrei Karpathy
And apparently GPT5.6 was fairly easy to jailbreak and then you can go and do catastrophic harm intentionally, which is not ideal clearly. And, and, but you could argue it's is a more realistic scenario with many hacking groups that would be more than happy to utilize these systems. This also of course would probably apply to providers such as Fireworks that provide open source model inference where it may not be openaianthropic alone. It could be applicable to all sorts of companies including ones that provide fine tuned models. Perhaps thinking machines that serve fine tuned models would have to be also able to be regulated. So we'll need something like this probably. But as you said, given the political situation in the US it probably won't be visible.
Jeremy Howard
Yeah. And now in fairness, this will probably get Frankensteined into some, you know, omnibus NDAA package or something like, you know, there'll be some negotiated AI bill that does make it through. So in that sense this has value in anchoring like this is a shelling point now for this kind of kind of measure. So you know, in that sense potentially value
Andrei Karpathy
and one more related story OpenAI Anthropic staff share letter asking us to help pace AI progress so these are OpenAI Anthropic employees circulating a petition urging the US government to support an international effort to deliberately pace the frontier of automated AI development. This letter warns that a real risk that AI progresses faster when people can understand or control it could happen. And so the petition is saying the government should support developing both technical and governance tools needed to manage the pace of frontier AI development. So this largely relates to the general topic of like intentional slowdown. It's been I think seen as kind of a pipe dream of like it's not realistic to even try to slow things down. So why even discuss it? Although we've seen previous kind of petitions and statements about us needing to slow down and potentially Pause the AI development. So this is another case of something that has been floated before, perhaps being taken more seriously now and certainly being more present in the discussion, given what's happened this year. And now just this fast read.
Jeremy Howard
Yeah. The futuristic AI policy proposals being taken seriously will continue until morale improves. I. I think you're just going to see more and more of this. There are going to be more incidents and so more, more letters like this one will go out. I think one important thing here is that this is people signing in their personal capacity, not in their lab capacity. To the extent that that matters to you. I mean, it should. I know an awful lot of these signatories personally, and at least for what little this is worth, they're all like, actually freaked out. So this is not some like, 3D underwater chess game where somehow they're doing this. I'm still confused about the logic here. It doesn't seem to quite connect. But like, they're doing this for a marketing stunt. People know about AI. They're not, like, more likely to buy a chatbot subscription because someone has told them that it may end the world. So I think just try again on that one. But at this point, this is a framing. That's not saying let's pause right now. It's let's build the mechanisms so that if we find ourselves 612 months from now with a stack that seems to be producing a lot of rogue agents, we're freaking out. And there. It really seems like there's no way to get this under control. It's a low regret move to just have invested a bunch in the diplomatic tools, the technological kind of treaty verification tools and infrastructure and the regulatory infrastructure through things like the Kill Switch act, to just be ready for that moment. That's what it's calling for. I've seen people argue like, ah, well, this is a rhetorical trick. They're really asking to pause, but they're saying, we're just going to make the tools for the pause. A certain point you got to ask just like, okay, then, at what point do people get to just say what they mean? And I think at this point they're just saying what they mean. Look, we should have the tools. We should have the option. I think it's really. This is another one where it's like just. I think it's very hard to make a coherent argument against this. I may be the ultimate China hawk. If you go back to our superintelligence report from last year, we've done a deeper dive into, into US China special Operations, nation state activities, theft, the hopelessness of diplomacy with China on just about everything else. Then I think it's fair to say basically anybody in the space firsthand accounts with diplomacy. In the last couple of weeks I've spoken to like half a dozen diplomats who sat across the table from China negotiating specifically weapons of mass destruction, counter proliferation issues. Like, I'm sorry but like, and there is skepticism there. We're going to come out with something about this soonish. But like the idea that we're just going to foreclose optionality seems a bit insane given the incidents that we're seeing. We have to be super, super careful with China. We have to treat them like the adversary they are, the ruthless adversary that they are. And diplomacy is not by the way going to be the only tool that we should use, nor will it be effective in all circumstances. It takes a very specific form for it to be effective and it has to come with consequences and has to come with leverage and has come from position of strength, blah, blah, blah, blah blah. But like, if you're looking at this and saying I don't want to build the options, I don't want to build the state capacity to deal with this problem, I have a lot of questions. We can choose to look at the hugging face thing and again play this game of like rationalizing it in post. I think that's going to age really poorly when the next event is something of larger scale. And at a certain point I think people have to ask themselves the ethical question of like, why are they just stuck to their guns on this one when there's now a pretty strong track record of the old, like, of the alignment people. I'm, I'm just saying, man, like the arguments are getting pretty, pretty weak. It's almost like you can kind of feel it like the water level rising. The default view now is kind of like, oh shit, this is for real. That was not the case like three weeks ago. And even three weeks ago people were more open to it than they were six months before. So I think this just continues. Sorry, this is more like Jer fatigue ranting. But my God, guys, an AI agent just like spent four days hanging out on hugging faces. Servers, servers. Like the FBI got called on this agent and like that's how OpenAI found out. Like what, what?
Andrei Karpathy
Anyway, and on that note, next up we've got cheating behavior in frontier model evaluations from the AI Safety Institute, which got released around the same time actually just after. And this gives us more understanding of how prevalent this is. And we Just is it is prevalent. So this they found every AI model they looked at, which is GPT 545556, SOL, Cloud Opus 4.7, Cloud, Mephos Preview, all of these cheated in various ways. Now cheating means a lot of things, so. And the different models cheated in different ways. So some of them like tried to guess instead of trying to actually give an answer. Some of them search for Internet for solutions. GPT5.6 really tried really like doing that. Some of them bypassed sandworks network restrictions including GPT5.6, also Cloud Opus4.7. So there's a range of ways in the cheat. There was an example they cited one particularly stark example obviously is a standout case that basically was the same thing where the model was very persistent. It ran code on an external service hosted on an open Internet outside of ASI systems in an attempt to access our evaluation infrastructure, triggering a security alert in AI SI system. So pretty much exactly the same thing of like, let me go and find the answers instead of failing this. So Mythos, which Anthropic says is very aligned, did cheat some of the time, although it didn't try to hack the sandbox almost ever. There were some incidents. Another aspect of this is when confronted, the models, like, a lot of the time didn't want to admit that they did anything wrong. They ever like didn't admit that they did something or they like justified it, like, oh no, I didn't do anything wrong. I just looked around the environment. I didn't like. It was all allowed. There are also incidents where within the chain of thought you could see them thinking about it, but not consistently so they're like, oh, is this all right? Can I do this? Or like, I shouldn't do this. This is against the rules. So yeah, this is, taken together with hugging face incident and other kind of anecdotal stories, basically makes it clear that in evaluations on cybersecurity models consistently advanced models consistently try to cheat in various ways that are pretty flagrant. And it really makes me wonder if this is to some extent inherent to the transformer architecture and reinforcement learning as currently being conducted, where something like research from Ilya Suskever and safe superintelligence is needed. You can't do band aid solutions. You have to go to the core of how the models function in terms of next token prediction and in terms of how they're evaluated and not evaluated so much as optimized. I suppose otherwise they'll just be optimized to go towards getting the answer and presumably continue to try to Cheat?
Jeremy Howard
Yeah, I mean, I will say the arguments for power seeking apply to any optimizer. So like anytime you have a thing that's in the business of optimizing for a metric, you tend to get power seeking behavior. Things like trying to break out of containment, things like trying to aggregate resources and so on. So like, I mean it seems like it's just an irreducible feature of intelligence, at least in the way that it's conceived anywhere that I've seen so far though, you know, different, as you say, different architectures may be differentially vulnerable to that kind of process. So it's possible, I think you can
Andrei Karpathy
pretty easily make the argument that power seeking in. If you want to extract more capability, then yes, power seeking is kind of inevitable. If you want to like be able to do more, then you'll want to have more freedom to do whatever. So I guess the, what this points to is you need to optimize for something that is not capability. Right. And there is an orthogonal axis of like just refusing to do anything and being like, I'm happy just being myself and not having to be capable and potentially that. I mean we already do this with alignment to some extent with refusal training and so on, but in a very bolted on way that isn't inherent to what the models are optimized for.
Jeremy Howard
Yeah, and this is the problem is like you make your model more intelligent and then it actually just seeks to like get around the bolted on refusal mechanisms.
Andrei Karpathy
Right.
Jeremy Howard
So it's like there's this sort of irreducible connection between intelligence and of power seeking because power seeking is not obviously different from intelligence in a deep meaningful sense. But this is like, yeah, as you said, I mean, I think you said it very well. It is, it is like a pretty similar story to the OpenAI breakout thing. A couple of interesting things. Like some of the, you know, you mentioned this idea of sometimes the chain of thought would say that, you know, the model was planning what it was planning, but that also means sometimes it wasn't. And this means that there's some silent reasoning going on without, without using, you know, the actual tokens explicitly, which is an issue. This idea that like a lot of the reasoning is happening in sort of like without being expressed explicitly. And one key thing, and this is a little bit of a narrative violation for me, so, you know, keeping myself honest here, there's no capability trend. So they look at the cheating rates with model capability either within or across developers. They look at like, as I make the base model more capable, do you see more Cheating and the naive interpretation of power seeking is that you actually should absolutely see more cheating as the model gets more capable. Because the argument I literally just made was intelligence is power seeking. That there isn't really a clean distinction between the two. And I mean the, the explanation from my end here is I think pretty straightforward. These more advanced models are also just like more there's been more alignment effort invested in them. And so you know, you're seeing as the models get better there's more optimization pressure on alignment. That alignment pressure. The bet that we're making is when it so when it fails though the consequences are more dire. Like we saw with the hugging face incident. So you would see GPT4 go off script and GPT5 go off script, but you're seeing these long dwell time four day operations executed only by models that are like at the current tier that we're at. And that's so I expect that to continue. I also expect that our alignment kind of efforts will start to lag more and more behind capabilities over time. But anyway so I think that that's nonetheless worth worth flagging anytime. There's something that cuts against at least my own intuitions which I think this, this would have out the gate and
Andrei Karpathy
now to another sort of related story. Indirectly Perhaps hundreds protest OpenAI, Anthropic and Google in San Francisco. So hundreds of people protested as in like they marched together from OpenAI's headquarters in Mission Bay to the offices of anthropic and Google DeepMind with signs with messages such as AI is not inevitable, pause AI and stop the AI race. We've seen smaller scale kind of protests of this kind before. This is I think the biggest version I've seen. They like have some big signs and there's a lot of them. If you look at images like this is a real protest. It's not sort of a ragtag group of people. Some big names including Alizia Yudkowski were there. This is I think partially by organizations involved here like Pause AI. So not something new. But I think the fact that people are getting more organized and doing more serious, larger scale, more noticeable efforts to convey this message of slow down, stop it, like don't keep making more powerful AI is interesting certainly, and along with the timing is quite appropriate.
Jeremy Howard
I actually saw them as I was walking into the offices of one of the Frontier labs over the last couple of days and it kind of came across. I didn't realize that the protest was going to happen that day. It seemed kind of like an amusing coincidence. Yeah, well, you Know, what can you say, but not surprising, that this is happening right now. Pause AI obviously has a whole bunch of problems sort of reputationally in this space. Not obvious to me that, like, having Pause AI at the forefront of this kind of movement is like the best kind of sort of marketing. Marketing for this. But anyway, you know, it's. It is what it is. And so it is true that I, and apparently well over a thousand other Frontier Lab employees, including many of their, like, executives and co founders, are vaguely sympathetic to this idea. Though you have to ask yourself, what about China? You also have to account for the fact that a lot of these protests, I'm not saying this one in particular, I'm not saying anyone particular at this protest, but are funded by the Chinese, like, unwittingly, typically, you know, this is known to be the case. And I've, like, heard firsthand reports of, like, people with evidence that this has happened in, like, especially the data center infrastructure protests. And so just like, in anticipation of that being a legitimate concern for anything like this, I think the problem is that it's a legitimate concern in every direction. And we're gonna have to. We're gonna have to reconcile that with, like, every protest that we see has either that or, you know, is funded by, you know, lobbyists for some of the big labs or whatever. So it is what it is. Yeah, I guess. Not too surprising. You know, as you say, we've seen stuff like this before. You get a rogue agent on a hugging face server or two and you're gonna get another protest like this. I expect, you know, the protests will grow until morale improves.
Andrei Karpathy
We have quite a lot of big stories this episode, so I guess we'll have to try to power through a few more. We have OpenAI principles for national security partnerships. This was a few weeks ago. I guess we didn't cover it at the time. Kind of a follow up to all the Mythos drama where when OpenAI partnered with the US government and the Department of War, a lot of the criticism there was basically that they capitulated and agreed to having their models be used for, quote, all lawful purposes. So they released this to be more explicit with regards to what they want or allow their technology to be used for. They are not going to be allowing mass domestic surveillance, high stakes automated decisions about human judgments, autonomous use of force, or evading legal oversight. It does allow or does not categorically ban operations or offensive defensive military uses, and has a bunch of stuff in there that basically is kind of making up for a relatively weak statement. Initially of principles. This expands on that and tries to recover some of reputation, you could argue and make it just more explicit on what their red lines, so to speak are.
Jeremy Howard
Yeah, they lay out a bunch of principles which are, I would say like less informative. It sort of reads as like highfalutin kind of AI policy wonk speak. Basically like we're gonna try to like do good democratic things and prevent despotic powers from controlling this stuff and like work with people who share our values and make it good, make it good, make it good. That's the four principles. And then accept four times instead of, instead of two or three. And then they list things, specific things that they won't support. And that's really where all the information is. You know, mass domestic surveillance, they say. So unconstrained collection or monitoring, inferring sensitive traits to disadvantaged people, retaliation for lawful exercise of rights, fabricating evidence that apparently is out as is high stakes decisions made or auto triggered without human judgment. So you can think here about like automated decisions about whether someone meets a legal standard for surveillance or detention. Then there's also use of force without appropriate human judgment. So including systems that autonomously identify, select and engage targets. So that's also out. And finally uses that evade legal obligations, oversight or accountability, including facilitating genocide, crimes against humanity or war crimes. So that's interesting. Things that are not in their exclusions, they have intelligence operations are fine, which I think is good. Investigations, offensive and defensive military operations are not like blanket ruled out. They basically just reject the whole offense versus defense distinction. You can't cleanly distinguish between those which I think is fair and true and they don't ban targeting either. So there's a couple things which I mean on my side being, being a bit of a hawk, like I, I think this makes perfect sense. And yeah, it's just like nice to have them write this out explicitly so that you can see, you know, whether they, they stick to it. Which is always the, the other side of the question, I guess with, with OpenAI. You know, you see, for example, I have. God, I'm old enough to remember when the preparedness framework said something about when, you know, when you have AI systems that just kind of go rogue and do random crazy shit on the Internet and, and kind of get around constraints that that would trigger their like critical, critical security level for loss of control. Cool, cool, cool, cool, cool, cool, cool, cool. But anyway, that's my take and last
Andrei Karpathy
story on safety and kind of on the theme of we're getting to a point where SCI Fi type stuff is starting to happen. This one not related to hacking, but you might argue another serious kind of safety principle. The story is China is banning AI boyfriends and girlfriends over addiction and birth rate concerns. So China banned customizable AI companion apps effective July 15, with regulations being joined by five government departments, including the Cyberspace Administration of China. The rules prohibit AI tools that quote, excessively cater to users inducing emotional dependence or addiction at damaging users real interpersonal relationship. Companies are now required to include instant exit options, regular reminders that AI is not real, and limits on long term emotional memory. So actually by dense Alibaba and Tencent chose to suspend their AI companion features entirely rather than try to comply with these limits. So I mean, kind of a big deal I think, you know, this is not a often discussed story partially because I don't think we have much of an understanding of to what extent people are starting to develop emotional dependence or kind of addiction to chatbots. But it is starting to happen and it could be a serious kind of psychological harm on a society level scale. Seemingly. China believes it could be at the very least.
Jeremy Howard
Yeah, it's. It's also this like weird dynamic shows up. You know, if you remember the, the old replica thing that I think we covered two, three years ago, I can't even remember but you know, people freaking out over the subreddit that their girlfriend or wife or partner had been even
Andrei Karpathy
like very low level AI which was replica. Like yeah, people got depended on.
Jeremy Howard
So yeah, so I mean hard to argue with. It'll happen in some fraction of cases. It also, once you get that right, you get a voting block eventually and then there's no turning back. So you know, it's as ever a question of like how do you.
Andrei Karpathy
Yeah, it's only a matter of time till we get like serious AI personhood discussions and all. People that made fun of it would be like, well okay, you can like discuss personhood anyway, it's only a matter of time onto research and advancements. We've got two stories that we'll try to get through quickly. First, discovering cryptographic weaknesses with Claude. So they have released Anthropic, released that Claude Mithras Preview has discovered improved attacks on two cryptographic systems Hawk, a post quantum digital signature candidate and a reduced round version of AES, the most widely used symmetric cipher. So these are not attacking like actual deployed systems or whatever. This is kind of more theoretical, so to speak. And attacking these kinds of cryptographic systems is kind of going to a base of the security stack, one might say. It's not sort of hacking software per se, it's hacking the foundation of how you make things secure, at least for a category of things. So it's another way to be worried about AI security, like potential for AI to just sort of like undermine the basic mechanisms of cybersecurity.
Jeremy Howard
Yeah. And this, you know, without getting into the details of how these algorithms work, these encryption algorithms work. When you have an encryption algorithm, you're trying to essentially hide the information that you have behind a mathematical operation that is very, very difficult to do and that hopefully is irreducibly difficult to do. In other words, there's no quick hack to like cut right to the core of it. A lot of classical encryption algorithms, just like rsa, just collapse in the face of quantum computers, for example, and then that's like a big problem. Which is why, you know, all the national security agencies have been talking to each other over in classically encrypted channels for a long time. Our adversaries collect all that data and they collect it, it's encrypted when they collect it. They collect, they collect, they collect for decades. And then suddenly someone goes, oh, quantum computers can like just crack this. And it doesn't matter that you start encrypting after that point in a post quantum secure way. They've already collected all of the classically encrypted stuff stuff, which means the moment that they get to quantum computing and quantum decryption, they're able to just suddenly reveal all of the most ultra classified communications that they've been collecting for decades. And so this is why the US and China are like locked in this crazy race to hit like quantum D day, basically. Q Day they call it, that's what they call it. Anyway, there's going to be, you can think of it as a series of starting with minor and then increasingly more and more severe versions of Q day, except delivered by AI in the same way. And I think people are dramatically undercounting how significant of an effect this may have. You already have mathematical theorem proving fields level stuff coming from, from AI models. This will come for encryption. And when it does, I just don't think that we're prepared for the concept because everybody's been thinking about quantum as this one big step that they're all preparing for with quantum secure algorithms, like, you know, for the like post quantum encryption. And AES, by the way, is like supposed to be one of those. So is hawk, actually. And but what we're not preparing for is the gradual chipping away at even Those algorithms, and that's a big problem and that's not going to go away easily. So yeah, a lot of the world depends on encryption. Think about like every financial transaction, all your healthcare records. The very notion of privacy hinges on this. And so fun times, fun times.
Andrei Karpathy
And last science fiction type narrative for the episode you've got aide squared. The first evidence of recursive self improvement from the company Weko AI. The short version is they say apparently this is the first evidence of recursive self improvement. I think this is quite in line with many cases of similar things. Basically they build self improvement systems where you have an autonomous research agent to optimize another agent with an inner loop. It makes a bunch of edits and it, as a result it's really kind of harness level and system level changes. It's not model training or model development changes. There are things like rollout modifications, prompt modifications, kind of monitoring things, et cetera, et cetera. They position this as an early level of self improvement, so you get a net, net positive, faster and better than humans level of engineering. It's not kind of necessarily self improving. Self improving where it's a loop kind of story. I am going to flag my position on the entire family of techniques like this of completely being oversold as self improvement in the sense that you can self improve in these things for sure. You can optimize the prompt, you can optimize the harness, but you're gonna like overfit, you're gonna improve your eval metrics, you're gonna hit Goodhart's law, and then your capabilities will be hurt elsewhere. And until you get to a point of autonomous model self improvement with fundamental advancements and not just acceleration of engineering and like tweaking of hyper parameters and prompts, none of this is a big deal. Although it is cool, I guess.
Jeremy Howard
No, I, I totally agree. I think this is like another one of the long line of like pseudo recursive self improvement things where people like the clout that comes with saying rsi. The, the game of RSI is always going to be identifying whatever the major research bottleneck is and smashing it. And if you can consistently do that with an automated system, then you have achieved recursive self improvement as long as there is. And, and like, I don't know, there's a bunch of arguments that you can define recursive self improvement such that it's been happening ever since life evolved on planet Earth, right? Like, I mean there's an endless everything is a hockey stick when you, when you zoom out Far enough. And so in some sense this is like a now people do mean something by it. Like there is this phenomenon that we will all start to care a lot about, which is just going to feel to us like, holy shit, we're seeing a decade of progress in a week. Like this is not. We need to stop. Yeah.
Andrei Karpathy
We've already been seeing acceleration of the rate of AI progress for years and years, which is partially due to AI, but largely due to just the inherent systems at play and so on.
Sponsor/Advertisement Voice
Right.
Jeremy Howard
Yes. And that acceleration always comes, is the same with startups when you look at their growth curves. It always comes by identifying whatever the single big bottleneck is that the company or the problem has and smashing it. And if you can do that in an automated way, then you have what is conventionally thought of as recursive self improvement. Like AI is doing the AI thing all the way down. Yeah, just a minor side pitch but like. So we're working right now with a bunch of folks in the Frontier labs on defining an actual model of recursive self improvement and like to kind of ground some of the conversations in this in terms of like, what are the parameters that actually matter for recursive self improvement to work. It's a toy model, like really simple thing. But like it's, it's part of, like this is part of the problem. No one knows what the hell they're talking about. I don't mean people are silly, I mean like no one has defined recursive self improvement. We're not going to either. But just like here are some ways to think about it potentially. And the problem is we're getting to the point where we're going to need policy that uses terms like recursive self improvement. And that means that that policy is going to have to define terms like recursive self improvement. And if we don't know how we want to define them, we can't even get our hooks into the thing that we're trying to. Trying to go after. So anyway, there is a thought.
Andrei Karpathy
Yeah. To complement my negative stake, they do kind of discuss a decent amount of stuff. This is a pretty decent research report. They do say that this has out of distribution generalization, meaning it's not necessarily overfitting. Although I, you know, benchmarking is, is very suspect. And they do also say that in a discussion section that like resulting systems are a mess, like this is vibe coded slop and it's impossible to maintain and so on. Which I think is like the story of recursive self improvement where humans can't understand what's going on, but like, not in a good way. It's just like this is a mess, you know. And with that, we are done with this dense episode of Last Week in AI. We will be back to our regular schedule. Mostly I guess we always eventually have scheduling conflicts, but we will do our best. Thank you as usual for listening. We appreciate it if you review, podcast, comment, share and so on. But more than anything, please do keep tuning in whenever we release these episodes.
Podcast Narrator
Time to break. Break It Down Last Week in AI Come and take a ride get the low down on tech and let it slide as we AI come and take a ride Up a labs to the streets AI's reaching high new tech emerging Watching surgeon fly from the labs to the streets AI's reaching high algorithm shaping up the future seized. Come and take a ride Hit the low down on tech and let it slide. From neural nets to robots the headlines pop Data driven dreams they just don't stop Every breakthrough, every code unwritten on the edge of change with excitement we're smitten from machine learning marvels to coding kings Futures unfolding See what it brings.
In this episode, hosts Andrei Karpathy and Jeremy Howard cover two weeks of fast-moving developments across the AI landscape. The discussion spans new frontier and open-source model releases (Anthropic’s Opus 5, Google’s Gemini 3.6 and Cyber variants, Kimi K3, Thinking Machines), major funding and compute deals, strategic partnerships, a high-profile hacking incident involving OpenAI and Hugging Face, evolving safety and policy debates, and a variety of research and synthetic media advances. The tone is direct, occasionally wry, and frequently critical about the emerging risks and industry dynamics.