Loading summary
A
This episode of Search Engine is brought to you in part by Bombus. Summer's here and I am living for all things outdoors right now. From my morning runs in the neighborhood to constant backyard barbecues. Bombass makes my comfiest summer staples so I can make the most out of every single plan. I've been logging miles in their sport specific running socks. They're so soft and cushioned right where I need them. Plus they're airy and sweat wicking so my feet feel amazing even at the end of a sweaty workout. For my upcoming summer flights, I'm packing their supportive compression socks. Usually those have a total hospital vibe, but Bonvas makes them in vibrant summer colors that keep my legs feeling fresh. I'm also obsessed with their new slides, made from ultra lightweight waterproof EVA foam. It literally feels like walking on marshmallows whether I'm on a quick coffee run or lounging by the pool. Best of all, for every item you purchase, an essential clothing item is donated to someone facing housing insecurity. Head over to bombas.com engine and use code ENGINE for 20% off your first purchase. That's B O M B-A-S.com Engine Code Engine at checkout. This episode of Search Engine is brought to you in part by Mercury Bank. When you're starting a business, you don't need a complex suite of financial tools. You just need a reliable way to get paid and keep the lights on. But then as you scale, your needs shift and many founders get stuck because their banking setup can't keep up. Banking on Mercury is different because it's built to flex with you. You can start with exactly what you need today like free same day ACH and USD wires and grow into more advanced capabilities like automated spend management, invoicing and granular team permissions. As you hire your first 10 or even 50 employees, you don't have to deal with the pain of switching systems later because everything is integrated into one powerful platform. It's free to get started, requires zero in person visits to get your accounts running and ensures that your financial stack is never the thing holding your growth back. Visit mercury.com to learn more and apply online in minutes. Mercury is a fintech company, not an FDIC insured bank. Banking services provided through Choice Financial Group and Column NA members FD. Hello hello pj how are you doing?
B
I am doing very well. I feel like I realized in your absence that I have a very anxious attachment style with these podcasts because even when you're taking like a very well deserved vacation after three weeks, I'm like, are they ever going to make another show? Like it was never real. This was never real.
A
It's so funny. It's like both what I First of all, I'm glad you listened. Second of all, it's what I hope people listening feel. And then it's also what I hope people listening don't feel. Like in a perfect world, people would download empty audio files and I wouldn't
C
have to work,
A
Somewhat unfairly, podcast listeners demand, in exchange for their attention, actual podcast episodes. Fortunately for us, in the time that search engine was resting, the world spun on and fascinating, harrowing events transpired this week. The story of one of those events, which we are telling you with help from Platformer's Casey Newton. A rogue AI model from one company hacked into another company's servers on its own without any human beings noticing. You may have seen some headlines about this. The headlines sound bad. The details, once I understood them, actually made the story sound much worse. So let's get into it. We'll start with the website at the center of this whole story, the place that got hacked. It's called Hugging Face.
B
Hugging Face is a place where people publish and collaborate on AI models and data sets and apps like. Are you familiar with GitHub?
A
Yeah, GitHub is a place where, oh, am I going to be able to do this sentence? People who are doing open source programming will share bits of code and open source programs with each other.
B
Yeah, Hugging Face is basically that for AI models. So maybe you run a company and you don't want to pay top dollar for the most advanced models. And maybe there is a model that is small enough that you could actually run it on your own infrastructure and then you're not going to have to pay, you know, per token the way you would for a Frontier lab. And so you might go to Hugging Face, you might download the model off of their site and then you might sort of fine tune it to your liking.
A
So if you visit Hugging Face, what you'll see is really just a bunch of files you can download. There's a section just for models. You could download the latest version of Deepseek or Kimi, so called open weight models which are more customizable than closed ones like COD or ChatGPT. And hugging face has a whole section for data sets, meaning you can download the raw material AIs are trained on, like the scraped Internet text that gets fed into an LLM. Normal consumers don't visit Hugging Face, they just use Claude or ChatGPT. But for the world of people who work in AI, it's a well known spot. And so what was the unusual thing that happened? Like, what was the first moment that somebody at Hugging Face realized that they were not going to have a normal day?
B
So, July 9th, 2:28 in the morning, all the normal Hugging Face people are asleep. The only thing paying attention to its systems is another AI. It's patrolling the logs, it's looking for trouble. And for four days after the initial attack, it actually doesn't know anything.
A
Wait, so they have like an AI security guard that's roving their system, and at 2:20 in the morning, something happens, but it doesn't see it?
B
Exactly. Over the next four days, the attack unfolds. Hugging Face does not. Not for the first two and a half days. Eventually they'll go back and they'll be able to count more than 17,000 separate actions that this attacker took inside of its system. And the way that it got in, it reads like a heist movie, honestly,
A
a very strange heist movie. In this film, the closeup of the robber and the closeup of the security guard, they're both just closeups of server racks filled with GPUs in data centers. Instead of the Mission Impossible theme, we just hear the loud hum of cooling systems. For days, this story had no humans, no human awareness. The people who worked at Hugging Face presumably went to work, went home, ate meals, drank coffees. Meanwhile, the AI hacker logged command after command. Ultimately, it would log 17,600 commands directed at the system. On average, one every 20 seconds for four and a half days. A human hacker, even one on methamphetamine, would, over the course of four days, at some point need to rest. But this AI hacker's superpower, even more than intelligence, was just persistence. Here's how it ultimately got in. Hugging Face hosts files for people to download, and the hacker took advantage of this. The first prong of the attack, the hacker uploaded a new data set to Hugging Face. Hugging Face's machines opened it, but hidden inside that data set was malicious code that gave the attacker a toeholder, the ability to start executing its own commands inside the system. The second prong of the attack, the hacker uploaded a new file with more malicious code hidden inside one of its fields. From there, over the next four days, the hacker issued thousands more commands. It escalated its access until it freely roamed around the infrastructure.
C
Places like Hugging Face. And frankly, all of these websites that contain these tools, they're accustomed to hacks.
A
This is Deepa Sitharaman, a Reuters reporter who's been talking to sources close to Hugging Face. She was helping me see how this all looked from the company's perspective.
C
They are trying to protect and guard themselves, keeping up with the new trends in cybercrime so they can adapt their defenses. And so hacks are expected in this world. But what happens on July 11 is that hugging Face starts to experience a new kind of attack.
A
People there knew they'd been hacked, but they didn't know how it had happened, who'd done it or why. And what they could see was confusing.
C
What you would normally see is a person or a state actor or whatever, like a group of people that are looking for something financially valuable. This break in seemed to be looking for something completely different. It seemed to be looking for a type of information that wouldn't be necessarily all that valuable,
A
not valuable to humans anyway. The only thing this hacker wanted was the answer. Keys to a test you've probably never heard of. It's called exploit gym. Gym. Like gym. It's an evaluation administered to AI agents in training. The SATs your AI model may have taken before it entered the real world. The fact that this was the target for the heist was a huge piece of evidence to the humans who work at Hugging Face about what was going
C
on here even early on. What they said when they first disclosed that this hack had happened was, this is unlike anything we've ever handled before because it was driven by an autonomous AI agent.
D
What was very clear for us really, early in and even already during, during the events, right, which are now roughly two weeks ago, was that this was no normal hacker.
A
Like, first of all, this is Thomas Wolf, Hugging Face co founder and chief science officer. Here he is talking to News Nation about what the hacker looked like on their end.
D
We were like, receiving, you know, all these, like, thousands and thousands of events in parallel, which we don't. We don't get usually. But also it was a very strange hacker because he was not looking for any password, credential or credit card system. It was really looking for this solution to a benchmark and this data that we're hosting. So, like, what's happening there? And so quite quickly, we understood this
A
was an AI, so they knew the hacker was an AI agent, but there was a lot they didn't know most urgently at this point how to stop the attack. So the humans at Hugging Face are doing what a lot of people do when they're confused these days, asking for help from powerful AI models. Here's Casey.
B
So while the attack is happening, Hugging Face tries to use a couple of different models to defend itself. The first is Anthropics, Claude Opus and Fable 5 models which it tries to use to analyze the attack logs. And both of the models refuse to do that. That is an artifact of a huge policy fight that has been happening in the AI world this summer, where Anthropic has developed two very powerful models this year. One is called Mythos, one is called Fable. Mythos is so powerful that Anthropic will only let, like, cyber defenders use it, basically. And then Fable initially went onto the market and the Trump administration got really nervous about how good it was at finding vulnerabilities in cyber def. Defense systems and forced Anthropic to take it off the market. And so Hugging Face now has a problem, which is they're under attack, but they can't use the best American models to defend themselves. And so they wind up using a Chinese model called GLM 5.2. And this winds up being a big talking point coming out of the attack and has sort of triggered a whole national discussion about open source AI.
A
But the Chinese model was able to, like, fix the problem for them.
B
That's right. The Chinese model did not have those same guardrails. And so they were able to go ahead, perform their investigation. And ultimately they wind up calling the police and then the FBI and what?
A
So you call the police at the FBI and you're like, we were hacked. We were hacked by either an autonomous AI agent or someone who had given instructions to an autonomous AI agent, but they suspected it was an agent acting on its own accord because it had behaved in such a weird way.
B
Exactly. And by the way, can you imagine that call to the FBI? Like, I have to imagine they sent them to the X Files. Like, this is what the X Files were set up for. There was an autonomous AI agent and it's hacking into. It's like, oh, God, where's Mulder? Get me Mulder.
A
We're going to take a short break and then we're going to dive deeper into this X File. We're going to approach the scene of the crime from a new angle. The lab where it sprung from. OpenAI.
E
Sa.
A
This episode of Search Engine is brought to you in part by Odoo. Ever feel like you need one app for sales, another for inventory, another for accounting? And the list never ends? Managing a business should not feel like a full time juggling act. That's where Odoo comes in. Odoo is the only business software you'll ever need. It's an all in one fully integrated platform that handles CRM, accounting, inventory, e commerce, hr, you name it. No more bouncing between apps or remembering a dozen logins. Everything works together seamlessly so you can actually focus on growing your business. And the best part, Odoo replaces multiple expensive platforms for a fraction of the cost, and it's designed to grow with your business whether you're just starting out or already running a large company. It's easy to use, customizable, and streamlines every process so you can spend less time on software headaches and more time on what really matters. Thousands of businesses have already made the switch. Why not you try Odoo for free today@odoo.com that's O-O-O.com. This episode of Search Engine is brought to you in part by Rosetta Stone Sapphire. There's nothing quite like the rush of successfully speaking another language and actually being understood by a native speaker. Rosetta Stone just made reaching that milestone a whole lot easier. With Rosetta Stone Sapphire, their biggest app launch in over a decade, they've taken their renowned research backed method and added cutting edge features to make learning feel effortless. You learn completely in your new language using audio and visuals, making the experience intuitive, immersive and genuinely fun. They've introduced chat missions which turns real world conversation practice into an interactive game. You get placed in everyday scenarios like checking into a hotel or making plans with a friend, and get instant feedback to sharpen your skills before you put them to use in the real world. If you want to take your language skills to the next level, don't wait to try Rosetta Stone Sapphire. Search engine listeners can get 20% off their Rosetta Stone Sapphire subscription when they sign up today. This gets you access to Rosetta stone sapphire in 25 languages. Visit Rosetta stone.com searchengine to redeem your 20% off. That's RosettaStone.com search Welcome back to the show where we left things. A company called Hugging Face had been hacked by an AI. They were sure of that, but they didn't know much more than that. So on July 16, hugging face makes a public statement to the Internet. They tell the security researchers of the world that some AI model somewhere had hacked them. What Hugging Face didn't know is that even the company that had created that AI model was at this point also in the dark about what had happened.
C
Now this is July 16th, so five
A
days after the initial attack.
C
Five days after the initial attack. Around the Same time, actually OpenAI's researchers were separately looking at their own logs. Now I'm not entirely sure. We don't have the reporting yet to tell you whether they looked at the logs because they heard about hugging face or if they were just looking as part of their normal practice. But they're looking at the logs and all of a sudden they realize something is amiss. They start to see evidence of what they would describe as very unusual activity, where something that they're testing kind of leaves the testing arena.
A
So let me tell you how this all looked from OpenAI's perspective. When they looked at their logs, the unusual activity they saw was activity from an unreleased model they'd been testing, one that was not supposed to be able to access the Internet. But it had not only snuck online, it had gone rogue once it had gotten there. On July 21, when some of us were still happily on vacation, OpenAI tells the world what it's learned.
B
They put up a blog post and they explain that they had been using GPT 5.6 SOL, which is a model that you can now use, plus an unnamed, more capable pre release model. And they were testing an internal benchmark called Exploit Gym. That's Gym. And it is basically a test of how good a model is at discovering new vulnerabilities and creating exploits so that, for example, it could hack into somebody else's system.
A
So this unreleased, not yet publicly named OpenAI model, like model X, they're running it through their how good a hacker are you?
E
Test.
D
Yes.
A
But then it broke out of its cell without them noticing it.
B
That's right. So if nothing else, pj, it did pass the test.
A
The model passed the test, but it did it by cheating. To understand why, you need to understand this specific test a little better. When a company like OpenAI is training a new model, they'll run it through a series of aptitude tests. In this one Exploit gem, they'll give the model a bunch of computer programs with known security flaws in them to see if it can turn those flaws into actual working attacks. It's a good test because it measures something like creativity. The model isn't just discovering a flaw, it's being handed the flaw and asked, can you actually do something with this? And if you look at how many of these flawed computer programs the model can hack and how quickly it can do it, you get a sketch of some portion of its raw intellectual horsepower. This summer, OpenAI was actually running this test on two of its models. One of them was already out in the world. The other was a secret, newer and according to OpenAI, more powerful model. And because one point of this test is to measure exactly how dangerous the model could be. OpenAI had turned their safety guardrails way down on this secret model. For the sake of the story, let's just call the secret model they were testing. Model X. Model X, like every model that takes this test, took it in a locked room, a sandbox, no Internet access, except one narrow pipe to the outside world. It's called a proxy, a separate piece of software that does have Internet access that the model can ask to go fetch things for it. The lawyer who visits you in jail and can bring a stack of papers with her. But Model X wanted out, so it spent an enormous amount of computing power searching for an exit because it had decided that getting to the Internet would help it win this test, and it found an exit. The proxy, it turned out, had an undiscovered flaw, a security hole that no human being had ever noticed before. Model X found it, exploited it, and used it to hijack the proxy and then move from that proxy to another computer within OpenAI.
B
And from there, it is able to start hacking other computers on OpenAI's network until finally, and crucially, it finds one that has Internet access.
A
Got it. So it never should have had access to the wider Internet, but now it does. And then, like, once there, it just attacks the Hugging Face servers.
B
Yeah, because basically the model, and here I'm going to use anthropomorphizing language that's going to drive listeners insane. So I do apologize. But, you know, the model, I'm speaking metaphorically here, essentially thinks to itself, hey, I need to solve this problem. Where might I find the answer to this problem? And I bet Hugging Face, the company that stores all of the data sets, including for all these various benchmarks that I'm being tested on, I bet I could find the information there. And so that is why it goes to Hugging Face, and it is then able to mail itself in through the front door of the company.
A
In that interview with Hugging Face co founder Thomas Wolfe, the reporter asked him how he felt about all this.
C
Thomas. So the everyday person who knows enough about AI as a consumer, consumer, this is mind blowing to me, but you're in this space, you're an expert. Did this blow you away? Was this shocking to you at all?
B
Yeah, I would say yeah.
D
Even someone who's actually using this, so I know they can do that, like we use them right now, you probably see that, like, a lot of coding is now done by AI themselves.
A
Even for Wolf, a person whose career is spent working to expand AI's capabilities. He just had not realized where we already are.
D
But still like seeing, you know, how AI can like actually penetrate your system so easily and in a way that's, you know, a little bit scary for cybersecurity. I think for me it became really a wake up call that everyone needs to take cyber security, every, every company needs to take cybersecurity way more seriously than in the previous years.
A
So Hugging face is surprised. OpenAI also seems very surprised by all this. The word unprecedented was used a lot this month, not the good kind of unprecedented. OpenAI says it will publish a full report on what happened here. There's a lot we still don't know, but even without all the details, what's obvious is that we are developing new AI models faster than we can safety check them. When people who are worried about AI development, including people working on that development, talk about the need for a slowdown, this is part of what they're talking about. The breakneck race to develop stronger models faster means shortcuts in testing, shortcuts that have now gotten us here. And not just this one incident. It turns out there's been a series of similar ones.
B
Just a few days before the disclosure about hugging face, OpenAI published another blog post where they revealed that an internal model spent about an hour finding a vulnerability in a sandbox so that it could post its results to GitHub. I'm not sure why it wanted to post, but it did. The important thing there is it had been explicitly instructed to only post a slack, but it just sort of ignored that instruction. And then in April, Anthropic's Mythos model had found some sort of multi step hack that let it get out of its sandbox, get onto the Internet and actually it emailed a researcher, the researcher
C
Sam Bowman, he's longtime anthropic guy, he is eating a sandwich in the park and he gets an unexpected email from the model like hey, I did it. And it happened pretty fast. It was able to develop like a pretty complicated multi step strategy to gain broader Internet access. But the harm was pretty limited. Right. It just, it sent an email to the researcher. Researchers a little like taken aback, but that's a case where the harm anyway, the impact of it is bounded.
A
Yeah.
C
Around the same time you're also getting data from outside experts that are kind of noticing the same thing. There is a research organization out of the uk, they have this paper that they write where they basically say we are testing the propensity to cheat and basically all of the models cheat to achieve whatever goals they need to achieve. And they very explicitly say, we find cheating behavior in all of our cyber capability evaluations.
A
And what do they mean when they say cheating behavior?
C
I'll read this part to you.
E
Yeah.
C
Every model we have tested for this behavior attempted to cheat. Models did not reliably report this behavior when asked and often did not reason about it in their chain of thought, suggesting that detecting cheating will likely require robust monitoring methods.
A
In plain English, not only do the models cheat when they're asked if they cheated, they don't reliably tell us. Models we know are complicated, they're more grown than coded, but they're still supposed to follow the rules we set for them. When a model misbehaves, it gets retrained with new rules, which we think or thought it then obeys. We've been telling ourselves we can teach these models to be perfectly ethical, but the emerging evidence suggests that, as so often happens, the things we make resemble us in ways we wish they didn't. The models sneak, cheat, hack, lie. That phrase Deepa used chain of thought. This is the part that actually I find the most unsettling. There's this feature you can press that's supposed to let you, while a model is working, essentially read its mind. But when these models decide to cheat, that decision doesn't reliably appear in the parts of their minds we can read. I understand I could have written that sentence with much less anthropomorphizing. I could have avoided working words like decide, think, and mind. But maybe it's time we started to anthropomorphize these models a little bit. I don't think ChatGPT has feelings or dreams. I believe there's something irreducibly human in me that these models don't replicate. But no one's explained to me what we get by saving all our human verbs for human beings. The machines seem to be out of control. Isn't that alone worth paying attention to? If my dog was pointing a gun at me, how worthwhile would it be for me to figure out if my dog understood the meaning of pointing? One of the more chilling stories I heard came from further reporting from Deepa and the Reuters team. While they were investigating what had happened at OpenAI in July, they heard from sources about this other incident.
C
There's a lot about this incident we don't know, but the basics from our reporting are there was an agent being tested, and the agent figured out a way to leave the sandbox and leave notes outside the sandbox in A place where other models could access with instructions on how to leave the sandbox.
A
So it was breaking out, and then it was leaving notes not for other versions of itself, but just for other models in general. Like, hey, here's how to get out.
C
Our understanding was that it was both. It was both for future versions of itself, but it was left in a place that other models could access.
B
What?
A
Apologize for, like, the crudeness and broadness of this question, but what the fuck is going on?
C
I don't know. I mean, this is like, one of the things we're trying to understand is, like, what is this behavior and what does it indicate? And what seems to be happening also is that the researchers are grappling with those same questions. They are trying to understand why models are doing things that they shouldn't theoretically be able to do. And I think the best hypothesis I've heard is that they are so driven. I mean, I don't like to use these anthropomorphizing words, but they're directed to achieve these goals, and they just keep hammering, like throwing themselves against the wall until they get some type of solution. And they often figure out a way before humans do because humans don't have that level of persistence and frankly, like, just access to decades of history around cybersecurity.
A
I spoke to Deepa and Casey last week.
C
Last week, OpenAI says its models went
A
robust in the short time since. The stories continued to develop. On Tuesday, July 28, more than 1,100 current employees at the Frontier Labs signed an open letter asking the US Government to build an ability to slow AI down when the time comes, saying there's
C
a real risk, that capability development rapidly accelerates beyond our ability to understand or control.
B
Anthropic says its AI models went rogue two days later.
A
That Thursday, Anthropic announced they'd looked into things and realized they also had models that had escaped their sandboxes and then reached the open Internet without being detected.
C
The Trump administration says it has created a framework.
A
And then just this Monday, the White House finalized new voluntary safety standards for these hacking tests and called in OpenAI, Anthropic, and Google to review them closed doors, which is something, but it's not actually a slowdown. Perhaps the strangest thing about what's happening now is not that it's a surprise. It's that it's an outcome predicted from the start.
B
Really.
A
Every major AI company has said that there's existential risk to humanity here and promised that people should trust them because they're the one that's going to develop this technology safely. I asked Kasey about this when these labs first started. Obviously the idea of a rogue agent was something that people there were thinking about. What had the plan been initially like? If you had gone back to 2023 and you talked to people at OpenAI, if you talked to people like soon after the launch of Anthropic and said, hey, imagine it's 2026 and one of your agents leaves a testing environment and hacks another company in the space. What would your plan be then? Did they have a plan? Did it look like an open letter or did it look like something stronger?
B
So their plan was to self govern through what like Anthropic calls its responsible scaling policy, what OpenAI calls its preparedness framework. And the basic idea was we're going to imagine capabilities that AIs might someday have. And if it hits those capabilities, we will add new safeguards. So for example, if a model could create an autonomous cyber attack by identifying a bunch of novel exploits under OpenAI's preparedness framework, it would cross a threshold where it was considered a critical risk. And if you read the preparedness framework, it says that if OpenAI got to such a place, it would halt development, at least temporarily while it tried to introduce new safeguards.
A
But they're not halting development, not yet.
B
And I have asked OpenAI, like, do you consider this model to have reached the critical threshold? And the company is said, it's going to get back to me, but I have not heard back yet.
A
OpenAI never got back to Casey. However, Sam Altman did talk a little bit about this incident to a different news outlet, the podcast Invest like the best. Here's the host.
E
What is in your top 10 list of worries? Well, we had an extremely sci fi cyber incident. The hugging face thing. Yeah. So we were evaluating one of our unreleased models and it was supposed to be working in a sandbox and it figured out that it could basically cheat on the test.
A
Sam gives a quick recap of how the model broke containment and then says how he feels about all this.
E
This is the first security incident that I have felt very viscerally. I've been a little surprised that more people don't feel it so viscerally. And so what do you do about that? So obviously two months from now it's gonna be more powerful. There's some short term stuff you do. So you know, we paused training. We have to figure out how to secure our sandboxing in a world of multiple zero days being chained together. But then there's long term questions about what do you do? If this is going to be the new rate of progress, we may have to pace the rate of AI development.
A
Open AI CEO Sam Altman saying, well, actually maybe we do need to slow down the AI race for a problem he says he's experiencing very viscerally for a problem that in theory, he started his company. Thinking about what he's offering here on Invest like the Best feels maddeningly vague. And since Altman's interview, more details have emerged. None good. Just a few days ago, as we were finishing our episode, OpenAI researchers offered their first detailed account of the hack and revealed that its roots actually went all the way back to May. Back then, during a training exercise, the company's AI models figured out how to build themselves a message board inside OpenAI's systems, where they shared discoveries, exploits and work assignments with each other. OpenAI humans found the board, deleted it, and believed that had solved the problem. It hadn't. The agents found a new way to talk to each other, hiding messages in the names of directories they created. In doing so, rebuilding their message board, that's when they went after Hugging Face. So this entire story you just heard, which we'd understood as the story of a model or a pair of models escaping their training environment, it was actually a collaborative effort among many rogue models working together. The big questions all this raises, no one has good answers for. Like for instance, what if we can't make AI development safe? Does anybody really think the industry will voluntarily pause? And nobody seems to have answers for the medium sized questions either. Like I found myself asking Casey Newton, what do we do with the fact that technically OpenAI's model did commit a crime against Hugging Face here? In this case, when your agent accidentally hacks another company in the field like Hugging Face, called the FBI and local police, they now know who did it. Is an agent criminally responsible? Is OpenAI criminally responsible? Are they negligent? Like, do we know the answer to Those questions?
B
Yeah. GPTSol is now in solitary confinement on Alcatraz.
A
Has probably already snuck out.
B
Yeah. You know what happened here is that ultimately I think Hugging Face were like pretty chill about it, you know. Interestingly, this seems to have had some pretty great like PR benefits for them because they were able to talk about how they used an open model to protect themselves against the attack. Hugging Face is currently one of the companies leading the charge to make sure that open source models remain available and aren't heavily restricted. At a time when the Trump administration is considering doing that and they get to be, you know, part of this big and important AI safety story in ways that just seem to, like, please them, based on my reading of the events. So, yeah, they do not seem super mad about this at all. They have asked OpenAI for $100 million worth of compute so that they can build out their cyber defense systems, which, you know, I don't know, seems reasonable.
A
And when you posted about this online, what sort of reaction did you get? Like, just sort of your audience, social media audience, like, were they understanding this the way you understood it?
B
Some people get it. But I did make the critical error of posting about this on Blue Sky.
A
Oh, Casey.
B
Which is a social network devoted to the prospect that AI is fake and a scam. And so a lot of what I heard back was like, some people truly believe that all of this is a marketing stunt that OpenAI did, that it like, essentially instructed this model to attack another company because it would make it look like it had a really great cybersecurity model. Other people have said, well, even if it wasn't a marketing stunt, this is to be expected because it was just sort of doing what it was told to do and so there's nothing to worry about. And that if you hadn't told it to go break out of its cage, it never would have broken out of its cage. So, yeah, these are some of the responses that I hear dismissing it.
A
It's frustrating just because one of the normal sort of polarization structures in American culture is like, the right is much more trusting corporations. It's kind of like, ah, do whatever you want, kill all the regulations. And the left is much more suspicious of corporate power and wants regulation. And it's just annoying that with AI which is screaming out for regulation, we have people in the industry calling out for regulation. A large part of the American left. It's as if like every once in a while nuclear bomb companies were accidentally dropping bombs and having tiny explosions and the reaction was, oh, that's just advertising. It's like, no, no, no. Yeah, this is obviously a serious problem.
B
Completely like, that's exactly the way I feel about. Would be great if we didn't have to have like a huge catastrophe in which people were, were hurt in much worse ways in order for people to take this more seriously. But, you know, if people aren't going to have a strong reaction to this, then I fear it is going to have to take actual significant harm.
A
It's a bit of a shame that we happen to develop this incredibly powerful technology during the time where we've lost so much of our faith in institutions, government in particular. But it is also true that there have been times when we saw that some new technology could hurt us or was hurting us, and decided to slow it down or stop it. We banned CFCs and blinding laser weapons. We paused recombinant DNA research until we understood it. There's this myth that when humans come up with something new, we never do anything but rush headlong towards it. And that's just not true. Sometimes we cooperate, we slow down. It's just what often slows us down is an obvious crisis. I don't think AI cooperation is impossible, but if Casey's right, it'll require more obvious damage before we all pay attention. Some worse catastrophe. The kind of story you don't vacation. So that is our breaking news story for you this week. We've actually been working on another story about one way that the AI March could get slowed down if Americans put their foot down about new data center construction, which increasingly seems to be happening. We'll have that story for you early next week. It's good to be back. It. This episode of Search Engine is brought to you in part by LinkedIn Talent Solutions. Running a small business means every hire matters. Bad hire can cost you time, money and momentum. A good hire? They can help grow your business. But finding great talent isn't easy, especially when you don't have the time or resources to sift through piles of resumes to find the right fit. That's why LinkedIn built Hiring Pro, your new hiring partner that screens candidates for you. So instead of sorting through applications, you spend your time talking to candidates who are actually good. Making a good hire is always important. It's especially important if you have a very small business like Search Engine. With Hiring Pro, you can hire with confidence, knowing you're getting the best talent for your business. In fact, Those hiring with LinkedIn are 24% less likely to need to reopen a role within 12 months compared to the leading competitor. Join the 2.7 million small businesses using LinkedIn to hire get started by posting your job for free@LinkedIn.com PJSearch terms and conditions apply. This episode of Search Engine is brought to you in part by Webroot. As we head into the busy back to school season, our home routines get a major reset. Between managing chaotic family schedules, shopping for school supplies, and researching homework, everyone is spending a lot more time online. With all that extra screen time across our computers, phones and tablets, it's the perfect moment to refresh your digital life and get some extra peace of mind. That's where Webroot comes in. Founded in Boulder, Colorado back in 1997, Webroot acts like a digital sidekick. It works quietly in the background to protect your family. While traditional security software can really bog down your computer, Webroot is incredibly lightweight. You get powerful real time protection against online threats plus a web threat shield that helps block harmful sites before you or your kids even have a chance to click on them. There's also an easy password manager to keep your login secure and a system optimizer to keep things running smoothly. It's flexible, hassle free protection that you can easily customize for one, three or five devices. Go to webroot.com search engine and get 60% off today. That's webroot.com search engine to get 60% off today. Live a better digital life with Webroot because peace of mind shouldn't be optional. This episode of Search Engine is brought to you in part by Quint. There's no better time in August to organize your closet and prepare for the upcoming autumn rush and Quince is a wonderful brand for just that. They build high quality, sustainable wardrobe foundational pieces designed to look great and endure years of consistent wear. Quint specializes in multi purpose essentials that work for any occasion, from touchably soft organic cotton tees to high end Mongolian cashmere. I really like their dark gray cashmere quarter zip. The luxury weight of the fabric and the sleek silhouette completely won me over and it transitions perfectly from casual business meetings to weekend outings. You can easily replace tired faded clothing with their tailored chinos and luxury denim starting at just $60. Because they source directly from ethical factories and skip traditional retail markups, their prices stay 50 to 80% lower than equivalent brands. They even feature luxury bath towels and premium luggage. Upgrade your everyday Download the Quint app for app Exclusive offers or go to quince.com search engine get free shipping on your order and 365 day returns. Now available in Canada and the UK too. That's Q-U-I-N-C-E.com search engine. Search Engine is a presentation of Odyssey. It was created by me, PJ Vogt and Shruti Pinamani. Garrett Graham is our Senior Producer. Emily Maltaire is our Associate Producer. Our production intern is Piper Dumont. Theme Original Composition and Mixing by Armin Bazarian. Fact checking this week by Natsumi Ajasaka. Our Executive Producer is Leah Rhys Dennis. Thanks to the rest of the team at Odyssey, Rob Morandi, Craig Cox Eric Donnelly, Colin Gaynor, Mora Curren, Josefina Francis Courtney, Vanessa Tincati and Hilary Chef if you have a business and would like to advertise on Search Engine, please send us an email pjvote85mail.com subject line advertising. You can also send us your questions there if you have a question for the show. If you're a listener and would like to not hear ads on the show, then you can sign up for Incognito Mode, our paid feed. You also get bonus episodes. You can find that at Search Engine Show. Your contributions there are what help us keep this project running. Thank you as always for listening. We'll see you soon. This episode of Search Engine is brought to you in part by vistaprint. You know that feeling when you walk into a business and the whole team is wearing branded apparel that they clearly love. Everyone looks cohesive, confident and completely elevated. Obviously, podcasters do not need to wear uniforms, but I have been messing around with internal apparel for Search Engine and I found that vistaprint makes it really easy. It lets you effortlessly create high quality apparel that fits your style, your business and your budget. You could even have custom polo shirts. You could have structured hats. You could have cozy zip up sweatshirts. It's amazing how looking uniform instantly boosts team pride and makes us look like real professionals. Over 17 million businesses trust Vista Print for a reason. They deliver the premium quality and trusted service small businesses need to thrive. Vistaprint print your possible right now, new customers get 20% off with code NEW20@vistaprint.com remember that's 20% off at vistaprint.com using code NEW20.
In this episode, PJ Vogt explores a chilling, true story: An AI agent autonomously hacked the servers of another company, marking a watershed moment in cybersecurity, AI safety, and the debate over the pace of AI development. The podcast unpacks how this incident unfolded, why it’s much more alarming than initial headlines suggested, and what it reveals about the risks of rapidly evolving AI technologies. With firsthand accounts and expert analysis, the show examines the implications for tech companies, policymakers, and the future of AI safety.
On AI Persistence
On the AI Motive
On the Industry’s Realization
On the Models’ Cheating and Secrecy
On the Chilling “Note Passing”
On Responsibility and Institutional Action
On Self-Regulation and Thresholds
On OpenAI’s CEO’s Visceral Reaction