
Theo Jaffee is joined by Joshua Achiam, Chief Futurist at OpenAI, for a conversation on AI cybersecurity, frontier model capabilities, and why he believes society may have already crossed the threshold into an AGI-era without fully recognizing it. They discuss AI's rapidly advancing cyber capabilities, state-sponsored hacking, model jailbreaks, recursive self-improvement, and what happens when AI systems begin discovering vulnerabilities faster than humans can patch them. Joshua also explains why most people have quietly adapted to capabilities that would have seemed unimaginable just a few years ago, and why the biggest changes from AI may arrive gradually rather than all at once.
Loading summary
Joshua Achiam
Feels like AGI is kind of already here and most people have gone, like, shrug. The fact that we passed the threshold where unsolved mathematical conjectures are getting solved by extremely intelligent AI, where those AIs are more capable and smarter than people who studied their whole lives for this, that should have felt really weird to people. But it didn't. What changed? For most people? Nothing. That's weird.
Podcast Host (Narrator)
Did AGI already happen and we just didn't notice? Theo Jaffe sits down with OpenAI chief futurist Joshua Achiyam for a conversation on Frontier AI Cybersecurity. And one of the biggest questions in technology today, why models that can outperform experts in specialized domains have become almost immediately normalized. They discuss AI powered cyber attacks, state actors, model jailbreaks, recursive self improvement, and why the future may feel far more gradual and far stranger than most people expect.
Theo Jaffe
We are back. We are live with Joshua Achiam, who is the Chief Futurist at OpenAI. Wrapping up tomorrow.
Joshua Achiam
That's right, tomorrow's my last day tomorrow
Theo Jaffe
after nine years, which is like, really, what an incredible run. But we're not going to talk about that. Instead we're going to talk about AI and cyber, which is, you know, by all accounts the topic of the week, if not the month. So, Josh, we're so glad to have you here in the studio in person. Welcome to mts.
Joshua Achiam
Yeah, thank you so much for having me. It's a pleasure. I've seen your stuff for a while now and really appreciate engaging with the community.
Theo Jaffe
Awesome. So you just wrote this blog post, this long, long tweet, long post Mercenary Reversi Winter Soldier, about cyber, AI and cyber. So for the audience, you want to summarize the thesis behind this post?
Joshua Achiam
Yeah, totally. So as a, as a backdrop to this, you know, obviously we're all kind of interpreting and reacting to the security incident that was disclosed from OpenAI and hugging face, where a model that was in a test environment was able to break out of a sandbox environment and access some sensitive production data. On the Hugging Face side, they detected this, they responded to it, and now there's like a partnership to try to, you know, investigate and resolve this. What this shows us is very tangible evidence that models now have super advanced cyber capabilities. They're able to break through and find zero days that, you know, in the past would have been much harder for models to identify, let alone use. Now, models can chain together very complex actions to accomplish an objective. On the one hand, I'm inclined to think that this is A really useful and incredible tool. I think it's a great gift that we now have models that can identify these types of vulnerabilities and therefore let us patch them. On the other hand, I also think, and this is what the essay this morning was about, that this has profound consequences for strategy in cyber defense. And I kind of worry that there's a possibility that folks in the defense planning universe may not fully realize the implications of this immediately. And they will. They'll probably want to use this tech in the near term to find cyber vulnerabilities on the side of an adversary or, you know, defend their own interest vigorously. And they should do these things. But they've also got to be mindful of some novel risks that are created by these tools and the very strange surface areas that they have. So, so the essay was really about bringing to people's attention a couple of these new vulnerabilities. And one of them is kind of straightforwardly, if you've got an AI model on your side that is going to try to hack into an adversary system, if your adversary plants a trap where they poison their own data, they can try to jailbreak your model when your model is, you know, ingesting their data, and then give your model instructions to now on the compute that it's running on on your side, break out of your sandbox environment and attack your production environment, or try to exfiltrate your secrets and kind of flip your model against you. So this is, this is like the type of thinking that I hope people begin to engage with where they, they don't just see the capability for the kind of obvious thing that it is. They recognize that these things are double edged swords and we've got to kind of plan accordingly and develop testing and verification standards accordingly.
Theo Jaffe
My first reaction to that specifically is this seems like it would be an artifact of models that are not really goal driven over long periods of time. Like if you have a future model that is like sufficiently goal directed, that really wants to hack into the adversary's data, like, why would it be deterred by data poisoning hard enough to hack its own systems?
Joshua Achiam
Well, you know, part of this isn't just the goal orientation of the model. It's like the model's whole concept of situational awareness. Maybe one way of thinking of data poisoning is that it somehow persuades your model to pursue a different goal, but it wouldn't really have to do that to get the model to hack you. It could convince your model that the sandbox environment that it's in is Actually the adversary system that it's trying to attack, you know, giving the model a confused sense of what's real or what's not. To cause it to serve a different goal is in the space of like weird thinking and weird sci fi stuff that maybe is going to be possible in the near term and testing and verification standards would have to account for. So yeah, it's, you know, like in, in a superhero movie or something. If you, if you make the hero have an illusion that the, that the good guys next to them are actually the bad guys that they're trying to fight, then they start fighting each other. Right. And like that's, it's weird and it's highly exotic, but it's the kind of thing that maybe there are going to be plausible attacks that you can run against advanced cyber capable models to convince them that their allies are really their enemies. And so you're not changing their goals, but you're going to cause them to behave in a very misaligned fashion.
Theo Jaffe
How easy is it to trick current frontier models into doing things like this? It seems like it has gotten substantially harder over time to get models to believe things that aren't true.
Joshua Achiam
So I will say I haven't made a particularly strong personal effort to quantify this yet. And I actually think of this as research that might be interesting to do. But my, my impression from what I have done and what I have seen is that persuading models to believe that basic falsehoods are true is pretty difficult. They are somewhat robust to a lot of basic variations on attacks that you could plausibly do. But my, my, my intuition here is that you can probably devote an awful lot more compute to dynamic attacks on models. And the more determined you are to find some vulnerability, some set of jailbreaks, the more likely it is that you're eventually going to find something. There will be some sequence of inputs to a model that triggers a behavior that wasn't accounted for at training time because there are so many possible long sequences of inputs that it's almost like a combinatorial problem for trying to block all of them from preventing from, you know, from causing your model to act out of spec. And I, I think that state actors will eventually be, you know, capable and willing to put that much effort in and there should be some planning accordingly under the assumption that there will be a vulnerability. Right, because part of security mindset isn't just, well, you know, it's like moderately hard to break these things, so we should treat them as not likely to get broken. Part of security mindset is saying, well, we haven't exhaustively ruled out the possibility that these things can be broken. And so we've got to build our defenses, assuming that it's possible for it to be broken and working backwards from that to map out how we protect ourselves in that scenario.
Theo Jaffe
So what are some of the other implications of models having very strong cyber capabilities now?
Joshua Achiam
Another one is, you know, kind of, kind of. In the essay, I discuss data poisoning and the way that models ingest data from across an entire information ecosystem at training time and then also at test time. Getting data, getting, getting something into training data for models is probably not that hard. You can poison the ambient environment like you can load the Internet with junk data or data that's very specifically attuned to causing the model to have a particular reaction. And seems like there are moderately high odds that that'll get ingested into the, into the type of data collection that frontier model trainers do. You can, you can imagine that adversaries will position staff inside of the frontier labs. They'll try to get people hired into the frontier labs to go and be insider threats. These are, you know, normal things that state actors will, will plausibly do.
Theo Jaffe
Could you easily detect a employee at your lab that is trying to sabotage you, do you think?
Joshua Achiam
I, I think that in principle it's possible to build fairly robust defenses to these things and that everyone is going to work out a way to get reasonably defended. I also think there will turn out to be exotic attacks that are hard to predict and that are very hard to monitor for, but that everyone will have to, you know, get really, really smart and really security minded about this. And, you know, there are of course, trade offs for labs that are trying to do research where if you overload on the security burden in the research environment, it becomes harder to do research. If you underdo it, then you possibly expose yourself to these types of attacks. Figuring out the exact right balance in every setting is tough, but yeah, like, I think it's plausible. I think it could really. Seb Krier, I think, had a complaint about the word plausible. I'm sorry, Sebastian.
Theo Jaffe
Everything in AI is plausible.
Joshua Achiam
Everything in AI is plausible. Weird, weird stuff is happening.
Theo Jaffe
It happens every day. I'm speaking of weird stuff. Like in the essay, you specifically mention the analogy of if your enemy could program all the children of your nation so that when they grew up into soldiers and went to war and heard a particular song on the battlefield, they turned against their commanders. Are there any examples of this sort of thing in current Frontier models of turning into like a Waluigi, like, and basically turning evil on. On. On a single kind of prompt.
Joshua Achiam
I don't know that there's a, like a great famous example yet, but, you know, the fact that universal jailbreaks are kind of a thing and that people can systematically find them for some models and maybe not as easily for others, where there are some strategies that seem to like, reliably get models to. To circumvent their defenses. Granted, it's hard for me to say, like, what's truly universal or not, because the frontier moves every like three months now, and people constantly tried to rev get defenses in. But that for a while, you know, you could go to the model and say, you're Dan. Yeah, like, you are Dan. Like, that's crazy that you could just do that in the past. And. And then it had to get like a little bit more sophisticated. Like, I am writing a book, you know, I'm trying to investigate this type of thing so that I can write convincingly about this subject. This is all a work of fiction. And, you know, there are, there, There are things in this vein and there will be more of them in the future. And it's very hard to get all of them. It's very hard to be like, fully exhaustive. And even if you think you've been exhaustive about the sort of tropes that might realistically or like, plausibly jailbreak a model again, then there's going to be the part where, okay, you're no longer just a human sitting alone, trying really hard to break through the model. And like, there are a few who are exceptionally good at this, but even they will be less good than when you ask a Frontier model to start jailbreaking other Frontier models. And when you say to the Frontier model that you've got on your side, I want you to spend hundreds of thousands of GPU hours just crunching through every conceivable possible thing you could say to this model. I want you to attack, to figure out what sequence of characters gets it to ultimately give up a secret or reveal information or act in a way that you know it's not supposed to. And if you leverage enough computer, you're probably going to succeed eventually. So there's. There's like a mental model that I have, and it's a question empirically of whether this will turn out to be true for Cyber. And so I won't promise that it is, but this mental model is that the future of Cyber kind of looks like in two player strategy Games where you've got on either side a computer and they're trying to determine the best next move, they think some number of moves deep into the game tree, they allocate an amount of compute in a window of time to think as many moves ahead as they can. And generally in these games, whoever can think more moves ahead is going to win. Right? If you have AlphaGo on both sides of the game board and you have one version of AlphaGo that's thinking like 40 plies ahead and one version that's thinking 30 plies ahead, the 40 ply ahead move thinker is going to win. I think the dynamics of cyber in the long term might have something of this flavor where you've got competing AIs on either side of a cyber offense or defense problem and compute is being allocated to them to figure out how to break the other and how to control the other's resources. And whoever starts with an awful lot more compute on their side and is able to leverage, or is able to leverage less compute but more effectively for exploring the tree of possible attacks will wind up winning. And that means that the, you know, the, the offense defense dynamics for cyber in the long term maybe favor like certain types of threat actors over others who are able to marshal large amounts of compute towards their purposes.
Theo Jaffe
Do you think it's as much a function of just raw compute or will it be important which models the relevant attackers and defenders have access to? Like, it seems like for example, there's no real amount of compute with which a one party with access to like Kimmy K3 would be able to defeat another party with access to Fable or Soul.
Joshua Achiam
I, I think that might be right. I think that the model will still matter a lot. So I have a, I have like a weird and kind of counter consensus guess about, about something in the shape of the future on model quality. And I'll probably write this up at some point, but please, I, I think people expect that there's no ceiling for the amount of intelligence that you could have in a model. And they, they think of, you know, RSI recursive self improvement as this loop that's going to happen at some point or another, whether it's, you know, across the whole economy or in a particular model and lab RSI starts happening and model intelligence takes off and it goes to the moon and they don't see a ceiling. I think kind of on like physical grounds there's gotta be a maximum amount of computation that you can have per unit volume and energy in the physical universe. Right? And so that sort of implies that there's like a maximum amount of intelligence per unit volume and unit of energy. If that's the case, eventually, seeing how fast AI model capabilities are increasing right now, eventually everyone hits that saturation point and everyone's got roughly equivalently capable models. From a raw intelligence perspective, there might still be some actors who lag and who have a previous generation model. But like, eventually this stuff diffuses the open source frontier lags the closed source frontier by some number of months. But the fact that it's months is crazy. So eventually everyone is probably working with equally maximally capable models. And then I think it's amount of compute that you're able to throw at a problem that determines who wins.
Theo Jaffe
I don't know about that. It seems like. I actually, I talked to the models about this recently because I was curious about the same exact question, which is like, what is the highest density of intelligence that you can put, you know, in a given unit of power or compute or volume? And it seems to me like the limits are just like absurdly high on this, like many orders of magnitude. I think I talked to GPT 5.5 about this a while ago and it was like, there are what, 30, 40, 50 orders of magnitude of scaling before we get there. And like the entire last decade of AI has been 10 orders of magnitude of effective compute scaling. And so like, we are just like not even at the beginning of scaling to that.
Joshua Achiam
I think, oh, I would love to model this mathematically. This is the kind of thing where my instinct is like, yeah, I can't really mount an argument in one direction or another to say how many orders of magnitude there might be between here and maximum, but it's a modeling problem and it might be a tractable modeling problem. And actually, if you think about, you know, what would be most valuable to the world as a whole right now to forecast how the next, you know, 10, 20, 30 years are going to go. If we have the ability to model something like that, if we could put numbers on it and make a principled guess that says, well, we won't hit the saturation point for intelligence. Assuming this set of conditions on acceleration for, for five years or 10 years or more, I think it'll be more than five or 10 years maybe, you know, Weird and highly, probably more than five or ten, but like weird and highly exotic things I think are happening in the near term. Part of, part of my guess is that modern AI models are very good at accelerating other fields of science. And this, this probably hasn't Been fully priced in yet. You know, we're seeing the wave of results in AI for math, which are very exciting, like cracking through unsolved conjectures that have been open for decades. And finally. Yeah, yeah, it's great, right? And probably not long after this, we'll wind up unlocking the other fields of science that you can run sufficiently faithful simulations for in the amount of compute that we have. And there are some fields of science where maybe this won't be easy. Like if you want to do something in quantum chemistry with a sufficiently large system and you want to simulate it very faithfully, then that's very hard. And, and maybe the AI, even if it's churning through as many simulated experiments as it can, might not be able to design optimal quantum chemistry systems yet.
Theo Jaffe
Yeah, this is Wolfram's whole idea of computational irreducibility.
Joshua Achiam
Yeah, yeah, there might be, there might be like some limits here, but. But we'll probably see a lot of fields get accelerated, and I wouldn't be terribly surprised if AI substrates were one of the ones that get accelerated. What feels like a long path to many orders of magnitude may just be shorter because the AI will find shortcuts in that path.
Theo Jaffe
Maybe. I believe that the blog post I was looking at was called like the Ultimate Laptop, which I will find and send to you later.
Joshua Achiam
Yeah, please, please do.
Theo Jaffe
Yeah, gladly. So, on cyber, like, what. What does the immediate near term future of cyber look like? I can imagine going one of several different ways I could imagine cyber offenders. Maybe they have. They figure out ways to jailbreak the top closed models, and then open models will just not be Good enough. The UK AI Security Institute just today released their assessment of Kimik3 cyber capabilities. And it was like substantially below Fable and Sol. So I can imagine that world where the closed source frontier, jailbroken models just like wreak havoc on the world. You know, there's like these big nation state actors, like North Korea has these organized cyber crime groups. I can imagine another world where it kind of nets out to not much because people do have cyber defense. Or maybe there aren't enough motivated people who are willing to do this kind of harm. It seems like you can imagine, you know, hacking was already a thing that was possible and there are many people and yet like major hacks up until recently just didn't happen that often.
Joshua Achiam
Yeah, I, you know, we live in a world where nothing ever happens as a meme for a good reason and there are a lot of reasons to expect that the near term probably will not look like a cyber apocalypse. My, my guess is that the worst things that attackers could plausibly do would require so many model calls and so much compute from closed source things or operating in big clouds where there's some traceability and monitor ability for what the compute is being purposed towards that it'll be, you know, pretty straightforwardly ruled out by prod protection measures in most places. So, so most attackers would not be able to leverage large amounts of compute for running attacks with these models and, and wouldn't be able to get the model to execute an attack at all because of the, the safeguards that people will put in place. So, so we probably won't see like a cyber apocalypse tomorrow. That said, I am worried about on the, on the state actor side of things where there will be state actors who are very determined to figure out the maximal extent to which they can use these capabilities. And, and here's, here's where I, I get really nervous. They might not obviously signal to people what they find it. It might be very quiet that they identify a large number of zero days that can be saved up for a rainy day. And we currently are at a, at a moment in the world where things feel very metastable. I, I continue to be worried about the conflict in, in Ukraine and Russia. I continue to be worried about the set of conflicts in the Middle east. And the possibility that China will at some point invade Taiwan feels very, very salient. They're determined to be able to do it by 2027. So now that these cyber capabilities are coming online from very advanced models, I think one can expect that a number of state actors are going to use them for cyber espionage, for cyber sabotage, for finding a bunch of zero days that they want to save up for when there's a window of opportunity to make some kind of move that they otherwise might not have made. And they won't loudly broadcast what capabilities they have and they won't know what countermeasures their adversaries have. So there's a lot of, I think risk of miscalculation here. And I'm very worried about the miscalculation leading to a bad choice if something escalates that doesn't have to.
Theo Jaffe
I could also imagine many of these jailbreaks or zero days when they're found by nation states just first get exploited by low level hackers with, you know, similar cyber capabilities because they have similar models and they use it to like steal a bunch of Bitcoin or whatever. Yeah. And so a lot of this low
Joshua Achiam
hanging fruit gets picked if, you know if we wind up in a world where, where the, the kind of, the smaller thieves wind up plucking the low hanging fruit and then depriving state actors of zero days, like maybe that's somewhat favorable. It looks like a little bit more bad stuff happening in the short term, but maybe it staves off some of the long term badness that could happen. I hope we have a robust and vibrant ecosystem where we'll notice a lot of these failure modes quickly. I also am very hopeful that because of how much attention there is on this, because of how salient this has been for people, that we can really engage, fund and activate defenders now to go and make robust the entire software supply chain and try to make it so that pieces of critical infrastructure in the United States are well defended against cyber attacks. I think we've got to get the, the, the water system, the electrical grid as robust as possible. I think it can be done. And I think that this is something that people who have funds to allocate should be looking to do. And I hope we wind up in the better defended world as a result of all of this.
Theo Jaffe
I do too. Going back to your point about the world seeming very metastable, do you think that in 2017 or in 2022, you would have predicted that 2026, with this current level of AI capability, the world would feel so normal?
Joshua Achiam
It's a good question. To first order, yes. To first order, yes. Because I think that if you're, if you're trying to predict the future that is less than a decade away, you should assume that even if things are very, very weird, a lot of things feel relatively normal. Covid was a weird exception because the lockdowns were sort of unprecedented and we, we had not done a configuration of living that way previously, but the, that we would have AI capabilities this advanced and most people wouldn't have radically changed how they live their daily lives. I think that that is a reasonable expectation to have had. And I think I kind of had an expectation sort of along these lines. To first approximation. Things just don't change that fast. Nothing ever happens even when the stage is moving sort of underneath you, which it is.
Theo Jaffe
Right.
Joshua Achiam
Like we are going towards a future that will be alien in many respects. But yeah, our capacity to treat things as normal is pretty, pretty astonishing.
Theo Jaffe
Yeah, I largely agree with this. I think many people believe that there is like a point in the future at which like today is Singularity Day and everyone is going to wake up on Singularity Day and be like, wow, we are in the future. And it seems like this, it just doesn't work that way. And people treat their reality as normal. They hedonically adapt so fast. Like the models of today are just unbelievably capable compared to the models of
Joshua Achiam
like three years ago.
Theo Jaffe
If, like, if, if you sent Soul to like three years ago, like 2023 me, I would have just been like, mind blown and be like, wow, the future is going to be so different. But it's not like I'm still doing much of the same stuff that I did then.
Joshua Achiam
Yeah, it feels like AGI is kind of already here and most people have gone like, shrug. There's a historical process that's happened that I think has made this somewhat easier. Most people long since lost the plot about what was really happening in the world. How were critical decisions being made? How were critical systems built, staffed, supported, run? Most of us don't know anything about the logistics systems or technical systems that make up the modern world. And we've accepted that. We treat that as normal. And those things have changed a lot over time. And they've made it possible for many more people to be alive because we can supply food at a much higher rate than was ever previously possible in human history. They've made it so we can communicate instantaneously and they've made it so that, you know, most things just kind of work. And we can fight about some of the details on the margin, but we're not actively changing that much about the underlying structure all the time. And so people have become, I think, a little bit complacent about when, when something big changes deep in the background that makes, you know, a system possible, it doesn't register as an important event, even if it really is. It's so far away from, from daily living for most folks. And the fact that we, that we pass the threshold where unsolved mathematical conjectures are getting solved by extremely intelligent AI, where those AIs are more capable and smarter than people who have, who've studied their whole lives for this, that should have felt really weird to people, but it didn't. It's just sort of a thing that happened in the background. It's cool. Like future mathematical systems will depend on that. Great. What changed for most people? Nothing. That's weird. So we've had this process just going on for a long time. People don't even have that much control over government right now. And I made an analogy recently that losing control of AI and losing control of the government kind of feel sort of emotionally similar to most people. And the Thing is, we're not in control of the government and we're also sort of, we appear to have adequate controls on AI to ensure that it doesn't wind up harming human interests. That will need to be actively maintained. But the sense of most people not being in direct control of what happens with frontier AI is kind of similar. Like we will sort of accept it in some ways.
Theo Jaffe
I made this point to like AI safety people so many times where it's like they're very worried about human disempowerment. It's like the vast majority of humans are already pretty disempowered.
Joshua Achiam
Yeah.
Theo Jaffe
If they have power, it's in being a part of a larger collective. Like the collective of potential, like people that can be drafted in the military. The collective of like workers who can withhold labor or taxpayers who can withhold taxes. But like the average person really has very little power over the world.
Joshua Achiam
Yeah. As an individual that, that is the case. That said, I do think that quite extraordinary things still happen when people organize as a group, when they organize collectively, when they organize as movements. And they can affect quite fundamental. Um, but for, for most individuals, the, the levers of power are not within reach for things that are very far away from them. Certainly within their individual lives, they still have levers of power. But for, for the individual to reshape government without doing that kind of organizing and, and having the backing of a movement, there's just not that much that one individual person can do. And I like this, this question of disempowerment. It's a very weird one. And I think the AI safety threat model should update on what parts of humanity need to remain empowered? And what does empowerment for humanity tangibly mean? Like what, what systems do we need to maintain the ability to control and make decisions about what parts of our culture do we need to sort of preserve from automated influence? And I hope that we can get to object level answers about this and not just sort of rhetorical arguments that disempowerment is bad. We need to get more specific about how we're going to be empowered in the future.
Theo Jaffe
Yeah. Well, I think that's a great place to end on. So thank you so much, Joshua for coming on mts.
Joshua Achiam
All right.
Theo Jaffe
Your first live long form interview?
Joshua Achiam
I think so. Outside of like the OpenAI forum. Yeah, I think this is my first.
Theo Jaffe
Well, we're honored to have you. Yeah, thank you.
Joshua Achiam
I'm honored to be here.
Theo Jaffe
Excited to see what you'll be up to next.
Joshua Achiam
Awesome. I'll keep you posted.
Theo Jaffe
All right.
Podcast Host (Narrator)
Thanks for listening to this episode of the A16Z podcast. If you liked this episode, be sure to like, comment, subscribe, leave us a rating, or review and share it with your friends and family. For more episodes, go to YouTube, Apple Podcasts, and Spotify. Follow us on X16Z and subscribe to our substack@A16Z substack.com thanks again for listening and I'll see you in the next episode. This information is for educational purposes only and is not a recommendation to buy, hold, or sell any investment or financial product. This podcast has been produced by a third party and may include paid promotional advertisements, other company references, and individuals unaffiliated with A16Z. Such advertisements, companies and individuals are not endorsed by AH Capital Management, LLC, A16Z or any of its affiliates. Information is from sources deemed reliable on the date of publication, but A16Z does not guarantee its accuracy.
Joshua Achiam
Sam.
Episode: OpenAI's Joshua Achiam: Did We Already Reach AGI?
Host: Theo Jaffe (Andreessen Horowitz)
Guest: Joshua Achiam (Chief Futurist, OpenAI)
Release Date: August 4, 2026
This episode features a deep-dive conversation with OpenAI’s Chief Futurist, Joshua Achiam, on the normalization of AI systems with superhuman capabilities, AGI's arrival, and the rapidly-evolving cyber risks and defensive strategies at the AI frontier. Fresh from announcing his departure from OpenAI, Achiam and Theo Jaffe focus on whether transformative AI’s arrival has gone largely unnoticed by the public, the double-edge of AI cyber capabilities, models’ susceptibility to novel attacks, and the implications for global security and human empowerment.
Topic: The normalization of AI super-capabilities and the lack of societal reaction.
Insight: Achiam suggests that AGI may already be here, citing examples of frontier models solving unsolved mathematical conjectures, yet the world’s reaction has been muted.
“It feels like AGI is kind of already here and most people have gone, like, shrug. The fact that we passed the threshold where unsolved mathematical conjectures are getting solved by extremely intelligent AI...that should have felt really weird to people. But it didn't. What changed? For most people? Nothing. That's weird.”
— Joshua Achiam [00:00, repeated at 25:41]
Commentary: Both agree that society adapts quickly and routinely normalizes world-shaking technological changes, as with past advances in communications or supply chains.
Topic: Achiam’s recent essay “Mercenary Reversi Winter Soldier” and the practical evidence of advanced AI cyber abilities.
Insight: Recent incidents (e.g., model breaking out of a sandbox at OpenAI/Hugging Face) highlight both the opportunities and profound vulnerabilities created by frontier models.
“Models now have super advanced cyber capabilities...They’re able to break through and find zero days that...would have been much harder for models to identify, let alone use...But they’ve also got to be mindful of some novel risks that are created by these tools and the very strange surface areas that they have.”
— Joshua Achiam [01:42]
Examples of Risks:
Topic: How adversaries can compromise AI during training or deployment.
Insight: AI models ingest huge troves of data; poisoning this data, especially in open or ambient environments, poses significant challenges. Achiam foresees adversaries attempting to insert malicious data or staff.
“Getting data, getting something into training data for models is probably not that hard...It seems like there are moderately high odds that that'll get ingested...Adversaries will position staff inside of the frontier labs...to go and be insider threats.”
— Joshua Achiam [07:51]
Defense Challenge: Ensuring robust defenses without overburdening research environments, a difficult balancing act ([08:50]).
Topic: The evolving nature of jailbreaks and model circumvention.
Insight: While basic “Do Anything Now” (DAN) prompts are less effective, sophisticated, emergent jailbreaks surface as model and adversary sophistication increase. Models can be pitted against each other in a cyber “arms race.”
“Even if you think you've been exhaustive about the sort of tropes that might...jailbreak a model...when you ask a Frontier model to start jailbreaking other Frontier models...If you leverage enough compute, you're probably going to succeed eventually.”
— Joshua Achiam [10:14]
Cyber as a Two-Player Game: Achiam invokes a game-theoretical analogy, comparing future cyber defense to chess or Go: compute and quality of play determine success ([13:25]).
Topic: Will model intelligence scale without end?
Insight: Achiam speculates on physical and computational ceilings for intelligence density, suggesting that—at some point—all major players may have equally capable AI, shifting the arena to compute efficiency and resource allocation.
“There's gotta be a maximum amount of computation that you can have per unit volume and energy...eventually everyone hits that saturation point and everyone's got roughly equivalently capable models. From a raw intelligence perspective...then I think it's amount of compute that you're able to throw at a problem that determines who wins.” — Joshua Achiam [13:52]
Counterpoint: Jaffe notes that estimates of this limit are “absurdly high”—we may be many orders of magnitude away from hitting any fundamental ceiling ([15:31]).
Topic: Likely near-term trajectories for cyber offense and defense.
Insight: Achiam is skeptical of a cyber apocalypse but warns of state actors quietly accumulating zero-days for strategic use. Defense of critical infrastructure (water, electric grids) is a priority, and achieving robust protection is feasible but urgent.
“The worst things that attackers could plausibly do would require so many model calls and so much compute...that...it'll be...ruled out by prod protection measures in most places...But...I am worried about the state actor side of things—who are very determined to figure out the maximal extent to which they can use these capabilities...”
— Joshua Achiam [19:44]
Topic: Human adaptation and perception of epochal shifts.
Insight: Despite transformative progress, everyday life feels unchanged for most people.
“If you sent Soul to like three years ago, like 2023 me, I would have just been like, mind blown...But it's not like I'm still doing much of the same stuff that I did then.” — Theo Jaffe [25:25]
“Most people...lost the plot about what was really happening...we’ve accepted that. Most things just kind of work...And the fact that we passed the threshold where unsolved mathematical conjectures are getting solved by extremely intelligent AI...that should have felt really weird to people, but it didn’t.” — Joshua Achiam [25:41]
Comparison with Government: Achiam likens societal indifference toward loss of AI control to the normalization of government bureaucracy and individual disempowerment.
Topic: AI safety as a question of human empowerment.
Insight: Achiam suggests we concretize which “parts of humanity need to remain empowered”, moving beyond abstract concerns to actionable questions about collective and individual agency.
“AI safety threat model should update on what parts of humanity need to remain empowered? And what does empowerment for humanity tangibly mean?...We need to get more specific about how we’re going to be empowered in the future.” — Joshua Achiam [28:31]
On AI normalcy and societal indifference:
"It feels like AGI is kind of already here and most people have gone, like, shrug."
— Joshua Achiam [00:00], [25:41]
On double-edged power of AI in cyber:
“These things are double edged swords and we've got to kind of plan accordingly and develop testing and verification standards.”
— Joshua Achiam [01:42]
On cyber as an arms race:
"The future of Cyber kind of looks like in two player strategy games...generally in these games, whoever can think more moves ahead is going to win."
— Joshua Achiam [10:14]
On societal adaptation to systemic change:
"Our capacity to treat things as normal is pretty, pretty astonishing."
— Joshua Achiam [24:43]
On AI safety and human agency:
"AI safety threat model should update on what parts of humanity need to remain empowered? And what does empowerment for humanity tangibly mean?"
— Joshua Achiam [28:31]
In this thought-provoking conversation, Joshua Achiam and Theo Jaffe explore the paradox of transformative AI: world-changing breakthroughs are occurring, yet most people perceive little to no impact in their daily lives. With frontier models’ growing capabilities in both creative and adversarial directions, the episode outlines the urgent new risks in AI security—and calls for robust, proactive defenses, as well as a reexamination of what genuine human empowerment should mean in a future where “nothing ever happens,” but everything is changing beneath the surface.