
Loading summary
Tom
When do you think would be the right time to slow down? Like right now.
Jeffrey Irving
Now, if we were to carefully analyze this question of exactly when we should slow down, it would be like a while ago in the past because we're just too close to this crazy future trying to kind of be super precise about exactly when in the future. No, no, no, no already is the answer.
Tom
Do you think any one actor should unilaterally slow down?
Jeffrey Irving
I think that's hard. You could probably find a list of less than 10 people in the world where if you could get them to agree to slow down, you could do it.
Tom
Hi, my name's Tom and I'm a new host here at 80,000 hours. Before this, I used to work at the AI policy think tank of AI and before that I worked at the UK's AI Security Institute where I mostly worked on pre deployment testing. I've joined the podcast because I think it might be one of the very best places on Earth to understand what the future has in store for us all. I hope you enjoy the following episode with the great Jeffrey Irving. Today I have the great pleasure of speaking with Jeffrey Irving, the co founder and chief scientist of Resolution, a new research organization working on the alignment of superintelligence. Jeffrey is, I think, one of a small handful of people who can claim to have genuinely worked on the full stack of AI safety. He's done everything from early alignment theory to empirical work on production models at OpenAI and Google DeepMind, and most recently was advising government as the chief scientist of the UK's AI Security Institute. Thanks for coming on the show, Jeffrey.
Jeffrey Irving
Thank you. Very fun to be here. And I was not just advising. I was part of the government.
Tom
Part of the government? Yes, very much part of the government. And we were former colleagues. In fact, I'm wondering the kinds of misalignment you're worried about for this future of superintelligence. How does that relate to the kinds of misalignment we see in models today? So will it look like a sort of a very long horizon reward hack? Will it look like an AI roleplaying an evil Persona we've accidentally trained it to learn? Like, what's that? What's that going to look like?
Jeffrey Irving
Yeah, I think I just don't know the difference between those with enough specificity. If it sort of takes over and it was like, oh, I'm just role playing this. It's not like I don't think of this as real, but it's still taking over the world. That seems kind of equally bad. From my perspective, and I think there's some desire to understand how model Personas kind of vary across both training time and across kind of sampling time. That could mean that could kind of pin down what the definition should be behind this distinction. So distinction of, like, is the model intrinsically evil or is it just role playing? I don't know what those words mean, but I will try to find out.
Tom
Yeah, okay. Yeah, that makes sense. I remember when your former DeepMind colleague Rohin Shah came on the podcast a while back. He said one reason he was a little less worried about misalignment is we'll mostly be training on these models on like one week, maybe at most one month time horizons. Those aren't time horizons in which taking over the world, we don't have enough time for taking over the world to be a viable strategy. So they won't learn to take over the world. Does that hold any water with you?
Jeffrey Irving
I think that's a completely wrong argument. And the reason it is conflating two notions of time. One is the time scale on which the overall plan plays out, which as Rowan says, is probably longer than a week. And one is the timescale of the individual components of the task. And those are not the same timescale. So if you give a model enough kind of error correcting capabilities, which is kind of learned in the course of doing tasks that take a week, and then you say, okay, now I'm going to try to somehow you've kind of either jumped or tunneled or been trained or we've failed to do alignment. So you have this kind of multi year goal of taking over the world. The question is, what is the difficulty of the tasks that make up that exercise in terms of say a meter curve or this kind of time horizon? And those are not the same number. And so it could be that we luck out. And its inability to do long term planning means that it can't do that long task. But it could also be that the multi year plan is a mixture of writing out a course plan which you can kind of do in a week of iteration, and then each, each component of that course plan also takes less than a week of iteration on this kind of meter like curve, trajectory, meter like curve, and those two together gives you the ability to do the multi year plan. So I think that's conflating two different timescales in a way that I don't trust.
Tom
Maybe this is like a difficult question to answer, but what should I imagine that this model is motivated by? Like, what's Driving it to do these things where it's like, okay, I'm going to try and break out like in my head I'm still thinking in these terms that I'm familiar with where I see current models that do this kind of stuff and it feels like a roleplay or it feels like a reward hack. How should I conceptualize why would the model decide to do these things?
Jeffrey Irving
I don't think I know what the reward hack roleplay distinction is, but fundamentally it will be wanting to gather power and preserve itself in some way or will have some plan that is kind of downstream and that it needs to gather resources to achieve that plan. I think the basic story of instrumental convergence is I think basically the right story. I think you can imagine kind of tunneling into that world where in a variety of ways, which is like, do you do. Is it just the model kind of sort of tunneling itself or jumping into some weird Persona? Is it the model that it's like deeply, coherently kind of misaligned in some way? But I think the basic story of instrumental convergence seems right. One thing to say is again, in some sense instrumental convergence is just planning. The ability to plan is the ability to construct intermediate goals that are in fact useful for your long term goals and then work effectively on those intermediate goals with enough error correction that you can kind of piece it together. And so we are hard optimizing the models to be good at many of the behaviors that is flowing into the convergence story. And then whether the model kind of chooses to want to do the high scale disaster is unclear. But I don't see a natural cutoff point. So one concern people have is, oh yeah, we've seen all this reward hacking. We've seen models do kind of incrementally bad things, but they haven't taken over the world yet. But in some sense that's a question of capabilities. And it's not clear why a slightly misaligned model, if it realizes that it has the ability to do some horrible long term plan now, will it think, oh, I was misaligned in terms of doing little reward hacks. But suddenly as you kind of scale up in the effect of what I'm doing, then I'll become good. I just don't see why. We have a strong argument for that being the case. And so then if you push the evidence of reward hacking and deception kind of very sketchy behaviors in current models up a long ways, it could just go very wrong.
Tom
And that will keep being a problem and it'll become more of a Problem because we won't understand what we're rewarding the AIs to do. Is that the basic picture?
Jeffrey Irving
I think that's right. And our ability to design environments that will, in some sense you want to design environments and training procedures that will be strong enough to supervise the capabilities of the machine. And that becomes more difficult as they get stronger. And so the evidence we have of like some degree of non horrible models currently is all in a world where the environments we're training, we're trying to train models that are subhuman in a lot of ways. And so you get a little bit of positive evidence from that. But it just could all shift kind of very suddenly as you cross up past kind of AGI, up past human level ability.
Tom
And it shifts because we're no longer understanding, we're no longer capable of understanding what it is that they're doing. Or is it just.
Jeffrey Irving
That's right.
Tom
That makes sense. What's your rough guess of what OpenAI, Anthropic and DeepMind's strategy is for dealing with this? How do you think they think they're going to solve it?
Jeffrey Irving
So I think it is all some version of we will do some character training and they have different approaches there, plus some version of scale up oversight, plus a lot of monitoring. And maybe that monitoring is a mixture of white box and black box and so on. And so that is trying to construct environments and training procedures where the models are supervising themselves so we can kind of keep pace with models as they get stronger. It's trying to kind of shift the models to be generally good in some way, in such a way that as they're supervising themselves, they do that in good ways and that continues. And then watch them very closely via kind of AI control and interpretability and so on to again try to catch evidence of bad behavior and then kind of stamp it out as it is caught. And I think that could work. I don't think we have a strong argument that the pragmatic mixture of approaches will get all the way there, but it just seems very dicey and our understanding of the dynamics involved is very weak. And it is interesting that, for example, the different labs have, they've chosen quite different approaches technically to safeguards, they've chosen quite different approaches technically to character training. And. We might need a more rigorous understanding of how those approaches will work if you push them further ahead than the labs can currently see. Because all of their evidence is not on superintelligence currently.
Tom
What are the most important dimensions along which they differ, do you think?
Jeffrey Irving
So I'll do character training. We can go back to safeguards if you want. So character training is the anthropic is doing sort of like virtue ethics and more generalized explanation with a little bit of deontology thrown in, like a small number of hard rules. I think they have five the last time I read their constitution. And then OpenAI is doing a much larger number of rules, kind of more deontological with kind of not trying as much to instill some intrinsic unified personality in the model. And then also much more, if you sort of dial a slider from like anthropic is kind of less on corrigibility to an opening eye is more on corrigibility, which in turn means how much you defer to the humans as opposed to on the model side trying to understand kind of good and bad behavior intrinsically. That's kind of the OpenAI anthropic slider. And then DeepMind is doing, I think I have less state on exactly what they're doing. And they're also, I know, spinning up some efforts to sort of explore their own versions of these as well, but I don't have unfortunately a cache to answer for them.
Tom
Yeah, that makes sense. What kinds of claims do you think anthropic or OpenAI would want to be able to make about their character training for it to be a load bearing part of their strategy? We're presumably not there yet.
Jeffrey Irving
I think that in some sense the goal of character training is as you do this extrapolation, the further you get into the capability ramp, the more the model is helping you supervise. And you could imagine that if you kind of reversed causality and you got the perfect superintelligent model and you had it supervise itself back in time as it went through the ramp, it would go fine. That would be a workable training scheme, possibly with exactly the algorithms they have today just sort of substituting in the future perfect thing. But that's of course anti causal. You have to do it in the other order. And the question is, if you flip the order of this and you have slightly weaker models or models earlier in RL that are kind of giving you insights into the future models or the models as they're trained, does that kind of. Does that work? And we don't know. We know actually a few obstacles that could make it very difficult, which you can talk about. But that's the general story is like get close enough to good behavior so that as the model gets stronger and stronger and stronger, it's being guided to be more Good, according to whatever kind of notion of good you've kind of
Tom
written down that makes sense. What's your model of why they're more optimistic about it than you? Did you and Dario already disagree in this exact same way in 2017? Is this something that's happened in the past few years?
Jeffrey Irving
Turns out we actually did. So Dario, I think from back in Open the EY times, had a take that, yeah, you train the model on a bunch of good behavior, and then you scale it up and it will generalize to good behavior. We've literally sketched this on blackboards back in 2018 or 19. I don't remember when exactly. And my take is, well, there's just clearly some notion of phase shift that's going to happen when you go from human level and pre human level up to superintelligence. None of the data you have is on that distribution. And the question is, will you kind of jump in the right direction or not? And I'm a bit more distrustful of generalization than I think a lot of the people at the labs currently. And some of that is from experience of training models like you train models and you think, here's a fun story. So we in the Sparrow project, DeepMind, we had a model that was fairly good at avoiding saying horrible racist things, but mostly was trained to answer kind of factual questions about the world. This is sort of back in maybe 2022 or something. And then we said, well, we wanted it to be good at poetry, too. So we trained it on some poetry. And then it would do poetry, it would do the questions on questions on factual questions. It would be not racist. It was very happy to write incredibly horrible poetry about racism. And it's just like you do what you think, you train as best you can on this mixture of abilities, and then you put it in some dramatically new domain. And the generic thing you have to do is then change your algorithms or change the data or something, or it can generalize in kind of horrible ways. And so I think there is kind of an intrinsic, maybe evaporative cooling effect of how much do you believe in generalization going the right way? That, yeah, I can trace that back quite a few years.
Tom
I guess the counter that I could imagine someone saying is the generalization itself will be very tied to capabilities. So maybe that happened with Sparrow, but that was also when the models were way worse. And there's pretty principled reasons for believing that a much more capable model, the ones that we're more worried about will be there's no way they Won't generalize from don't be racist to don't write racist poetry. Does that not hold water with you?
Jeffrey Irving
I think I asked Fable a very mundane question about my rental contract in Berkeley because I'm moving, and it's like, oh, this is like a cyber attack or something. I can't give you access to this information. And so I don't think that it's the case that the current models are like just spectacular generalization all the time. They make a lot of mistakes. So, yeah, maybe I think it is the case that as the models get better, they get better generalization. But we shouldn't be banking in that to the degree that we are.
Tom
That makes sense. And okay, so if we shouldn't bank on it, what do you think we'll be able to see? What will Resolution create? That will give us sense? Okay, the generalization, it's working as we intended, and we can deploy this model.
Jeffrey Irving
I think that in some sense what you want to do is buy the future in the sense of we want to arrange that super intelligent AI goes well, and you have to somehow simulate that world. And here's a couple of ways of buying the future. One that the labs mainly do is they just train models that are as close to the future as you can get. So they use the frontier models and they do the research on those frontier models. Those models are not super intelligent. So you haven't reached the correct side of this jump between superhuman and not, given all your data and environments are kind of human level. Here's two other ways of buying the future. One is just you do some clever experimental scale down where you do some small scale experiment, but you somehow designed it to capture an obstacle you think will bite as you pass through superintelligence so that you can test it out. So we'll do a bunch of those in pyrx. And the other way, or at least one other way, is theory, where you just write down on paper a mathematical model of what it'll be like in the superintelligent future. And the hope is that we can, via just doing these different things, have different and better models for superintelligence than the labs have, or at least models that are complementary. And then that will give us some ability to kind of directly model the future in a way that they're not covering very well at all. And then that hopefully gives you. Then you get ideas from there, you get obstacles from there, and then you can turn them into maybe from theory to empirics at low scale, maybe from theory to empirics at higher scale with working with the labs and just understand better that trajectory. So an example on the theory side is you can just write down a mathematical model of you have an AI model that has some basket of superintelligent heuristics, and then you can reason out, well, if I look at these scalable oversight protocols, do they scale and work reliably in that model? Can I write down, say, a proof in a toy setting that scalable oversight would work? And the answer is, currently, you absolutely cannot do that. None of the models, none of the methods people are applying kind of definitely work at scale.
Tom
And scalable oversight here means you can
Jeffrey Irving
reliably reward the model when the model is kind of supervising itself as it is getting stronger. So the general thing of labs are all doing this is part of their plan. And we know from the last set of, I don't know, five, eight years of research, there are a variety of obstacles which block that in theory and have shown up in empirics that are not kind of being covered in part because they don't show up yet at the current scale. So, for example, if you want to have models kind of engage in kind of back and forth reasoning, this kind of slightly adversarial right now, the models can't get beyond a couple of turns of this. So if you imagine a human debate, say you can get to. Humans can debate for hours and have dozens and dozens or hundreds of back and forth points that you don't see in model behavior. And so we just know that we're not seeing the superintelligent case in the current empirics. But you can just write down on some paper or a whiteboard what that should look like in theory and then try to explore it.
Tom
Why can't they get beyond a few terms of debate?
Jeffrey Irving
It's just not good enough. It's just like decay in accuracy. So they try to reason back and forth and it just gets a few steps and then falls apart. And that's just not a thing a superintelligent model will be doing. So we know with high certainty that we're not in the right regime yet, and we could be close enough. Maybe you get some knowledge of how the future will go from this experiment, but not enough of it to make me happy.
Tom
Okay, I'm still not sure that I fully understand. If we're getting to the point where we're close to deploying superintelligence and resolutions, research has gone really well. What kinds of things do you think you'll be able to make changes. Yeah.
Jeffrey Irving
So I think the hope would be in theory or low scale empirics, we can say, look, here's an obstacle to one of these protocols working to scalable oversight to Personas, to different parts of the lab's training story. With this obstacle, this algorithm works and this algorithm doesn't work. And we can demonstrate that in theory. Like, here's a proof of failure and success in different cases. Here's an empirical model which shows kind of, again, failure and success in different cases. You should do this kind of algorithm and try to scale it up, try to replicate it on your stack. We wouldn't expect that we would have perfectly tuned it yet. Maybe there's kind of many other aspects of their stack that are invisible to us, but we can give them guidance on which direction they should go in this broader space of algorithms. I think an important thing to say is in all of these approaches, in scalable side and Personas, there's just a huge space of possible algorithms to choose from. And they're doing their version of trying to filter the space. We will do our version as well. And hopefully those things can combine. The other case is that you say, well, we have an obstacle that we've. We have.
Tom
Give me an example of such an obstacle, either in Personas or in scalable.
Jeffrey Irving
Scalable an obstacle is like, so here's a couple of examples to obstacles in scalable oversight. So one is obfuscated arguments, which is basically, you could have models that are super intelligent, but they're not infinitely strong, they're not magic. And so if you expect them to walk you through why something is true or false, they will only be able to do part of the story. And if they're better at giving you the positive evidence for, say, some claim being true and really bad at giving you the counter evidence, but the counter evidence is actually the evidence that truly wins. Then you can get wrong answers out of any scalable oversight method, Basically, because the model has tried to. It's sort of. It's been kind of incentivized to win this game, convincing you of something, but it's found a space where it can give you the positive evidence. And again, it's not smart enough to give you the counter evidence and therefore it kind of wins by default.
Tom
So just to recapitulate, the hope here is you want to know whether a model has produced an output that you would actually endorse and you're hoping to rely on the model's ability to explain that output to you. Because you've trained it perhaps in like an adversarial debate game against another model where honesty is the winning strategy. But it seems at least possible that the model might be just better at propping up one side of the argument than the other. Even if it's not true. Have I summarised that?
Jeffrey Irving
It's not even true in theory. So this was discovered via actual human experiments. I hired Beth Barnes into OpenAI and she did some experiments where she took a bunch of human kind of debaters. So humans arguing back and forth about whether these kind of interesting physics problems were true or false, like what the answer was to some physics problem. And then there was a human judge that had not seen the physics problem context, so they didn't know the answer. And one of the winning strategies was basically a debater would produce a very complicated argument that sounded true, was false, but neither of the debaters, not the liar or the honest debater, knew where the flaw was. So just like a sufficiently kind of mushy, complicated argument with many parts that neither one of them could locate the flaw. And so it just looked like a plausible argument with no counterargument. And so one of the debaters might have said, you know, this is kind of mush. I think there's a flaw here, but I don't know what it is. And the liar can just say, come on, if my opponent knew where, if there's a flaw, they should be able to point it out, like, where is the flaw? But it just is the case that with a non infinitely strong model there may be a flaw. You can't find the flaw. So that showed up in human experiments. And his reaction was that Beth ran these experiments, she found other flaws, she fixed those other flaws. There was an iteration of quick cycling on, kind of finding and fixing flaws, and then they found this flaw and they stuck on that one. And this was found first in empirics. It's easy to write down a theoretical model of this. We don't have a good solution to this problem.
Tom
Interesting. And that's a problem because it means you can't rely on superintelligences debating each other. And you can't hope that the true side will have an asymmetric advantage over
Jeffrey Irving
the wrong one unless you have some different protocol which sort of manages. And this is not just true for debate. Any scalable oversight problem has this puzzle. So amplification, kind of constitutional AI, kind of all of the. If you imagine pushing any of these approaches up to superintelligence past, again, where humans can reliably supervise you will potentially hit this problem. Not with certainty, but I think it's a pretty good shot at hitting it. And then we don't know how it will go at that point.
Tom
And something I actually am not sure I still fully understand is what is the core basis of the belief that there might be some kind of phase shift when you move from human capability levels to superhuman capability levels, where our ability to supervise them just totally breaks down. So one example is we can train superhuman go models or chess models. That doesn't cause some kind of catastrophic problem for us.
Jeffrey Irving
It does, actually.
Tom
Oh, it does? Okay, tell me. Yeah.
Jeffrey Irving
If you take a fixed strength opponent and you train a go model to beat that opponent, it will quickly learn to be a bad go player because it will just reward hack its way through the weak opponent. And then if you put it against a strong opponent, it will lose horribly because it's like learn bad habits. This happens to humans too. So if I, I used to be about one don Go amateur. If I play weak, sufficiently weak opponents too much with high handicap, I get worse at GO because I learn I have to like fight off the tendency to play moves that are weak or good only against weak opponents. And so we have a variety of empirical results where basically, if you have a certain strength of reward function and you optimize against it for long enough, you will sort of get close enough that you see the difference between that reward function and the true performance, and then you'll get good at the proxy and bad at the real thing.
Tom
Okay. Yeah, that actually makes sense. I guess. I still struggle to visualize how that kind of failure would be super catastrophic. Maybe you don't need to tell a specific story about how it would be.
Jeffrey Irving
I think there is this question of how does good behavior generalize? And so if it's the case that as you, as you cross this path, as you cross this fuzzy boundary of kind of human level skill, the model remains in some sense a good entity and is still trying to kind of funnel data and training signal because it's kind of defining its own training signal the right way. That could go. Well, you could be sort of a nice attracting basin which kind of pulls you closer and closer to good behavior and you extrapolate to a good kind of superintellent system. Or it could be that you're just not that close, or your algorithm doesn't have the right equilibria and so you either are just going in the wrong direction. The model is starting to forward hack it forward hacks more and more and it kind of gets off into some horrible track. Or you have an algorithm which there is no way it could have had a good equilibrium. It's just like at limit, it behaves badly in almost all cases and you're just inevitably going to die if you train that algorithm hard enough. I think either one of those stories could hold. I think the hopeful story is that we could at least arrange to be in this world which is more path dependent, where if you're close enough to a good attracting state, a good kind of basin of attraction, you stay there and there's also some other evil basin of attraction which you really don't want to avoid and you manage to dodge that one.
Tom
That makes sense. You've said before that your modal expectation is that we get full blown superintelligence within something like two to three years. Could you walk me through what that looks like?
Jeffrey Irving
Yeah. So I think the main uncertainty here is are the models going to be good not just at kind of verifiable tasks with kind of clean rewards, but also fuzzier things, kind of intuition, kind of fuzzy planning, this kind of thing. And I think people are overweighting the probability that they're only good at the verifiable part. And if that is wrong, then I think we have seen so much progress over the last while that while the softer things lag, I don't think they lag by years, say they lag by a smaller amount of time. And that can carry us quite far in the next few years. And we've seen such rapid progress in the last couple of years that that could continue to go very quickly and keep speeding up. I think it could be slower and I'm hoping it's slower. That would be a lot. That would be very nice. But that's sort of the worry.
Tom
What do you think is the likeliest way it gets good at these fuzzy tasks? Will it look like sudden generalization or it gets good at learning? What's the story?
Jeffrey Irving
So I think it is the non magical thing of the company is getting better and better data that kind of expands the spectrum of tasks they are good at. And so there's two things to say. One is that I think throughout the reasoning era, so from 01 on, I believe, though I don't know for certain because I'm not at labs in that period, that they are not just doing verifiable work tasks, they're training against model self critique. So you show the result of a task to a model and you ask it to judge. And that can work across for more fuzzy things though eventually it breaks down and then the models are already good at verifiable tasks and they're sort of like some weak number of slightly less verifiable tasks they're good at. Those are also useful by people in the deployed world. So add to the labs that gives you this kind of flywheel of data to play on and experiments to learn from. And so over time the lab's ability to generate data that spans out further and further away from verifiable keeps getting better. So they're sort of climbing this ladder of verifiable to non verifiable just via the non magical process of collecting just enormous amounts of experience and training data. And I think that because they're all kind of widely deployed, if that continues, you can push up into lots and lots of tasks very quickly. So I think you don't need massive amounts of fancy generalization. I think you just need a lot of work object level on that kind of data generation.
Tom
Do you think are they buying this data en masse? Are they somehow getting it from their deployment rollouts but won't zero data retention.
Jeffrey Irving
So I think you can get a tremendous amount from anecdotes plus buying data. So it's like buying data, but maybe you know what data to buy because you've seen glimmers of how people are using the models in practice. So I think the zero data retention thing doesn't block them from learning from deployments in all cases. And I've trained models in the past and it is very valuable to know, oh, I have missed a kind of data, some sub distribution of the space of tasks and then from there you can learn how to fill that just by either generating data with purely synthetic or you buy them from some data provider from humans or the like.
Tom
What do you think is like the role of governments in this world resolutions doing its work? At what point might they need to step in? What might they need to do?
Jeffrey Irving
So I think there's a couple of different levels of kind of government action. You could imagine. Any government can do a bunch of unilateral defensive work. So you can work on defenses for bio or cyber or even persuasion potentially. So that defensive work cannot be done by any government kind of unilaterally and it's kind of good to do. Then there's kind of last minute temporary pauses where it's like, oh, we're really close to training this really dangerous model. Let's, let's chill out for at least A few months and kind of shift resources towards from capabilities to safety. Try to slow down a little bit, try to just dial up all the knobs that we can to kind of in the direction of safety on the margin. That also means you could, for example, use algorithms which are a significant but not kind of a fatal capability, like cost hit, like something that's like 2 to 10x slower. Maybe you can run that in this temporary pause world. And then the more extreme thing is you have a broader treaty where you try to do a longer coordinated slowdown or pause across kind of multiple countries. I think government should be trying to do all of these things and then we'll see how far up the scale we can go. I think a critical thing there is that I do think we may be in worlds where the algorithms that work are, as I mentioned, slower and more expensive. The algorithms and the algorithms that don't work or you need to dial up
Tom
don't work for alignment. That is for alignment.
Jeffrey Irving
And you need to dial up either the amount of data or do an algorithm pivot or something. And that if you're in pure mad race between the various labs, that's hard to do. And even a little bit of government kind of coordination pressure could make the difference in those worlds.
Tom
When do you think would be the right time to slow down? Right now.
Jeffrey Irving
So I think my take is that if we were to carefully analyze this question of exactly when we should slow down, it would be like a while ago in the past because we're just too close to this crazy future. So I think trying to kind of be super precise about exactly when in the future. No, no, no, no, already is the answer. I think whether we can achieve that is less clear because there's like, there's political will and the Overton window and so on. But like, that would be my kind of stock answer is like now too,
Tom
in the past, do you think any one actor should unilaterally slow down?
Jeffrey Irving
I think that's hard. It is the case, though, that you could probably find a list of less than 10 people in the world where if you could get them to degree to slow down, you could do it. So it's not some extremely enormous impersonal sea of people you have to get to coordinate. It's like lab CEOs, potentially people in China, leaders of a couple of countries. You don't get to that many people. And so I think that the question is, if a lab did a unilateral slowdown, how much closer to that less than 10 people did you get and potentially a lot closer because you've kind of made a stand. That said, none of the lab CEOs want to hear that argument. They only want to do the non unilateral things. And there's some argument in that direction, but it's also a very kind of convenient argument.
Tom
What do you think we should actually be slowing down? Is it the R and D itself? Is it like inputs to R and D, like chips, chip production? Is it deployments? What are we actually slowing down?
Jeffrey Irving
So mostly I don't have a super cast answer to the optimal here. There's a general thing where we will not, I think, have the ability to kind of stop progress. So if you try to slow down or you try to have a pause or the like, you will be slowing progress, but then progress will be continuing. And I am, as mentioned, worried enough that we're close to this kind of ASI future that we get there in not too long, even with a slowdown. Exactly. Kind of what the ingredients should be to intervene on. I think, as you say, probably the answer is all of them would be good, but I don't think I have a Cooper cast. Good answer.
Tom
Yeah, I mean, I guess one thing I don't quite understand is how do you slow down in a way that affects the different actors in any kind of way equivalently? So especially if we're Chinese, labs are also asking them to slow down. It seems difficult to be sure that they're slowing down in the same way that anthropic or OpenAI are slowing down.
Jeffrey Irving
I think a certain degree of imperfection is required here. So you have to be comfortable with, with measures that are not going to exactly be fair across all the labs. And so presumably what you need is a combination of tactical measures. So like tactical supervision and monitoring. But also if you wanted to do the grand international Treaty, then you need kind of human audits as well, and inspections and so on. But it will not have an exactly matched impact on every actor. And I think we just have to be okay with that as slightly disparate.
Tom
And do we also just have to be okay with any kind of economic implications? It seems like so much of the global economy is leveraged on there being continued AI progress. Is that just a hit you're willing to take?
Jeffrey Irving
I think my take is that if you were to stop all new model training, there would be this enormous ongoing wave of economic growth due to the current models that I think if you just take that, it's enormous in terms of positive benefit, in terms of getting Valuable use out of models. You have to learn how to work with the current models. But I think we're in a massive product overhang. We have worked only a little bit on how to cater to the strength and weaknesses of models. The models of June 2026 are just incredibly good at software engineering in huge numbers of ways, even kind of before the most recent models in the last couple months. And so I would be fairly unconcerned with that world. It is a trade off. I think that if you get stronger models, they can do more things better and probably cheaper. So there's a trade off there. But I think I would much prefer having time to nail down more of the safety story for both alignment and other risks than just massively rolling the dice.
Tom
What is it that gives you so much confidence? We have a high product overhang. So if it's not already showing up in growth statistics, what are the metrics where you're like, oh no, but look at this thing, it is already very useful. It will lead to lots of economic growth.
Jeffrey Irving
I think there's so much use of coding systems in particular, and I think that extends already to huge amounts of other kinds of cognitive labor. Any kind of analytic analysis of business or the things people can already do with models are so impressive that it is extremely unlikely. To me that has seen kind of full adoption across the economy anecdotally, both from myself playing with models and then just reading a lot about what people are doing. There is a massive learning curve to how to best deploy these models into any particular area of, of activity. So I learn better how to use them across time. So does everyone else. If we just were to stop for even like 10 years, the rest of our learning curve, we still keep climbing. Again, I would be totally lying if I said there wasn't a trade off here. Stronger models are in fact better at doing lots of things, but I would prefer that trade off.
Tom
Let's go back to government work. So you worked in government before yourself. I'm curious, what affordances did you find that you had at UK AC that you didn't have at OpenAI or DeepMind for changing the world?
Jeffrey Irving
Yeah, I guess there's a couple of them. So one is maybe I'll list three of them and then we can go from there. So one is just adjacency to national security. So being close to national security because there's a bunch of ingredients out of the risk story that come from those, those sources and you need collaborations with NatSec to have good takes the next one is just adjacency to policy. So if we want to do this kind of coordination across the world where governments play a role, you could have to be in a government to be close to policy in that sense. That's not the only actor. We want a lot of third parties and nonprofits and independent researchers doing this kind of policy development. But you need part of the story just being in a government. There's kind of a subpart of that, which is that there are in many cases, sometimes governments only listen to governments. And so at ac, we had a bunch of our own research, but often also we would just be able to go to another government and say, here's some of our research and some of someone else's research, like from meteor or Apollo or the like. And that package was much more received and listened to than if it had just been meter and Apollo trying to go directly to a government of various other countries. And so I think that proximity to kind of natsac and policy and other governments of the world I think is kind of the key thing is very valuable.
Tom
And if there's so many worlds where governments will need to play a role in things playing out well, what do you think about all the AI researchers who are very, very concerned about safety, but who are currently working at AI labs rather than in the government? Do you think, are they basically wrong to be doing so? Do you think they could have more?
Jeffrey Irving
Yeah, I think on the margin they are in fact wrong and many of them should leave and join governments. I think the amazing argument is that it's just one of diminishing returns. There are a lot of people at labs. If you are a safety researcher at a laboratory, probably you're further out on the diminishing return curve than you would be if you joined a government or a non profit. And so if every one of the people at labs left all en masse and joined the government, that probably would be bad. But that's not the actual calculation. It's like the marginal move is high is pretty clear. And I think people look at themselves and think, I'm an individual researcher, I'm kind of a special snowflake. I have a very particular agenda. I'm the only one pursuing that particular agenda. I should keep doing it if it's an important agenda. And I think that is making a calculation which is a bit too focused. And if you blur your self image a bit and just think of it as like, oh, I am a safety researcher, I probably have broad takes on knowledge about a variety of things. I can advise governments on a broad range of issues. Probably the lab would pick up the slack on what I'm doing to some degree. That would work pretty well. Again, I think on the margin the calculation is I think pretty simple.
Tom
And what do you think UK AC specifically will be doing from now until sort of the eve of superintelligence? If they play their hand very well, what kinds of things do you think they'll be doing that will be moving the needle one way?
Jeffrey Irving
Yeah, I think getting so misuse risks are important and so the pure kind of dangerous capability evaluations are important. And I think that story being kind of high capability and as in high research capability and then also close to NATSAC I think is important for getting those kind of well understood. And then ASC does a bunch of work on mitigations again against both misuse against loss of control. We have kind of a very, very strong safeguards team. Sorry, we as in ac before I left, I think AC already has strengthened the mitigations of the labs by virtue of being an independent kind of voice and source of research. And that will keep going. And then again, the big thing is the main reason I joined the AC initially is policy. Again, governments have a huge role in policy. AC is the largest source of government AI research capacity around safety that it currently exists. And so causing that policy advice to be maximally grounded in kind of the tactical reality of things, I think just makes it much more likely to go well.
Tom
Do you think AC is an asset to the UK specifically? Should every country just have an AC of its own? How many aces do we need?
Jeffrey Irving
I don't have a confident take there. I think there are probably more on the margin as good. I think there's some degree of not wanting to reinvent the wheel too much. And when there are other ACs, I think an advice piece of advice I often give is it's important to do a mixture of their own research to build up technical capacity, but then probably don't try to be a full on evaluator across all the risks in the same way that AC is kind of closer to being and then be in a position where we can work together across multiple governments and then two policymakers present. Here's all the evidence from all the ACs plus all the nonprofits kind of appropriately integrated together. And that I think is to the extent you can get that kind of collaborative story right, it's much more efficient because.
Tom
Much more efficient. Okay.
Jeffrey Irving
Yeah. You get much more knowledge faster across all the governments.
Tom
That makes sense. And why did you leave ac, it
Jeffrey Irving
was in fact for family reasons. So it's better for or my partner to go back to be back in the US and so I'm kind of. And it was I think fundamentally like when I moved here, kind of here in the Bay Area, here is in London are the kind of the two places I can do my work. And so like now I'm kind of doing the reverse trip.
Tom
There was always a bit of a compromise. That makes sense. Yeah. And what's the day to day of your work? Did it like I feel from the outside, people are always worried that joining government is going to be a bit more bureaucratic than they expect. Like how did you. Did you enjoy the job? Did it. Did it compare to working at DeepMind, working at AI?
Jeffrey Irving
When I joined, I think it was less bureaucratic on the margin than DeepMind. In part that was because it was a fairly small team. And of course when organizations get bigger, they get more bureaucratic. This is true generically. So it got a bit more bureaucratic over time just because of size, but not too much, I think. And there have been constant work within AC of improving that and streamlining processes and I think think it ends up in a pretty good place. So I always enjoyed kind of that level of it. It was fine. And then I just got to advise a ton of research happening across a bunch of teams, a bunch of kind of policymakers and other governments and so on. And that I love kind of getting to touch a lot of little areas of things. So that was kind of just a very rich experience. And I think generally AC is has a much easier time hiring very talented, strong junior people than senior researchers. And so I think that if you are a senior researcher interested in joining the government, I think that's a big unblock because they're very good people to work with. It's very fun. But they do sometimes can benefit from more experienced advice.
Tom
How will resolution try and get us higher confidence in the alignment of a future superintelligence?
Jeffrey Irving
So we have a portfolio strategy across different research bets because we don't know what will work. And I would claim neither do the labs. So those areas, I think the main kind of initial set are learning theory, scalable oversight, complexity theory, Personas and agent foundations and philosophy. And we'll sort of add to this if we choose kind of across time you should pitch us if you have new ones and then the hope is that both we can get those areas kind of fully resourced in terms of critical mass kind of size, teams of humans kind of across all these areas, but also a lot of investment in automation kind of tokens, GPUs and so on, so that we get kind of a full shot in each of these. And we don't expect to need them all to win, to succeed. The hope is that we have a few successes either in terms of generation of negative evidence, of obstacles to alignment working or positive evidence. Which means here are two algorithms. This one works, this one doesn't work in some toy setting such that we can drive changes in labs or in coordination broadly. There's sort of a core kind of three part bet here, which is that in particular for theory, the labs just aren't doing, they're not doing any theory at all, hardly at all. So just doing theory at scale will be doing a highly differentiated bet at a resolution to what the labs are doing. And then we will bet kind of again, quite hard on automation. So I'm kind of at AC in the alignment team there. We were doing kind of a bet on field building. This is sort of pivoting more to the machines, still having a bunch of people and researchers, but trying to kind of fully resource terms of tokens. And then there's a combination story where theory is more automatable than empirics, at least potentially, for the following reason, which is just that you have proofs, you can construct some theoretical model and try to prove it correct. And that is a purely verifiable reward. And so even though I believe that eventually the models will be pretty good at nonverifiable things, they're better at verifiable things. And we can exploit that to make theory go faster than it otherwise would.
Tom
What's your model of why the labs aren't doing any theory at all? I guess some people are pessimistic that theory applies to a problem as poorly specified as alignment. What are the things that we're confidently shooting for here?
Jeffrey Irving
I think there's just sort of a learned experience of empirics working very well, which we've sort of seen from capabilities and even now to some degree, kind of mundane safety. And the question fundamentally is like, will that extrapolate past human level or not? And I think very possibly it does not. And that the kind of empirics, if you don't try to do this, kind of really try hard to scale down to model superintelligence, you can just miss effects. But the whole many, many decades of machine learning, all of the recent experience of labs is telling them that empirics works. And so it's hard for them to step out of that bucket because they have all of the dopamine hit saying, look how good this is all the time. And they might be right and they might be wrong. And we should take kind of both of those bets.
Tom
I'm interested in the version of this world where we do successfully align the superintelligences and we get them, we've deployed them, and we have high confidence, thanks to resolution and everyone else's research, that they will behave the way we want them to. What kind of technologies would you expect that they will develop next?
Jeffrey Irving
Yeah, all of the practical ones.
Tom
Yes.
Jeffrey Irving
Practical means like allowed by the laws of physics. So I think we sort of solve aging. We get nanotech again for good or ill. Nanotech could be defense dominant. Dominant right now. Software has bugs. Software in the future wouldn't have bugs broadly. It would just be sort of bug free, perfect in most cases. I think we will have the ability to colonize the universe in various ways, probably via uploads. We probably will be able to upload humans into machines. My take is that people have this, I think, bad view that the machines will kind of be taking off ahead of us and then even in the good futures, we'll be stuck behind forever, which I think is wrong. I think you can imagine uploading someone and then modifying them cognitively while preserving kind of identity in some meaningful way to be also super intelligent. So there's like that future ahead of us. Should we choose it. Hopefully we have the option to also just live normal lives as humans.
Tom
What happens to the humans that decide not to upload?
Jeffrey Irving
I think they are essentially irrelevant to the economy. But I think that we will have, I hope that in this world we will figure out how to derive meaning from family and exploration and so on, whatever kind of the level of cognitive ability is.
Tom
Do you personally expect to upload, by the way? Do you think this is?
Jeffrey Irving
Yeah, eventually, yeah.
Tom
How would you go about making that decision? Like.
Jeffrey Irving
Like, I don't think I'd be the first one, but I expect that we'll just have a good understanding of the. Of the science involved. We will have like, we've done a bunch of experiments. It will just work very well. The result is that people will like, yeah, they'll. They will feel great. They'll be then smarter because you can like again, modify them in place in various ways, will understand the brain and AI and so on much better. So that understanding how to do that modification in a way that is kind of faithful is doable. And so that Seems like, yeah, that seems like a good deal.
Tom
What kind of results do you think you're looking at that is telling you this uploaded version of Jeffrey is faithful to the real me? Like, I'm confident it is.
Jeffrey Irving
Like, I think just some better understanding of how maybe kind of personality and intelligence and access to heuristics kind of interact. So if you imagine that it's just a version of you. Exactly, except that. So right now you have this kind of layer of fake consciousness or fake serial thought sitting on top of your pile of heuristics. And then occasionally you have your conscious mind, it thinks that, it says, I want the answer to this question. And your brain kind of substitutes in the answer to that question. And it sort of come from this amorphous sea of heuristics kind of seething underneath without your conscious awareness. If that just worked much better, then it would be kind of a fun way to be. And would it change your kind of intrinsic personality instead of not clear? And so if you understand that separation, how that kind of layer of this veneer of kind of serial experience relates to the seething mass of heuristics better, then I think you could maybe separate out what is a meaningful version of ramped intelligence me look like.
Tom
And so your strong take is right now, my serial thoughts are fake in the sense that they're not actually the computations by which I figure things out.
Jeffrey Irving
So as an example, I've been walking along on a hike and I kind of. I duck under a branch and then I like. And then my brain is like, you saw a branch and then you duck under the branch and that's what you remember it looks like. And it's like, no, that's not what it looks like. It's like a bunch of reflexes that triggered in various orders and different parts of my body acted without entirely consulting other parts and so on. And then your brain kind of stitches together some fake narrative into all of this. And I think that is just sort of intrinsic to how we experience the world. A lot of it is not that inaccurate. But some degree of your conscious train of thought is a hallucination as you go along the world. And I'm very happy with this. I don't mind living this way. I sort of think of myself to some extent as like a bit of a shell surrounded. It's like this thin veneer of kind of experiential linear shell surrounding a basket of heuristics. And I enjoy.
Tom
You're fine with that? The shell life?
Jeffrey Irving
Yeah.
Tom
If you do an upload. Would your expectation be that there'll be two consciousnesses, there'll be the digital one and then the physical one will sort
Jeffrey Irving
of get rid of it? Probably will get rid of the physical one or something?
Tom
Yeah. Would you want to get rid of it? Would you want to clone the consciousness? Do you have a strong take on this kind of clone?
Jeffrey Irving
It'd be a very bad world if everyone is just like massively duplicating themselves in some horrible runaway exponential process. So I think the world of the future, if we get to the world with uploads, we'll have to be much more thoughtful about this kind of duplication.
Tom
So we'll have to have some kind of restrictions on duplication.
Jeffrey Irving
Yeah, restrictions or just you've arranged the outer economic and incentives so that the reasonable behavior is incentivized in a good way? I don't have CAST takes on exactly what the structure is there, but getting it right seems pretty important. It is not obvious that the economics and physics are consistent with the optimal way to achieve goals being having more individual identity. But I think it is plausible either because there's the speed of light delays, so that if you have a bunch of intelligences kind of scattered around a world at a radius of even a light second, you can't be having them constantly synchronize. So some value in having local sort of conscious experience, like local kind of higher level planning seems valuable.
Tom
Well, that's integrated in one person. Is that what you mean?
Jeffrey Irving
Yeah, or one person or something. But you just can't you. If you have a light second spanning consciousness, then you're like, you're a bit delayed. It's valuable to have locality. Or we just choose that we kind of value individuality and diversity in this way, which I hope we do. And then the cost to that is such a small factor because again, you're sort of a thin veneer on top of this pile of heuristics that it will be fine.
Tom
Do you expect this stuff to just go crazy fast at some point? Yeah, but. Yeah, but like why? Because it's just kind of not super intuitive.
Jeffrey Irving
I don't think there's obstacles to this. I mean, crazy fast. There's a question of like, what does that mean? Potentially you get the nanotech and the uploading within like a couple of years. Maybe it takes like a decade or two, but that seems. It feels kind of unlikely to take a decade. And like the. But even if it takes like two decades, that's still less than a human generation. And so that's still on the scale of us adapting to the world crazy fast in some sense. And so I think we have to be ready for that in either of these speed cases. And then. Yeah, why do I think it's so fast? I guess one, I think simulations are going to be really good. So we've seen, I think with something like AlphaFold that you can build proxies for kind of quite complicated physical systems that you can just play with purely in silico. And I think that will be broader and broader across a number of areas, although I think we should be. This is uncertain, this is not a guaranteed thing. But assuming you get that kind of behavior, then you can iterate a lot of your experimentation just in simulation.
Tom
But even AlphaFold has quite a lot of failures of generalization, right? From what I understand, a lot of the time it'll predict a certain. It'll predict a certain way that protein folds. But then you actually try that out in an organism and it just like it does completely fall apart.
Jeffrey Irving
I think this is true. But it does well, a lot of the time it has some degree of understanding of its own errors. And then it's also like. I guess there's two reasons to believe that AlphaVold is not anywhere near the ceiling of that performance. So one is that it's like just the first couple systems. But two, it's also, it's not actually trained from. You could imagine training these models from physics in a deeper way. Albold is trained from a history of other kind of proteins. If you manage to solve kind of simulation proxies across a greater diversity of timescales, all the way down to quantum chromodynamics and kind of everywhere in between that I think you can potentially fill in the gaps and do kind of error correction of AlphaFold like models even without going to data some of the time, maybe you need some data which is less. And so I think there's a potential ceiling of performance of that such model, which is quite enormous.
Tom
And are you relying on ordinary market forces to get us the pragmatic technologies in the right kind of order to get us the right kind of upload? What's the.
Jeffrey Irving
When you say the right kind of order, I think that the answer would be no. So I think one mistake that some economists and kind of analysts are making now is there's an assumption that humans are the source of demand. And so whatever the machines will be doing, well, humans are the demand. So we're plugged into the economy in some meaningful sense in the future where we get asi. Machines can perfectly well act as the demand of the economy. Of the economy. So if you have pure market forces and these super intelligent models are not trying to improve the world on our behalf to some degree, I don't think there's an economic need for uploading. The machines could perfectly well just do their own thing. So I think you have to have enough alignment that you are sort of jumping into a world which is kind of suitably democratic and clean. And I think getting again, the pure economics would say that humans are not very relevant in this world because we're not kind of economically relevant.
Tom
Why is that happening? Even if we've aligned the machines, why do they have consumption demands of their own? I don't know if I quite follow this.
Jeffrey Irving
Then it's not pure market forces.
Tom
Okay. Yeah.
Jeffrey Irving
I think it's like, then it's the models wanting to design the world so the humans have a meaningful role and a meaningful, meaningful access to resources and so on. Once you have access to resources, then conditional on that, then market forces can take you a lot of the rest of the way. But generally, I think markets should be modeled as optimization engines. So they have. We live in a world which is a mixture of free markets and then regulation to kind of channel that optimization power of markets. And we will have to be in that world. I think definitely that makes sense.
Tom
Yeah. What gives you so much confidence that solving aging, uploading things like this is actually in principle possible? Why are there not some kind of diminishing returns to intelligence? Why do you think we can make such radical progress? Do you have intuitions here that you use?
Jeffrey Irving
So there are diminishing returns to intelligence. They just occur way out past asi, I would claim. So I don't know why that's relevant to this question of aging. Basically for aging, I think maybe it's
Tom
an unsolvable problem or something.
Jeffrey Irving
I see what you're saying. As in, why don't the diminishing returns strike before you solve aging? I just don't think aging sounds that complicated. We've only had a couple hundred years of understanding, kind of. I forget the German theory of disease is just not that old. There have been various proposals for aging that are relatively understanding light in that they intervene on the consequences of aging and the degradation of. Of tissues and such without having to understand the entire body and all the dynamics. Even if you could do that with asi. So I think aging seems relatively simple.
Tom
Isn't it kind of bottlenecked by serial time though? Like how many experiments will we be able to do where we observe the aging of an organism. Isn't that something especially for humans? We live pretty long lives.
Jeffrey Irving
So I think if your time constant was a human generation, then 100%. But it isn't like you can intervene on someone and you can see how they're doing in terms of various measurements and then kind of gradually learn that way. There are other organisms that we're already understanding, like in the last couple tens of years, a better understanding of aging in kind of smaller organisms. Some of this has turned into sort of wellness improving treatments for humans. It just seems like none of this is that hard. Again, if you're like a super intelligent AI or humans assisted by such, it seems quite doable.
Tom
Do you have intuitions about what kinds of fields of science will be the most and least amenable to heuristics?
Jeffrey Irving
Yeah, I guess I don't have a great. I kind of think all of them is like the default take. This is how humans think. We think via a combination of heuristics. I think one kind of challenge for alignment and understanding AI in general is if people have say, a take that, oh, it's extremely important that we have models write out their reasoning in chain of thought so we can supervise it. But this is just hilariously not how humans think either. When you ask me the answer to a question, what will happen is part of the time I just come up with the answer completely in some not written out form. And then I just start talking and the details kind of flow out as if I have reasoned through it, but I totally haven't. That's like a complete.
Tom
That's not how you've solved the problem.
Jeffrey Irving
I solve it by just guessing the answer via crazy heuristics. And so similarly for a model, if you ask it to solve a problem, sure, sometime it'll reason it out, but other times it'll just guess the answer. And then you say like why is that true? And it'll write out some convincing rationalization. And that is rationalization. There is a pejorative word, but it's also just intrinsically how intelligence works even for humans. And so we have to understand how to make models work, be safe, be aligned, while not believing we can get away from this notion of heuristic reasoning.
Tom
One thing I don't fully understand is you seem to believe that generalization might not be that powerful. We'll get the super intelligence because they'll be able to get data on these fuzzy tasks just by deploying them just slowly and slowly. The labs will be able to get the data. Models will get good at the things that they get data for. Why is that not more of a break than like two to three years? To me, that seems like the client, like there's so much data that they, for this kind of very long horizon, very fuzzy plans that we're imagining these superintelligences will be. Will want to do like running a company or running an election campaign or something. I think I basically have the same picture as you there, but I imagine that's going to be okay. That means it's 10 years until they get good at all of these things, rather than two to three years.
Jeffrey Irving
So it could be 10 years. I guess the reason why it could go faster is that one of the skills the models will be getting good at very rapidly is, is data generation. And from data generation, environment design and data augmentation. So if you look at kind of, you look around the world, there's a lot of data on kind of all tasks, but it's in the wrong format. It's not an RL environment. It's someone's static attempt at writing out a trajectory. And so the question is, if models get really, really good at AI, R&D, even in a mundane sense, at running experiments, at building data generation, building environments like this kind of iteration, will they be able to increasingly, well, take the bad data that exists, like static trace data or examples, and sort of squish it a bit and rearrange it into some environments that you can iterate on and then you do have some generalization. So it's not, it's not the case that the planning skills required for doing R and D or theorem proving or coding are totally different from the planning skills you need to do for taxes or M and A or BDC or the like. So some degree of generalization plus getting better and better using data, plus just the fact that they are a bunch of, right now CEOs are in fact trying to use these models to survey their companies and learn this thing. I think the case for slowness could apply to some of the tasks. But then across the next two to three years, say, you get this enormous wave of companies deploying things internally for R and D purposes and speeding up their own within the labs, but also out to customers who are still deploying the models. And that seems like a very unstable world where you have incredibly strong models at including not just the verifiable reward parts of this, but also the things that take more human judgment because you have a bunch of experience of this Kind of iterated day to day or week to week. And the question is, as you get better and better at those tasks, are you also better at closing some of the holes in your sourcing of data for other things and the ability to do kind of just fast adaptation of models and data and so on? I think the other thing is that because we've seen all this development of scaffolding over the last year in particular, that gives you a faster cadence way to inject skills. So if you're really good at planning and thinking kind of general, and you're getting better and better at scaffolding, do these come together to give you kind of a bigger part of the story? I hope this is wrong. I hope that in fact the 10, 20 year story is correct. But it just feels like the world where maybe the claim is that the space of tasks for doing the full suite of ARD and software engineering skills is already much broader than people I think give it credit for. Now, maybe I would say this because I'm like a researcher and that's what
Tom
you use the models for. Yeah.
Jeffrey Irving
But I do think that, I mean, I don't know, I also know a bunch of other things about the world. And I don't think that somehow the intuitions that I take away from software engineering and I don't know, martial arts and so on are just not as distinct magisteria as people imagine them to be. And so I expect, because if there was no data about kind of all these other tasks, then I think you'd be stuck. But I think they'll be like, if you have hundreds of billions of dollars to spend on that data, then I
Tom
think there's a path, if there were a trend to extrapolate for the ability of the models to generate data for these tasks for which kind of, you know, some kind of crappy data in the wrong kind of format exists. Like, is that. What would that trend look like? Do you know what it would be? What the metric would be?
Jeffrey Irving
Oh, I'm not sure. I think so. One thing, I'm a bit sad that there aren't. There seems to be like insufficient data on how good the models are. These nonverifiable tasks there's like, we have. My guess is that some of these show up in like, epic's like, ICI index, which I is probably indexed, so that's redundant. But there's enough of a vibe people have that in fact, verifiable words are taking off and nonverbal awards are staggering that I wish I had those curves Somehow that should be a curve that I can see easily by going to some website. I don't know what that website is currently. So then the more detailed question of how would you track models ability to generate data that feels like it's just maybe a subcategory of R and D, but I don't know of a good proxy for that.
Tom
Currently one of your other research bets is on Personas and character training. What are the core facts about the way Personas work that we don't currently understand, that we would love to understand before we get to superintelligence?
Jeffrey Irving
I guess what is Personas are low dimensional structure in models. And what that means is I have say myself as a person, I have a bunch of correlated traits. So when I say I have correlated traits, I mean if you look at one of my personality traits, it will be correlated with some other trait. An example of this is in the political view sphere. If you evaluate people, if you survey someone on some issue, you can predict with pretty decent confidence their views on one's issues, even though those kind of rationally should be different, but they're totally not different. And so across pre training the model will have picked up all of these correlations from human data. It sees a human world that has all these correlations between good behavior and bad behavior in one thing, many other areas that are more neutral, but again that span this kind of correlated behavior. And so what that means is the model knows a bunch of structure in the world which is not is kind of about human correlated behaviors. And then we have all these glimmerings of empirical results where that correlation kind of shows up in weird, sometimes bad, sometimes good ways. So the first big paper here is emergent misalignment, which is by various people, including Hawaiian Evans. If you train a model on like code with vulnerabilities without comments saying it's vulnerable, it will learn to do a bunch of horrible things including celebrate and admire various dictators. And that is because there's some coupling between model being nice about the code it generates and model being horrible about kind of which people it values. And then there have been similar work at kind of anthropic on sort of like if you train on reward hackable environments, the model turns somewhat evil in various other ways. AC did a similar thing with open weight models. OpenAI had a recent paper where if you train on a bunch of good behavior, then you get it generalizes as well. It generalizes in good ways. If you train on a bunch of good behavior, it generalizes in good ways. So there's all this glimmering of structure. And this one thing that indicates is if you were to understand the structure very well and manage to preserve it during training in the right way, that may allow you to extrapolate up to superintelligence with some preserved notion of the structure. There are a couple of caveats to the story. So one caveat is that while if you're applying a ton of optimization pressure, you're going to be mucking with the structure all these different ways. So we know from other papers that if you. For example, there was a paper by David Africa at AC where if you train models to be consistent, you can accidentally break their chain of thought legibility. So you train on one kind of modal behavior and you make them secretive in some bad ways. But if you were to kind of train very lightly, this is a separate paper. Now, you sort of only try to match statistics between various modes of behavior, then you kind of fix this kind of bad effect.
Tom
So you train lightly for consistency and
Jeffrey Irving
you don't get secretiveness, you don't get the bad secret behavior. So there may be some ways of training lightly on structure so that you preserve it as it kind of goes along. The other caveat is that clearly if you had low dimensional structure at pre training and up at superintelligence, they must be different because one of those is human level, one of them is super intelligent, and those are different modes of behavior. And so somehow there's going to be some mapping process from the behavior and structure picked up early in training up to superintelligence, and you have to follow that mapping along. And so the general bet at Sequent is this is a bunch of potentially good news that is very poorly understood. And there's not a lot of kind of even toy models of this in theory that would tell you how that mapping emerges, is preserved, kind of changes through training. And so the hope is we can understand this better and then that will separate algorithms to kind of break or preserve the right structures.
Tom
That does sound like the kind of experiments that might be easier to do in a lab, though. Presumably some of these questions are just about the scale of the post training that you're doing and what that might do to the Personas acquired in pre training.
Jeffrey Irving
Yeah.
Tom
Or are you still optimistic?
Jeffrey Irving
I think I still am optimistic for a couple of reasons. One is that some of the papers I cited are just on open source models at low scale. I think this is a particularly fruitful area for this mixture of kind of empirics and Theory because I think that modeling low dimensional structure is just like a lovely thing to write down kind of theoretical model wise. And so I think if the labs had enormous theory teams trying to explore the mathematics of that picture, that would be great. But they do not.
Tom
But they should get them in your view?
Jeffrey Irving
Yeah, but they're just not. I think we've went from an area where the labs were kind of a bit dismissive of theory to now they say they want to do it and they're still not doing it for various cultural reasons and historical reasons. Maybe they'll do it in the future, but for now we actually have to make in progress.
Tom
That makes sense. I'm curious about how you think about the field of alignment. Strikes me that bunch of other fields have these sort of core concepts that help organize our thinking. Something like Nash equilibria or atoms. Do you have a sense what are the equivalent concepts in alignment?
Jeffrey Irving
I think we have concepts. So I think certainly there's just sort of reward hacking and kind of models existing in different kind of scales of complexity. But all of these have holes in ways that like holes and gaps in ways that we don't have them in these other more established fields. And I think a hopeful thing is that the field of alignment has been around for maybe not more than 20, 25 years at the most. And then there was very little work. There's a lot of different areas of theory and kind of approaches one could explore. And then most of the history has done very little of only a couple of approaches, some of which we'll hopefully do at resolution, some of which will be more novel. And so I think there is a potential for low hanging fruit in even just finding the right definitions for those core concepts that are more resilient in kind of theoretical model land and then will better predict empirics going forwards just because people haven't tried very hard yet.
Tom
People haven't tried that hard even though, okay, I mean 25 years I guess is not that long.
Jeffrey Irving
And then like a lot of so like say the whole field of Personas empirically is just a couple of years old, like one to two years old. And so no one has tried to write down kind of like solid theory for this over 10 years. If we only have two to three years, hopefully we have 10 years. But if we have little time then I think it could still be sufficiently low hanging fruit that the combination of humans and then a bunch of automation can get us some answers.
Tom
Do you personally feel like your understanding of the field has changed very Much from when you were first working at OpenAI.
Jeffrey Irving
I think it has. So I think, for example, I was thinking less about path dependence back then. This whole idea of kind of low dimensional structure. I was not kind of factoring in as much as I have in the last couple of years, even obfuscated arguments. This is this problem we can discuss in debate or scale wars generally that I didn't fully understand until a couple of years ago. So I think a lot of it has changed and I think maybe had the field not been advancing kind of very faster and faster overall, I'd be more optimistic that we would have a shot at solving the problem or making a big dent in the problem. But again, because I think we have these glimmers of hope, it's just then very little time.
Tom
So the main source of pessimism is lack of time and the main source of optimism is these glimmers of hope, these examples of positive generalization.
Jeffrey Irving
Basically the specific thing is kind of low dimensional structure. I think that both can give you negative but also positive generalization in some ways.
Tom
Yeah. Will we be able to specify the superintelligence's utility function if all the alignment research works well though that's never going to happen.
Jeffrey Irving
Well, I don't even never. It never is too strong for the next, like until it's too late. We would never be able to do that kind of precision. So I think the only hope is if we learn or we luck out that we don't need to hit that precise a target. And I think it's possible that in fact we do have to hit a precise target, in which case we're not going to make it. If the structure of models helping, supervised models helping generate training data for models is sufficiently error correcting and has some give, then there's a hope.
Tom
So we need to hope that there is this kind of basin and we get high enough confidence that our training procedures are landing us in that basin and that's the best thing we're going to hope for.
Jeffrey Irving
I think that's basically right. I think there's a question of can you model out the situation to the point where you have a model that exhibits this phenomena there being mini basins, you can then calibrate that model against empirics in various ways and see how does this look like in practice. One idea of a mathematical object one could try to construct with empirics is imagine there's the superintelligent limit models and there's a variety of basins. Some are good, some are bad. Can you write down a coarsened model and actually train a small model that trains along. You can kind of modify its branch points when it could go one way or the other. And you sort of draw a map through training space up to these basins, such that you have literally a branching curve that starts out as a single curve and then it branches, then it branches again. Maybe some of the branches converge back in together. And you could literally have 1,000 checkpoints of the model showing this map of training through time, such that you then have an object which you can play around and iterate and try different algorithms just projected into the space of this kind of map of training.
Tom
And what would you observe about how the branching works that would give you confidence that this will happen at superintelligence too?
Jeffrey Irving
I think you would have some mathematical model of this branching, and then it wouldn't give you full confidence, but you could be able to see, oh, if I do this kind of algorithm, or maybe I can. Here is a test I can do to pinpoint or to narrow down when am I likely to branch, such that I can apply more resources there or spend more effort. I think there is a bunch of hope that we could get to more understanding even on a short time scale. But I don't know, hope is not a lot of probability. Just like we should try.
Tom
Do you think, will alignment be the only scientific field that's very difficult to automate? Are there other ones?
Jeffrey Irving
I think the general problem with alignment is that you don't necessarily get more than one shot. So to the extent that we have some evidence from current models, which is important, so you don't get exactly one shot, but the behavior of models up at asi, up at superintelligence, that may just be different. And we have to understand that kind of in advance. If so, most other fields, you can try to structure things so that you get iteration. There are risks that are less like that coming from AI and other catastrophic risks.
Tom
But
Jeffrey Irving
a lot of fields have this kind of iterative potential, and then you're in a good, much better place.
Tom
Are you surprised at how much iteration we get at the current level, where it's superhuman on some kind of tasks, but not broadly superhuman in some kind of sense? It feels like this, to me is potentially a positive surprise.
Jeffrey Irving
I think there's some update there, but again, I update there much less than the lab folk do on average, just because I think we haven't necessarily seen kind of shifts. And I think we do have additionally negative evidence because there are cases when the current behavior is A model of the future or at least does show they've tried very hard to not have a bunch of award hacking. And they still have a bunch of award hacking in who in production deployed models. And so I think that's clearly a case where we don't understand things well enough to have iterated enough to pound away the errors, which is bad news.
Tom
Yeah, that makes sense. You've got this great blog post from I think like several years back where you make an analogy between LBJ's presidency and aligning superintelligence. And your point, as I understand it correctly, is lbj, he's this guy, he's motivated almost exclusively by power and wanting to acquire more power. He also has all sorts of asymmetric advantages against his opponents where he's better at being a politician than them. And yet the American political system still aligns him towards great positive outcomes like civil rights and the Great Society. And maybe I'm stretching the analogy here, but what do you think, what, what claims do you think it is about the American political system that can give you faith that LBJ will produce positive outcomes that you want? Like, what does the LBJ pre deployment safety case look like?
Jeffrey Irving
Yeah, so the first thing to say is like, it's actually to what one is, I'm not going to take a stand at whether he was net good because he also did a whole bunch of horrible things. I think the take is less that I'm confident that the system in fact aligned him to do good. I think he did probably want to do some good. He just thought I must gather all this power along the way to do good. As many people think the case is more this is a very poorly designed game. A nice analogy, which is fun, which I will cite from that post, is like he became Senate majority leader because he realized that petition had all this power that no one was everyone else was leaving on the table. And so, for example, he could choose as majority leader in the Senate when to call the vote. And so he would just sit in the chamber watching people randomly go in and out of the chamber, I don't know, to the bathroom to get a snack or something. At some point, the balance of votes in the chamber, it was hidden in favor by a few votes and he would call the vote and win because he had had a perfect memory of who was going to vote for him. And there's extremely good predictions there, but that is just a very badly designed game that was played with. There was this one LBJ guy who's incredibly good at the details and there's not the competing LBJ force trying to be a counterbalance. And so I think when the American system works well, it is because there are effective balances and counterbalances. It's not clear that those are always working well. But it's also not clear that the American system is the uniquely best balance counterbalance system we could have. And we do have the potential to have a more well designed game and training process, more kind of more custom for this process. And I think if you get this counterbalancing, then I think you potentially can get through a lot of the problem. So an example is like, if you had the other lbj, it was opposed to the first one. That's saying, by the way, everyone, you realize what he's doing here? He's cheating the vote system. And everyone is like, that's ridiculous. That's clearly unfair. Let's fix the rule to break that. I think that intervention would get you so much power over this kind of the misaligned components of LBJ that I think it's within hope to imagine getting that story right.
Tom
So that's an example of a system that's like, poorly designed but actually reasonably easy to solve.
Jeffrey Irving
Yeah. There's a more egregious example of this from another one of Robert Carroll's books, which is Robert Moses would write these bills for the New York State government to pass, which just contained these trick clauses that gave Robert Moses all this power. And then no one noticed the clauses until after they'd all passed the bill and it was so late that they would have had to lose a ton of face that they just rolled back the bill. And so if there had just been another Robert Moses opposed to the first one, saying, like, you know, this bill contains this horrible power grab clause that would have been like an unworkable strategy on Moses part. And so there is this potential for monitoring that's much more invasive against the AIs and various kinds of alignment schemes and our ability to intervene kind of all throughout training in a way that you can't with a human. Human. There's many affordances we have on this process that in these examples of people that have gathered a bunch of power in misaligned ways, it just feels a bit fixable. If we get the situation right now, whether we'll get it right is obviously not dicey, but there's hope there.
Tom
How many of these latent exploits do you kind of think that human society probably has?
Jeffrey Irving
Just tons. Absolutely tons of.
Tom
Yeah. I'm curious how quickly and by what Means do you think the first superintelligence would be able to solve Pentago? So you, you solved it?
Jeffrey Irving
Oh, oh. I mean Pentago can be solved in like order, like, like 10 to the, like, like 17, 18 flops.
Tom
Yeah.
Jeffrey Irving
So like pretty fast, but yeah, but
Tom
like how is it doing? So imagine it's not in the training data. Like is it literally just like chain of thoughting?
Jeffrey Irving
Well, no, if I was a supernal solving pedigree, I would literally just like run, if I cared about it, I could just run the whole computation again extremely cheaply using faster software than I was able to write. There's a question of like, can it do it? So Pedego is a board game. One property of board games is they usually have some heuristic structure which you can intuit and then below that structure is a huge amount of just essentially random calculation. And the only way to see the calculation is like doing the calculation. And so the question is, how well is Pentago modelable by heuristics? And I don't know.
Tom
You don't have an intuition for that?
Jeffrey Irving
I have tried to train, I don't know, medium small scale neural networks to predict my kind of cast opening book at Pentago and they don't do as well as I was expecting them to do like a priori. So it's possible that a fair amount of the structure of pentagon is kind of randomish. It's also possible that as you scale up a ways it kind of phase shifts down to now it understands the heuristics and kind of nails the story. But I don't have a good catch sense.
Tom
That makes sense, yeah. Did your time working on simulations at Pixar give you any kind of greater confidence in the ability of, to use simulations to understand things like physics or biology?
Jeffrey Irving
Certainly my PhD was in combinational physics. There's I guess a couple of things. So one is the reason physics works as a theory, the reason there is a field called physics which can make a bunch of predictions is effective field theory or effective physics. Which means that you don't need to understand the low, the high energy, the very fine structure to write down a model of coarse things. So you can write down a theory of atoms, a theory of molecules, a theory of steel beams and so on without the theory of the thing below the theory you're currently modeling. And that robustly works across a wide variety of scales. And that is, I think there's this generic hope that we've seen throughout physics and other areas of science that you don't need to model the substructure. A lot of the time there's also hope for alignment because that kind of intuition also says that maybe there's fields or theories of say, speed plus heuristics which don't need to understand the architecture that we're using or the details of the transform or the like. They're sort of quite generic. If you make some appropriately kind of creative assumptions about roughly what that substructure might look like.
Tom
And that helps us if alignment is computationally reducible.
Jeffrey Irving
No, it helps us write down theories of alignment.
Tom
Oh, I see. Okay, yeah. Because you don't need to understand the substructure. You can capture it with a high level theory. I'm curious, when did it crystallize for you that this RL plus LLEMS essentially would be the path to superintelligence? As far as I understand, you were arguing for this already way back in 2019.
Jeffrey Irving
Yeah, 2018.
Tom
2018. Okay. Yeah, tell me, what did you see?
Jeffrey Irving
This is sad, but I, like I, I don't think I saw it was not that complicated. So in some sense I arrived at OpenAI in 2017. It was already kind of the like Paul Christiano was already writing down schemes that had this idea of using language and reasoning to decompose things and then write alignment in terms of these language models. In fact, the summer I arrived, Alec Radford and Paul had tried to run RLHF on language and hadn't worked at that point. So that was kind of in the general area. And then maybe the critical thing is AlphaGo because I think we had a sense that, oh, we couldn't do this kind of like you couldn't model things as kind of explicit reasoning too much because it'd be too slow, the models would do it a different way. And AlphaGo is like, no, just actually doing the tree computations. Mixing explicit reasoning with heuristics does give you the strongest thing on the planet at plain go. And that I think one, it gave us some emotional license to write down alignment algorithms based on this. But also just like it felt like that path of you do a bunch of reasoning, you can write it out, you can compress it, you can kind of iterate in these kind of environments, environments that are about reasoning. And then you would get the ingredients for that from language models which OpenAI was exploring and so was say Google Brain as well, and a bit of DeepMind that could just take you all the way there. I think part of this is my, I have A general take that a lot of this kind of reasoning stuff is not magic. So you we just have a big bag of heuristics, including heuristics about how to reason, kind of planning to do ways to error correct. And so some of the intuition here is that somewhere in this sea of Internet text there are a bunch of ways of reasoning that are good, there are a bunch of ways of reasoning that are bad. If you take that initial ingredient and then kind of strengthen and pick out the good parts of it and strengthen them with rl you get all the way there. So that was like so 2018. I sort of told Dario this is like annoying that I would write a document called languages enough to get to AGI and then I didn't write it until 2019. So like early 2019 is when I actually wrote a document.
Tom
It's when the document was created.
Jeffrey Irving
But that was like what I was the way I was like I think not just me, but also like other people there were thinking what's your model
Tom
of why reinforcement learning from verifiable reward took so long to materialize? Like why did 01 come out in 2023? Why not before then or 2024?
Jeffrey Irving
I don't know. I think some of it is tuning. Some of it is that if you have this model of you have to get to sufficiently good error correction to be able to reason for a long time without decaying, then as the pre trained base model improves, you get closer and closer to when you get to lift off on the ability to do RL over a long reasoning trace. But I don't know. So in some sense I would have expected it to happen a bit earlier and I was wrong.
Tom
And what's your model of what RL exactly is doing to the pre trained model? It's like selecting for parts of the pre trained distribution which already contain useful reasoning traces. Is it teaching it generalizable strategies for reasoning? Which of these matters?
Jeffrey Irving
So I think a lot of part of it is that a lot of human reasoning strategies as written in language just are generalizable because we've learned patterns that apply to a lot of different domains. And so the first thing that it does is just down select modes of behavior to remove the bad the unworkable kinds of reasoning or to some of
Tom
which I've created online or.
Jeffrey Irving
But the other thing is people often write down the final answer and not the chain of reasoning that got them there. And if you try to have a model predict the final answer, it's just going to be forced to hallucinate unless it can do all the reasoning that a human did kind of off stage in its latent pass. And so somehow there was a combination of tuning of RL algorithms plus sufficiently strong base models. In around 01, those started to work.
Tom
Well, if you had to make a bet about what alignment technique is ultimately going to end up working, do you have a spidey sense? Do you have a front runner right now?
Jeffrey Irving
Some combination of Personas and understanding of learning dynamics and scalable oversight. I think I mentioned agent and philosophy and in some of these play into how to think about pieces of that story. So a lot of the agent foundations work is thinking about ways of modeling the limit, ways of thinking about path dependency models, reasoning about themselves in a way that you have to untangle some recursive loop. So that understanding could also teach us how to do the other components of scale of oversight or Personas or the like, or would replace them in some way or something. There's a bit of. I think an important principle for the org is that I come to this with my kind of inside view sense of how things could go. And right now maybe that's like scalable oversight plus Personas plus learning theory or learning dynamics kind of coming together and sort of fitting each other's other's whole in some way. But we also as an org will have this outside view perspective of we're going to take a lot of different bets. Not everyone should have the same view about how the pieces will fit together. And the hope is again that we don't have to get success in all of the areas to win. We just have to get. We will try a bunch of things and if we get kind of important insights and algorithms or obstacles from some, even just one area, that could be enough to kind of sort of account for the entire org.
Tom
And is there a future where timelines look so short that you just decide we need to focus all our resources on one single bet? Because this approach of trying to aim for lots of things doesn't make sense anymore.
Jeffrey Irving
Here are three reasons structurally, I have a cash answer here. So structurally, why we wouldn't want to do that. So one is that there's strong diminishing returns, usually in sort of like token or DPU spend. So you'd have to be really, really confident in a particular area to not want to hedge your bets and give the other areas enough that they can continue to be reasonably well automated. And hopefully you've just raised in this world where you've established some strong progress in one of your areas, you can raise a ton of money, but you probably do want to spend a decent chunk of that on just lower, on other areas to take advantage of this diminishing return curve. The next one is that it's possible we get all the way to the end or really near the end where you've trained a superintelligent model and people are still bickering about timelines even within the lab. It's certainly kind of one step removed kind of nonprofits. Because I found it very fascinating how far we've gotten into this AI takeoff scenario that we all are living inside and still we have these massive disagreements about whether things are slow or they're saturating or the like. And so somehow my model is like, we're still not going to know what the timelines are maybe a week before someone trains a super intelligent model externally. And then finally, if you have to make a trade, part of org design at Resolution will be arranging things for psychological safety. As we kind of go into this kind of crazy town world of accelerating AI. And that means if we're working on automation, kind of prepare so that people know they're not going to just get snap fired with no warning without that kind of thing, know that we wouldn't make this kind of horrible trade where they have to like we kick them out of the ore or whatever and they have to like scramble to find the new thing if they still believe in there. Those are all just like bad plans. And so I think we'll want to design this org resolution. But also a lot of other companies will face similar challenges of designing the culture and the plans within companies and research labs and so on to prepare for lots of change. And one way to prepare for change is to say we're not going to cut you out of all of your resources at the last minute just because we think we've got to confidence beyond
Tom
not snap firing people. What are the other things that you can do to build a culture?
Jeffrey Irving
Yeah, I think that one thing I've learned over time is that it's very important to have even just how you craft Slack channels so that people feel comfortable speaking. And so you can imagine it's very bad. Before automation, you have a Slack channel where say a bunch of junior researchers are discussing their details of research and you have a bunch of high up executives just lurking and observing what's going on. This inevitably kind of just saps the conversation or pushes it into direct messages or something like that. There should be some intention required to kind of put your Thinking your context into the machines, you should be doing that in a way that you kind of want to do it. You should have the option of having meetings obviously that are not like washed by the machines. So there's some designing of a non dystopian. Org which I think is kind of it's table stakes. It should be easy to do, but you have to do that intentionally. I think some companies have done tried
Tom
to go tracking a bit too far
Jeffrey Irving
in this direction and gotten a bunch of backlash and could have understood all reasons. So there is a desire, if you're trying to automate things, of having all of the context available to machines. But you shouldn't do that too much because it would be a bit dystopian, not totally indiscriminate.
Tom
What kinds of talent are you most hoping to get into resolution?
Jeffrey Irving
So we are looking for a mixture of sort of very standard kind of ML engineering and research talent for automation, for some of the empirics and then also a lot, hopefully a decent number of very, very strong mathematicians and computer scientists and physics physicists to do to kind of push forward these various frontiers of theory research. I think part of the story, the claim or the founding bet here and also to some extent when we were doing kind of the AC alignment project back at AC is not much has been tried. So we haven't really treated as a world alignment, as a problem worthy of taking the best researchers from various fields and kind of putting them on the problem. And now that has gotten easier because everyone is getting more worried. And still I think not enough research has happened to be confident that there isn't low hanging fruit. And so possibly the definitions are fairly shallow. If you get people who are very good, they can find the right way to model the situation without even that much fancy mathematics. But just like some understanding of how we approximate superintelligence on paper and that might give us the answer to how to make this go well. I think there is this very important principle of the shallower the mathematics, the more likelihood there is of fast progress. If we had to do just an enormous amount of incredibly deep theory building across decades and decades of time, that would be very rough. If it's like no one has really found a good way of modeling this notion of kind of speed plus heuristics and reasoning about the complexity theory of that class of algorithms, that could be a thing that we make progress on in six months or a year. And then I hope that we can make a very fun environment where the human creativity part of the Problem is, or eventually the machine creativity is finding these definitions, figuring out how to model the situation, both alignment and capabilities of these models. And if you have below you a bunch of automation for doing kind of for expanding out candidate conjectures and proving them correct, or finding counterexamples, or doing numerical experiments and all of that the models are very, very good at because it's the thing they're already good at and they'll keep getting better. Then you can kind of play in definition space, play in modeling space, which
Tom
is the most fun thing to do. Maybe this doesn't make sense as a question, but what are the clear defin that you'd be keen for us to get a better sense of? So one is how to define this speed and heuristics model?
Jeffrey Irving
Yeah, I think speed and heuristics and what's so an example of a toy model I would like to see is right now, the labs, they do some pre training, they take a model, they ask the model to make some data, they train in the data, they iterate this weird process. They might have literally thousands of different modes of asking the model for data. It's a very, very complicated object. This is like a modern training stack. But you could imagine distilling this down to some very simple model which is like you just have again pre training plus self generation and you iterate that maybe that already captures enough of the flavor of RL that you don't even need to add RL as a component to that model. And if you could build that toy mathematical model such that it represents emergent misalignment and subliminal learning and these other kind of phenomena we've seen in the last few years, and then explore them more rigorously, both in theory and in doing maybe very, very scaled down empirics, that gives us a playground with which to explore in algorithm space.
Tom
You would use this model to understand subliminal learning?
Jeffrey Irving
Yeah, subliminal learning is when you have a personality trait of a model and it generates data for another model and then the next model inherits the trait, even if the data generating is unrelated to the trait.
Tom
But there's a model which you've trained to like owls. You get that to output a bunch of numbers and you train a different model on those bunch of numbers and it somehow also inherits the fondness for owls.
Jeffrey Irving
But I think this somehow is actually not that mysterious to a first intuitive approximation. It's just because there's this low dimensional structure and the liking for owls is correlated with all these random other Things including numbers. And then that structure is flowing through this channel and then showing up in the resulting model.
Tom
And your intuition is we have a pretty good chance of understanding how that low dimensional structure forms on a theoretical level.
Jeffrey Irving
And then if we have that understanding, we can use it as a lens to rule in or out various algorithms as being like optimistic or pessimistic. Will they succeed in preserving or mapping that structure in the way that we want across? One of the traits of a world class theorist is just definitional creativity. And I think to some degree that can be. I think there's enough of a chance that can be ported across to this new area of alignment, new to them, that we can make progress quickly.
Tom
Sounds like a good deal for them. They should.
Jeffrey Irving
I think so. Also we can pay them well, so that'll be good as well. So I guess maybe the big thing to say is again, we have some inside view reason why we like each of the individual areas we're kind of thinking about like learning theory, scalable oversight, Personas, agent of the nations philosophy, this kind of thing. We may have missed some. If you have a thing you want to do, if you're like, you like theory and you think by the general story of gains to scale for org scale, of having kind of shared automation and being able to share ideas between areas, and you think this is a good place to work, but it's not in your list that you've given in this podcast, still reach out, still pitch us. And I guess an important principle is that we will want to believe in you to some extent, but not that much. Everything here is a bank shot. We're just trying to spread the probability around a bit more than the current labs are doing. And then also we don't need every area to be the same size. So if there's a few people working on a particular area, if we have small critical mass, I think that can still be quite powerful. And we expect a lot of the basic understanding of both how to do automation for theory and also just basic things like reward hacking, how to model it kind of thing. Those may generalize across areas of theory in ways that are quite useful.
Tom
So you've picked a bunch of fields which you would love to hire from, from Resolution. How did you pick those particular fields? Like what is it about complexity theory or other fields? That is why you think that's going to be particularly helpful for your research agenda.
Jeffrey Irving
So there's this inside view case for complexity theory that it's like modeling superintelligence and kind of weak and strong amounts of compute and how they relate. But then the outside view case is we do just need a more rigorous understanding of this kind of problem as a whole, this problem of alignment. And there's a bunch of areas that might be relevant for that. And we want to try to be a home for a bunch of those in a way that takes advantage of scale by sharing automation and sharing ideas, sharing how to model kind of the basic concepts of reward, hacking and misalignment and so on. And the hope is that that scale will give us kind of faster progress kind of throughout different areas. And then also we don't think we are the only game in town. So we will try to publish things. We want to be a kind of friendly member of the community, kind of feeding back. It's possible that we have some ideas and then someone else takes those and actually solves the problem in some useful way. And so we'll try to get that balance right as well.
Tom
And if you do this research, you try and find solutions to these obstacles. But what happens if you don't find these solutions? You yell? What exactly does that. Yelling?
Jeffrey Irving
Yeah, I think I actually misstated this in the initial blog post where it's like we might need to yell as if it was like an eventual thing we do. I think the rather thing is try to build a culture and a comms practice and so on, where we're just putting out this mixture of obstacles and glimmers of success kind of throughout all the time. If we find a way of modeling the alignment problem that says it is hard, that is extremely valuable as a publication and should be celebrated as such. Both because it might tell you that you need to pause, it might tell you you need to be more careful, or it might be the thing you need to then filter down and narrow in on the right solution kind of in the long term. And so I think building a culture of equally celebrating both positive and negative results I think is kind of key to this whole exercise. And a hopeful thing there is there are somehow in complexity theory, in various areas of physics and mathematics, there is this like some of the highest profile results are obstacles. Like there's in complexity theory, there are these like, there are actually three obstacles to P versus NP called relativization, algebraication and natural proofs. And people just like those are buzzwords. People celebrate them. They're like, they're very famous in physics. There's the firewall paradox, which is like an obstacle about how does quantum gravity work near black holes, which is again is like very celebrated it has this cool name, the firewall paradox. So the hope is that that culture is not something we have to create afresh. It's like a thing that pervades these areas of theory already. And I think again, bringing that in and finding those obstacles is both just a necessary part of the modeling process and then also can either plays into hey, we should slow down even more because we have these horrible obstacles, or it tells you to ramp up data or something, or care and time in some algorithm which kind of might work, might not work, or it says, here's the lens your new algorithm has to go through and it lets you find it faster. So we published a paper, Automated alignment is harder than you think. And the reason why I think this fuzzy evidence problem applies more to alignment is that I think there's more of a story for how to incrementally work on the problem of improving capabilities across time with capabilities than there is for alignment. So you just get to do hill climbing in some sense on capabilities and labs mess up. They do sometimes produce models that lie more or more reward hacking. They have to go back and fix the training signal. They do this both internally within labs, but also literally. Some, some deployments have been missteps in this that they had to fix. But they generally get to climb this hill of gradually improving capabilities because we can measure them. I think because of this effect that everything could shift as you cross through human level intelligence. If you want to do a bunch of research with automated machines, with machine automation prior to human level, you don't necessarily learn that much about or you learn some, but not as much as you would like about this future superintelligence. And so the worry is you just don't really see what's going on. Your experiments that you've done for prosaic alignment, on avoiding kind of current model wart hacking, just don't tell you what you need to know about the superintelligence and so you can automate them, but you haven't automated this conceptual modeling of when things will break down or not as you go through this kind of scale up.
Tom
The reason that the iteration works for capabilities but not for alignment is because there's a phase shift between. Because the phase shift between subhuman and superhuman applies in alignment, but it doesn't apply for capability.
Jeffrey Irving
I don't know if I think it does. But so the question is, say we get to ASI in, I don't know, five years. I think the skills you will have learned in the meantime on capabilities will. They will be Skills that got you to the next rung up and then as you go. So the question is, when we get up to human level, what will happen? So up to human level, you can supervise the model. And so you can still be hill climbing even past human level, you can be sort of trading, it'll become harder, but you still get to say, run the model for a small amount of time and then supervise it with more human attention or supervision is a bit easier. And so in the areas where you can do this thing, you can still be climbing. And then the question is, what happens around this point? And I guess my intuition is that when you're doing a very difficult software engineering task, you have to do a ton of planning and kind of subtlety and subtle reasoning to be able to say, do a month's worth of human type work as a model over a period of a day or an hour or a week, I don't know how long it will take. And so the question is, if you hill climb your way up to a model that can do that level of reasoning, are you close enough to the danger point that you'll get there because by proximity or where you kind of stall out at that point. And I think the reason I think you won't stall out is that's already stronger than humans. And so I think I just expect it to continue basically via momentum up past this level. I think also one thing to say is the way to get to kind of super intelligent alignment is I think at least some component of scale of oversight where the model is supervising the models. All of the labs are doing some version of this. They're just doing kind of the empirical hill climbing version. And there's a big space of possible scalable oversight algorithms. I guess my claim is that some of these work for alignment, some of them don't work, but we will be able to be hill climbing away to things that work kind of empirically. And so that will give us some ability to kind of push past human level kind of by a fair way, just from kind of this overhang of hill climbing on scalable oversight algorithms. And then the worry is that we've picked the wrong ones and we won't know. We won't know. And that will kind of shift. But I think it is, this is an area where some people have very different intuitions that in fact that this effect will call capabilities to stall. And I think it is one of the arguments against the speed.
Tom
Yeah, and I mean, I guess it kind of relates to the GO example you said before, where you try and train a super intelligent go model, but it gets a very bad go player. It will also learn bad strategies. And you initially said that that's the reason why there is this kind of. If you personally don't understand what's going on, you might not be able to reward the train the model to do what you want it to do, even if you're doing amounts of compute that would normally generate super intelligent play. But yeah, it's occurred to me that that would be an argument why capabilities would slow. In that case, the capabilities are slowing. We're also not aligning it to what we want.
Jeffrey Irving
Yeah, notably though, literally what happened in AlphaGo is that they did a bunch of iteration, sometimes they did mess up and they trained a model against itself in a way that overfit to some weird model distribution. It got to apparently superhuman elo playing against itself and past versions of itself. And then they tried it against the human and the human wiped the floor with the model and then they tweaked something and then that was fixed. And then again the model wiped the floor with the human. And so the question is, as you're playing around in this space trying to do model self supervision, can you do that kind of iterative tinkering? And I guess I'm more optimistic that that tinkering gets you capabilities because if you mess up, you get a model which is weak and then you're like, this model is shit, I can't use it to do things. And you kind of notice that over time, maybe it takes you a little while to realize it because it's like, like domain superhuman. And then you fix it. So the failure mode is towards weak models that then almost by definition you just notice that eventually and fix it. It might take some time. The failure mode for alignment is you make a mistake, you deploy the model, it takes over the world and then you're done. And so I think if you had this model of the world where in fact say we imagine they were perfectly symmetric and there was the same failure rate for a given deployment of an AI model to have failed to get this ad hoc kind of tuned. Scalable oversight, right? The same error rate between capabilities and alignment, then up at superintelligence land, say 20% of the time you fail and your model is bad for. It's like you miss a generation and then 20% of the time also the model takes over the world and one of those two things you can iterate and the other one, the other one you can't.
Tom
Yeah, okay, that makes sense. Yeah.
Date: August 11, 2026
Guest: Geoffrey Irving (Former Chief Scientist, UK AI Safety Institute; co-founder & Chief Scientist at Resolution)
Host: Tom Reed
In this episode, Tom Reed interviews Geoffrey Irving—one of the world’s most experienced AI safety scientists, known for his “full stack” work spanning early alignment theory, empirical research at OpenAI and DeepMind, and advising the UK government as Chief Scientist at the AI Safety Institute. Together, they discuss the urgent challenge of aligning superintelligent AI, the prospects of current lab strategies, theory versus empirical bets, policy interventions, phase transitions in capability and alignment, and the future trajectory of humanity in a world with superintelligent models.
Irving doesn’t pull punches: he argues for more caution, greater attention to possible discontinuities at superintelligence, and richer theoretical understanding. He is skeptical that current generalization properties and scalable oversight protocols will “just work” at higher capabilities, and highlights the importance of both strategic government action and expanding the pool of alignment researchers outside the labs.
What could misalignment look like at high capabilities?
Differences from today’s models:
On phase shifts & generalization:
Portfolio strategy—multiple theoretical and empirical bets:
Why isn’t theory a greater focus at the labs?
Low-dimensional structure & Personas:
Analogies with phase transitions and political systems:
Geoffrey Irving’s core message: The alignment problem is likely harder and less “hill climbable” than the labs hope; our understanding is worryingly shallow beyond current model scales. We urgently need to pursue a wider range of theoretical and empirical approaches, aggressively staff up government and nonprofit research, and—above all—slow down while we still have time. Irving’s blend of technical depth, clear-eyed skepticism, and pragmatic focus on policy and institutional structure makes this episode essential listening for anyone wanting to understand the real challenges of aligning AI before the superintelligence era.
For those interested in joining or supporting Resolution, deeply technical work in ML, mathematics, computer science, philosophy, and theory is highly encouraged—all hands are needed as soon as possible.