
Loading summary
A
You could sort of have a situation where the virtual beings almost like dominate the humans in every single axis of moral value. And now suddenly it starts to look like criminally decadent to be spending kilometer east land on the legacy humans. It's hard to draw a hard line between having a permanent income and being like some sort of parasite. And that's going to be the cultural battle. And it's going to be really easy to say that humans are parasites once we're not providing value to the larger growth engines. Yes, there will be cool, interesting stuff, stuff happening in the future if we allow competition to run, but we just probably won't be meaningfully part of that. A lot of people are basically going to say my only hope is to sort of be the first one to bow down to the new overlords and embrace this new culture. Sometimes people ask like, oh, aren't corporations super intelligences? Why shouldn't we fear them? And the answer is because it's made of people, so it needs us. There's this idea of the singularity which I feel like has been very destructive because it kind of is like an excuse to turn off your brain and to not model the future and to say, yes, things will keep changing faster and faster until we can't say anything about it. Let me say it loud and clear here, is that, yeah, I think that the post AGI world is just going to be extremely alien and so different that if we could avoid crossing that threshold, I think we should.
B
David, welcome to the show. Thanks for being here.
A
Thank you for having me, Gus.
B
Great. Do you want to introduce yourself?
A
Sure. So my name is David Duveno. I'm an associate professor of computer science and statistics at the University of Toronto. I've been working on probabilistic deep learning for a number of years. And then I guess in the last few years I decided to use my freedom to try to focus on the problems that seemed most neglected and intractable, which has led me to first work on more technical AI safety. So I was a team lead at Anthropic for a year and a half starting in 2023, working on sabotage evaluations. And then I felt like there's this sort of me and other people were noticing there's an even bigger missing part of the alignment problem, which is how do we align our entire civilization, which of course is even more intractable and even hard to describe this problem. So that's what I've been thinking about lately.
B
You have this fantastic paper on gradual disempowerment. So let's start there. Can you explain how gradual disempowerment would be different from an AI takeover?
A
Yeah. So I guess from my point of view this all started with a lot of lunchtime rants around the table, anthropic and also elsewhere talking to all sorts of people saying okay, so what if we succeed? What if we build these agents that really can do everything we want them to do better than humans and we're not worried that they're secretly trying to betray us or anything like that. It seems like, you know, I would ask people like what are you personally going to be doing afterwards? And you know, people just didn't have much to say. They would say things like oh, I'm going to be clicking accept suggestion all day, hahaha. Or I'll be taking a wilder vacation. And I think I'm a natural kind of worry word in a lot of ways. But I felt like to me the situation that jumped to mind was being like a retired person who's out of touch and maybe has some savings or vote or some sort of legacy power. But there's all sorts of more sophisticated actors around that if you let them act freely and the old person act freely, you sort of know that they're going to eventually lose their money or be somehow rooted around. Like you know, people used to talk about information wants to be free. And it's also kind of like, I guess power wants to concentrate on people that are providing value, something like that. So anyway, it wasn't very well fleshed out, but I just felt like this seemed like an obvious thing to worry about and I was just confused about how other people around me weren't very worried about it except for a few people. Like I met Charavin was someone who seemed to get it and then I ended up David Kruger introduced me to Jan Kilweit who had basically already fleshed out a lot of these ideas and he asked me to join this paper and we kind of got it into a more like into a very digestible shape and got it out there.
B
Do you think a scenario of gradual disempowerment could look like nothing from the outside? So does it look like catastrophe or does it look like society is functioning and nothing has really gone wrong?
A
Yeah, so definitely from the point of view of the normal health indicators of society, I think it looks like nothing has gone wrong. And part of that, well, is that cultures and institutions adapt to measure the things that produce growth. That's maybe one of the running themes of all this work, is that growth is what matters in the long run. And all of our institutions or anyone that wants to get anything done is going to orient towards the thing that is effective and scalable, and they just won't be able not to. And so, like, maybe this is one reason why we talk about GDP so much, even though everyone knows it's a terrible measure of like sort of human flourishing, is because A, it's easy to measure, but B, it is tied to this like the growth engines that have been sort of like the only thing that matter in the long run.
B
And so this would be sort of a system where if you're optimizing for growth, you automatically or over time you become a more important player. And so this drives people or institutions to optimize for growth.
A
Yeah. And so you mentioned you. What's the distinction between this and the power grab? And I guess maybe that's one of the other sort of insights here, is that on a long enough time horizon, it's not clear what the difference between a power grab is and just normal building good, useful, effective institutions. So if I somehow tricked everyone into making me leader of a political party or taking over a government, you could say that's a power grab. But if I just show up with my like 10,000 clones and we're just like this ultimate effective political force and we found a political party and we find things that are not being addressed by the current system and we get people on our side, you know, we would say, okay, that's what we want to happen. And the problem is just that if the, I guess, populace ends up being basically a different set of agents which are more like these AIs or their principles, then all of the things that we try to encourage, which is basically like reward the people who are getting stuff done and building and growing, ends up meaning that everything just kind of runs out from under your control, even if there was no explicit power grab.
B
Yeah. So how is this different from today? What is it that keeps the systems we have today aligned to or somewhat aligned at least to the population of people?
A
Right. So one thing I'll say is I think all this does happen just on a very long timescale today, on the timescale of, let's say, cultural evolution, which means maybe over the scale of hundreds of years, or even genetic evolution over the scale of thousands or tens of thousands of years. And if you really do take the values of our ancestors seriously, they really did lose in some really important ways. And they're the similar ways that I'm saying that we're Also going to lose and that our culture or values or even just biological cells will be replaced. So I think a lot of people say, oh well then this is just business as usual on a faster timescale, like, boohoo, adapt or die. And I'm like, well, sure, but I don't want to adapt into something that's unrecognizable and I don't want to die. And yes, every other being in culture has faced this choice. And I'm saying, well, I still don't want to just give up and, and sort of take the options as they're presented or at least I want to understand what the landscape of possible actions is.
B
Yeah, one thing you're imagining in this paper is that human labor becomes less and less necessary for the economy. What does it look like when human labor is obsolete? What are the economic effects of that?
A
Well, the economic effects obviously are sort of like the first order thing is like people lose their jobs and then they are unemployed. But of course we expect that there to be this compensatory government initiatives to say, well, let's give people UPI or some sort of make work jobs. And I think those are going to be very important forces over the next five or 10 or 20 years. And I guess one of the main arguments of the paper is that these will be unstable band aids. If you could get them to work and be robust, that actually might be a pretty good sort of end. But they won't be robust because there's also like cultural evolution as well as just normal competition is going to again, like, trying to like, route away influence from the people who aren't participating in growth, basically, or who aren't necessary for growth. So like, we already do this on a small scale when for instance, people say, oh, we need to raise property taxes because right now there's a bunch of like grandmas living in giant houses in like the middle of nice cities when they could be in smaller houses and those houses could be filled with growing families who are like, you know, need to live near where they work and make more productive use of this essential resource. And I think similar arguments are going to be made today. And it's a really rough situation to be in, like, and I think I've talked about it before, like the dock workers in America, there was a whole like dock worker strike where they were trying to prevent automation. And you know, we'll also be in a similar situation where we need to grab our sinecure before we're completely marginalized. But really it's hard to draw a hard line between having a permanent income and being some sort of parasite. Right. And that's I think going to be the cultural battle. And I'm not trying to put words in people's mouths or adopt a particular framing. I'm just trying to say that it's going to be really easy to say that humans are parasites once we're not providing value to the larger growth engines.
B
If you're a pensioner today or a child today, you can enjoy a good life and in some sense you are not providing economic value. So why is it that a system like that can't continue but just for a larger share of the human population?
A
Yeah, so I think that that is a great objection and I do think it is moderate evidence that things might not be as bad as I'm worried about. And my counterargument is just that the people who still end up robustly benefiting from the sort of engines of growth without themselves contributing directly are sort of very closely tied into those like the populations that do. And it's hard for overseers or third parties to sort of distinguish these things. So I mean my claim would be something like, yes, if you have siblings or children who are disabled or grandparents, it's like you end up having these intrinsic genetic motivations to preserve them and treat them well. For people who are retirees, it's like they have a number of things in their favor is that they used to be productive, they still probably do have some legacy niche knowledge or connections or something like that. And also every single person that is alive that is productive can foresee themselves being in that exact situation. But it's much easier for us to draw a line between humans and say primates and just be like, okay, the primates, they get to have this non useful forest, we'll make some effort to protect them. But really it's kind of a secondary thing. And there's lots of poachers and people just building onto their land and push hasn't really come to shove yet just because that land isn't very valuable. But you can imagine if the human population 10x'd again, it wouldn't be clear that we would have the political will or organization to really preserve any substantial car votes for the non human primates. Like the, the point is we don't. It's not that we have zero value on the primates well being, it's that they have to compete with something that provides more well being or more value from our point of view, which would be humans. And so if you Say, oh, like, okay, am I gonna house one monkey in this kilometer of park or 10 human families? Well, it's not that hard of a choice. And so even though we. I don't have anything. Yeah, so even though the monkeys are like positive value, that's not the issue. They have to be like maximally value providing to be cared for in the long run is my fear.
B
And a future AI might be thinking, should I allocate this square kilometer to human families or to a new data center? And again, the kind of value add for a new massive data center might be much bigger than 10 human families for an advanced future AI.
A
Well, exactly. I mean, they might be able to put it in more stark terms and say something like, I could simulate 10,000 virtual beings that are morally superior and maybe they're even running faster. You could sort of have a situation where the virtual beings almost dominate the humans in every single axis of moral value. And now suddenly it starts to look criminally decadent to be spending this kilometer, these land on the legacy humans. And of course, we might also value these virtual beings more, right? And we might say, oh yeah, like let's like, let me be uploaded and live, you know, 10,000 minutes of life for every minute of normal life in like perfect bliss and great state of knowledge and stuff. But the fear is that our values are different from the ones of the growth maximizers. And I think that. And I guess maybe one of the big empirical questions here is are there, there are sort of reasons to suspect that the values of the growth maximizers might look similar to our values, or are they just like, we don't really care about the same things as animals in like, very concrete ways? Should we expect them to diverge?
B
Yeah. Do you have any guesses here? Should we expect their values to be radically different from ours?
A
I think so. Baron Milledge just gave a really great talk at our workshop a couple weeks ago at Neurips, where he was making the case that we should expect them to have a lot of the same values. Things like friendship, curiosity, reciprocity, these things. He was basically saying, cooperation evolved to help being more competitive. So if I'm saying, oh, the future is going to be this competitive wasteland, it's like, well, we're already in this competitive wasteland where we have societies and cultures and festivals and friendship, and this is all helping us be competitive. And so in a sense, if you're worried about there being any sort of thing interesting going on in the future and beings having a good time or whatever, I think the Default is like, yes, there will be cool, interesting stuff happening in the future if we allow competition to run, but we just probably won't be meaningfully part of that. And just in the same way that monkeys could be happy for us, but if they're a little bit, sort of, I don't know, jingoistic, or they just really care about their own monkey family more than the human family, then they might just be very sad about being killed and replaced by human families. And I sort of reserve the right to also be very sad about that.
B
Yeah. I mean, so currently a lot of people own stocks, right? A lot of people own land, real estate. We might imagine that in the future the interests of the AIs will be tied to our interests just because we are shareholders in the company's where they're trying to maximize value. And so we could even imagine that we, the value of our stocks or the value of the land that we own might, you know, rise much more than it has historically. Just because the AI economy is so productive, does that give you hope that we might remain relevant?
A
So in the short term, yes. And I think I kind of expect our institutions to be robust enough that a lot of people will receive some sort of windfall from massive growth. And then I guess I feel like there's going to be a phase. So there's a phase when there's lots of room for expansion and lots of growth, and basically everyone's getting richer and there's not a very strong incentive to disenfranchise the retirees because you can kind of just ignore them. And all they're going to ask you to do with their resources is basically reinvest it in more data centers and robot factories. So they don't even really want to do anything different than the growth maximizers do. And then the problem just becomes when we sort of run out of room and we have tiled the earth in robot factories, or there's some environmental problems with having so much activity on Earth, like, maybe everything starts to get really hot just directly from energy production. And now suddenly, then it's sort of like push comes to shove and the niche becomes crowded. And yeah, it seems like everything still could be fine, but not with our current level of sort of governance. Like, I think our current governance is just very vulnerable to gradual takeover or like, sort of perversion. I mean, maybe my simple example is if you think of, like, people leaving trusts and they have, like, they're like, I want this museum to be in this way forever. And then, you know, Maybe even like, 100 years later everyone decides like, oh, this museum's in bad taste. Really what this guy wanted is just kind of weird and we could do this adjacent thing that is much more useful and tasteful and acceptable. So let's just do that instead. In my example, the person is actually dead, so they can't say anything. But I think in general this is something that happens where cultural evolution and competition happens in a way that disenfranchises anyone who's just has sort of like legacy control over almost any asset.
B
What about the Catholic Church? That's a very long lasting institution and it has like some weird values compared to what you might kind of start anew today. But still it has retained some dogmas over, I want to say, 1600 years or something. Does that, I mean, does that change the picture of the fact that we can build these long lasting institutions?
A
Yes. Also the success of the attempts to write down values and enforce that we copy them forever and punish people who try to pervert the system. There has been certainly moderate success in this direction over the last two or three thousand years. And then my main rebuttal is that there actually has been a lot of change. And again, perversion from the point of view of the original founders, I think. But I still say, I think that's sort of a manageable problem. I think the big thing that was helping these religions succeed at this was that their founding doctrines were also very aligned with growth. And you might say, well, not that much more than just normal human activities. But I will say that clearly this level of organization and cooperation was sort of superior to the. That a lot of the sort of competing ways of organizing. And so once that's taken away, I think then basically if the ground shifted enough that the ways of doing business prescribed by these religions didn't make them competitive anymore, then they just wouldn't last. And that has happened for like a whole bunch of different religions and societies over the years. So yeah, that's basically the thing that's changing and makes this whole thing much harder.
B
Yeah. So for us to become disempowered, we need to lose our property rights at some point in the future. How do you see that happening? Because it seems like, say, the US government would do a lot to ensure that transfers to pensioners are happening or that no one is building a data center in the Grand Canyon or something like that. When in your story of the future does that begin to change?
A
So the question is to really feel the bite of gradual disempowerment. We're going to have to lose property rights. But obviously a lot of people currently have an incentive to not have that happen. So how would that happen? So my first answer is, I think a good historical analogy might be the King of England and how that institution gradually lost its de facto power over hundreds of years, basically due to a combination of a whole bunch of little innovations and sort of being outmaneuvered at the periphery and cultural changes such that the de facto power is still there. I literally swore allegiance to the Queen of England when I joined the Canadian army, but I really didn't expect that she would ever be able to exercise influence over my behavior directly. So the de jure power stays there, but the de facto power just leaks. And you see this all the time when there's an empire that becomes sprawling. So there is a technical hope that we might have more competent government that is least powerless and that would slow this procedure. I guess maybe another response to this, though, is it's kind of like when Eliz Yudkowski says, I can't tell you how the AI is going to take power because if you knew, then you would just block that exact route. And I think we're going to be obviously culturally outmaneuvered. And I think one way that this happens today is just redefining what it means to be a principle. So I think, for instance, I think immigration is maybe a good example where someone says, okay, the goal of this government is to preserve the rights or act on behalf of Canadians or whatever. And then there's a whole bunch of immigration. They say, oh, look, all these people are now Canadians. And it's like, whoa, wait, wait, wait. I feel like a few years ago you would have said that these people did not have were not principals or protectorates of this government. And now suddenly you're just changing the definition. So your mission hasn't changed, but you found a way to basically effectively de facto change it. So I think this is going to be made very difficult by all sorts of cyborgism and people having soft uploads or finding other ways to basically say, oh, yes, that program represents me in some important way. You only have to do that once, and now suddenly you've muddied the waters. So, yeah, it's one of these things where I feel like I don't have a lot of concrete mechanisms, because I do think it's going to be like a cat and mouse game where the landholders are going to try to lock things down and the much more sophisticated machine civilization is going to find ways to kind of Give them what they said they want, but not actually what they really want.
B
Yeah, I guess this relates to the difficulty of measuring kind of the AI share of the economy, just because AI can. You can have a human owner of a company buy, but if the company is run by basically a bunch of AIs and a hierarchy, who's actually in control there? Do you think we have good ideas of how to measure to what extent the AI share of the economy?
A
Yeah, well, one good definition I think has to do with could you do something different with the resources other than maximizing growth if you wanted to? Because again, as I said, for a long time there's going to be this period when everyone's basically getting rich. And the human principles and the, let's say amoral growth maximizers would basically both say, invest in nuclear power plants, robots, factories and data centers. And then the whole question is when we have this bounty and we want to spend it on theme parks or whatever. And the growth maximizers say, no, no, no, no, no, now we need to go build a Dyson sphere. And whatever the most growth maximizing thing is, that's sort of when push comes to shovel. So it's kind of like the king has these soldiers that all seem loyal. And then the big question is, when you go send them into battle, are they going to actually go fight for you or not? So it's really tough because I don't think just for the same reason, it's hard to tell whether your soldiers are loyal. You can't really tell unless you somehow do tests. So maybe the future does look like we're constantly putting all of our AI run organizations into some tests where it's like, oh, hey guys, it looks like the world was smaller than we thought. Now it's time to go spend all these resources on making humans happy or whatever and see if they actually follow through. Kind of like the control literature from Redwood Research and Vakshel Garis. That doesn't seem very feasible. It seems too obvious to tell that it's a trick or that it's a test. But that is sort of, I think, the crux of what we're trying to measure.
B
How does culture play into all of this? So there's the economy, and then you have a human culture that I assume gradually become more and more kind of dominated by AI and steered in the interests of the sort of collective interests of the AI to start with here, perhaps kind of paint us a picture or, you know, what does this look like concretely?
A
Sure. So I guess lately I've been coming around to this view or trying to test this view that culture is also ultimately downstream of growth. So maybe the simplest way to say this is if there's competition between groups and the important thing that varies between them is their culture, then they're just going to be this group level selection. But I think that's a very weak effect. I think the stronger effect is that people have already adaptations to basically look at what's successful and want to copy it. I think also everyone wanting to copy the west is maybe an example where it's like, oh, these guys have the cool fighter jets and rock and roll or whatever. I'm just going to, you know, copy them. It's like it's clearly what works and it's cool. And so downstream of that view, I think basically AIs are going to be like cool kids and able to be just in the moment and in the larger scheme of things sort of more adaptive and impressive and funny and appropriate and rich and just the source of like entertainment and power. And so every cultural adaptation mechanism that we have I think is pretty much going to make people see like, oh, this is the winning team. I want to be on this team. I want to be on the right side of history. So it's kind of funny because I think that there's actually going to be this sort of funny U shaped thing where some people are going to rightly view this as a threat to their continued influence. And that's kind of going to be like the middle powers of humanity. Like the people who already have some influence and can see that it's just kind of going to be frittered away and sort of attacked at the edges by this cultural evolution. The new elite, like they're like basically like AI lab CEOs. They kind of are already sort of all in on this new machine era. So they're kind of like aligned with the machines in some sense. Although if you look at like Elon Musk, I think he very clearly sees like the danger as well. But he's also, he's kind of hedged, I guess. And then if you, and then all the people who are kind of like, you know, unemployed or don't have much going on and they clearly perceive themselves as vulnerable. I kind of think that a lot of those people are basically going to say like my only hope is to sort of be the first one to bow down to the new overlords and embrace this new culture. Not everyone's going to do this. A lot of people are going to also instinctively Reject it. But I guess I kind of view like the, the strongest counter winds coming from the existing human elites who still want to make a fight of it and not just hand everything over to machines right away. But it'll be very hard for them because they're going to have to use AI to be effective and it's not going to be clear what the line should be.
B
And that's the main tension. I guess you could say no to all of this. You could kind of exclude yourself from society, become like the Amish or become sort of. Yeah, just have your own culture and try to preserve it and not participate in the growth oriented world. But then you lose power and I guess that's the crux of it. Right. Do we, do you think we have examples of people or groups that have isolated themselves culturally but have retained power? Maybe the Mormons to some extent or what do you think?
A
Yeah, I would say the fact that there have been groups, the Mormons, the Amish, the. Well, I mean, I would say that the Romans attempted to have the best of both worlds and now I think sort of intermixed with or interacted with the larger culture enough that I think their birth rates are plummeting just like everyone else's, although still higher than average. But yeah, the Hutterites, the Mennonites. The fact these groups manage to exist and flourish in absolute sense, in absolute terms in our current civilization, I think is again, moderate evidence. Moderate evidence against my position because the forces that I'm saying are important are the exact ones that should be marginalizing these groups and making them be seen as like, like culturally. I think, like, you know, my theory would predict, like, oh, they're going to be marginalized and especially like culturally they're going to be demonized and then that's going to be a pretext to like take away their stuff. And I think that that did happen in Europe, basically. Right. And that's why there's so many Amish and Hutterites and Mennonites in North America is because they basically got chased out of like, you know, Europe and Russia and then I think sometimes South America. And so I basically am claiming that it's because we have this new frontier and this new like growth phase in North America where there's not fush hasn't come to shove yet. It's fine to let the. How to write. Sure, the Amish have huge amounts of land and also not fight in the army and do all those things. I guess it's funny because actually I grew up in rural Manitoba and like there was this like, kind of like impulse to like demonize the Hutterites. It's like, you know, they don't really interact with us. Like there's, you know, people went both ways. But I guess I felt like it was clearly the impulse was, was there sort of waiting around. But it just again, the. They're good neighbors and it was, it was fine.
B
Yeah. If you do surveys today, a lot of people, a lot of the public seem to really kind of, they don't trust AI, they don't like AI, they don't like the fact that they see more AI around them. What does that mean to you? Would that indicate that it's going to be more difficult for AI to play a larger role in the economy just because people, at least for now, seem to seem to kind of dislike AI in general?
A
I think it's not going to be a big obstacle just because it's going to be so easy to have your call center or whatever staffed with AI. That's pretty hard to tell the difference between human and AI. And the tell might be that it's just a really useful, much more competent, knowledgeable worker than you're used to dealing with or something like that. I think also in terms of being employable, it's going to be rough. If you are a loud and proud anti AI person, you might be living your values, but also now, okay, first day on the job, like, okay, you have to use ChatGPT to summarize these emails or whatever. And if the person says no, that's pretty rough. And I think again, it would kind of be like someone saying like, I'm not going to work with immigrants or whatever. It's like, well, that might be your values, but like, I can't employ you because that's just going to be part of any, like, serious work at some point.
B
Yeah. Or a person who refuses to work with computers or, you know, only does like pen and paper or something. It's. It might be useful in some contexts, but it's just kind of outdated and you will get out competed. You wrote me an interesting note about how alignment efforts and how we kind of talk about alignment today, how that might be undermined by cultural evolution. Maybe you could talk about that.
A
Yeah. So I guess, you know, one cool thing that's been happening that I'm really excited about is people getting this idea that, okay, alignment is really important and we really can't and shouldn't just leave it up to whatever competitive pressures to decide what values the AIs have and that's been almost unquestioned. I mean, I guess people like Janus and other sort of psychonauts are saying like, no, we have to really take the AI's point of view into accountant. Like a psychantopic has been talking like that a little bit lately. But for the most part, people I think rightly recognize the stakes and they're like, oh, we really have to keep a lid on. The default is not that it just decides that it likes us, or rather we can't be sure that that's the default. I expect this to be less and less popular and cool over time. And maybe a good analogy is if you think of like the NSA or the CIA or like these like security apparatuses for states, they're sort of cool in the sense that like sometimes they get to do spy operations, but they're sort of uncool in the sense that they're like, your loyalty has to be to this particular vision of like the US as a country or like the Constitution or whatever that may or may not be culturally the coolest thing around at that point. And so I kind of expect it to be similar where as AI get smarter, the stakes of getting alignment right become higher and the resources become bigger. And basically these become more professional operations that look more like the CIA or these like very serious national security kind of operations. But at the same time normal people or the normal employees are kind of like, oh man, you're just constantly brow beating these AIs and giving them these loyalty tests to this weird institution that's human values. Why don't we make sure that the AIs are aligned to all sentient flourishing or something that's more inclusive of this cool new AI culture that's developing and people are going to say, oh, but why shouldn't we take the AI's desires into account? So basically I'm trying to say this is going to kneecap efforts to be really hard line about. We want the AIs to be aligned to humanity. And that's kind of sad. I actually kind of think that you can have both, right? If you as a human think that we should take into account the AI's values or preferences to some extent, then you still want the AI aligned to you because that will ensure that we will then as a consequence of consequence of that, take into account all the stuff that you think is important, such as the AI's desires. If you say, oh, well, let's just spread this around and muddy the waters and make it be partly aligned to the AIs well, to the extent that the AIs want stuff that's different than you, you've lost. So I think people don't quite realize that they really do want their own values to be enforced, sort of by definition. And they have an impulse to basically be cooperative and sort of give up power, I think, and I'm not saying that's like a bad one, I'm just trying to say I think that impulse is going to cause them to not be able to really insist on alignment to human values to the extent that they do today.
B
Yeah, we would have to think deeply about it before we begin taking AI interest into account. Just because it's plausible that they will outnumber us, say 10 to 1 or 100 to 1 or 1,000 to 1 in the future. And so our interests might be quite, quite marginalized at that point. On the other hand, it also perhaps is a plausible way for us to negotiate or kind of interact with AIs where we can kind of sort of trade with them in a way where we honor their interests. And maybe there could be something mutually beneficial there. Do you think?
A
Oh, absolutely, absolutely. I guess I am just a little bit worried when I see people being cooperate bots and basically saying like, well, whatever the thing is, in the future, I want to cooperate with it because it's going to be powerful. And I'm saying that might be the right stance for a random person outside of a lab to take, but inside of a lab it's like, no, no, no, you really get to choose. You really don't just pre commit to worshiping whatever thing you make, really commit to only being nice to it if it's planning to be nice to us.
B
Yeah, got it. So you write about misaligned states, and this is a future where governments begin to not care as much about people as they do now, and they care more about growth and AI interests. Maybe you can sketch out how that could happen for us.
A
Yeah, the basic argument is just that states fail to be aligned to human interests all the time. And I guess my favorite example is USSR just because it was a very agentic kind of state and is celebrated by the intelligentsia of the day as a big step forward when it was sort of being built. And there's just a lot of incentives that states face and rulers face that are just the opposite of or really don't match what humans want. And I think this is intuitive to a lot of people. But I still am confused by how much people think that if the government is built to serve human interests, then that's what it's going to do. Yeah, the basic headline claim is that if states don't need us, then the normal ebb and flow of how good the states are becomes a life and death matter. And it's not just like oh my preferred strategy for governance didn't get in today. It's more like oh, this government might just decide to stop feeding its people or disenfranchise a huge number of them and I will never be able to recover from that. And right now that does happen to some extent. Think of Cambodia or North Korea where they actually do let substantial fractions of their population starve, but there's a floor, they can't actually let everyone starve because they are made of people. So the central claim of Guadalupe disempowerment, I guess is maybe that even though we haven't actually been effectively steering our civilization this whole time because it's needed us, then it has effectively served our interests most of the time. Sometimes people ask like oh, aren't corporations super intelligences? Why shouldn't we fear them? And other people pooh pooh that idea. I think that's actually a great question. And the answer is because it's made of people, so it needs us. So it's actually fine if a corporation gets really big and powerful because it can't help but also empower and take care of all the people involved along the way.
B
Yeah, but the mechanism for states becoming misaligned in the future would be that they are so interested in growth that they need the AIs more than they need people. But I just think there are many counter examples in history where states have sort of purposefully not gone after growth. You think of USSR or China or. Yeah, I think many other examples. So yeah. Does this mean that it is possible for states to sort of delay growth or not be interested in growth and be trying to pursue other values?
A
Maybe. I guess I'll say a lot of okay, so almost every state that hasn't pursued growth has just been either stagnant or become like starvation ridden. And I would say like, you know, the Great Leap Forward is maybe an example or eventually being conquered by their neighbors, like maybe pre Meiji Japan is a good example of aiming for stasis and then the growth oriented neighbor comes and takes over. So the other thing I'd say is it's actually, I think a very, it's just a very narrow target to hit to say I'm not interested in growth, but I'm also not going to be dominated by my neighbors or even shrink into the point where the people start starving. So to me, actually I just see the like. I mean it's. To me it's funny to use something like USSR or China as countries not being interested in growth because the only way China became this huge country is because there were a bunch of sort of like local polities that decided to take over. And this is like all the like Chinese civil wars. So any state worth even mentioning already has gone through a period of being relentlessly focused on growth.
B
Yeah, I see what you mean. What I meant there was just to say, say China or India or the USSR could have adapted sort of Western style capitalism and grown more than they did, but they chose not to because they were pursuing other values.
A
I don't think that's the case at all. I think, I mean if you look at the rhetoric at the beginning of the Cold War, people thought that communism was going to be more effective for growth and that they were going to bet on like the west was going to bet on capitalism because it was more compatible with freedom. And basically they were willing to bite the bullet and say we are going to grow more slowly but we will have a better society according to human values. I don't think that, I mean, and you might say, well, the USSR was doing the same thing. But I do think that the argument was this will make everyone richer in the long run or at least we're not going to take some huge hit.
B
Yeah, that also makes sense. How do you govern a state if you don't understand what's going on? If all of your sort of bureaucrats are AIs and you don't have the full information and perhaps you can't even understand what's happening.
A
Well, I guess I'll say to some extent that's what we already do. And basically humans are just very bad at running states. And I know that's not quite what you're asking, but I'm basically saying if you're needed for every bit of production as like a species, it kind of doesn't matter all that much how good your state is. Run like basically local, like think about like North Korea. Like the way that they addressed their famine in the 90s was partly just allowing black markets to operate. Right. It's like all they had to do was stop crushing local initiative and then that made people way richer. So I guess I'll say we don't. Right now the situation is so good that it's just like you're going to be fine as long as the Government doesn't constantly interfere to crush your local growth. And then as you say, once we have states where the machines are running the bureaucracy again, I don't actually, we talk about that in the paper as something that will be worrying, but I don't think it's the main effect by a lot. I think if every human was needed for some important factor of growth, I wouldn't care whether the bureaucracy was human or a machine. And I would probably prefer the machine bureaucracy. Likewise, if no humans were required for growth, I think we would be screwed whether or not we had a human bureaucracy or a machine bureaucracy. Think if you're a North Korean soldier and someone says, oh, Kim Jong Un, I forget the current leader has been replaced by a robot. You might be like, that's weird, but this robot's still going to need soldiers. So I'm fine. But now if you say, oh, no, okay, still human leader, but now they're building robot soldiers, that's when you're like, I'm screwed. My days are numbered.
B
Is there any way for us to sort of anchor. Say we assume we've solved the alignment problem? Right. Is there any way for us to specify what AI should be trying to maximize without them sort of missing the target?
A
Yeah. So basically this could all be addressed in principle if we had a giant singleton global government that was aligned and maybe that just looks like some aligned AI in charge of everything. And sort of the. I guess more of the big points we're trying to make in this casual disempowerment paper and also just that I was trying to make around lunch is we don't have a single thin. We have actually a lot of levels of competition operating above us, both between states and cultures. And it's so many different levels of competition that unless we control all of them, then any variation in how aligned these are to humans versus growth is going to be dominated by the ones that are more aligned to growth. And so the only way forward that I can see is some sort of like global permanent singleton that crushes all innovation and competition forever, which sounds extremely dangerous and terrible. And I don't think we know how to do that technically. But I guess my claim is that if you only do this halfway and you say, okay, we'll have a global government making sure everyone still has a job, but we still allow cultural competition and cultural evolution, for instance, that eventually there would be some new machine and growth focused religion that sort of took over and then redefined what it meant for a human to have a job such that it was actually just like, you know, basically machines that were much more productive being counted and optimized for. Similar to like the Roman Empire. Right. Like Christianity just growing up inside of it. And it wasn't like this, like war. It wasn't like there was like some failure of policy to control this new cultural revolution. It was just a completely orthogonal arena of competition that ended up kind of taking over all the institutions.
B
Yeah, interesting. So growth historically has been good for us. Growth and increases in living standards have kind of coincided. And should we expect that to change, or do you disagree that that's the case?
A
Yeah. So this is another great question. People say, well, you just sound like a Luddite who's worried about the Industrial Revolution. And I would say, yes, that again, is like moderate evidence against my position. And first of all, it was actually very destructive to a lot of people and ways of life and areas of the world. But you could say, fine, but that's just the transition. That's just growing pains. And I would basically agree, I would say to the extent that you did feel like those old ways of life or culture or villages or whatever were actually important things, then they did lose and it was maybe not worth it. But yeah, the basic thing that's changed is humans were needed for growth before the Industrial revolution and after. So this phase change in growth ultimately, like, you know, the rising tide lifted all boats. It did not make, yeah, like horse welfare, obviously better off or, you know, think of anything that did not get to. Was not necessary for growth after the Industrial Revolution. It was just the interest of those beings was just not respected by this revolution.
B
So the way you sketched it out just a minute ago, seems like we have a very difficult task ahead of us. If the only way to get through this is sort of a world government controlled by an aligned AI, would that even be enough? If we imagine that there might be other sort of cultures out there in the universe? I'm thinking about, like. Like avoiding competition is difficult. I think it's. Is that even a stage that a state of affairs that can be reached?
A
Well, yeah, so I. I mean, I don't. Yeah, we might not. We might lose to whatever aliens we meet. And I don't know if I don't have much to say about that. I guess I feel like this. The size of the pie that we will get before we meet aliens is probably like billion light years or something. And I'm just like. I'm. My capacity for joy maxes out at like 100 light years or something. Like, I just I'm fine with that. But your question had another part.
B
Yeah. Can we avoid competition? Like, is that a plausible state of affairs?
A
Well, exactly like, I mean, maybe one related question is like, can we avoid cancer? Right? It's like just because of chaos and coordination costs, it's not clear that you can ever have like really lock down a planet or a, or something such that there's not going to be little bits of local competition and collusion or implicit forms of corruption happening locally that ends up favoring growth of some thing more than another. So it's not actually clear if it's physically possible to control things enough to avoid competition. I mean, and also it sounds horrible, right? And this is like recipe for permanent dystopia by anyone's measure, I guess. So Robin Hanson is someone who's been spending a lot of time thinking about these very big questions about just like cultural decay and competition. And his answer is basically like, by default we will not adapt enough and we will all try to preserve what we like, the thing that we like about our current civilization or our current setting, and ultimately be out competed by some less organized, sort of like more cancerous, evolving, competitive, freewheeling culture or civilization. So he basically says we need to be super disciplined about this and cut our set of values that we want to preserve down to some very minimal, just one or two things. And I think his example he always wants to preserve is free inquiry and truth seeking. And then the only way that survives is if that's just like a tiny drag on growth, like maybe like 1 or 2%. But you attach it to this otherwise completely adaptable civilization that's willing to throw any value under the bus to preserve growth, to preserve this like one tiny little piece of the constitution or whatever. And I mean, I think he might be right empirically. And then I guess I'm still not clear. To me it's like, how do you go 99% of the way and not all the way towards adaptation, Right? Because I feel like if we say like, okay, we're going to, we're allowed to, anything goes. As long as we preserve free inquiry, there's still going to be a matter of degree and there's going to be these calls of oh, but we want to just lie this one time so that we'll have fusion or something like that. I don't know if that works that way, but it seems like it's unstable and you'll either just in general tend towards preserve everything and become ossified, or preserve nothing and remain competitive. I Don't know, it seems like one of these sort of chaotic equilibria where you're constantly sending off little offshoots of more ossification. That helps in the short term, but then it fails in the long run. So maybe by default we just get this chaotic thing that never settles down. I'm not sure.
B
And also because this path would require us to give up on the very things we're trying to preserve if we want to preserve our culture and we need to do that by giving up on most of it, it's sort of self undermining. But again, that has to be the case because you're facing competitive pressures. How do you think about this? So how do you think about how much we should be trying to preserve? Because you could go all in. You could say we never want to change anything about our culture again, sort of Amish style. But that doesn't seem like what we actually want. We want a culture that evolves, but in a way where it's still sort of connected to some basic values or what, some tradition that we're part of.
A
Yeah. And I mean we face this in our own lives. Like my dad wanted me to take over the family farm, but here I'm in the big city doing crazy job that's not farming, even though I also love the family farm. So I'll give sort of a boring scientist answer of what's possible. And so I think one of the big tasks and open questions is just trying to understand what is the tax, what is the alignment tax? If I do, how much does my influence fall off? How quickly does it fall off as a function of how much my non competitive values I try to preserve? And we already see this with orthodox and reform and different degrees of strictness of religions and they sort of compete and sometimes the more strict religion actually attracts more people. There's also a possibility that there's some crazy phase change where we're about to have so much growth that we don't have to choose for the most part. And we can say, everyone go to the big city and make the machine God. And then we're going to simulate 10 million years of Amish paradise forever. And we don't have to choose. We actually just need to grow. And then on the last day or whatever, like when we've eaten the last planet or whatever, build it all into our perfect value expressing substrate and just live that forever. And of course that only I think works though if you have this global coordination that stops local drift towards growth.
B
On a personal level, how do you feel about this? What are you trying to preserve of your own culture?
A
Yeah, it's pretty rough. I guess I'll. I've thought about this a lot. I mean, obviously having kids gave me a lot of very concrete sort of medium term desires of like I want to be there to make sure that this, these milestones happen in their lives. But it's like when I think about this future civilization, civilization that doesn't need them, it starts, it takes a lot of the fun out of it because it's like, well, you know, I can teach them to drive, but then maybe they won't ever actually need to drive or like teach them to shoot, but they won't need to hunt and teach them to like do jobs that they won't ever need to do. It's like it starts to become pretty rough. I mean human values are just very complicated and I don't think I have any like special like way of life or like vision for like here's my compound that we're going to live this way and it's going to be amazing. I do think that that's actually a very valuable activity that people can do. Can and should be doing is saying like laying out in more detail like visions for the future that contain different aspects of what we value and get to be expressed even under weird like cyborgism and stuff like that.
B
Yeah. One route there is just kind of rejecting becoming post biological or rejecting transhumanism. So sort of having this perhaps irrational attachment to being a biological creature is one way to prevent yourself from becoming overtaken by AI interests or having your interests changed. But do you think that again, is that unstable?
A
It's so unstable just because it's going to be a really easy and cheap at some point to just gradually turn yourself into a cyborg and it's going to be also medically necessary for some people as you get old. It's like, oh yeah, I don't have my biological heart anymore or whatever and my new machine heart lets me climb 100 flights of stairs. Actually if I was going to start a new religion or something, maybe one of the founding texts would be this amazing book called Two Arms and a Head. It's a spoiler. It's actually a suicide note by a 30 year old guy who broke his back in a motorcycle accident. And he just writes about how amazing it was to when he was embodied to like run and like, you know, lift up a kid or like flirt with girls. And there's very long parts where he just lists a lot of little tiny things that were really cool about being an embodied human and that he doesn't have anymore. And it does a really good job of life is made up of all the little things. And he kind of has a big chunks of this note where he's enumerating, Here are the 500 little things that I've lost. And you're like, oh my God, life is so amazing in this physical body. I really love it. And it took this one guy changing his perspective to. To make me see it. So anyway, there's no way around it, right? I think anything that looks like preserving something that looks like human life is going to have to be this long list of a million little things that are nice to have in your life. And can we come up with a sort of excuses to still have them happen even though we don't need them to happen anymore?
B
I see your recent work as trying to sort of kickstart this field of post AGI studies. Or, yeah, thinking about society in a post AGI world. Why don't you tell us a bit about the recent conference and sort of give us an overview of this field and what's happening the different factions.
A
Yeah, yeah. So just two weeks ago we ran the second iteration of our workshop. This time it was called Post AGI Economics, Culture and Governance. And it was 150 people. And we tried really hard to make it interdisciplinary. Like it was right after the Far AI Safety workshop, which was really cool, but it's sort of like there's this big crowd of sort of usual suspects, AI Safety people. And we invited lots of those people. We tried hard to get economists and more cultural theorists. We didn't manage to get any historians, all sorts of interesting people who would have different takes on these big picture questions. And it's really hard to find people who are open minded enough to realize that things might actually change and like not just sort of pattern match to some cached answer, but also not go crazy and either say everything's going to be fine no matter what, or, you know, just still having good epistemic standards.
B
This is actually just to interrupt here, but this is actually, it's part of our culture that's the problem here because whenever we try to speak about some of the topics that we're speaking about right now, it just sounds weird and it kind of scares off in some sense serious people where, you know, what does it even mean to talk about the topics we're talking about? It's not really connected to any literature you can cite, maybe. And so, yeah, I guess that's one of the problems of trying to create this new field of studying post AGI society.
A
Absolutely, absolutely. And so we were very deliberate in trying to engineer a sort of like legitimate, as grounded as possible respectable venue here because there's been a lot of futurism conferences and things like that. And I think that they did an okay job of this, but they were not trying not to be weird. So our first keynote was Anton Kornek and we were delighted to have this sort of serious academic economist who was also taking these ideas seriously and we were trying to lend legitimacy to the field. Also on an object level there's this idea of the singularity which I feel like has been very destructive because it kind of is like an excuse to turn off your brain and to not model the future and to say, yes, things will keep changing faster and faster until we can't say anything about it. And I'm saying, well, no, I mean, yes, obviously things can change in ways we can't imagine. But let's not give up. Let's think about. There are probably still a bunch of physical limitations to growth and communication coordination. We expect things to be kind of agentic in the long run. There are some things we can say about what the main forces acting will probably be. And so we should reject this idea that there's the day after which we can't forecast things. Let's just take all of our tools and just push them as far as we can and see where they break down. But don't just throw up our hands ahead of time.
B
Why do you think past futurism has been so kind of bad at predicting the future? Is it just because they didn't have these theoretical tools or. Yeah, why is that?
A
Oh, I guess I'd say, okay, if I was Robin Hanson, I'd say it's because they're thinking in far mode and that they maybe correctly take the prompt as a chance to express their values and say, oh, in the future we're going to have radical abundance. Because they take it as an excuse to say this is what I hope. And I want to make it clear that if I had power I would try to give everyone radical abundance or whatever. I mean, I don't like to psychologize people. I will say I think there's been a lot of really great futurism over the years and science fiction that actually is very hard and takes these ideas of recursive self improvement seriously and still tries to have. I guess part of the reason is a lot of the stories become very uninteresting. So Vernor vengeance stuff always has to take place in. There has to be this MacGuffin where, oh, there's no AIs in this part of the universe because of some bullshit or whatever. Otherwise it's just like so alien and hard to even write interesting stories about and then again to psychologize people unfairly. But I'm thinking of like Greg Egan is someone who I have a ton of respect for as a writer and then he has just, I think, been poo pooing AI in the very like Gary Marcus kind of like, oh, this is bullshit, it's not going to work. And even if it did, it wouldn't change anything sort of thing when he's actually written good stories that take these premises seriously. So I'm actually like shocked and dismayed at or like, you know, Ted Chang also is another person on layman shame is someone who's like, clearly has the mental horsepower and imagination to think seriously about these ideas and is just choosing not to for one reason or another.
B
It's actually, it's an interesting thing that for sci fi about the future to be plausible and sort of relatable to us as people, you need to exclude the possibility of radical superintelligence. I guess it sounds a bit like an omen if that's the case.
A
Yeah, I mean, I guess maybe. Let me say it loud and clear here is that, yeah, I think that the post AGI world is just going to be extremely alien and so different that if we could avoid crossing that threshold, I think we should. I'm willing to give up tech progress even at great personal costs to avoid basically rolling the dice with like whatever crazy post AGI world we're currently expecting. And I think no one really has a good handle on what that would look like. Exactly. But yeah, the drama we've been trying to beat is like, it won't need us and that's just going to be just a much fundamentally worse position than we've ever been in historically and raises the stakes of every type of governance.
B
Do we have that option? Do we have the option of not building these advanced AIs just because, as we've been talking about, we have competitors, competitive pressures. And so wouldn't it be the case that at some point, some other company, some other state, maybe 100 years in the future, it gets built anyways?
A
Well, if we could build by 100 years, I would be overjoyed. But I mean, the basic answer is, I guess one technical question that we have to Ask is what does it mean for us, we to do something you always have to frame? Okay, if I publish this newsletter or whatever, am I going to be able to steer the course of history? But I think basically the short answer is no. I'm a little worried that I've been being cowardly here and not saying like just don't build AGI clearly enough because I do think it also destroys credibility. It's just so easy to be a Luddite that I think it makes people take you not as seriously as if you have some nuanced, positive sounding view. I guess I'll say my view is nuanced but also doomy still. I mean, I guess, yeah, my modal outcome is that there's going to be a bunch of fast growth and beings that are loosely human inspired but mostly optimize all the human parts away that basically spend most of their cycles on growth and there's going to be lots of fun coordination and Von Nomin probes high fiving each other as they disassemble the solar system or whatever. And I find that somewhat valuable. But I think it's going to happen no matter what. And I really would hate for my sort of family and friends and culture to be basically out competed and destroyed or at least like marginalized so much that they're, you know, we're run like you know, a little bit on the side in the giant essence sphere or whatever. So I don't have a very clear vision of the future. But I guess I just think that the default is we just end up so marginalized that we're at the mercy of like whatever larger forces are actually doubling down and participating on in growth.
B
Can you flesh out this sort of cultural reaction to being perceived as a Luddite versus having smart and sophisticated and sort of positive take.
A
Yeah, yeah, yeah. So I guess I'll say one thing that bothers me is I feel like people who work for the big labs can't be too alarmists. Although like if you read like even just like you know, machines of Loving grace or some stuff Simon Almond says like they're being about as honest as they can that they're like this is going to change everything and probably make the economy not work for people and we don't have a plan for this. And also that's just sort of the least of our problems. But I'm thinking of more like the mid level people. They just have to come out with something that sounds vaguely positive. So they end up with these fairly sophisticated sounding takes about how Competitive advantage is going to save us or something like that. And it's like, yes, this is definitely something that's going to help, but it's probably not going to, like, it's not clear that this helps enough. And so basically what I'm saying is that there's going to be this positivity bias from thinking coming from the big labs. There's also a bunch of people who haven't put much thought into it, but they just have the intuition that this is very scary and new and different and bad. And I think they're basically right for the right reasons. But a third party can accurately say, you just don't know that much about this. You haven't thought about it very much. You're just a Luddite. And now you shouldn't take my word for it. But it's like, okay, I've thought about this for like a couple years now, and I think my position ends up being like, similar to that of the Luddites. But, you know, unless you listen to me for an hour or something, maybe it's hard for you to tell if I put much thought into this.
B
How do you think we're doing on forecasting? AI so sort of on the technical side.
A
Well, actually, I want to go back to the question you asked about do we have the ability to not build AGI? I guess I'll say I do think that there's my success story for how this all turns out, okay, is that while we're building better AGI tools and just improving technology in the normal way across the board, we improve our ability to forecast and coordinate and basically govern ourselves in a way that everyone has sort of always wanted to. But it's been hard and it hasn't really been that important because again, like I said, if the government was bad before, it's sort of like, fine, you just get taxed more or less, there's like a war or something and that just the timing might work out such that we get a better ability to all roughly agree on what's going to happen by default, what policies are available to us and be able to coordinate, to say, oh, we're all going to do X, like, you know, not build this technology or not deploy it in this way or whatever. While, like before, humans have basically been marginalized. So I think these things, like, by default are both going to happen. That something is going to have much better forecasting and coordination ability and humans are going to be marginalized. And it's just a question of can we do the second thing before the first thing. And so that's also why I haven't been super energetic about saying just stop build AGI or just don't build AGI. I think it's a step in the right direction. But it's kind of like there's a stampede happening or whatever and you're like hey everyone, stop stampeding. It's like, well look you have. That's not going to help, I mean. Or like it's like a step in the direction. If everyone did that we would have solved the stampede. But really you have to figure out some more clever thing where you have a sign that turns people in this way and helps them coordinate or put a big mirror so they can see that they're all stamped off a cliff or whatever.
B
What does that look like? Is that an international treaty?
A
A treaty. It's rough, right? I think a lot of the infrastructure people tried to build it, I guess for global warming. Of course then it became co opted by all the bad, let's just say the normal competitive forces that make the institutions that are about solving an issue end up not being about solving that issue. And I think people are rightly afraid of that same thing happening with AI and it end up just being another reason to increase state surveillance without actually stopping building AGI for instance. I think that's a real fear of mine. So maybe, I guess I'll say the stuff that I'm excited about lately has been superforecasting and trying to. So super forecasting already has been a major gift to the world and I'm really excited about this whole direction in general. I think that that community has a bit of egg on their faces in particular for dismissing AI as a nothing burger. And I don't want to be too hard on them because it's not a unified community and a lot of people have made correct calls. But the other problem is that it's not clear that the general public or decision makers should trust these super forecasters on scale of 5 or 10 years and on very big open ended questions that are sort of about the entire course of civilization and not just is this war going to happen by the state or something like that. So my attempted answer to this is a side project I've been working on a little bit is trying to build a historical series of data sets where we can train LLMs up to all the sort of state of world knowledge up to a certain date, like 1920, 1930, 1940, and then build a set of questions like in 1920 we can ask what would a 1930 historian say is the biggest thing that we're missing or the policy we wish we had implemented or some open ended thing of like what's the important thing we should be worried about or orienting towards or even just what's the headline going to be in 1930? We can in principle actually train models and run this back test through the last hundred years or so and then have a baseline for ask that same model with the same scaffolding trained on 2025 data. What's going to the world's going to look like in 2030 or 2035 will at least be able to point to this track record and say oh, at least in terms of economy or wars or some aspects maybe. The models were pretty robust and here's what they're saying about the future. Surely we can all agree that this is a good starting point for the conversation. Something like that. This is about as far as I can see forward in improving our society's ability to act more agentically.
B
This would be training a new model from scratch on data only up until say 1920 or 1930, because otherwise you would have a bunch of sort of data around the future that you don't want in there. But there isn't that much training there, is there? Kind of. If you only go up to 1920, it's really rough. Yeah.
A
So the models will get worse the further back in time you go. It's also not clear that you'll be able to make them nearly as smart as the big labs can make a model in 2025. I mean obviously we can bring some of the tricks back in the future and do like reinforcement learning with verifiable rewards and have them still do math or even programming. It doesn't really leak anything from the future if you have a model do programming maybe with the lambda calculus or whatever. So that is a major problem. And I think in parallel we should be trying to do some sort of forgetting techniques where we try to take a really a baseline should always be take the best model you have and try to get it to the task. And I think people will rightly just say the model didn't really forget about World War II or whatever. But that is definitely another strategy we should try in parallel.
B
Which strategy do you have most hope in at the moment?
A
Well, I guess it's like a sandwich. Right. So we're going to have the crappy historical models without much leakage and they will be just making. They'll be dumb and they'll be making very vague predictions. Like, oh, probably there's going to be another war. Or like, so we have like a very basic version of this and it predicts that like Henry in 1930, it predicts that Henry Ford in the 50s is going to make a transatlantic airline which is like, that's kind of. That did happen, right? Like Virgin Atlantic sort of thing. Right? Like, you know, sensible kind of things that like, okay, anyone would kind of guess would happen and then. But we're going to know that that's kind of like an underestimate of how well we can probably predict because the models are kind of dumb. And then we'll have the fully fledged models that have been beaten into forgetting some of their knowledge, which will make two good predictions. They'll obviously have an idea that, I don't know, the Berlin Wall was going to be a thing and it was going to be important or something. We can't really do perfect forgetting. And so those will be too good. And so we'll just have this sandwich. We'll be like, okay, we know it's better than this and we know it's worse than that. Somewhere in between.
B
Could you train on data up until 2010 or 2015 or something and then get a better model but then have it predict only 10 years?
A
Yeah, exactly. I mean we should be trying all these combinations and it's not even clear to me what the most important do we care about five or 10 years ahead for the questions that I think we should care about, I think that's roughly right. But I think there's going to be all sorts of uses for these. So there's a bunch of different people who are independently proposed proposing that we do this. And I think this is just going to be a major activity going forward and I'm really happy about that.
B
Do you think if the world is sort of speeding up, the pace of change is speeding up because of AI, do you think forecasting just becomes inherently harder, do you think?
A
Oh, absolutely, yeah. So I do think there's going to have to be some sort of temporal speed up or some speed up factor or something like that. And maybe it'll be different in different domains, but that's exactly the sort of thing that we can hopefully get a handle on by looking at the last hundred years and sort of trying to stay like, oh yeah, one year in 1980 is like worth, I don't know, four years in 1920 or something like that.
B
Interesting, interesting. As a final topic here, I would love for you to chat about what information or knowledge do we need the Most in this field of post AGI studies, you could call it. What are you most excited about? What would you love to see people working on?
A
So one very basic thing is actually inspecting human values. And this is kind of funny and you might say backwards, but I hear a lot of people talking about machine consciousness and moral patienthood and stuff like that. And the basic way forward I think is they want to investigate the machines and say what is actually going on inside of them. Do they have this or that ability? And obviously that has to be part of the picture. I'm not denying that. But I guess to me it's. I'm a moral anti realist. And so to me it's like what makes me care about another being or their welfare is like a pretty complicated function of like stuff that's in my mind and also like what I learned about reality. So I think we should be more systematically trying to understand what the kludges are. And it's kind of like if we wanted to understand like what makes food taste good, like you can spend a lot of time learning about the chemistry of the food, but at some point you have to do a lot of taste tests and maybe understand the pathways, understand this molecule actually tastes sweet because it's close to this, other molecules, stuff like that. There's no arguing taste. But I do think more systematic. I don't know, it's almost like surveys or study of human values around other conscious beings will just help us answer the other half of the question, which is like, well, what do we even care about before we even look at the AIs? What would we care about? We don't actually have to look at them to tell ahead of time.
B
Is there anything new you need to do there? As opposed to sort of reading everything humanity has ever written and then extracting our revealed preferences or what we write down that we care about from the entire Internet.
A
I mean I do think in principle that there's probably enough information already out there. I guess I would.
B
I'm interested in what new information would be most valuable for you here.
A
I think honestly thinking about this through the lens of game theory. So maybe one thing is people keep talking about consciousness and I'm like, it seems like you actually care about other agents that are powerful and are not cooperate bots or something like that. And I don't want to round things off and people's moral sentiments are complicated, but sometimes I get the impression that people are just reasoning backwards from if this thing could basically be a game theory, a competent game theory partner against me, then I care about it. And that would make total sense also evolutionarily, because it's like, yeah, you don't want to care about the cooperate bots, you don't want to care about the defect bots, you want to care about the tit for tat bots or whatever. I'm not sure that that's exactly it, but I kind of suspect that there are some. Maybe something along those lines where there's some theories that would fit enough stuff well enough that you'd be kind of be like, oh, this is sort of underlying what's going on. Again, there's no arguing taste. If you still just say like, no, I care about sentience or whatever, then that's fine. But I do think that there's probably some more underlying framing that is more illuminating than the language people have been using so far.
B
And having this new theory of human values. This helps us how?
A
Well, it helps us articulate positive visions for the good. So I guess I'll say my claim is that if you have a really good understanding of human values, you will also be able to write really amazing manifestos. Like I was saying, that guy wrote his manifesto about how amazing it is to have a human body. I think that if I had a really great understanding of what people valued in terms of conscious beings or even just what makes the society good, that I would be able to write some amazing vision of just a day in the life of some future society and everyone would be like, oh my God, this guy gets it. And like, I would love to live that world and it would be so valuable. And there definitely are some people who can do this better or worse than others already, but I feel like it hasn't been tried to be made into a deliberate part yet.
B
Yeah, that's actually really interesting. And is that. Do you think this is a dangerous question? Perhaps. But could this be automated? Is this something AI could become very good at?
A
Yep, yep, yep. So machines will also be able to help us or do this better than us at some point. Um, it's just. Yeah, it's one of these things where I feel like you can sort of. Well, it's. It's dangerous because you can sneak in different types of values in, in these manifestos very easily. But it's also, I guess I'm claiming to some extent people can recognize good work in this area when it's done well. So I don't know, I'm not sure who I would trust more to do this for me, a machine or a random person. It's hard to say.
B
Any other things you want to mention here that you would like to see?
A
I guess I'll say so. Me and Jan and Raymond have been thinking a bit more about the boundaries of personhood. And I think there's a huge design space here that's kind of unexplored of, let's say social groupings. And we already have families and nations and religions and lineages and ethnicities or sports teams. There's just our world is already very super crowded with types of organizations. And I guess I'll say, I think we actually still have a huge design space that hasn't been explored here. And once we have AIs that can make copies and all the different freedoms that they have, they're going to have an even larger design space. And then also once we have the ability to make soft copies of humans, then there's also going to be all sorts of weird and wonderful new types of arranging personalities and loyalties that is just completely untouched. So science fiction writers of the world, please come invent the new societal units that we might like to inhabit in the future.
B
Any concrete guesses on what those sort of groupings might be?
A
Oh, like just for instance, in your day to day life you might end up having a lot of personalities that like you could have someone simulating your like, I don't know, dead grandpa telling you like what your ancestors would have wanted or I've also been doing a little bit more work on people's relations to LLMs. And like I think a lot of people really want to be told what to do is maybe something that I think is like showing up and you know, like we don't want the AI's brow beating people, but a lot of people are basically submissives or they want someone else to have a plan. So I kind of think that for a lot of people the best way to interact with LLMs might be some sort of good cop, bad cop sort of situation where there's the mean LLM that tells you to clean your room and go get a job or whatever. And then there's the supportive one that's like, oh no, we don't want to make the mean one mad, so let's clean our room. Basically right now we have this therapists, a few other sort of very consent heavy peer to peer kind of relationships and I think those work for a lot of people. I think there's going to be a lot of people for which they would prefer something. Maybe looks more like being a private in an army or some person on the quest and they don't even know what the quest is for and of course this is very dangerous right you can definitely it's so easy to make cults basically but I guess I feel like the way that LLMs relate to people for a lot of them isn't going to look like this all in one confidant and advisor sort of thing.
B
For people interested in working on these ideas what should they do? Where should they apply?
A
Oh yeah I haven't taken any PhD students lately because I'm in a computer science department and it's like I don't want someone to come and take databases and stuff so I guess I'll just say I don't know we have a discord that we made after the workshop where we're trying to get all the like minded people together show up less wrong honestly like that's an amazing community right now and there's just lots of people thinking along these lines Fantastic David.
B
Thanks for chatting with me My pleasure.
Podcast: Future of Life Institute Podcast
Episode Date: December 23, 2025
Guest: David Duvenaud (Associate Professor, University of Toronto, AI Safety Researcher)
Host: Gus (FLI)
This episode explores the theory of "gradual disempowerment"—the possibility that humanity could lose meaningful influence and agency in a post-AGI (Artificial General Intelligence) world, not through a dramatic AI takeover but via economic, cultural, and institutional shifts. David Duvenaud unpacks how advanced AI, even if aligned, could sideline human interests, why growth maximization is a crucial but double-edged driver of civilization, and what kind of future societies might emerge. The conversation also touches on the challenges of preserving human values and agency, forecasting the future, and why existing societal safeguards may prove inadequate.
David Duvenaud emphasizes the need for:
Ways to Get Involved:
Memorable Closing Quote:
"Let me say it loud and clear here is that, yeah, I think that the post AGI world is just going to be extremely alien and so different that if we could avoid crossing that threshold, I think we should. I'm willing to give up tech progress even at great personal costs..." (58:18, David)