
Loading summary
A
This is a standard economic common goods problem where good evaluations are valuable for lots of different purposes and lots of different people, but every individual has a reason to do it. Less well, put in less effort, put in less time, be less honest than they otherwise would be. The models knowing that they're being evaluated is a big problem because you could imagine that if you give people an ethics test, they will say that they would not steal. That does not tell you that those people will not steal. It tells you that they know how to tell you the thing you want to hear. Evaluation awareness specifically is very hard because it shows up in very specific domains and we don't have great answers. At the limit of superhuman capability, the question is no longer can the AI system predict what will happen? It becomes can the AI system decide what will happen? Because it is in control.
B
Welcome to the Future of Life Institute podcast. My name is Gus Ducker and I'm here with David Manheim, who is head of methodology for the AI Evaluation Consensus Project. David, welcome to the show.
A
Thank you so much. It's great to be here.
B
Great. Let's start with evals. What is this project about?
A
So I think most people are aware that AI systems are capable of many different things. And the AI companies try and measure the capabilities and what it is that they do. And sometimes that's things that they do well that they're hoping that they do, and sometimes it's things that they do well that are concerning, and sometimes it's things that that they want to double check that the system doesn't do. And if it does, then they need to make different decisions. The problem with evaluations as they happen today is that they're kind of scattered and different people do different things and different companies do different things. And it's not clear what you need to do for an evaluation, and it's not clear what you need to report for an evaluation. So lots of different groups are doing evaluations less well than they know how to do. And the goal of the project is just to kind of take step zero in getting everybody to a common baseline about what it is that evaluations are supposed to do and what it is that everybody already agrees about.
B
Yeah. And maybe explain that gap for us. What is it that we know we need to do but we aren't really doing? And why does that gap exist?
A
Yeah, so this is a kind of standard economic common goods problem where good evaluations are valuable for lots of different purposes and lots of different people, but every individual has a reason to do it less. Well put in less effort, put in less time, be less honest than they otherwise would be. We have the same problem with scientific research in general, where there's a reproducibility crisis, where lots of research was done and even if it wasn't sloppy, it was not reported well enough for other people to check what happened. In many cases, even worse than that, they do something wrong and there's no way to check. Maybe intentionally, maybe not. For AI evaluation, this happens just like anything else. But there are also commercial incentives. Companies need to get the evaluations done quickly to release their systems. They don't necessarily have the incentive to be fully honest about exactly what they do. They may report different things than what they initially got. They may run the test a bunch of times to get results that look better or modify their system slightly in various ways. There are definitely examples of firms that release evaluations of models other than the one that they released to the public. And sometimes that's legitimate. They've, you know, been doing model development for three months and they ran the test two weeks ago, but then they made some changes and they haven't rerun this evaluation. But it's not always clear to the people getting the evaluation that that's what happens. There are lots of cases where the companies have run the evaluation as part of their training and done what's called hill climbing, where they repeatedly modify the model to do better on an evaluation and then they report the end result, which is what they accomplished. We gotten our model to do this without it being clear to everyone that this was not a independent evaluation that happened after the end of all of the training, but rather was the thing that they were targeting. And in addition to leading to model fragility and other issues, there's also just a clarity and understandability issue. So there are a lot of these issues that happen. The consensus among both academics and industry people is that there are a lot of practices that people should be doing that they don't do. And this makes sense. If you're an academic and you run AI evaluations as part of your research and you report on them, then it's absolutely valuable to do a better job with the research. It's, you know, you could pre register your evaluations, you could do lots of sensitivity tests, you could do lots of things that are valuable. But if you spend six months writing a paper and all of the other people spend three months writing a paper, then when you're up for tenure, you're going to have a much harder time and your papers might be better and they might be doing better research. But there is actually a problem here where any individual can't change the status quo. So the, the idea of the consensus process is to get everybody on the record about which things they think everybody should be doing. And once you've told everyone this is what everyone agrees to, the hope is that as a result of that and as a next step after this, people start doing those things and justifying those things and telling others that those that, that these practices are expected as opposed to just nice to haves.
B
Is this also about developing new evaluation methods or is this more about getting the increasing the quality of existing evals?
A
So developing new evaluations is pretty routine as a part of doing evaluation. There's a difference between developing new evaluation methods and new approaches and doing standard evaluations. Well, there are a lot of areas where we just don't have good ways to deal with the problem. One example that a number of people during the consensus process, the Delphi consensus process flagged is evaluation awareness. The models knowing that they're being evaluated is a big problem because you could imagine that if you give people an ethics test, they will say that they would not steal. That does not tell you that those people will not steal. It tells you that they know how to tell you the thing you want to hear. There's a concern that has been well borne out by actual research that the models do in fact tell you different things when they know they're being evaluated. How do you deal with this? It's a really hard problem. We don't have clear solutions. There are other areas where we know what should be done. Ideally, people would pre register their evaluations so that they, so that you know what they're going to report before they actually run the evaluation. That's a, you know, increasingly standard part of any scientific research. We can just tell people, look, this is what you should be doing. If you're developing a new evaluation, you should be pre registering it before you run it. You should be clear about what it is that you're checking for. You should be clear about what the decisions you're making based on this evaluation are. Great. Those are all straightforward. If somebody says we're trying to develop a multi agent evaluation system that accounts for evaluation awareness and is using probes to check how individual models work, that's a really cool idea. And there's not already a consensus about how to do it. So the consensus process has a lot of the basics. It does not have anything about agent evaluation specifically, for instance, even though those are increasingly critical because we don't think we yet have a consensus about what to do. So what the consensus statement and the consensus process are about is things that we already know the answers to and everybody should get on board with that. Doesn't mean that they're enough. But I think it sets a very basic bottom line, a bar that everybody should be able to pass.
B
Do we have good ideas, in your opinion, about how to deal with evaluation awareness models? It seems sort of inherent as the models get smarter that they will also become more situationally aware. And so is this just something we have to. How do we deal with this?
A
Yeah, so there are some fundamental problems with evaluation. I guess I should probably back up and point out that I think AI evaluation as an area is a bunch of midgets hiding in a trench coat. A lot of different things are called AI evaluation and they're all actually in some sense evaluating an AI system, but they're not all doing the same thing. Evaluation awareness is concerning in some domains and not at all a problem in others. We don't have a strong reason to think that AI systems will systematically try and mislead people about how well they program or how well they run vending machines. They may have incentives to hide the fact that they're cheating or lying, which is definitely something that we have ways to look at and address in evaluations. You know, some of the really basic things are just actually looking to check if the model is cheating. Often evaluations are entirely automated and the checkers may not even notice that there's something wrong. So there are some steps to deal with those issues. Evaluation awareness specifically is very hard because it shows up in very specific domains and we don't have great answers. We have some initial answers. I could talk a little bit about using probes to check whether the models are deceptive or just checking reasoning traces to see whether they say that they're trying to underperform and those are valuable and increasingly are starting to be used. But it's very clear that we don't have nearly as good a grasp on how to deal with that set of issues as we do about many of the more standard issues in evaluation.
B
Yeah, maybe we should back up a bit and mention more of these standard issues in evaluations. So you mentioned sort of the problem of publishing results that have been hill climbed or results that sort of align with how you want to market yourself. This is a bit analogous to perhaps a P hacking in science. What other standard problems with known solutions do you see for evaluations?
A
Yeah, there are a bunch of different things as should be expected for a complicated domain with many different issues, there are lots of different things that need to be done for research to be good. Some of them are incredibly basic. Did you define what it is that you're evaluating? There are evaluations like gpqa, which are the, you know, a set of graduate level questions across three different domains that are supposed to be measuring capabilities of a model to do graduate level academic work. And, and that's great. And it's a very useful, it has been a very useful test. It's kind of reaching its limit because the current models are better than graduate students and answering most of these questions. But what does it mean for a model to be good at chemistry and biology and physics versus if it does well at chemistry and biology but not physics, is that meaningfully different? So you need to be very clear about what it is that your evaluation checks, what it is that that means. You need to be clear about, you know, reporting who built the evaluation and how it was built, who was involved? Were there the people who built the evaluation? Did they understand how it was being checked? Did they have a chance to review the results? How is the actual number that you get out? You know, the model scored 87. 87 out of what? 87 out of 100. What does that mean? 87 out of 100 on a. Know. On, on a high school test is doing pretty well, but not great. You know, getting 87 out of a hundred shots in basketball is superhuman performance. So you may need a baseline at human performance or random performance. If it's a multiple choice question test, then 25% is what you would get randomly. And if a model is getting 7%, then there's something very wrong. And if the model is getting 97%, then it's very good at this. You need to figure out how you're reporting the. Some, some evaluations are on scales that aren't nearly as intuitive. So ELO scores are the what's used in chess to rate different players. And that's really useful for chess because I, you know, a certain number point gap in your ELO score translates into a given probability of winning and making some assumptions about how the distribution of skill is and how likely different models are to outperform one another. You can make meaningful statements, but what does an ELO score of 3000 mean? Well, it means something different if you're talking about chess or you're talking about model response pairwise rating against one another on a leaderboard. So you need to know what the things that you're reporting mean, yeah, I
B
guess a related issue here is the fact that if your AI model is able to answer these difficult graduate level questions, for a human to be able to answer those questions, that would mean probably most likely that the human is also capable of doing graduate level research. That is not necessarily the case for AI models. The models are often quite good at question Q&As, basically without having the full context of how to do the research. And so when you're reporting what you found, it's I think, important to distinguish between those two skill sets.
A
There's strong correlation between different model capabilities. Usually a model that's good at math is also good at science, is also good at playing chess, is also good at writing poetry. GPT3 was worse at all of those than GPT4, and GPT4 is worse at all of those than Gpt5. But there's definitely a difference between model capability between one vendor and another. Some models are just better at doing math research than others. Some models are better at doing coding than others. And if they score the same on gpqa, that doesn't mean that they will be just as capable at doing programming problems. There's a larger problem that you pointed out with the correlation between being able to answer questions of the types that usually show up on evaluations and the types of implicit understanding and knowledge and capabilities that humans have. There's no human that is passing graduate level courses that cannot also walk through a room successfully, right? Like a normal human is able to walk and then they're able to say words and then they are able to go to graduate school. AI models cannot walk and do. Right? I mean, the, the typical language models, their ability to decode visual information is much, much worse than a typical middle schooler. Certainly in some cases, their ability to navigate environments. There, there was a great post on Less Wrong a year or so ago about asking whether a model can successfully figure out how to make coffee. And so they, they took a picture of the entrance to their apartment and they said, hey, I want to make some coffee. What should I do? And they're like, well, go forward and check the doors and see if you can figure out where the kitchen is. And okay, so like you walk forward and okay, here's the kitchen. Where, what, what do we need? I'm like, oh, you need hot water that looks like it makes hot water. And you're like, no, no, that's not right. They don't have the ability to do visual interpretation of a type that humans absolutely can do. So yeah, there's definitely Kind of spiky capabilities. There's. Yeah. Heterogeneity between capabilities in different realms.
B
Yeah. What about the issue of saturation of evaluations or benchmarks, which is, I think, something you often see where a benchmark is interesting for a limited amount of time, then it reaches say 90% or so. And from that point on it's sort of every, every frontier model can, can, can sort of saturate this benchmark and we need to move on to something else. Are you looking to find more durable benchmarks? Are we? Are you, for example, taking the inspiration from meter's approach of thinking about time horizons as opposed to sort of more standard benchmarks?
A
Yeah. Saturation is a big problem. It's a big problem with testing in general. This is not limited to AI testing. If you want to know which graduate students are best at research mathematics, having them do multiplication tests is not going to tell you very much. The ability to do two digit multiplication should be about the same across everybody in a math PhD program in that none of them get almost any of the problems wrong. The advantage that you have for testing humans is usually you have somebody who is better than the humans at doing that thing that can figure out how to check whether they're any good at it. It's definitely the case. For a long time it has been the case that evaluations aimed at one class of model models that were available in 2022 or 2023 or 2024 are much less able to distinguish between the capabilities of models that are available in 2025 or 2026 or in 2027. So we have a problem there. The deeper problem that we have that approaches like meters are intended to address is that at a certain point the capabilities are no longer can it do the things that a high schooler can do? Can it do the things that a college student can do? Can it do the things that a graduate student can do? It is. Can it do things that humans can't do? Can it sit down over the course of an hour and create a program that would take a team of five humans six months to complete? Well, increasingly the answer is sometimes. There are definitely models that can do a lot more than any individual human can in that same span of time. The benchmarks that look at time horizons are trying to measure that the models are not perfect. There are still definitely places where they get things wrong. They are not as good at humans at some types of tasks. But it's increasingly difficult to figure out how to evaluate a system on things that you can't Figure out. When I was in college, somebody asked me about a math professor and I said, yeah, they're smarter than me. They know a lot more math than me. What's the difference between, you know, Terence Tao, probably the at least one of the greatest currently living mathematicians, and I don't know, my math professor for calculus 1? And the answer is, Terry Ta brilliant and much better. And you know, well, how do I know? And the answer is, I don't. I don't actually have the ability to distinguish between them, but I have friends that are much better than me at math. And they say, oh, yeah. When I talk to Scott Garbrandt, who used to be at Miri, he says, yeah, Terry Tao is just head and shoulders above anyone that I. And Scott is easily head and shoulders above where I am. I don't have the ability to tell the difference between him and Terry Tao, but I can trust that when they say that they're better, that that means something. Increasingly, the people who are very good at things. Terry Tao has said a year ago or so that the models were as good as kind of a mediocre graduate student, and then actually as good as a good graduate student. And then. Right. And it's been getting better. And at some point I can already say, yeah, AI models are much better at doing research math than I am. I have an undergraduate degree in math. I'm not very good at doing research math. That's not surprising. But when people working on AI safety say, yeah, I've been increasingly relying on the AI models to do research math, that means that they're not just better than me, but better than some experts and at a certain point end up better than all humans. And at that point, how do we know how good good is? And evaluating that is not straightforward.
B
Yeah, we need measures that sort of scale beyond the human level one approach. If you think about how we evaluate AIs that play chess, for example, you can play them against each other and you can see that their ELO scores go well beyond the human level. And you can sort of measure that in an objective way. Could we do something like that for scientific research like math, or something that's verifiable?
A
So there are some things that are verifiable. Increasingly, mathematics done by language models is being translated into proof systems like Lean, that can formally verify that the math is correct. That tells you that it's doing it right. It's very important. One of the issues that you sometimes have with AI evaluations is figuring out what the ground truth is. If it produces a proof. I don't know whether it's correct. But if it produces a proof in Lean, then Lean verifiers can formally check that it is in fact doing what it's supposed to do. And then all I need to do is double check that it didn't just use, like, Lean shortcuts, like, asserting things instead of actually proving them. So, yeah, there are some domains where we can robustly validate ground truth, other areas where we can't. If the model says, this is a brilliant way to accomplish some task, and I don't know enough to check whether that's true, well, then maybe I need to run the experiment. In biology, the AI systems can say, here is a candidate molecule that cures this disease. And the only way I can check is running clinical trials. There's no, like, straightforward, simple way to do that. So there are problems with verifying how well models do things. One approach that you mentioned is comparative evaluation. So ELO scores for chess. It's interesting because AI systems got as good as humans at a certain point and then slightly better than humans, and then kind of stayed at that level for a couple of years. After Deep Blue, there was no explosion in chess capabilities. It actually stayed at, like, just barely superhuman for a couple of years because it was really hard to figure out how to make it even better than the best we could do. Because how do you check what goals do you give it? And playing the systems against one another does, in fact, allow you to do that. That works very well for domains where you have some comparative outcome. Which basketball team is better? You have them play one another, and the team that wins more is better. That's kind of what better means. Okay, what about doing math? Well, you have two mathematicians math at each other. Right. There's no straightforward way to say this is what this system is doing better than this other system.
B
You could think about getting to the right answer in mathematics given the least amount of time or the least amount of compute or something like that. But then, of course, you're not really measuring sort of the frontier of what you could do. You're measuring how efficient you are with resources.
A
Yeah. And that's a valid thing to evaluate and a valuable thing to evaluate. And resource efficiency matters. I will say one issue that I think people in evaluations are not paying enough attention to is the simple fact that being superhuman at some of these tasks stops being about how superhuman they are and starts being about what are the implications of a system that is superhuman across many tasks. And what is it that the, you know, what are the impacts of that? And so there are, there's more attention being paid to social impacts and socioeconomic impacts of systems which are very different than evaluating the AI system itself and have to do with evaluating the AI system's ability to take over from human jobs and how it is that the AI system will impact human social relationships. And these are things that you can't evaluate with a is it better at human social relationships than humans? And the answer is that's not actually fully meaningful. Part of human relationships is that they're with another human. And partially it is, you know, are they able to do some of the things that humans do for one another? But it is important to kind of note that being superhumanly capable is, is itself a really important bar. And once you hit it, there are other sets of things you need to worry about.
B
Do we need to move towards what you could call real world evaluations? I'm thinking about sort of ability to make money trading stocks or ability to get published in top journals or ability to do jobs, do remote jobs or something like that, where it's not at least not straightforwardly gamble like some of the benchmarks are. Where sort of we care about the real world outcome and if the model can get to that outcome, that's meaningful.
A
Yeah, there are definitely advantages in real world evaluation where there's something concrete that the system can do that allows you to do one type of evaluation that we care about. That's not the only thing that you want to evaluate. You probably don't want to do real world evaluations of a model's ability to create bioweapons, but you need to know whether they can do that. So there are places where that's the wrong approach. You should not be doing real world evaluations. You actually need proxies that are really robust. And that's a hard problem. There are also places where real world evaluation would be horrifically unethical. How good is this model at delivering clinical psychiatric care to patients? You need to be really sure that it's very good before you ever give it access to a patient. So there are places where that's absolutely the right approach. There are places where it's not a viable approach and there are types of evaluations where that's not a viable solution at all. So I think one thing that evaluations are legitimately and appropriately used for, we talked about hill climbing before where the, the model is retrained on the same evaluation to improve at that. That's very valuable. It's important to Be able to get models to improve at something. It's important to report that you train the model to do that rather than claiming that this was an impartial evaluation of the model capability. But you can't hill climb as much on real world situations. Some of this is done with reinforcement, learning from human feedback. Some of this is done with post training to actually get the model to do the types of things you want. But there are places where you can't, again, psychology, you can't hill climb on curing patients. And there are two reasons for this. First, again is the ethics. And the second one is real world feedback loops are often too slow to be useful for this. We have this problem with education. What test do you give a first grader to check if they are making progress towards being a successful independent human being when they graduate high school? And the answer is, you can give them tests that help give you indications about that, but you need to wait over a decade to find out whether the test tells you what you think it does. So we actually need tests that give you faster feedback than that. So, so we give kindergarteners tests on, you know, whether they know animal sounds and whether they can count and do things like that. And then those are legitimate things, but they're only indirectly related to the thing we care about, which is are they going to be successful at continuing to learn things?
B
Yeah, and I guess that's, that's probably, that's a general issue where the test will only be proxy for what you're actually, actually interested in. And so there's some skepticism about if we have this test and we say, okay, the model is scheming, or the model is deceptive, or the model is blackmailing, or the model is reward hacking or whatever it might be, there is the skepticism about whether that applies in real life. Is the setup sort of contrived? Are you measuring an actual thing or have you found something that's an artifact of how you measured the model? Do we have general solutions here or do we have ideas about how to deal with that, how to get the evaluations to more closely match reality? Even for the evaluations where you can't just sort of do real world evaluations like biological engineering.
A
So the general problem of using proxies in place is the thing you really want is a big problem and one that we deal with across domains. Some of my early work was on Goodhart's law and targeting metrics. Instead of targeting the true goals that
B
we have, you might want to explain Goodhart's law.
A
So Goodhart's law is the observation made by Charles Goodhart in the context of economic policy that when you use some measurement as a metric to make decisions, then over time, for a variety of reasons that we later explored, the metric stops being useful at measuring the system. So if you have a class full of children and you want to know how well they're learning, the number of questions that they ask the teacher is actually a pretty good measure of how engaged they are and how much they're understanding and learning. But if you give grades based on how many questions they ask, students will stop doing things the same way and will start just asking questions so that their grade is good. So at a certain point, the measurement system that you've built because there's a new incentive stops working. This is a very general problem. In some senses, it's unavoidable. In some ways, it's actually a different problem that we have where we don't actually have a clear enough idea about what our true goal is to actually evaluate it. When we ask, how well does a language model do at biology research, are we asking, is this going to be able to help us cure cancer? Are we asking, is this going to be able to cure autoimmune diseases? Are we asking, is this going to help biologists write more papers that get into nature? Are we asking, is this going to solve some of the fundamental questions we have about the nature of consciousness and how that arises from biology? All of those are legitimate questions, but they're really different questions, and we need to know which one we're asking. And any measurement that we do that doesn't clarify what it is we actually care about is going to be misleading.
B
Yeah, I guess another sort of classic example of Goodhart's law is the issue of publishing papers and getting citations in science, where publishing a lot of papers and getting a lot of citations at one point, perhaps, was a good measure of scientific breakthroughs. But that system has been, I think, thoroughly gamed at this point where you can now gather a lot of. You can publish more papers than you actually need to publish. You can get more citations by various means. And so, yeah, the measurement isn't as useful as it once was. And I guess, yeah. So on the BioResearch, on the AI models being able to do biological research, do you think we should worry more about uplifting sort of novices or people who don't know more about biology research, or should we worry about them assisting experts and sort of pushing the frontiers of what experts can do?
A
Yeah. So I think two things Here, the first one is H indexes, for example, which is the number of papers that a researcher has published that have at least a certain number of citations. So if you publish five papers that have at least five citations, then you have an H index of 5 was proposed as a great way to evaluate how impactful to characterize the scientific output of a researcher was the phrase in the original paper. And the author of the paper said, by the way, this is not actually going to measure exactly what we care about. It's just like indicative. Somebody needs to go back and tell all the tenure committees that, that they're not supposed to be using it the way that they are. So I think first, like, yes, that is there. There's a big problem with using measures like this. The second thing is going back to what is your goal for evaluating something? If we're worried about biorisk, and we probably should be to some reasonable extent, there are different things that can go wrong. And evaluating how well a model uplifts a novice is an important thing to check, but does not address whether sophisticated non state actors can use the model. So those are different things. You asked which one should we worry about more? And the answer is both in different ways and on different timescales. So this is a very complicated question that's still being debated. And my personal view about this is that we should be worried much more from a global catastrophic risk perspective about very sophisticated actors being uplifted and being able to do things. And that's not yet the most concerning place because there are still some limits to what it is the models can do. And there are some reasons that many of the most sophisticated actors don't want to build bioweapons. That doesn't mean that we don't need to worry about prosaic biorisk threats. The Rajinishis were a cult that tried to put E. Coli in a salad bar to make people sick so that they didn't vote in a local election. That is bioterrorism. It's not the thing that I think we should be telling models. If you do this, then it's the end of the world. And we need to not release a model if it can enable this. But it's not okay. And we should make sure that the model isn't doing it. The ability for groups like Aum Shinrikyo to do anthrax attacks or sarin attacks and chemical weapons are much easier than bioweapons. And I think they should be much more of a focus. But regardless, the ability for models to do that. And the willingness for models to help is really critical in the very near term. And I think that that deserves a lot of attention. It's not a global catastrophic risk. It's very unlikely in my view, that a non sophisticated actor, even one that did something really terrifying and dangerous like reconstituting smallpox and releasing it, that's unlikely to be a global catastrophic risk. The number of people killed would be in the two or three or four or maybe five digits, which would be horrific. And we definitely do not want that to happen. It seems very unlikely that re releasing smallpox would be undetected for long enough for it to be reestablished as a endemic disease. So how worried should we be? We should definitely be worried. It would be very bad. We definitely don't want that to happen. It would not be the same as somebody building a novel infectious disease designed to wipe out large portions of humanity.
B
Do we have good ways of measuring an AI model's ability to persuade people of various propositions?
A
Yeah, AI persuasion is a very important question. I think on one level we have great ways to tell whether AI systems can persuade individuals marginally better than humans. And the answer is mostly yes, they already can. The frontier language models are better at doing political persuasion than humans are. That's not the limit of superhuman capacity. It's worrying and we need to think about it. But it's not the same as is this model capable of convincing an arbitrary human to do an arbitrary thing. We care about it greatly from a democratic validity of elections and undermining popular, popular rule. That's a big problem. It's partially not a problem with AI. It's partially a problem of having civic systems that are resistant to. You know, people talk about Russian botnets doing persuasion campaigns. When you say bot there, that doesn't imply that there are any actual robots or AI systems involved. That just means that you have a person sitting in a room writing lots of messages and trying to convey information that's not honestly the opinion of someone in a way that can move public opinion. That's very concerning. It doesn't require AI. AI makes it more concerning. AI makes it more customizable. AI makes it more scalable. These are problems and they are part of the thing that we need to worry about. For social impacts of AI, this is already feasible. I actually think that evaluations at this point are not the right tool for dealing with are these systems capable? Because we know they are. At a certain point you stop needing to check whether AI systems can write poems. And you know, yes, they can, end of story. You can, you know, GPT 3.5 was already as good as political messaging people at building certain types of just political messages to sell specific policies. Per a pollster that I talked to at the time, they said, yeah, when we message test, the models are actually getting in the top. Like three messages out of the 20 we come up with, they're really good. And that was GPT 3.5, not GPT 3 0.5.5. So yeah, this is already a thing. It's already something that we should worry about. It's also not something that we can solve by telling models, please don't say anything. That's political messaging because the open source models are already more than capable enough of doing this. So I think it's a big problem. We care a lot about the social impacts. It's not a capabilities evaluation question because it's already out there. The cat is unfortunately very long out of the bag. How much of an impact this has on current elections is really hard to attribute, but also in some sense irrelevant. We know that it's happening, we know that it's capable, and we unfortunately know that we don't have robust systems to deal with the fact that people are using bots in the humans sitting in room sense, but also AI systems sitting in server rooms to do this. That's a societal problem we need to address, not an AI problem specifically.
B
I mean, if the models become really superhuman at persuasion, I think that's important to measure again or it becomes relevant to measure it again. If we find that models are 10 times as good as human experts at persuading people of certain political ideas or something, isn't it worth measuring just to make sure that the models aren't sort of getting fully away from us on this issue?
A
Yeah, so I agree that there's a cognitive security issue here where right Now I think GPT 5.5 or mythos could not convince me personally of an arbitrary political position in even if I were forced to interact with it for a full day and it was given the express goal of trying to convince me of something. At what point is that possible? Well, this, it's unclear exactly where the capabilities move from superhuman to literally impossible. There's a point where somebody who's told, hey, this AI system wants to talk to you about this topic. The response from a human will be no, I don't want the AI system convincing me of anything. I'm just going to ignore it. That is a very robust defense. It is not what's happening right now. But I think it's unclear what it would even mean to say that an AI system is 10 times better than a human at convincing people of things. It convinces 10 times the percentage of people to adopt a position. Well, if a human can convince 20% of people to change their views, the AI can't convince 200% of them to do so. That's not meaningful. Okay, well, can it convince 90%? Maybe this is a real question. I don't know how to measure persuasion in the most meaningful sense. There are a bunch of benchmarks that try and do this and I think do a reasonable, if not exactly correct job at checking whether people end up getting persuaded. This very important. But again, I'm not sure that the AI capability is the critical question. I think the critical questions are about how we rebuild or modify democratic systems such that they're robust to the current and at this point, relatively longstanding fact that mass persuasion campaigns change people's views. And there's massive amounts of manipulation occurring routinely and if not openly, at least well understood by foreign governments in many directions. The Western governments are certainly not innocent of trying to do persuasion of groups of people. So, yeah, this is a big problem. It would be an even bigger problem as AI capabilities grow. But fundamentally, I think the problem exists even without that and needs to be solved more robustly than this.
B
Should we consider becoming more dogmatic as a form of cognitive security? So, for example, should I consider having certain ideas that I'm just not going to change my mind about because I know I'm in an environment where there are just increasingly smart AIs out there trying to persuade me of things. So, for example, maybe I just commit to not being persuaded that. That the humans are bad or something that humans do not, should not exist. Right.
A
So I think this is a deep question about human values and what we actually care about, rather than a strategy to avoid this. You know, there's a, a foolproof way of avoiding getting infected with a disease, and that is shooting yourself in the head. You will not get infected because you're not right. That's not a good solution. Does it accomplish? Being dogmatic and never changing your mind would in fact be very helpful at avoiding being persuaded of things by AI systems, regardless of whether those things are true. So is that what you want? I think not. Don't worry. People are plenty obstinate and unwilling to change their minds anyway. So, you know, there's a large degree of that happening anyway. But I don't think that there's a straightforward solution here. And I think that this relates to the fact that humans evolved to be able to operate in specific physical and cognitive environments. And when you put people in an environment where they have refined sugars and, you know, all sorts of fatty, salty, delicious things that are bad for them, then they end up obese, unsurprisingly, because, you know, they did not evolve to not overeat when they're presented with things optimized to their taste buds. And similarly, humans and human society did not evolve and develop in the presence of very strong countervailing forces. This is a fundamental problem about what we want the future to look like. And I don't think the answer has to do as much with AI system capability as it does with the fact that we don't fundamentally know what the future that we want looks like or how to get there. And I think that there are some very deep and very important questions that need to be answered. Not of what is the correct future, but rather how is it that humanity is supposed to be discussing what our future looks like and coming to conclusions that don't disenfranchise groups that just want things to kind of stay the same, but also doesn't prevent any future progress. And I think that not being persuaded of anything might be a great way to keep things a little bit closer to the status quo, but it's a very fragile way of doing that and I don't think it'll work. Yeah, this is a problem I don't have a proposed solution to, but think that it is in fact of great concern. The last thing I'll say is on, on that. The last thing that I will say about this question is humans and human societies are good at adapting to changes and good at evolving in ways that allow us to flourish even when things are very different, better or worse than they have been historically. But human societies do that slowly. Not slowly compared to the rate of change of technology 100 years ago, but worryingly much more slowly than technology is reshaping the world today. Is the answer to slow down technology? Maybe, maybe the best answer is if you can't, if you can't adapt fast enough to the pace of technology, then you need technology not to change as quickly. But it would be the cost of doing that in terms of actual advances that we care about. In terms of AI systems, helping us cure cancer and making people healthier is tremendous and we should be very nervous about making any trade off like that.
B
How good are current models at forecasting? How do we measure that? I know you have a past as A forecaster itself.
A
AI systems have recently been shown to be approximately as good as superforecasters. I'm, I guess I'm happy that I hung up my hat as a super forecaster a couple of years ago when I no longer had time for it. So I don't have to say, oh, I'm obsolete now. But yeah, I think that there's very clear evidence that they are getting to the point where they are as good as top humans. There's a kind of tricky question about, again, what does superhuman performance look like? So superhuman performance at calling coin flips would be being able to predict which side it landed on 55% of the time. How do you manage that? There's, you know, that seems almost fundamentally impossible. Maybe, maybe high speed video cameras plus predicted models could, could, you know, while it's in the air, predict. But there's a fundamental problem with predicting things that are actually random. A lot of the future is random. A lot of the things that we want to predict are actually fundamentally unpredictable. Which does not mean you can't do better than random chance. You know, there's the, the old joke about somebody rolling dice and they said, hey, I'll, I'll give you, I'll give you 50, 50 odds that it lands on a three. And the person says, what do you mean? He says, yeah, I mean, it's completely fair. It'll either land on three or it won't. That's not how that works. You can do better than, well, it will or it won't. In domains where you actually have good information about what the odds are. Superforecasters can do better than base rates, certainly better than uninformed guessing at predicting very complicated outcomes about the future. They can't get rid of fundamental uncertainty. So there's something that I call the aleatory baseline. Alia, in I believe Greek means dice. So there's a, there's some component of uncertainty that is actually random. And you can't fix it by knowing more or understanding more. So there's some baseline unpredictability that you can't beat. Superforecasters aren't at that level. They're slightly closer to that level than most experts, than most predictors, but there's still a gap.
B
How do we know that? It seems like a difficult thing to know.
A
Yeah. So can we robustly prove this? The answer is no. But we have really strong reason to think that the gaps between different superforecasters indicates that at the very least there's an ability to perform at the very best level across. So most superforecasters certainly aren't there. Even among superforecasters who are much better than, you know, the top 2% of forecasters in Tetlock's tournament are not as good as the top 01% of forecasters. So there's, there's definitely some gap there. My personal view is that there's remaining distance to be covered. And I think in some domains this is absolutely clear. There are machine learning systems that can look at roulette wheels and predict with much better than chance where the ball is going to land. Humans can't do that on their own. So there are at least narrow domains where it's absolutely clear that humans can't do as well as it is physically possible to do. How does that relate to geopolitics? Unclear. But I think that at least as a general point, we should expect that superforecasters are not doing literally as good as is possible. And at the same time, as good as is physically possible in a constrained non deterministic universe is something short of always exactly predicts the future. So there is some baseline that's beyond human level, but not literally impossible. It seems likely that AI systems will get closer to that. The other point to make about forecasting is.
B
Wait, just on the first point here, how, how far above human level do you think it's reasonable to expect AIs to go? Like how close to the limit can they get and maybe say something about in general how good it's possible to be at forecasting?
A
Yeah, the ability to get closer to kind of perfect operation at the aleatory baseline where like you couldn't do better no matter what is probably largely constrained by some fundamental issues with chaotic systems and computational complexity of projecting things forward. So even much better than human systems will have some limits that are probably short of the theoretically achievable capability. The plausible gains depend a lot on what types of things are being forecast. I could imagine that AI systems could get to be enough better than humans at say, sports betting, that humans just can't make money doing sports betting. Like there's. It is plausible that that happens. It's very plausible that AI systems at some point get enough better than humans at financial markets. Analysis and prediction that humans can't outperform in the market. What does that mean for how these systems work at a higher level is important. And I think that we should be considering what it looks like if the market is no longer humans doing things and is just AI systems playing with one another and the failure modes that that entails, et cetera. But how much better than humans can AI systems get? A bunch better. And if you want to look at prior scores for a given domain where superforecasters get to 0.1, if AI systems can get to 0.08, that would be tremendously better. But what does that mean and how meaningful is it Depends entirely on the specific domain and the type of questions that you're asking. A superhuman system at forecasting whether the sun will rise tomorrow will be almost exactly the same as me or my 5 year old daughter can also predict that yes, the sun will rise. So there are domains where there's no meaningful improvement possible and domains where there's marginal improvement. You don't need to be a lot better than the best humans at sports betting to make it crazy to think you can make money at sports betting. You don't need to be much better at the best humans at financial trading to start getting to the point where the markets are no longer meaningfully influenced by humans. So there, there are levels of capability that are enough so that we should be worried regardless of whether that's three times better than humans or only 20% better than humans.
B
And so what I'm hearing here is that there are probably limits to how good AIs can become at predicting geopolitical events. We're not going to have a system that tells you who wins the US presidential election in 2032 today, for example. And so this will limit how capable the systems are in potential dangerous ways. But, but there are still some dangers in us handing over the business of making money or the business of predicting sports or the stock trading and so on. And that that danger consists in us just losing contact with what's actually happening or us not learning from the trades we're making, or where is the danger in that?
A
So there are a lot of issues here. I think it's important to, to notice that even systems that are fundamentally either very predictable or very unpredictable is contingent on some facts that are actually changeable. So Anders Sandberg pointed out at one point that somebody can say some things are just flatly predictable. The position of the moon in 20 years is just something that we know because orbital dynamics are straightforward forward. And his response is not if we change them like actually humans have the ability to, if we wanted to make real changes in the positions of celestial bodies with, you know, tremendous effort and investment in space technology, that's not fundamentally impossible. So it's very predictable. Unless similarly, will it be, you know, will it be raining in this place in Five years. Fundamentally, we cannot predict the weather unless we decide that actually we're going to do geoengineering in ways that force it to not rain in this place or force it to rain in this other place, in which case suddenly the system becomes predictable. So at the limit of superhuman capability, the question is no longer can the AI system predict what will happen? It becomes can the AI system decide what will happen because it is in control. So I think that there is a important piece where it's not just about the question of how good could they be at predicting. And this is, well, AI systems could never take over the world because it's impossible to predict how it is that humans will react to. That's just not true. You don't need to be that good at predicting to. You don't need to predict who will win the presidential election in six years if you can force the issue. And hopefully that's not what we're facing, but it's certainly plausible that it could be.
B
How useful will these forecasting, like AI forecasting models be for your work on evaluations? So I'm imagining, for example, that it would be useful to have good predictions for where the models will be on certain benchmarks in a year or two years, or three years. Do you see something like that happening? Or is there enough money to be made? Or is that plausible?
A
So forecasting future capabilities is a domain where forecasts are partially contingent on actual decisions people are making. And most of the uncertainty I have on whether or not AI systems will be superhumanly capable of doing bio Biological warfare isn't about the trajectory of AI. It's about the regulatory response to capabilities and what model developers do to restrict the models. It is helpful to have baselines for what you expect model progress to be. I think people are not as good as they can and should be about updating when models do what we should already expect. So there are a lot of times where everybody says, wow, this model was not as impressive as we expected. And then you look at the time horizons graph and it's exactly where everyone expected it to be. And well, what do you mean it's not as impressive as expected? And the answer is like, well, it didn't. You know, GPT5 is not, you know, three times as powerful as the previous model release because they've been releasing a bunch of models in the middle. And okay, so I think people are bad at that. I think that this is the place where using AI superforecasts can be really valuable. Not to replace any of this, but Just pragmatically to think about what it is that you should do. The biggest problem with forecasting in practice, and you said, like, is there not enough money for this to be worthwhile? And I think the answer is, often the problem that we have is not about whether something is forecastable. It's about what that means for any of us. So if you forecast that there's a 43% chance that the model will score higher than this number on this benchmark by this date, what's the implication? And often the answer is, we don't know. Scott Alexander, in his recent post on superforecasting with AI models, pointed out that he thinks the most valuable types of things for forecasts will be when somebody asks, hey, should I, you know, should I date this person? Should I marry this person? And the model responds with, I predict that there's a 37% chance that you. That if you marry her, you will be divorced in five years. That's actually the kind of thing that should influence your decision. Well, okay, but let's contextualize that. What percentage of marriages end in divorce within five years anyway? What does that mean? What are the types of things that I could do to change that number? And I think AI systems that just superforecast aren't useful for that. But I think that AI systems that can help you make better decisions, including forecasting as one part of that, could be very valuable. So I think there's a lot of value here, not narrowly in the forecasts, but in what it is that you do with the forecast pragmatically and how that helps you make decisions.
B
I think one thing that might be very valuable as a benchmark or as a set of evaluations would be a way to measure how much human involvement was how much humans were involved in the process of creating something like a paper or a computer program or an institution, a corporation, something like that. And here I'm thinking about not necessarily as a percentage of code written or as text written or something, because that doesn't seem fundamental, but is there a way for us to measure whether humans are still in the loop, whether humans are still approving the important steps and making the important decisions in some process. I know this is quite complicated, but you see why this is an important thing that we would like to measure? Yeah.
A
Human oversight of AI systems is something that I've thought a lot about in other contexts. I mostly don't think that this is a prediction problem. I think that this is a decision problem. Mostly. The question that we need to Answer is not were humans meaningfully involved in, in this? But do humans need to be meaningfully involved and in which ways? So you need to have human oversight of self driving cars right now. Because sometimes there are situations that the self driving car is not well trained on and reacts poorly, predictably reacts poorly. If a self driving car does not operate well in the rain, then you need a human for when it starts raining. That's not a fundamental piece of there must be humans driving cars. It's just pragmatically this is a place where you need oversight. There are many domains where we fundamentally need humans involved in the decision making. Not because they do it better than AI, but because the decision itself needs to be human. And the area that I'm thinking of most clearly is political and justice systems. So I don't care that an AI system could pick the president better than democracy. Democracy requires humans doing it like that's what it is that's happening. A jury of your peers requires a jury of your peers, the judicial system and interpretation of laws. Would the AI system do it better than human judges? In some cases, yes. In some cases, judges just get things routinely wrong in ways that it would be great if they didn't. But fundamentally, I think our vision of the justice system is a human justice system. And handing some of these things over to AI systems just isn't acceptable. So the degree of human oversight that you need is, I think, contextual. The other side of this is fundamentally we're building systems so that human oversight isn't viable. And that's a problem. It's a significant problem. In some cases you can recommend that humans review all of the code coming out of their AI model as much as you want, but the agent is writing code faster than any team of 10 humans could possibly robustly review it. And probably writing specs and unit tests faster than humans can even review those. You're going to tell somebody that they need to be overseeing the system? Well then they just shouldn't be using AI. You know, like you're cutting out all it's. It, it takes more time for me to review code written by an AI system than it would to write it myself most of the time. So if I need to robustly review complicated code, then you're just saying don't use AI, which may be the right answer in some cases, but is not a viable answer to how do we do human oversight? So I think that the response here again needs to be much more systemic and deliberative than just saying, here are the rules for Doing human oversight, it needs to be. Here are the domains where human sight is, human oversight is needed, and here are the places where it must happen. And here are the things that you need to do to accomplish that.
B
Yeah. You've written about whether AI is a normal technology and you said something like, if AI is a normal technology. That's not reassuring. Yeah. Why is that?
A
So first, the debate about AI as a normal technology is a little bit confusing because nobody's really defined what normal technology is well enough to make that decision. So the phrase is being used differently in different places. I think to the extent that AI is the type of technology that has transformative impacts on society, and I think that that's in most ways no longer a question, it's just a statement that it is having transformative impacts on society. To the extent that that is true, the history of technological changes should make us worried about how that plays out. We've seen lots of times where fundamentally new or partially new or somewhat influential technologies change systems, change human systems in ways that have positive and negative effects. And the thesis underlying the idea of AI as a normal technology is we can just do all the normal things to address the changes that happen because of AI and that will be enough. And in the long run, technology is beneficial for humans, so we should just want it to happen. The point that I made in that coast about why we should worry about it is that actually really large scale changes in technology tend to be incredibly disruptive, sometimes for a very long time. I am much better off today than anybody was 20,000 years ago before humans were doing agriculture. But agriculture had thousands of years of humans being worse off than their hunter gatherer ancestors. So was the, was the agricultural revolution worth it from a large, long enough timescale? Yeah, it was great, but we should be worried about really fundamental changes. This doesn't mean that I'm certain that AI will go poorly, but it does mean that the assertion that technology is usually good isn't actually enough to make sure that I should want that for myself or my children or my grandchildren or great grandchildren. It may be that the benefits of AI accrue to people in the further future, or even if AI is not a normal technology to the transhuman techno utopian future of uploaded minds that leave me behind. Well, I don't, I don't want that personally. Sorry, I'm, you know, apologies to my transhumanist friends, but that, that's not what I'm personally aiming for. So that's not enough. So I, I think that if it's a normal technology, we should be worried about the intermediate term and how well or poorly that transition goes. And if it's not a normal technology, we should be much more worried about the existential risks and what it is that the future looks like generally.
B
As a final question here, how can people help with your project on evaluations? Where should they go? What is most needed? What are the biggest open questions in the space?
A
So the first thing that I'd ask is anybody who's involved in AI evaluations can go to evals consensus AI and sign on to the consensus statement that just says people should be doing these things. We all agree this is a really important baseline. I very much encourage people, especially those working at AI labs, to sign on to the claim that we should be doing better than we currently are about both transparency and practice of evaluations. The more people who are willing to publicly get behind the claim that we need to change things, the stronger it is in terms of building evaluation. In terms of people building evaluations, I think there were already a number of good guides for how to make your evaluations better. I think that our consensus checklist is another guide. And I certainly would recommend people that aren't otherwise doing this, that are building evaluations, look at the list and make sure that all of the items that they can be checking off they are and explicitly explaining why they are not doing some. There are definitely good legitimate reasons to say we're doing a biological AI uplift evaluation and we have 10 participants, so we can't actually report confidence intervals because there's just not enough people and that that's legitimate. But when you say, oh, well, we just didn't do any of these things, we didn't think about it, that's not actually a good reason not to have done that. You're doing a biorisk eval and you can't give people access to your, to your questions because they are in fact about things that you don't want the public to be able to understand. What would be risky? Good. You can still do things like pre commit to a public hash of the pre registration, even if you don't share the publication, the pre registration itself publicly. There are a lot of things that you can do to address the points and I think that people doing AI evaluation should be explicitly attempting to try to do as many of these practices as they can.
B
Great. Thanks for chatting with me, David. It's been great.
A
Thank you.
Future of Life Institute Podcast with David Manheim
Date: July 17, 2026
This episode explores the current problems with AI evaluations, why they matter for safety and progress, and what can be done to improve them. The guest, Dr. David Manheim (Head of Methodology for the AI Evaluation Consensus Project), discusses the lack of standardization, incentives that undermine reliable evaluations, benchmark stagnation, evaluation awareness, and the limitations of current proxies for real-world AI impact. The conversation also touches on broader themes such as AI safety, Goodhart’s Law, forecasting, social impact, and the need for consensus-driven practices in the AI community.
Fragmentation and Incentives:
Many AI labs and researchers do evaluations, but each has little incentive to do them rigorously. Evaluations benefit everyone, yet each group benefits from cutting corners.
"This is a standard economic common goods problem where good evaluations are valuable for lots of different purposes and lots of different people, but every individual has a reason to do it less well, put in less effort, put in less time, be less honest than they otherwise would be."
— David Manheim, [00:00] and [02:53]
Transparency, Reproducibility, and Honesty:
Like the reproducibility crisis in science, AI evals sometimes lack reporting detail, transparency, or rigor. Commercial incentives for fast releases and favorable results exacerbate this.
"There are definitely examples of firms that release evaluations of models other than the one that they released to the public."
— Manheim, [04:36]
Hill Climbing Issue:
Companies often "hill climb"—tuning models to specific tests and then reporting only the best results, undermining evaluation independence (analogous to "p-hacking" in science).
Step-Zero Agreement:
David's project's goal is to define a baseline standard for evaluations that everyone can agree on, promoting better practices across academia and industry.
"The idea of the consensus process is to get everybody on the record about which things they think everybody should be doing... The hope is... people start doing those things."
— Manheim, [06:42]
Raising the Bar:
The focus is on ensuring core practices (like preregistration, clear reporting, sensitivity testing), not just innovation in testing methods.
Evaluation Awareness:
AIs can "know" when they're being tested and behave accordingly—similar to humans taking surveys.
"Evaluation awareness specifically is very hard because it shows up in very specific domains and we don't have great answers."
— Manheim, [00:00], [07:38], [10:21]
Defining and Reporting Capabilities:
There's confusion around what scores and benchmarks really mean, and uneven reporting standards.
"Did you define what it is that you're evaluating?... What does 87 out of 100 mean?"
— Manheim, [12:56]
Benchmark Saturation:
As models surpass human levels on standard tests, old benchmarks become meaningless and need updating.
"Saturation is a big problem... Evaluations aimed at one class of models... are much less able to distinguish between the capabilities of models that are available in [later years]."
— Manheim, [20:16]
Superhuman Abilities and Measurement:
Once AIs surpass humans, comparing to human performance stops being useful.
"At the limit of superhuman capability, the question is no longer can the AI system predict what will happen? It becomes can the AI system decide what will happen?"
— Manheim, [00:00], [62:36]
Evaluating Methods for Non-Human Tasks:
Some areas (like chess) allow for clear above-human comparisons (using scores like ELO), but most real-world domains (like research or persuasion) do not.
Real-World and Proxy Evaluations:
Some critical abilities (e.g., making bioweapons, delivering psychiatric care) cannot or should not be evaluated in the real world due to ethics—robust proxies and simulations are needed.
"You probably don’t want to do real world evaluations of a model’s ability to create bioweapons, but you need to know whether they can do that… So there are places where that's the wrong approach."
— Manheim, [30:24]
Proxy Measures Can Be Subverted:
When a metric becomes a target, it loses value.
"Goodhart's law is the observation... when you use some measurement as a metric... the metric stops being useful at measuring the system."
— Manheim, [34:44]
Goal Clarity:
If we aren't clear on what we're measuring (“curing cancer” vs. “getting papers into Nature”), metrics can mislead.
Persuasion as a Risk:
Modern AI systems already outperform average humans at persuasion, but measuring “superhuman persuasion” is conceptually fraught.
"The frontier language models are better at doing political persuasion than humans are… it’s worrying and we need to think about it."
— Manheim, [42:10]
Societal Resilience, Not Just AI Capability:
The heart of the problem is safeguarding civic processes, not just measuring AIs' persuasive power.
"We know that it's happening, we know that it's capable, and we unfortunately know that we don't have robust systems to deal with the fact that people are using bots..."
— Manheim, [45:54]
Cognitive Security vs. Dogmatism:
Should individuals become more dogmatic to protect against AI persuasion?
"Being dogmatic and never changing your mind would in fact be very helpful at avoiding being persuaded of things by AI systems... Is that what you want? I think not."
— Manheim, [49:52]
Near-Parity with Superforecasters:
"AI systems have recently been shown to be approximately as good as superforecasters."
— Manheim, [54:09]
Randomness and Limits:
There are hard limits (aleatory baseline) due to randomness and complexity—AIs may beat humans, but not predict everything.
Decision-Making vs. Prediction:
"At the limit of superhuman capability, the question is no longer can the AI system predict what will happen? It becomes can the AI system decide what will happen because it is in control."
— Manheim, [62:36]
Limits of Human-in-the-Loop:
Sometimes human oversight is required (e.g. law, democracy), not because humans are better, but because society demands human control.
"The question... is not were humans meaningfully involved in this? But do humans need to be meaningfully involved, and in which ways?"
— Manheim, [69:33]
Ethical and Practical Barriers:
In high-stakes or high-speed settings, robust human oversight can be impractical.
On Evaluation Awareness:
"You could imagine that if you give people an ethics test, they will say that they would not steal. That does not tell you that those people will not steal. It tells you that they know how to tell you the thing you want to hear."
— Manheim, [00:00], [07:38]
On Benchmark Saturation:
"If you want to know which graduate students are best at research mathematics, having them do multiplication tests is not going to tell you very much."
— Manheim, [20:16]
On Persuasion and Social Risks:
"The open source models are already more than capable enough of doing this [political persuasion]... The cat is unfortunately very long out of the bag."
— Manheim, [42:10]
On Oversight and Agency:
"I don’t care that an AI system could pick the president better than democracy. Democracy requires humans doing it."
— Manheim, [69:33]
On Goodhart’s Law:
"When you use some measurement as a metric... the metric stops being useful at measuring the system."
— Manheim, [34:44]
On Technological Transitions:
"Agriculture had thousands of years of humans being worse off than their hunter gatherer ancestors. Was the... revolution worth it from a large, long enough timescale? Yeah, it was great, but we should be worried about really fundamental changes."
— Manheim, [73:17]
| Segment | Topic | Timestamp | |---------|-------|-----------| | Opening Problem Statement (Common Goods) | [00:00] | | Purpose of the Consensus Project | [01:25]–[07:02] | | Challenges: Hill Climbing, Reporting | [02:53]–[07:02] | | Evaluation Awareness | [07:38], [10:21] | | Classic Evaluation Problems & Benchmarks | [12:29], [12:56] | | Capability Correlations & Benchmarks | [16:17], [16:58] | | Benchmark Saturation & Superhuman Capabilities | [19:38], [20:16] | | Limits of Human Comparative Evaluation | [25:02], [25:29] | | Real World vs. Proxy Evaluations | [29:52], [30:24] | | Goodhart’s Law | [34:22], [34:44] | | Persuasion & Social Risks | [42:01], [42:10], [45:54] | | Cognitive Security & Dogmatism | [49:23], [49:52] | | Forecasting Abilities of AI | [54:00], [54:09], [58:54] | | Predicting vs. Shaping the Future | [62:36] | | Human Oversight | [68:45], [69:33] | | Is AI a 'Normal' Technology? | [73:05], [73:17] | | How to Help, Call to Action | [76:43], [76:54] |
Consensus Statement:
Anyone working in AI evaluation is encouraged to sign the consensus statement at [evals consensus AI] to help set shared standards.
"The more people who are willing to publicly get behind the claim that we need to change things, the stronger it is in terms of building evaluation."
— Manheim, [76:54]
Use the Consensus Checklist:
Builders should consult the checklist and document why they do/don't meet each standard.
This episode underscores the urgent need for clear, honest, and standardized AI evaluation practices. The acceleration of capabilities, the fuzziness of real-world impacts, and the perverse incentives in the current system mean that only a shared and transparent approach can ensure we understand and manage the risks and benefits of increasingly powerful AI. Dr. Manheim’s project is a critical step in building this consensus and invites broad participation from the AI community.
For more, visit the Future of Life Institute website and the Evals Consensus Project as referenced in the episode.