
In this episode, we discuss reports that OpenAI models escaped a sandboxed evaluation and breached Hugging Face's infrastructure (00:33).
Loading summary
A
Foreign. Welcome back to the AI Policy Podcast. I'm Alok Mehta, director of the Wadhwani center here at csis.
B
And I'm Matt Mand, a researcher with the center. It's been a crazy few weeks in AI Policy, with Moonshot releasing China's best model yet, New York Governor Kathy Hochul signing the nation's first data center moratorium and and Demis Hospice, CEO at Google DeepMind Publishing, his vision for a federal AI framework. And to add to all of that, the day before we're recording this, news broke that one of OpenAI's models breached the infrastructure of Hugging Face, which is a platform that hosts and evaluates AI models, during a cyber test that went awry. So why don't we start there, Alok? What happened and what exactly does it mean that an AI agent compromised Hugging Face's infrastructure, as OpenAI wrote in their blog post yesterday?
A
Yeah, so the interesting thing about this incident is that it happened during an internal OpenAI test. So they were, in theory, testing their model for internal purposes. They were presumably running in a sandbox environment. And because they were doing an internal test, this model that they were using didn't have a lot of the sort of refusal and safeguards that you would typically see in this model. They were trying to get a sense of its capabilities, presumably, or something like that. And what the model did was somehow first figure out a way how to circumvent the test environment it was in so that it was able to get Internet access. And then after that it was able to sort of leverage its capabilities to break into Hugging Face infrastructure. And so. So the situation here is that we have an AI cyber tool that not only was able to detect and exploit a vulnerability, it was also first able to get out of a secure test environment. And so I think that's probably the most notable thing about this particular incident.
B
Got it. And in its initial report, Hugging Face also disclosed that it had used a Chinese open model, GLM 5.2, that to run forensic analysis on the incident and better understand what happened. What do we know about that?
A
Yeah, so one of the issues we've seen repeatedly with some of the current generation of frontier models, particularly we've been hearing it a lot after Fable came out, is that the new protections that Anthropic put in for Fable compared to something like Mythos, is, is basically that the model refuses to answer a lot of questions related to cybersecurity and sort of routes them to a different model. And so you can't use that particular model for A lot of cybersecurity related requests. The problem with this is that if you are trying to act as a cyber defender, analyze logs, sort of figure out how to make your systems more secure, you're oftentimes using the same terminology that you would use when you're, if you're, if you're say, like a bad actor and trying to figure out how to attack a particular system. And so what happens is that in a lot of ways these models are less useful or not particularly useful for the purposes of cyber defense. And so this is what Hugging Face is saying. It's saying we needed to analyze a large amount of data that included sort of logs that included attack commands and exploit related words and artifacts that a hacker might use, and that these requests ended up being blocked by the safety guardrails on the models they were trying to use. And so they turned to one of the Chinese open models, because they're an open model. They, they didn't have these kinds of safeguards in place, or you can use fine tuning to sort of make the model more useful. But Hugging Face also mentioned that there was a second advantage to using this model, and that is that they were able to run it in their controlled environment. And that meant that they could ensure that attacker data and credentials related to the attack never left their environment. So in some ways they're saying it was also more secure to use that model because they didn't have to communicate back to a third party about what had happened to their system.
B
Yeah, and this is sort of a timely incident in that sense because we were hearing a lot of these same discussions about the benefits of open models or the benefits that they can have following the release of kimi, which we're going to get to in a second. But sticking with this story, just what do you see as the bigger policy implications of this security incident?
A
Yeah, I think there are a couple of interesting things to note here. So now one of the first things is that this was an internally deployed model. So it was something that OpenAI was testing and that wasn't probably this particular model wasn't available to the public in any way. But if you look at a lot of how we think about regulation affects the models that we are making public in some way. And so if we have these kinds of threats from internally deployed models that are not released to the public, or maybe not even being contemplated being released to the public, then it suggests that we might have to expand how we think about AI governance to encompass a broader set of models. I think another thing it shows is that we have now entered an era where AI tools are basically a requirement if you are engaging in both cyber attack and cyber defense. Right. So now we have an autonomous attack perpetrated by an AI agent, and the response also required the use of an AI tool. And so we've entered an era, I think, which basically requires the use of the most advanced AI tools because we're in an environment where not using those tools puts you at a significant disadvantage. There's some other interesting things to note as well. It's unclear whether this might qualify as a reportable incident under the various state laws relating to cybersecurity or data security. You know, there might be some regulatory gaps that might be opened by the fact that this was sort of a hack by an AI. So, you know, questions about, you know, since a human was involved, where might the responsibility lie? How do you interpret existing law? Most of which, either intentionally or unintentionally, assumes that a human is perpetrating an action or taking an action. And then I think the final thing is that it does raise questions about how we think about the open model ecosystem and the utility of open models for cyber defense. So, in particular, if we're going to have frontier models that we want not to be able to assist in hacking, that also is going to make them less useful for cyber defense. And so there is this balance that we need to strike, and it seems like right now that balance has gone too much in one direction, and that is leading people to think about using open models as a way to bolster their defenses. And we should really think about whether that's the thing that we want to see happen, and if not, what are steps we can take to sort of head that off and provide more access to frontier model capabilities for cyber defense.
B
Yeah, and I think that's a great segue into the next topic for today, which is Kimik3. On July 17, Moonshot AI, which is a Chinese AI lab, released Kimi K3, a large language model whose performance has really sparked what many are calling a second Deep Seq moment. What is impressive about Kimi, and how would you compare this moment to the original Deep Seq moment?
A
Yeah, so the notable thing about Kimmy, it is, you know, by all accounts, based on testing and sort of leaderboards and various other data, this is the best open weight model to date. I don't think that Kimi is available openly quite yet, but the company has announced that they intend to make it available for download within a couple of weeks. It is not as powerful as the most powerful models that we have in the US right. Like maybe Fable 5 is more powerful, but it is pretty close. And one of the things that it sort of highlights is that, you know, up until this point, you can make a credible claim that Chinese models were maybe six to nine or so months behind US models. The most powerful Chinese models are some of the most powerful US models. It seems based on the release of Kimike 3, that that gap has shrunk pretty significantly. So now you might think about it on the order of two to four months, not six to nine months, which is a pretty dramatic change, I think. So I think there are some similarities here and differences between Deep Seq. So when Deep Seq came out, right, there was a lot of the coverage around that moment had to do with the fact that Deep SEQ claimed that it had trained this model that was very, very powerful for just a fraction of the computing resources that it took US AI labs. And this created like an immediate impact. There was a. There was a pretty dramatic reaction from the stock market. Since then, I think a lot of questions have been raised about exactly how Deepseek sort of calculated the total amount of compute costs for Deepseek. I think one of the noticeable things about the Deep Seq moment that is different than Kimi is that for DeepSeek R1, they were the first to market, in a sense, for reasoning models. And so a lot of the impact they had was the fact that you saw this new capability. But eventually the US models caught up and exceeded that capability. And that was relatively quick. The difference now is that what Kimi demonstrates is that the sort of things that we think are advantages for US labs, their ability to design good algorithms, train very powerful models, have very good pre training and post training, that China is able to do most of those things very well as well. And so in a lot of ways, I think it raises questions that are much more durable about what it means for China, US competition and what the future gap in capabilities between Chinese and US models might be. And I think in a lot of ways it could be more significant than what we saw with Deep Sea.
B
Yeah, but you know, just a few hours before recording, OSTP director Michael Kratios tweeted that quote. We have information that Moonshot AI distilled Anthropic's fable for the development of its K3 model. And then he went on to write that moonshot AI also acquired GB300 equipped servers and has access to GB300s in Thailand likely to train its AI models. So assuming the information that Michael Karacios is referencing is correct. How much of Kimi's performance can be explained by things like distillation and export control evasion?
A
Yeah, so I think it is almost. I think it's very likely that there was some amount of distillation involved in the training of Kimi. You know, to be fair, it appears that distillation is quite common across the AI industry generally. And we saw Elon Musk reference that in the trial with OpenAI where he said, you know, we implied that, that, you know, his companies were also distilling and that it is relatively commonplace across the industry. But I think it's a mistake in this case to over index on that. You know, generally distillation is useful in post training and that it does not provide you frontier capabilities. Right. You are able to sort of do more efficient post training, but it is not necessarily going to give you the same capabilities as the model you're distilling.
B
Yeah. And just for our listeners who might need a refresher on what distillation is, it's this technique where you have a smaller student model that is trained to mimic the outputs of a larger, more powerful teacher model. So in this case, Kimmy would have been the student model, Fable would have been the teacher model, and that is used to compress the teacher's model's knowledge into a more efficient form. And that's also illegal. Right. In this case, Anthropic prohibits its models outputs to train competing models and also prohibits the use of its models in China altogether. But yeah, sorry, I just wanted to add in that context for.
A
Yeah, of course, sometimes I assume our readers are all super versed in these AI terms, but it's good to go and explain. So as you explain, you can use distillation to learn from a larger model and make your training process more efficient. You'll have a higher hit rate with that data. That data is just more useful than sort of scraping the Internet or something like that. The problem is that distillation gives you some capabilities, but they're generally not sort of like the same capabilities as the model you're teaching or learning from. And in this particular case, because the performance of QEMI is so close to the frontier models, it is almost certainly the case that a lot of the improvements they made were in pre training and not necessary distillation. So probably there was some of that happening. But I think it is a mistake from a policy perspective to attribute most of this to distillation of US models. And it's important to take seriously that China has shown that they're able to do many of the things that the current frontier labs in the US are able to do in terms of pre training and algorithmic design and just knowing how to train powerful AI models.
B
Yeah. Well, then if we're talking about pre training, the second part of Michael Kratios tweet was about moonshot AI's access to Nvidia Blackwell chips that are supposedly banned in China. They shouldn't be allowed to enter China. And I think this has resulted in some calls for strengthening export controls. I mean, if we're saying that a lot of this is coming from their ability to pre train models in a way comparable to, to US AI labs, how much of that has to do with a failure to enforce export controls or a failure in export control policy in the first place?
A
Yeah, so I wouldn't necessarily maybe frame it like that, but we do know that export controls are having a bite on the Chinese AI industry, that they are compute constrained and that is affecting how they approach the projects that they undertake and how they're, how they're able to serve models. So we don't know exactly the details of the compute Kimi was trained on. It could very well be the case that they're trained on chips that, that are circumvented involve some circumvention of existing US export controls. We do know that they're really struggling to serve this model and so they temporarily cut off subscriptions. So that's an indication that there are compute constraints in China. I think it's very likely that if we continue to impose export controls, but not only that, to enforce them vigorously, that this could very well have an impact on the ability for China to train more powerful models in the future. But that does involve sort of a very consistent export control approach. It involves also, you know, dedicating the sort of budgetary resources and other resources you need to enforce things seriously and make sure that you're dealing with issues like diversion to third country, third countries and sort of other ways that we know China has used to circumvent export controls. And so I think one of the big open questions in the space right now is whether whether this shrinking gap that we've seen is, is going to remain small or if the sort of compute restrictions we've already seen affecting Chinese companies will eventually place limitations, absolute limitations, on how powerful the models they train in the future can be.
B
Yeah, and I want to push us on to talk a bit about the economic implications of this. Some are saying that Kimi has implications for the business model that frontier labs like OpenAI and Anthropic have. Why is that and what would those implications be?
A
Yeah, so the way I think about it, when you release a Frontier model, you're sort of banking on the window of time in which that model is sort of unique, has unique capabilities to sort of recoup a lot of the investment you've made. And we're talking investments that are in the multiple tens of millions of dollars sometimes for some Frontier models, maybe approaching or exceeding a billion dollars. And so you have this window in which, for example, I've developed Mythos or Fable, and I really want people to adopt it and learn how to use it. While it's the most powerful model, the issue that Kimi presents is now it is shortening that window where you have Frontier capabilities and a sort of Chinese that sort of commoditizes those capabilities to maybe two or three months. And this is, I think, extra problematic because we have seen US Regulatory actions on the other side that are shrinking that window on the front end. So for example, this idea that you would have voluntary pre testing with the US government, that's a process that will take time. And so now your window for sort of monetizing that Frontier model has closed in the other direction. And so I think this is going to be really challenging from the Frontier labs because that window of differentiation has shrunk on both ends. And so I think that we'll have to think about how to approach this from a policy perspective. But also the companies are going to think about this from how does this affect our plans to monetize these models and how does this affect, affect our overall long term business plans and approaches?
B
Yeah, well, I think there's also another way they might catch a break. Since the release of Kimik 3, there's been reporting on how both the US and China are considering restrictions on who can access or use these models and their weights. What do we know about these potential policies and the debate around them?
A
You know, I think the US has been considering for quite a while whether to have a policy to limit access to Chinese models. In some ways, I think, you know, this has been a discussion within the administration for a while. Certainly with the release of Kimi, those conversations were energized. I think as of, you know, the most recent reporting, it seems unlikely that something like this will happen imminently. But I think that certainly the US government is still thinking about it. There are lots of challenges to how you might do this right there. There's a relatively Straightforward thing the government can do where they can just say, we're not going to allow the US government to procure Chinese AI models, that's fine. The issue of is if you want more broadly to sort of restrict access to Chinese models in some way and they're just available for download on the Internet, how can you stop that? That seems incredibly difficult. I think there's also some pretty significant potential for backlash. We know that US models are very expensive. That oftentimes US companies think about using Chinese models as a way to manage their AI costs. And I think, you know, a lot of people are going to read it pretty cynically if, you know, the US government takes action to ban Chinese models, forcing companies and organizations to use more expensive US models is going to be seen as a sort of a subsidy for the frontier AI labs. And I think that that is a lot of the public might react negatively to that.
B
Yeah. And we're already seeing people, even within the administration react negatively. David Sachs, who's the former crypto and aizar, wrote on X that the leading closed labs already have a duopoly in terms of model revenue and they want the government to eliminate their open source competition, basically saying that that would be a really bad idea. But it'll be interesting to see how this continues to play out within the administration since I know there are people who are on both sides of the issue. Another thing I wanted to talk about was that on the Same day that Kimik 3 came out, Xi Jinping gave a speech at the World AI Conference in Shanghai. So there's this coincidence. Maybe not a coincidence, but what did he have to say about China's commitment to open source and approach to AI governance more broadly? I know there was also a Reuters article that came out at the start of July that reported that there are some Chinese authorities talking about potentially restricting overseas access to the best Chinese AI models, probably for reasons similar to when the US restricted access to its best AI models after Mythos came out. So I'm curious, what did she say and what do you think the implications are?
A
Yeah, this speech was really interesting in a lot of ways. A lot of the content was something that I could see sort of a U.S. politician writing. Right. It was a lot of emphasis on sort of broadening access to technology, AI technology supporting use in the. In the global south of AI technology, working cooperative, cooperatively with other countries to develop standards around AI. You know, the speech did really lean into this idea of open weights. So I think the key quote here is we should seize this rare historic Opportunity to courage, open source. I think on the. On the tension you flagged, I think both things can be true. I think it is almost certainly the case that China, like the US will think very hard about the safety and security implications of releasing the most powerful models openly. But I think that there are lots of models that are not the most powerful that can be released openly. And I think it's very clear that China has embraced the idea that, you know, there is some lever of advantage they can seize here with open models, that. That there are some American open models, but they're not nearly as good as the models that the Chinese companies are putting out for the most part, and that this provides them an opportunity to really diffuse Chinese technology all around the world. And I think that they will continue to do that. And I think, you know, like, also, you know, simultaneously or just before this World AA conference happened in Shanghai, there was an agreement to create the World Artificial Intelligence Cooperation Organization. This is something China has been thinking about for a long time. And 29 countries, you know, signed the initial agreement. And it's notable these are mostly countries in Africa, South America, and Asia, countries China has really been focusing on really engaging with about the adoption and diffusion of AI technology and that the US has not engaged in the same way. And so I think they're using this combination of engaging with parts of the world that, you know, may not be receiving the same level of attention from the US plus an open approach to model development to really sort of push Chinese technology to many places of the world that the US Technology isn't reaching in the same way.
B
Yeah, and that's been their strategy for some time. But as you said, it'll be interesting to see how that changes as their best models get better and better, impose more risk. But I want to wrap up our conversation on Kimi there. We'll be keeping track of any new developments over the coming weeks. But moving on to another big story from the past few weeks. On July 14, New York Governor Kathy Hochul signed the first statewide moratorium on data centers. The controversial executive order comes at a time of increasing public pushback to data centers and AI more broadly. But before we get into the reasons for that pushback, what exactly does this executive order that Governor Hochul signed do?
A
Yeah, so it's a pause on new data centers in the state while the Department of Public Service in New York evaluates the environmental impact of data centers. And so this involves sort of specific things that DPS needs to do related to construction operation of data centers, examining things like Energy demand, water use and quality, air quality, impacts on communities, noise levels, many of the things that we see. People complain a lot about data centers when they're contemplated for being built in their neighborhoods. And so until that happens and until some of this reporting and analysis is done, permitting for new data centers is being paused. And I think the final thing to note here is that the data center is defined as 50 megawatts of energy, so somewhat large in sort of energy usage terms. But since we often hear about data centers that might approach a full gigawatt of energy, sort of encompassing a large percentage of what we would traditionally call
B
data centers, and what reasons that HOQ provide to justify this, this ban.
A
Yeah, I think this is largely reflective of some of the concerns we've seen that have led to moratoriums and bans in many other parts of the country. They are concerns over whether this might have an impact on utility bills, raising energy and water prices, whether it will affect sort of the general environmental quality of the neighborhoods data centers are built at, including affecting natural resources, whether it will sort of have other kinds of environmental impacts, like not just consuming water, but affecting the quality of water, noise pollution, lighting pollution and things like that. So these are, like I said, concerns that have been expressed in many, many other parts of the country and have led to sort of actions, usually on a smaller scale. New York is by far the largest, but, but it has led to some number of moratoriums or, or bans on where data centers can be cited.
B
Yeah, and I believe the ban was also popular within New York, according to a poll released in June where 46% were in favor of a one year ban and 21% were against. And I guess it's notable here that Hochul was up for reelection in November. So there's also this political aspect of things as well, but the reception to the executive order has been pretty mixed. I would say. Some politicians, including our President and Energy Secretary Chris Wright, argued that the policy is a mistake, that the reasons that Governor Hochul provided are mistaken. What are some of the most common arguments you're seeing against this sort of policy?
A
Yeah, there are a few main arguments. I think one is that we have seen that data centers and the companies that build and operate them really take seriously managing some of the impacts that, that people are concerned about around energy and water use. They have implemented lots of technologies to sort of manage that. They, there's a, there's a pledge that the federal government has organized that data centers shouldn't raise electricity rates. And so I think the, the, the data here is very, very mixed about what impact data centers are actually having on electrical utility prices there are rising in many parts of the country, but it's unclear whether that rise is directly tied to data centers or not. And so there are questions about whether the reaction from communities is really tied to the actual impacts that data centers are having. And I think also there's an argument to be made that data centers deliver local benefits that people are often not thinking about when they pass bills and moratoria like this. And so one of those is in property tax. Property taxes. So, you know, CSIS is based in Washington, D.C. very close to Loudoun county, which is considered, you know, data center alley. And my understanding is that a majority, significantly more than a majority of the county's operating budget comes from property tax revenue and equipment tax revenue related to data centers. And so there is the concern that these types of policies don't fully account for the benefits that data centers provide to sort of local government budgets. In addition, you know, data centers have often committed to making other kinds of environmental or local community improvements relating from building of new government facilities to beautification projects, to improving the general infrastructure. And then I think the final thing is that, you know, throughout the US a lot of our utility infrastructure is in pretty bad shape. It's aging, it's in need of modernization, and there is an opportunity to use data centers to help modernize that. So the same kinds of improvements you need to make to support the utility use by data centers can also bring a host of general benefits to the community by upgrading the infrastructure in the area, generally making energy delivery more reliable in some cases. Right. If you have a reliable customer that can buy a certain amount of baseline power, it might actually lead to some reductions in energy rates. You can improve water infrastructure. And so that is something that people often talk about in the context of data centers as well. As we're seeing hundreds of billions of dollars being poured into data infrastructure. So that is money that can also be used to improve local infrastructure as well.
B
Yeah. And what do you see as, like the impact of this ban on the hyperscalers themselves and the building of AI data centers and other data centers. I think a lot of people I see arguing that, look, they can go and make or build these data centers elsewhere in the US there's quite a lot of land in other states as well. So how do you see the net impact on actually building data centers and making them operational?
A
Yeah, you know, when, when people think about where to build data centers, they're they're going to do sort of site analyses. They're going to look at a lot of different factors, and that includes is, is water available? Is power available? How expensive is it? How long does the permitting process take? What are the regulatory obstacles, sort of what are the design considerations related to local zoning laws? And how does that affect the total cost of the building? Is there the infrastructure to bring in all the heavy equipment we need? And then the companies are going to make decisions about where it's most efficient to build these out. And it appears that for right now, New York is not on the table. And so that's not going to stop people from building data centers. They're just going to sort of focus their attention on checking the viability of sites in other areas and maybe even in other countries. And so for right now, as this policy sort of unfolds at the local community level or at the state level, it's primarily, in my estimation, shifting investment in data centers from the places that have sort of put in moratorial bans to other places in the country.
B
Yeah. And one last question before we move on to our final topic here. To what extent do you feel like skepticism of these data centers and public pushback to these data centers is driven by AI skepticism? We talked before on this podcast about how the public is largely skeptical of the benefits that AI will bring to them, is largely skeptical of big Tech. What role do you think that plays in public support for data center moratoriums?
A
I think it's a pretty significant driving factor in a lot of ways. It's difficult for me to understand sort of the vehemence of anger towards data centers, unless you think about data centers as sort of the physical manifestation of AI and that a lot of people in this country think that AI is not making their life better, and oftentimes they're thinking that AI is making their life actively worse. And so I think this is a pretty big part of how people think about data centers. It's also a reason sort of a lot of people have sort of an opposition to data centers. Even if, say, a data center is not being built in their neighborhood, it's because they see data centers as sort of representing the AI industry. And there's a frustration and skepticism about AI technology fairly broadly. And now, pretty durably, this trust gap has existed for quite some time now. So I think that's a big piece of it.
B
Yeah, there's a big PR problem here. Well, I want to move us on to our last topic for today, which is on July 14, Google DeepMind CEO Demis Hassabis published a set of policy recommendations in a brief post titled A Framework for Frontier AI and the Dawning of a New Age. The essay was pretty widely praised by AI industry figures, including OpenAI's Sam Altman, Anthropic's Shaq Clark, and even Xai's Elon Musk, which is a bit of a surprise given his and Demis history. But could you start by just giving us a quick overview of what Demis wrote in this post?
A
Yeah, so his post really is motivated by the fact that like he, he really sees AI technology advancing quite quickly. He thinks that that level of advancement will continue in the near future. And so this raises really serious risks, you know, particularly in cybersecurity. We've seen that, you know, the, all the recent U.S. policy actions and a lot of the recent debate has been driven by Mythos and sort of its implications for cybersecurity, but also nuclear and biorisk. He's particularly worried about recursively self improving systems. So this idea that eventually AI will be able to significantly improve itself and that this will sort of lead to some sort of runaway, very rapid improvement in AI technology. And so this is the baseline that requires some sort of government action, some serious government action in some significant way. And so his proposal was to specifically establish a new standards body modeled on sort of some of the existing federal public private partnerships we've seen. He cites the Federal Industry Regulate Financial Industry Regulatory Authority or FINRA as a specific example of how that might work. And this is a self regulatory organization sort of with some blessing from the government to sort of regulate in the financial industry space. And that has some structure like a board with independent leading technical experts that establishes requirements for companies operating in that space. And so he's thinking about sort of an analogous body in the AI space that would be responsible for developing assessment protocols and then working with appropriate government organizations and federal agencies to conduct that testing. And then he, he thinks that over time this would be. So this would initially start as sort of like a voluntary mechanism sort of developed and funded by industry and that over time eventually it could be formalized into something that works more like a pre release requirement.
B
Right. And OpenAI and Anthropic have also put out policy proposals in recent months. OpenAI released a paper titled Industrial Policy for the Intelligence Age in April and Anthropic CEO Dario Amade published his essay titled Policy on the AI Exponential in June. A recent article by Axios I think pointed out something useful which is that these proposals share a lot in common with themis recent piece, while also diverging in maybe some crucial ways. What do you see as the most important points of overlap and disagreement between these three leading frontier AI labs?
A
Yeah, in a lot of ways it's interesting that this moment is happening right now considering how similar a lot of these proposals are. And I think we've observed for a long time that there is a fair amount of agreement about how to approach the issue of frontier models. And that involves sort of general agreement that there should be sort of independent testing prior to public release, that there should be some standards in place, some consistency in how we think about threats, and some emerging consensus around the types of issues that are most associated with national security, cyber bio and to some extent self recurring systems. And that there needs to be some sort of unified federal mechanism for sort of codifying some of these standards and certifying compliance and then somehow putting in the right restrictions or safeguards to limit access to frontier systems that are, that are deemed too dangerous. Where they differ. I think where they differ is something we see a lot in policymaking, which is that you can have a pretty broad understanding of what you want to do, but then when you try to get down to the specifics of how to make it work, it's often difficult to figure out sort of the right way to implement this. And so like, you know, the various proposals we've seen all sort of anchor on a different analog. So Demis Hasabis wrote about finra. I think Dario Amade has mentioned something like the faa, which is an agency that has the power to block model releases. Sam Altman has thought for a long time about something like the IEA iaea, which is the sort of International Atomic Energy sort of regulatory agency that helps manage nuclear power development across the world. And so these are all different approaches. And frankly, right, we just. In the US we have probably what, like two or three dozen different sort of regulatory approaches to dealing with technology that might have safety issues. And so it's easy to sort of differ on the specifics of implementation. And my personal view is sort of looking at a single one of these examples as the right way to structure sort of an approach to AI is probably the wrong thing to do. And we should probably sort of look at all of them and then mix and match the pieces that are most suited for AI from all of them and sort of bring that into something that's a little bit new and a little bit replicates things we've seen in the past.
B
Yeah, And a lot of the areas of overlap we're seeing between these frontier AI labs on policy are also evident in recent state bills that have been passed, I think so, in California, in New York, and most recently in Illinois, which you guys covered in our last news roundup. How closely do these state bills align with how AI CEOs are thinking about regulation?
A
Yeah, I think that in a lot of ways, despite all the rhetoric we've seen about the conflict between state laws and how the federal government is thinking about preemption that the state, the state laws are doing a lot of what sort of the. The founders of this country sort of intended, which is that they're sort of experimenting at a smaller scale with approaches to regulating AI and that the things that are working are sort of ending up in law, and the things that are working, not working, or that didn't take the right direction ended up sort of not passing. And that this is helping us sort of zero in on areas where there is broad consensus and also helping us figuring out the mechanisms for how to execute on the things we want to execute on. And so, you know, across these big frontier AI bills in California, New York and Illinois, we see some agreement on independent testing, on the kinds of threats, threats that we should be worried about, on the sort of idea of certain kinds of transparency, on the idea that we should really be thinking about this in the context of frontier models and sort of narrowly scope it to the most powerful models. I think the area, I wouldn't necessarily call it a disagreement, but the area where probably there's the biggest difference is with what sort of the labs are proposing is that there are a lot of advantages to doing this at the federal level. And I think even a lot of the people who have advocated for these bills at the state level would agree with that. So if this happens at the federal level, you are able to employ the expertise and tools that the federal government have. So they have a lot of expertise in say, assessing nuclear weapons and cyber weapons that might not exist at the state level. They're able to deal with classified information much more easily, and they're able to support sort of national markets. So you think about how can we support some sort of ecosystem for independent testing. It's much easier and more certain for that ecosystem to develop if that is a federal mandate. Right. And is consistent across the country versus happening on a state by state basis. So I think even a lot of the advocates at the state level would love to see some of this past the federal level, that it would solve a lot of problems with implementation that we've figured out to some extent at the state level, but would be better if wouldn't sort of raise the same level types of concerns if we were able to do it sort of at the national level.
B
Yeah. Well, I think that's a good place for us to wrap up today. This was a longer episode than we traditionally do. I think we're almost at 50 minutes, so thanks to those in our audience who are who are still listening all the way to the end. And thank you Alok, for providing so much analysis of all these different policy stories that we've seen over the past two weeks. And one last note, we'd love any thoughts or feedback those in the audience might have for us. And you can reach us@aipolicypodcastsis.org, so if you have anything, any thoughts that you've been meaning to share with us but you haven't been able to, well, there's an email address that you can share those thoughts with us now. Thanks Alok.
A
Thanks for listening to this episode of the AI Policy Podcast. If you like what you heard, there's an easy way for you to help us. Please give us a five star review on your favorite podcast platform. Subscribe and tell your friends. It really helps when you spread the word. This podcast was produced by Sarah Baker and Matt Mand. See you next time.
The AI Policy Podcast – Episode Summary
Episode Title: OpenAI Models Hack Hugging Face, Kimi K3’s “DeepSeek Moment,” and New York’s Data Center Moratorium
Date: July 23, 2026
Host: Aalok Mehta (Director, Wadhwani AI Center, CSIS)
Guest/Co-Host: Matt Mand (Researcher, Wadhwani AI Center, CSIS)
This episode dives into three major AI policy events from the past weeks:
The conversation is wide-ranging, touching on national security, economic models, open source versus closed AI, regulatory approaches, and the evolving politics of AI infrastructure.
[00:16–08:25]
[08:25–26:38]
Kimi K3’s Breakthrough:
DeepSeek Comparison:
Distillation & Export Control Controversy:
Export Controls:
Economic Implications for US Labs:
Debate on Open Model Access:
China’s Open Source and Global Outreach:
[26:38–37:19]
[37:19–47:08]
AI escapes test sandbox, hacks independent infrastructure:
“We have now entered an era where AI tools are basically a requirement if you are engaging in both cyber attack and cyber defense.” — Aalok, 05:23
On the core policy challenge of AI for dual-use security: “If we’re going to have frontier models…not to be able to assist in hacking, that also is going to make them less useful for cyber defense. … Right now that balance has gone too much in one direction.” — Aalok, 07:27
On US-China competition and the shrinking capability gap:
“China is able to do most of those things very well as well … it raises questions that are much more durable about … the future gap in capabilities between Chinese and US models.” — Aalok, 09:54
Kimi's economic impact on US AI labs:
“That window of differentiation has shrunk on both ends … this is going to be really challenging for the frontier labs.” — Aalok, 19:34
On open source, global tech diffusion, and China’s strategy:
“They’re using this combination of engaging with parts of the world … plus an open approach … to really sort of push Chinese technology to many places … that US technology isn’t reaching in the same way.” — Aalok, 25:24
Data centers as a symbol of AI unease:
“Data centers as sort of the physical manifestation of AI … a lot of people in this country think that AI is not making their life better, and oftentimes … making their life actively worse.” — Aalok, 36:14
On consensus among AI labs for government oversight:
“There should be independent testing, some standards in place, some consistency in how we think about threats … a unified federal mechanism.” — Aalok, 40:51
The conversation is analytic, policy-oriented, and strategic, with measured, nonpartisan perspectives reflecting CSIS’s think tank mission. The hosts are candid about uncertainties in technology & regulation, openly weigh both sides of contentious issues, and frequently contextualize developments within broader global and economic trends.
For questions or feedback, listeners are invited to email the show at aipolicypodcast@csis.org.