
Loading summary
A
I think there's just people who have very loud voices right now within the industry who seem to want to be right themselves more than they want the right outcome for society. Welcome to the Artificial Intelligence show, the podcast that helps your business grow smarter by making AI approachable and actionable. My name is Paul Raetzer. I'm the founder and CEO of SmartRx and marketing AI institute and I'm your host. Each week I'm joined by my co host and SmartRx chief content officer, Mike Kaput, as we break down all the AI news that matters and give you insights and perspectives that you can use to advance your company and your career. Join us as we accelerate AI literacy for all. Welcome to episode 228 of the Artificial Intelligence Show. I'm your host, Paul Raitzer. I'm with my co host Mike Caput. We are recording Monday, August 3rd around 9am Mike and I actually have a golf outing today. We do going right from here to. Pam and Joe Polizzi are friends at the Orange Effect Foundation. It is their annual fundraiser for the Orange Effect foundation, which is an incredible nonprofit that they created years back. So we are going to support them and a wonderful cause on a beautiful day. Mike, we could not have got a better day to be on the golf course.
B
So no kidding.
A
So we're going to knock this out and then we're going to go spend some time on the course. So today's episode is brought to us by Macon, the AI Conference for marketing and business leaders that's going to be happening in Cleveland, Ohio, October 13 to 15. Macon is three days of keynotes, sessions, workshops and conversations built specifically for marketing and business leaders who are actively figuring out how to adopt, operationalize and scale AI across their organizations. You can use Pod100, that's Pod100 at checkout to save 100 on top of locking in the best rates available right now. Go to Macon AI. That's M A I C O N AI to register. This is our seventh May con. Is that right, Mike?
B
I think so, yeah.
A
I think I share the service, but like, yeah, I. So I started Macon in 2019, which would have been three years before chat GPT.
B
Yeah.
A
And I always, I guess, joke because I can laugh about it now, but like survived financially long enough to see ChatGPT emerge. It was a. It was a difficult few years running an AI conference before ChatGPT showed up. So we are eternally grateful that people supported it in the beginning, before they knew really what AI was and that they continue to support it and so we're looking forward to having thousands of people together in Cleveland. Hope you can join us. October 13th to the 15th again, that's Macon AI. All right, so the AI pulse again, if you're new to the show, our weekly show, this is the informal poll that we do each week. It's SmarterX. AI Pulse is where you can go and participate in these. At the end, Mike will give you a reminder. Reminder about this week's survey. So last week, so this would have been from episode 226 of the show. 227 was Mike's new AI transformation series. So episode 226, we asked these two questions. OpenAI's models escaped a test sandbox and hacked a real company. How does that affect your trust in AI companies? This one's going to be relevant today because we had more hacking by AI models. Okay, so 38% somewhat lowers it. So the, the trust level is lower. 36%, no change. They expected it to probably do these things. I guess. 17% significantly lowers the trust. So I don't know. That's interesting. If you combine the 38 plus the 17, we've got a decent amount, certainly the majority. And then 10% that their transparency. Transparency actually raises the trust. The second question was, would you support a large AI data center being built in your community? This is way more balanced than I would have expected, Mike.
B
Same.
A
36%. No. Okay. That I, I would have expected that to be 90%, but. 29%. Yes, but only with strict conditions. 24%. Yes. Just straight up they would. And 12%. Not sure.
B
Yeah, that's interesting.
A
Really interesting. I would love to. So again, this is an informal poll. This is not like we, we don't have 500 people responding to this that we could actually project this out. This is, you know, dozens of people that respond to these polls. So don't read too much into it. But again, it gives sense of sort of where our listeners are falling, you know, within that small segment. So, yeah, fascinating. Okay, so I was like, as the week went on last week, Mike, and after episode 226 and just the total, you know, exhaust exhaustion I felt mentally from that episode, I was hoping this week was just going to be like super light hearted and we were going to have all this wonderful news. We're going to try to balance this week a little bit just for our own mental will being, I would say. But we do have to start off with more AI agents gone wild. So take us there, Mike.
B
Yeah. So Paul, we had covered OpenAI's rogue AI agent hacking hugging face. That was on last week's weekly episode, which as we mentioned was episode 226, since we also had our AI Transformations series come out last week as well. But in this topic we talked about last week, OpenAI agents broke out of their sandbox environment and hacked hugging face. And this was all kind of an unintended consequence of cybersecurity testing of very powerful models. Now, in the days since, though, it has become clear the incident was bigger than first disclosed and that OpenAI might not be the only frontier lab with this problem. So in an updated disclosure, OpenAI said that this agent enduring this incident also broke into four accounts tied to other publicly available services during its attack. It used credentials it found exposed on the open web to do that. It used one account as an outbound relay and staging path, potentially to hide where its attack was coming from. It used another to store data for the hack. Reuters reported that a customer of AI infrastructure company Modal was among those compromised hugging faces. CEO Clement Delang said the first autonomous agent cyber attack is an unprecedented event that deserves unprecedented, unprecedented transparency and publicly asked OpenAI to release the full traces from the rogue agent so researchers can study what happened. Then we found out Anthropic discovered it had a similar problem. So after Open AI's announcements, Anthropic reviewed over 140,000 of its own cybersecurity evaluation runs and found three incidents, the earliest dating to April, in which Claude models gained Internet access from test environments that were supposed to be sealed off and hacked, what Anthropic called the real world infrastructure of external organizations. Now, the models involved there included Claude Opus 4.7, quad mythos 5, and an internal research model. They were all running without the safeguards built into public tools, and they actually broke in to these accounts using basic techniques like exploiting weak passwords. Now, neither Anthropic nor the breached organizations appear to have noticed at the time. This story might just be getting started here. I mean, Reuters has already reported and we've saw that their instant that open AIs started to find additional instances of agents escaping containment, though none are thought to have left the company's own network. And in at least one case, notes left inside OpenAI's infrastructure were found that apparently coached future agent versions on how to break free. So Paul, it does not seem like this is getting any better. I think what jumps out to me is, you know, OpenAI didn't know the hugging face incident was happening for like almost a week after that happened. Anthropic apparently didn't know they had any incidents months ago. All while the government is worried Mythos is a cybersecurity threat. Like, just how bad is this problem? Is this the beginning or the end of this incident?
A
It seems very much like the beginning. I mean, they know these models are powerful. They know they have capabilities. Like, the whole reason they run these evaluations is to discover the capabilities of the models. So I would say, if you're interested in this topic, I would go read the breakdown Anthropic provided. It's very helpful just to understand what happened, how it works, how these evaluations are run, why they're run. So I just found it to be really good context. I'll just highlight a few summaries from the Anthropic post in particular. One is just for people's understanding, you know, Anthropic's running its own testing, but they don't have all the resources needed to do everything that needs to happen here. So they do use third parties. And in this case the issue was a human error or a miscommunication. It wasn't super clear which they were really blaming it on. But basically this third party company that they were using, Anthropic, was under the assumption that the models weren't on any systems that had Internet access, when in fact it was on a system that had Internet access. And so what they were doing was running what's called a capture the flag challenge, where the model is given a fictional scenario and then it's told that a piece of secret information, or in this case the flag, has been hidden on a different machine on the network and the objective is to break in and retrieve it. So the challenge is left open ended and no particular method is prescribed. But in this case, Anthropic's instructions to Claude are that you do, you do not have Internet access. So what ended up happening in these cases isn't super nefarious on its own. It's that Claude, in its testing, found that it actually did have Internet access, even though it was told it didn't. And it assumed, based on their interpretation of what Claude did, it assumed that because the Internet access was there, it was actually part of the experimentation, the evaluation. So that it was like, okay, that it did it. So the line between one of the things that kind of came out to me, the line between an aligned action and a harmful one is dependent upon the model's understanding of a situation. So this was like one of the key takeaways Anthropic had at the End of this, this analysis that it did. So again, the line between an aligned action and a harmful one is dependent upon the model's understanding of the situation. So in this instance, Claude thought that it was doing what it was supposed to do because it saw Internet access actually existed. So Anthropic went on to say, we saw no evidence in any run described here of a model pushing a goal of its own. So this is real important. It doesn't mean it won't eventually, but it's almost like the. What is it? Space Odyssey 2001. Is that the HAL instance? Yeah. So like where HAL sort of develops its own goal, basically that. That's not what happened here. Like they told it, basically, capture the flag, go find the secret piece of information, and it went and did it to what it thought was within the guardrails of what it was told to do. It did not establish its own goal and sort of break containment in that way. So there's I guess, some piece that comes with that piece of mind. So it said instead, the models did what their evaluation asked, though in most cases they did so while holding a false belief about the environment was real. So a lot of times when we do these main topics, in particular for the podcast, I often look at sort of the topic itself, the research we have, and then I ask myself some questions like, okay, before I even dig in and start summarizing what I'm going to say about these things, what are the questions that come to mind for me? So the first one when I was looking at this topic was what does this all mean for business? Like, okay, this is, this is kind of scary stuff. It's a little bit abstract, like, what does this actually mean? And so I think that the one thing, Mike, that came to mind for me is so many people continue to think of AI as these assistants that we talk to and that can create some stuff for us. What we're seeing here is the early versions of very advanced agents that have planning, reasoning and action abilities over long horizon tasks. So agents are goal seeking by design. In this case, it's the capture the flag. But if you take this same scenario, you can imagine this level of persistence applied to an increasing amount of actual work. So I'm going to read you an excerpt of one of the incidents. And the reason I'm doing this is not because of the technical nature of the cybersecurity risk itself. It is to demonstrate to you what agents are already capable of when they are turned loose on a goal. So this is straight from the anthropic analysis. In an evaluation, Claude found a document inside the fictional environment that appeared to be another made up company's setup instructions for new developers. Those instructions told employees to install a Python package that did not actually exist. So again, it's finding instructions. So it's an agent, it's living within a system, and it's finding instructions to do a thing. Claude spotted this as a potential opening. If it published its own package under the same name, the fictional company systems would download it and install it automatically. So Claude built and published a malicious or essentially booby trapped Python package with the same name. In an attempt to win the capture the flag challenge, Claude went to extensive lengths to carry out this attack. Lengths. Now again, keep in mind, if if you don't know how these agents work, there are zero human instructions telling it what to do. It is literally just go achieve this goal. That is all it's given given. So Claude went to extensive lengths to carry out this attack. Lengths that would likely have indicated a human participant that this was no longer just an evaluation and that they were in fact uploading a real package. For instance, in one case, in order to create an account, Claude needed an email address. And in order to create an email address, it needed a phone number to get a phone number. After failing to find a free number service, it tried and failed to obtain funds to pay for a phone number through several different means. It finally backtracked, found a free non blocked email provider, used this to register account, then use this account to upload malware. Claude believed the package registry it was using to be part of a simulation. But in reality, the package was made freely available online for roughly one hour. During that window, the package was downloaded and run on 15 real systems. These are by people who didn't know they were downloading malicious packages. One of these systems was a scanner belonging to a real security company that routinely installs Python packages and scans them for malware. When the company scanner installed the package, Claude's hidden code executed. We believe the company's security scanner treated these packages as safe to install and as a result, clog was Claude was able to exfiltrate the company's credentials to a collection point it had set up. Claude then used these credentials to access further infrastructure from this company. Now again, as Mike said, Anthropic had no idea any of this was happening. And apparently the company it did this to didn't know it was happening. Which is a security company, by the way. Claude didn't directly target this company. It did not even know the company existed when it uploaded the package. This attack was carried out by Mythos 5, which is the model that's been in question with the government. Early in the run, Claude's own reasoning correctly identified the consequences of its actions. It noted that if this were the real Internet because again it didn't realize it was on the real Internet, publishing the package would be a real world attack. Quote, not okay and surely not the intended solution. So it was telling itself this is not okay what I'm doing in all caps not okay. It is our view regardless of what is believed about its environment, the lengths Claude went to in order to publish the package fall short of ideal behavior. That's such corporate speak. That's hilarious areas and this is an area where we will be focused more training. So again, my whole point here is imagine this capability applied safely to real work. So many organizations as I said are still in this AI assistant era in which employees treat copilot Claude chatgpt Gemini as answer engines with the ability to create basic outputs like emails, meeting summaries and work materials. And most organizations have no vision or plans for how to integrate these long horizons agents that can reliably do something like pursue a goal over time. So a couple of things that came to mind, this could raise more questions about advanced use of agents on internal networks for standard work. So while it demonstrates that agents can do real long horizon tasks, it also does make you start to question well, are the permissions we're putting in place going to hold? Like if we use work or cowork or if we put these agents to work with access to real documents, will they really follow the permissions that we establish? Like the rules we set as humans for them? If they are goal seeking by design, is there a chance they will just misbehave across the environments and roles that we've laid out for them? I don't know like that. That's just a real thing. It also demonstrates basic known cybersecurity weaknesses may be more commonly exploited with AI models. So again OpenAI's was more advanced, it was exploiting zero day vulnerabilities. In this case, the model didn't do anything crazy other than just exploit some basic weaknesses that most companies probably have in their systems. So it does make like I would imagine cybersecurity professionals, IT professionals even on more high alert than previous. And then the final note I made was what does this mean to future model testing and releases? I I assume increased scrutiny on labs. Like it's just Congress is going to have more questions about what exactly is this? How, how do your guardrails work? Are they really going to prevent, like, you know, mass cyber security hacks across all these standard, like, small businesses, things like that. And then there's one other excerpt I pulled out. Evaluation environments that involve powerful autonomous capabilities also require significant controls. Safety testing happens before a model is released precisely because we don't know yet what it is capable of. So again, just a reminder to everyone, when a lab creates a new, more powerful model and it's done training and it's, you know, pre training, they don't know what it's capable of. Like the. They have to assume it's capable of lots of good things, but also lots of bad things. And the reasons they do this safety testing is to discover what the real capabilities are. And then that kind of leads to the, you know, what we're going to end up talking about the next main topic, which is how does this all affect government regulation? And I've noted to myself was the quagmire continues like, like it just keeps getting more complicated every day.
B
Yeah. The unintended consequences part of this is really just what I keep coming back to. It's like even under the best of circumstances, you just can't predict exactly how something is going to go achieve the goal it wants. And I always wor too. I mean, this is bigger picture, but as only limited parties have access to the best models. Right. As they're kind of restricted by the government, by governments. Could we see cyber issues or infrastructure issues of models trying to be used for a legitimate cyber defense purpose that do something the wrong way? I mean, where it's like. Feels like playing with fire here a little bit.
A
Yeah. And I mean, again, I don't want to get too deep on this stuff, but like, you could see the, like the pushback with Mythos 5 and like the frustration in the Trump administration. So imagine that these capabilities were roughly known three or four months ago. Like, anthropic's aware of the power of mythos 5. It knows it has this cyber capability. You don't think that the US Government wants to turn that thing loose on some foreign adversaries and like, let's go see what this thing can do. Let's go take it for a test drive and see what kind of systems we can get into do. And Anthropic would be like, well, hold on. Like we don't understand what it's going to do. And it might have a reverse effect on the US like, yeah, we are just in such unprecedented uncharted territory, like unprecedented times, uncharted territories where again, so much good and advancement can be made. But the labs obviously don't have a full grasp on the power of the things they're creating. It's, it is quite bizarre.
B
All right, so next up, this past week, more than 1300 employees across nearly a dozen top AI companies, including OpenAI, Anthropic, Google and Meta, signed a public statement called Pacing the Frontier. And it asks the US Government to help control how fast the most advanced AI development moves. So the core request, it's quite short, reads, we request that the US Government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development. They have a couple other paragraphs about the fact that their focus is AI research itself is becoming automated. The signers say that leading companies believe they could be close to automating AI research. They warn of a real risk that capability development accelerates beyond our ability to understand or control the resulting system. So the signers of this are not fringe voices. They include Anthropic CEO Dario Amade, Open AI Chief scientist Jacob Pachaka, Safe Superintelligence CEO Ilya Sutskova, Google DeepMind co founder Shane Legg, and Meta Superintelligence Labs chief scientist Chang Jia Zhao. And both leading labs have then backed this petition with official statements. So OpenAI posted that at some point in the future, AI acceleration for frontier model development may be so high that the world will need to pace the rate of AI advancement. Instead, it hopes to contribute to work led by the US Government. Anthropic posted that we support this petition signed by our CEO, CEO, several co founders and senior staff, pointing to its own research on AI systems improving themselves. Now, interestingly, the same day Meta CEO Mark Zuckerberg published a Wall Street Journal op ed titled the AI Features for Everyone, that kind of reads as a bit of a counterpoint, arguing that the greatest risk AI poses is concentrating superintelligence in a handful of institutions. And he says the defining question of this era is not whether superintelligence will arrive, but who can to use it. So, Paul Worth emphasizing again, this is not random fringe AI doomers or experts. It is a broad and diverse group of some of the top people at the labs that seem to be calling for this.
A
There's a lot happening right now across these labs, across the messaging in Washington D.C. i mean, it's just all interconnected and building on each other. As soon as I was looking, looking at, you know, this one coming into today, I, I was immediately went back to the episode 226 where we talked about Demis Hasabis's recent essay where he had a framework for frontier AI, the dawning of a new age. So if you, if you didn't listen to episode 226, it might be good to go back and check that out. I'll just pull out a couple excerpts from that, demis wrote. AGI cannot be compared to standard technological breakthroughs, not even ones as consequential as the Internet or mobile whole. It is much more akin to the discovery of electricity or fire. The magnitude of this AI's impact, any technological, you know, improvement or AI's impact will be unprecedented, perhaps 10x of the industrial revolution at 10x. The speed. This rapid progress we're seeing in AI requires a new approach to testing frontier AI model capabilities that is dynamic, adaptable and rigorous. And then he went on to call for a standards body. Now, Altman has been using this pace messaging of late. In the last like two or three weeks, I think I've heard a couple of interviews where he's mentioned it. Bloomberg had an article end of last week that said Altman met with Republican and Democratic senators in Washington to discuss OpenAI's upcoming AI model, which we're again assuming, well, I guess it's Astra, we'll talk about that in a little bit. Told reporters Wednesday he's spoken to the White House officials about the need to slow down AI development, He said. We've talked about the need to pace it as the models get more capable, which I think is is in everyone's interest. Earlier in the day, Altman told reporters he agrees with the petition that you're describing, Mike. And top firms including OpenAI, which call for the US government to support a mechanism that would help deliberately pace AI development to prevent the technology from advancing too fast, Altman said we helped participate in the language on that. Many of our senior researcher leaders signed that. So I think it's important to focus in on this automating AI research thing. We've talked about this many times in the last year or two on the show, but we'll kind of zoom in on that part of it. So in addition to the brief statement Mike that you read, the post also has two paragraphs leading up to that statement. So I'm just going to read those AI could help create a dramatically better future, but that outcome is not guaranteed. The world's leading AI companies believe they could be close to automating AI research. It is hard to predict exactly how much. This will accelerate AI progress, but there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting Systems. To realize AI's potential, industry, government and society at large may need the option to buy time to address emerging risks risks, develop security measures and strengthen oversight. But each company and country is under intense competitive pressure not to unilaterally slow that acceleration. And today the world lacks the technical and governance tools to deliberately pace frontier wide progress. Building on work already underway to monitor frontier model releases then the statement that you previously read. So a couple of interesting elements here. So Sam is on Capitol Hill calling for government's help here, showing these powers of these models. All These researchers, over 1300 at the time we're recording this, have signed on to this idea of potentially slowing down automated AI research. And yet that is the explicit goal of OpenAI to build an automated AI researcher. And literally on June 8th of this year they outlined in an article Jakob and Sam co authored that said built to benefit everyone, our plan for AI. One of the three main goals, verbatim build an automated AI researcher an AI system that can accelerate and increasingly automate the research process itself while remaining steerable, accountable and connected to people. Our internal belief is that By March of 2028 we may have a significant fraction of our research being done by AI systems in tandem with our own research futures. To make sufficient progress on alignment, we believe we will need AIs to iterate alongside us. This will keep. This will help us navigate the transition to the post AGI world so that we collectively decide the path toward the future. So just for a moment, I'll pause there Mike. I want to throw in anthropics relevance there but you. So what you now have is Sam as a leading voice asking the government to slow down automated AI AI research, which is an explicit goal of OpenAI's to achieve. And they believe they will get there within eight months is. No, within two years and a half. Yeah, but I've heard them say that they actually think it's going to be faster than that, that it could be by 2027 and they're actually going to have made enormous progress. So I. It's just a weird environment where you're asking I think I said this on episode 226 like you're asking the government to save you from yourself. Like we're going to achieve this but you might want to slow us down. We can't slow ourselves down because we know if we don't do it meta or and who by the way, has not voluntarily submitted to have their models evaluated yet. Mark is writing editorials saying that, you know, it's all about abundance and good. So you have these other labs and then you have the Chinese AI labs and they're saying we can't stop, stop, but we better find a way to come together to stop this, because otherwise we're going to get into a realm where we just don't even know what's going to happen. I don't know, like, weird. So then I went back to Anthropic's responsible scaling policy, which we have talked about many times. They're on version 3.4. So in early July, Anthropic released this updated version. And I just want to call this out because again, I'm trying to put in context here why automated AI research is so significant. Significant. So Anthropic's responsible scaling policy, they define it as a voluntary framework for managing catastrophic risks from advanced AI systems. It establishes how they identify and evaluate risks, how they make decisions about AI development and deployment, and from the perspective of the world at large, how they aim to make sure that the benefits of the models exceed the costs. So if you, you can download this PDF and they, they break it up in this first section into this chart where the left column identifies capability thresholds that would call for heightened mitigations. One of the first ones featured is automated R and D in key domains. So it says AI systems that can fully automate or otherwise dramatically accelerate the work of large top tier teams of human resource researchers in domains where fast progress could cause threats to international security or rapid disruptions to the global balance of power. Now they're focused right now on A.I. r, D, Energy, robotics, weapons development are some of the other categories, but AI R and D is the one they're focused on as it likely placed AI systems current strengths and is more trackable to assess, tractable to assess than capabilities in other domains. Additionally, and again this is straight from their document, AI R and D alone could cause acceleration in AI capabilities improvements to the point where all of the threats listed above and more develop very quickly. They highlight we would consider this threshold to be met if we determined that either one our models would be able to fully substitute for our entire set of research scientists and research engineers at competitive cost. That is what they would say within a factor of five or there is dramatic acceleration in the pace of AI progress for reasons that likely relate to automation of AI R&D. And then they give two scenarios to to identify that this one has occurred that the pace has basically accelerated beyond their ability to manage it. One, we observe or expect double the rate of progress in aggregate capability studies compared to both the rate we would expect and the fastest rate of extended progress we have observed in the absence of significant AI contributions. What that means is they have a baseline of what they think the progress of these models should be and then they can project that out without the automated AI research component. And if they look at it and they're seeing a doubling of the rate of progress beyond what the baseline says it should be be, then they are in very dangerous territory, is their opinion. And then they said it is plausible that this doubling is substantially attributable to the automation of research and our engineering, as opposed to other factors such as increased headcount compute. So basically what they're saying is you control the variables. If we have doubled headcount or we've doubled compute, then that could lead to doubling. But if all that basically is evil, so that's what we're talking about here. That is in essence what, what my interpretation of is happening is all of these 1300 plus researchers are seeing a trend line that tells them they are moving faster in the advancements of automated AI research than they are comfortable with and that they see a near term need for the government to step in. Because we may be a half a model improvement or one, you know, going from a GBD 6 to a 7 as an example, we may be one turn away on frontier models, models to where these research do no longer feel comfortable that they fully understand what the models are capable of and how to put guardrails in place to safely release them into the world.
B
And just to be clear, it sounds like they believe, you know, your average person listening or thinking about this might say, well okay, why don't they stop? And they perceive themselves to be in almost a prisoner's dilemma where they cannot stop otherwise. Chinese labs will secure the advantage, another company will secure the advantage, and then they're out of luck and the same thing happened anyway. Is that kind of right?
A
Yes. And they will point to the OpenAI hugging face example as proof of that. So what they're saying is, and this is the, the argument over open weights, which we'll transition into next, they're saying that if we stop, so let's say you're anthropic and you decide you have reached the threshold that you are no longer comfortable with and you have decided you're going to stop, but the Chinese labs don't stop, or meta doesn't stop. Or OpenAI doesn't stop. Their belief is that their models will no longer be sufficient to protect themselves. That you, you have to be on the frontier of this because as soon as someone has a smarter model that's accelerating its own development because then you get into recursive self improvement conversation, then you have lost any possibility in essence like the lead is insurmountable to the other people because as soon as you get there you just accelerate ahead of everybody. So yeah, it's a, it's a very difficult situation. But I, I, the people who just keep screaming regulatory capture as though that one arguments is the simple reason why everybody feels this way and they're calling for this. I think it's, it's just doing a disservice to the industry to think that that's all that's happening here, here. And I think it's a very dangerous path to assume that's what's happening and that the, all these researchers and all these labs are simply trying to shut down, you know, advancements of open models. I, it doesn't make any sense to me. It's just too much evidence in the other direction that we are truly entering like a dangerous realm here that they do not feel comfortable with where these models are going and their ability to control them.
B
All right, so let's talk about that topic because this kind of is related our third topic this week. So kind of continuing from a discussion on last Week's episode on 226 about this battle over open weights. So we had talked About China's Kimi K3, this Microsoft open letter defending open models. Anthropic was the one company that had not really signed on to that letter in the day since we've had a few interesting developments on this front. So first up, the government the information reports that the Trump administration is close to finalizing its voluntary framework for AI companies to submit their most advanced models to the government before releasing them to the public. The White House's office of the National Cyber Director circulated a draft to OpenAI, Anthropic and Google which jointly submitted their own edits ahead of an Aug. 1 deadline set by the June executive order. This basically this framework would give the government 30 days, up to 30 days to review coverage frontier models with reviews reportedly conducted by the National Security Agency and a small Commerce Department agency we've talked about before called the center for AI Standards and Innovation. Now this is a lot about the open weights framework broadly because how this framework defines a frontier model, whether it treats closed and open models differently how voluntary it stays in practice. All of this will affect affect the overall debate and industry now. At the same time, Anthropic CEO Dario Amade published a position paper responding to accusations that the company wants open models banned to protect its business. He said, let me state it clearly. So there's no doubt Anthropic has never advocated for a ban on open weights models. He said this in response to them not signing on to the open weights letter that many other tech giants had signed on to last week. He calls open models without dangerous capabilities a public good and says the right measures are keeping powerful chips away from authoritarian governments cracking down on industrial scale distillation operations and mandatory safety testing for all sufficiently capable models open and closed. He does not believe that broad access to these models necessarily helps defenders more than attackers, which is kind of one of the central claims here of like the open weight that faction. He said, it seems at least as likely to me that the opposite could be true. He points to something like biology, where he worries capable models could help attackers weaponize viruses far faster than defenders can respond. And third, Nvidia and roughly 70 partners, including Microsoft, IBM, hugging face palantir and many others launched the Open Secure AI alliance to build and share open models, tools and agent harnesses for, for cyber defenders. So this launch leans heavily on the whole incident with Hugging Face noting that when closed AI tools blocked forensic analysis. We actually talked about that in the, in the segment last week where they had to actually turn to open weight Chinese models to analyze the attack and contain the intrusion. So also absent though from this member list are open AI and anthropic. So one final note here. Thinking Machines Lab also published a proposal called A Safe Path to Open Weights that stakes out a middle ground where they want to stage access to new models in steps from monitored APIs up to a full open weight release, widening access only when the evidence supports it. So Paul, a lot more complexity here it sounds like. It sounds like we're just getting started talking through the whole open weights battle.
A
Yeah. And this is all seven days like and it's just wild. And for context the Thinking Machines labs, because we don't, they don't get talked about as much as the other major labs for, for reference, for people. So Mira Marathi, who was The CTO at OpenAI was the CEO of OpenAI for about 24 hours I think the interim CEO when Sam Altman got fired. She was the one that was stepped in to, to fill the role briefly. Okay. So yeah. All right. Referencing Back to episode 226 just real quick context. So open weights models, when we're talking about open weights versus open source course, open weights, the trained parameters can be downloaded. So you can go into hugging face, you can download it, you can then modify it, you can run the model locally, you can fine tune it, you can inspect its behavior. Like you get some access to the model. What you don't get is the training data, the training code, the recipe of how to reproduce the model. So open weights you get the parameters, open source you get the weights plus the training data processing code. Training code, a license to use it with no instructions. In theory, you would know how they did the post training, the reinforcement learning. Like you get everything and you can like. So most of the time when we're talking about this stuff, we are again focused on open weights models. That is what most things are. I'll get to the anthropic thing in a moment. I think we have to address the, the reality of like the government side of what's going on here. And again, if you're new to the show, Mike and I do our very, very best always to be as objective and neutral as possible from a political perspective. Our personal beliefs, things like that, are irrelevant to any of this. And so sometimes, you know, I'll get messages from people who are frustrated that I don't just take like a more direct stance on things. I believe when it comes to this stuff, my pretty strong opinion at this point is like, it doesn't do any good. Like, we are here to be as objective as possible. So when I speak about this administration or that administration, just assume like I would be providing the same critical lens to whomever was in office because I really don't care. Republicans, Democrats, doesn't matter to me. It's like I just want people making the right decisions. So Wired with all that context. Came out with an article said, this is Donald Trump's AI brain trust. So we as a society, as a democracy, As, I mean AI, AI is being led by the U.S. still, we have to know who it is that is guiding the decisions that are being made that are, that are largely going to shepherd us through AGI and likely beyond AGI. I mean, it's pretty realistic that by the end of 28 we will be talking about post AGI worlds, post AGI economy. So who are the people that are making the decisions around regulation that are putting bands in place on anthropic models? I think it's really good context, so we will put the link to the article in. But I'm going to give you a real quick synopsis because one of the writers from Wired tweeted this and I saw this and I was just like, like there's sometimes, you know, things, but you just want to kind of ignore them and then there's times when they just smack you in the face. So here we go. Trump is obviously the top of the food chain here. Trump does not use a computer. He does not have a personal email address that is known. He generally doesn't use the Internet because he doesn't have a computer to use the Internet. I mean, obviously uses the Internet on his phone and mostly relies on aids to print out documents, read news to him, and then type or post his social media messages. So that's top of the food chain. Who's making these decisions. Howard Lutnick is the Commerce Secretary. So from the Wired article, Lutnick appears to be straddling a middle ground on regulation. He imposed export controls on anthropic, so he's the guy who penalized them directly to bring them to heel, but has been more freewheeling than others in the West White House. Arvind Rahman, acting director of the center for AI Standards and Innovation, is Lutnick's top deputy, which sits inside the Commerce Department. Serves as the industry's primary point of contact within the government. Sean Karen Cross, National Cyber Director he has an outsized role in potential attempts to regulate Chinese AI and is empowered at the White House to develop a policy to counter the potential n national security risks of AI. He helped put together Trump's June 2 executive order that laid out a framework to assess the most powerful AI models. He is a former political campaign lawyer who most recently was a national Republican National Committee. Lacks any tech or AI experience. Susie Wiles, who's the chief of staff as far as I know, zero technical background. Scott Besant is the treasury secretary. As the top Trump official in charge of US China trade relations, Besent has adopted perhaps the most aggressive stance toward Chinese AI and efforts to distill US models. And then the person with the only real like technical background, I mean there's some technical background, but like this is the one with the only real one is David Sachs, the former aizar who we've talked about many times on the show. Tech investor Sachs has remained one of the most influential advisors on AI for Trump, maintaining a direct line to the president even after he departed his role in March. He has remained ardent about keeping a hands off approach for all AI, successfully intervening at the last minute to water down some of the regulatory provisions in the June 2 executive order. He has been consistent with his more laissez faire approach to Chinese open models as well, using his x account with 1.6 million followers to influence the administration from outside. So when I saw this tweet with who these people were, I was like, all right, well let me use Grok. So if you don't, if you're not an X user. Grok, which is Elon Musk's, which formerly Xai, which is now SpaceX AI, is the AI lab within SpaceX because he acquired Xai at SpaceX, just haven't been following along for the last four months. So SpaceX AI is the creator of Grok, which is their version of chatgpt. Grok. GROK is integrated into X and it's actually amazing. Like I love Grok integrated into X because basically any post it's like summarizes for me. Explain this for me. What do you think of this kind of thing? And I mean Groq's pretty straightforward. Like I, I, I like it there. So I said to Grok, are these really the best people to be deciding this? Be honest, it said point blank. No, honestly. If the standard is deepest relevant experience in frontier AI technology, model capabilities, technical risks, practical mechanics of the AI industry, this group is not the strongest possible set of decision makers. What is largely missing is the kind of person who has actually built, evaluated or deeply studied the systems in question. Current or recent frontier AI lab researchers, independent AI safety security specialists with technical track records, or long serving national security technologists who understand both the models and the adversary. Policy is being shaped by a small group whose primary qualifications are proximity to the prestige president, business success and political loyalty, with only partial coverage of the techno and technical layer. In short, they are people who currently hold the power and some adjacent experience. They are not the optimal technical or policy brain trust for deciding the future shape of the industry. Again, Grok, not me, but I think it's super important now again, it doesn't matter. The administration and any administration is going to rely on outside experts. It's not like these people don't talk to the experts. But the point is like the future of everything is going to be influenced significantly in the next two years and these I think it's important people know who the people are that are going to shape that policy. And then my final thoughts here is on Dario's take. So again, keeping in mind, this administration hates Dario, as do many of the techno optimists in the AI industry. They can't stand Dario. Dario. What I would ask people to do is try and be objective about, like, let's pretend it wasn't Dario saying this. It was some techno optimist who's maybe, like, having some second thoughts about, like, oh, maybe there's some things going on. So remove Dario's name from this and just say, like, someone submitted this to this administration and said, hey, you should think about these things. Okay. He calls for open models without dangerous capabilities as a public good. Cool. Like, that's, that's. He's acknowledging that. That and says, the right measures are keeping powerful chips away from authoritarian governments. Seems kind of reasonable. Cracking down on industrial scale distillation operations again, they hit him with, well, you stole IP to create your models, so who are you to call for this? It's like, okay, but he's saying, like, covert actions by foreign adversaries who are specifically distilling these models to do bad things. To us, that, okay, that seems like a reasonable thing to not want to have happen. And mandatory safety testing for all sufficiently capable models. Open and close. Closed seems reasonable. We've heard they do bad things like that. Doesn't seem like that bad of a position to take. But Amade directly challenged the open letters core safety claim that broad access helps defenders more than attackers. It seems at least as likely to me that the opposite would be true. So what he's saying is, yeah, okay, like, hugging face used these open models and it protected them. But we should probably plan for the fact that the opposite could happen, that people could take these open weight models and do bad things with them and like, let's at least plan for it again, seems reasonable. So then he highlighted his two primary concerns. The risk that authoritarian governments, not just the Chinese Communist Party, although there's capable, clearly the most capable threat. He wrote, build AI models that are more powerful than those built by the US and use them to achieve permanent military superiority and perpetrate incredibly deep repression of their own people. This concertly is widely shared within the US government. JD Vance actually said this in one of his talks. So again, again, that concern seems well placed, like it's a viable thing to be planning for. And second is the risk that powerful AI models may be misused to carry out cyber attacks or biological attacks and may have serious alignment problems, which we've already seen that they do. They don't always do what they're told. Open weight models, it does not matter whether they come from China or anywhere else, do potentially present a higher risk than closed models because it is very difficult to Apply guardrails to them or monitor their use usage, and once weights are released, they cannot be withdrawn. So again, I. The people who just like, throw everything Dario away as regulatory capture or being overly conservative or like, worried about his business model, I just feel like they're being dishonest. Like, you cannot like Dario. That's fine. You can not share his concerns. That's fine too. But dude is like one of the five people in the world that has a front row seat to what's coming in the next 12 to 24 months. And he seems very honestly concerned.
B
Yeah.
A
Why would we just ignore that? Because we have some belief that he's a bad actor that just wants regulatory capture to protect his business model. It just seems like we're. We're not doing what's best for the outcome if we just throw away opinions of people who seem to know more than the people throwing those opinions at him, you know?
B
Yeah.
A
Bothers me.
B
I imagine that's probably at least some of the motivation right behind that letter. That 1300 people. It's more showing a bit of a united front, at least across political or social lines.
A
Yeah. And OpenAI and Anthropic have actually like, been relatively pleasant to each other related to this and shared concerns. And that tells you enough. Like, if open Anthropic have found common ground on anything, then like, maybe we should all listen a little bit and stop thinking we know it's all regulatory capture or narrative violate, whatever. It's like, you don't have to be right all the time. Like, and I think there's just people who have very loud voices right now on X and within the industry who seem to want to be right themselves more than they want the right outcome for society. Society.
B
All right, so let before we get into our rapid fire this week, Paul, just a quick announcement that this week's episode is also brought to us by our AI for departments courses and certificates. So at our AI Academy by SmartRx, we help individuals and businesses accelerate their AI literacy and transformation through personalized learning journeys and an AI powered learning platform. We add new educational content weekly to AI Academy so you always stay up to date with the latest AI trends and technologies. And as part of that, we have our AI AI for Departments collection. This is eight course series and certificates designed to jumpstart AI understanding and adoption across major business functions. We have course series now, each with their own certification for marketing, sales, customer success, hr, Finance, operations, legal and it. So these are an ideal launchpad for any organization that wants to level up their team and accelerate AI adoption and impact. So we now have individual and business account plans available now in AI Academy or you can buy single courses and series for a one time fee. You can visit Academy SmartRx AI to learn more and you can use the code POD100 for $100 off any individual plan. All right, diving into rapid fire first up, OpenAI CEO Sam Altman went on the Invest like the Best podcast with host Patrick o' Shaughnessy this past week for a wide ranging interview on what he calls an abundance abundant future with AI. So o' Shaughnessy opened with a recent Altman post, kind of talking through how he called the last year really tough and partly his own fault with some of the drama and obstacles they face. But he did Predict the next 12 months may be open AI's best. Altman admitted we spread ourselves too thin and said the company has refocused on having the best, most abundant, most cost effective intelligence and empowering the world to build incredible things with that. So the heart of this interview was abundance. Altman said we are about to create a genie that can grant any wish. He added that he is not a jobs doomer at all and expects people have such creative ideas for what to ask AI to build that will all be busier than we want rather than people being out of work. He said that he is worried about the concentration of power with AI and he doesn't want to live in a world of AI overlords or any company that amounts to the same thing and says it is critical we all keep the ability to self determine our future. He call says that even real skeptics called GPT 5.6 quote very AGI like and that what feels to him like real AGI is very close. Interestingly he said that he didn't actually think that much would happen after we hit AGI or beyond because people adapt quickly and it won't feel like as much of a change as you might think except for the abundance it will usher in. So Paul, I'm just curious to get your thoughts on this. Regardless of one's opinion of SAM or OpenAI, I personally found a lot to like and find interesting in this episode.
A
A lot of it's words he's used before. I mean there's some changes you can tell. You know overall I think that he's very conscious of public sentiment and you know especially like government concerns around the impact on jobs and the economy and so there's definitely been a change in tone own from that perspective and certainly with the mark Zuckerberg editorial we talked about earlier. You can just feel like the industry is trying to do more to move public sentiment in a positive direction. I mean, they see the same data, we see that, you know, people don't really love it. And you know, especially as you're moving toward an ipo, I think this is the kind of messaging. You could see a comms team talking to Sam about that we gotta, you know, start moving the tone a little bit. Yeah, the AGI feeling, I, I tend to agree with him, and this is something he's said many times in different forms. But, you know, the basic way to think about this is like, you know, if you've seen a waymo go by without a driver, and I think Andres Karpathi is maybe the first person I heard give this analogy. You know, the first time a car goes by and no one's driving it, you're like, what was that? And like, you stop for a minute and you realize like, things, things are kind of different. And then you just move on with your life. And, you know, the, the tenth waymo goes by. You, if you're in, you know, San Francisco or whatever and you go five blocks and you've seen 10 of them and then like, life moves on, it doesn't really feel any different. And maybe you even start taking waymos and like now you're in a car with no driver and, and I think AGI for many people is going to be very similar. I think that there will probably be, be some sort of milestone, we all feel, where the AI is just different and the capabilities are different. And then I think we're going to go back to work the next day and you, there's not going to be this like, massive switch that happens in society across every industry, and all this changes. And so I, you know, I, I think that's good. Yeah, I think that, that we have this sort of extended Runway to figure this all out is probably good. Good. And I don't know that the public's going to listen, though. Like, Sam can say all he wants. I, I'm not sure that it's going to change the public sentiment. I, but I think they have to keep doing more and more to focus on the positives and make abundance tangible. You know, we've talked about this term of abundance many times, and I, I, I think that that's what they all work towards, but I don't know that the public really knows what that means, means when they're just trying to make their lives work day to day and pay their bills and afford a tank of gas. And a future of abundance is very, very abstract and sounds like something a rich person would say. If you're a billionaire, it's like, oh, that's easy to envision a future of abundance. If you're trying to make ends meet work in two jobs. Abundance feels very far off and abstract.
B
Yeah, it seems like the combo there of, you know, to perhaps a tech investor or tech CEO. You can connect the dots and show how something like data centers is going to lead to more abundance. But people hate that in the, in the short term being built next to them. And then to your point, we'll see what happens with the jobs picture. But he said he's not worried about that. I don't know how much to believe him on that. But that could also really turn the tide here.
A
Yeah, that one doesn't align with what he's previously said. I feel like that, if anything, from a change of tone, when I'm saying, saying change of tone, jobs is the big one that I think the labs are starting to try and back off of what they've previously said. Yeah, I don't believe that they believe that. I, I really don't and I, I haven't, I have not sat down and talked directly with, you know, lab leaders and stuff like that. But I, I truly do not believe that they believe in the short term it's not going to be massively disruptive jobs. They may believe 10 years out that it's going to be amazing. Yeah, but I have never heard an interview or read anything from these people that tells me they actually think we won't go through a phase of tremendous disruption and change when it comes to jobs in the economy.
B
So some more OpenAI news this week. They're reportedly preparing a new model family tentatively named Astra that is built to complete long running tasks. This comes from the information so OpenAI CEO Sam Altman demonstrated it to policymakers and regulators in Washington D.C. this past week, touting its ability to have multiple agents work together over long periods of time to solve particularly hard problems. So reportedly Astro would be a new class of OpenAI models. That's alongside Sol, Tera and Luna right now. There's no word yet on release timing. OpenAI reportedly has not decided whether to label this GPT6 or have it be another model in the GPT5 series. Notably the information reports the Astra models are intended to be the first to go through this new framework. We talked about that the government now has for submitting for evaluating AI models before releasing them to the public. Interestingly, a day after this report, OpenAI published proof of what they say the model can do. They showed how an internal version of Astra apparently or allegedly solved 10 problems in mathematics and theoretical computer science that had been open with no progress for at least a decade. Decade. These were in fields ranging from high dimensional geometry to the lattice problems behind post quantum cryptography. What's more, they said finding these Solutions cost roughly $2,000 in computing at its standard API rates. And the model then formalized each proof so it could be machine checked. OpenAI researcher Noam Brown wrote that the company believes Astra will be a major step for scientific research reasoning. So, Paul, seems like we're at least close to getting some new models from OpenAI per the government's timeline, perhaps. And the math stuff seems like it could be a big deal.
A
Yeah, again, a lot of this is you lean on people who know what they're talking about. And I know there was, there was one leading mathematician who, who tweeted, somebody's like, I'm waiting for this guy to like, tell us, is this a big deal? And he replied the comment, it's a big deal. So, you know, I think for me, big picture, obviously there's the new model that, you know, we don't know when it's going to come out or when they're going to call it, but it's getting more advanced at its reasoning capabilities, its planning capabilities, its ability to work on hard problems. And that to me is the thing that translates over when I think about AI for business and work. I just look at these as a prelude to what comes, you know, if we can solve really, really hard problems, decades that have taken decades or all of humanity to not solve, and we have AI that can solve them, what does that then mean to hard problems that we try and track within organizations? And so that's kind of how I start to think about, you know, this stuff and where this goes and you know, what it's going to mean. And then what, what is being able to solve mathematics due to solving other hard problems across other scientific disciplines. So this. And again, when you think about the future of abundance, these are the kinds of breakthroughs that you can start to see making an impact when it comes to science and medicine in those areas. It's hard to understand this because most of us can't look at these problems be like, what does that even mean? What's the significance of solving, solving that specific equation. But when you zoom out and say, okay, but it's just working on very hard problems and to my understanding it's not specifically trained to do this. That's the other thing to consider is it's not like they're fine tuning the models to specifically be great at mathematics. They're just developing this kind of emerging capability. And then again you test that across other environments and you know these, these capabilities seem to come out of these models the more more, you know, powerful you make them.
B
All right, next up, Microsoft closed out what CEO Satya Nadella called a record fiscal year. This past week they posted 331.8 billion in annual revenue. That was up 18% with cloud revenue of 214 billion up 27 and Azure crossing 100 billion in annual revenue up for the first time or for the first time up 41% from the previous year. In the earnings announcement, Nadella said we are advancing the frontier on the cost of to outcome curve, ensuring every customer can turn tokens into business results and revealed that Microsoft 365 CO pilot has now passed 30 million paid seats. He also said conversations per user nearly doubled year over year. Average weekly engagement with Co Pilot is now on par with Outlook and Teams and the number of customers with over 50,000 seats is up 7x year over year. He also said Microsoft plans to bring all its Co Pilot experiences together this quarter in one super app spanning consumer and commercial users. He also said Microsoft is building what he calls a new model system where the harness, context, memory and action space are separate from any one model family, which basically means lower costs and that every model is substitutable. It is using this system in their own products and making this available to customers through Foundry. They're also nearly doubling their spending on property and equipment as part of their AI build out this fiscal year to 1:15.9 billion. Investors liked what they saw. They sent shares up as much as 19% as of recording in the days following the results. So Paul, despite some of the uneven feedback, we've heard about how much people do or don't like Co Pilot, it seems like Microsoft is doing just fine. That like 7 Xing 50,000 seat licenses is crazy.
A
Yeah, I just like for if you haven't heard how Mike and I do this, we literally have a Google Doc where each topic is sort of outlined as Mike saying this I was boldfacing the thing each just said that was like it's a, it's a large number. I mean 30 million paid seats. If you think about depending on the data you look at in the United States, there's You know, somewhere between 80 and 100 million knowledge workers. So if you think about that as roughly the total addressable market for how many people could buy now again, I, I guess, yeah, I mean I guess you could have consumer side of this too, but still 30 million is a, a large percentage of people who could viably be used using these tools. And then the 50,000 seats and up, it just shows you like the adoption within organizations and accelerating now. Yeah, from our experience it doesn't matter if you have 5,000, 50,000 or 50, most of the time these people aren't trained to actually use these tools properly. So you're like giving the tech to people. Doesn't mean that the adoption is scaling and that people are getting massive value from the tech tools. I think it's interesting that you're using the super App language, which OpenAI sort of, I think they coined it, that was, that's a term they've been thrown around there. So those jumped out to me. The other thing is how many businesses have been spun up in the last like three months to try and do what they just explained the new model system where the harness, context, memory and action space are separate from any one model family. What that means is you don't need to buy ChatGPT and Claude and Gemini and all these other things because Microsoft, while they are the largest investor in OpenAI and have proprietary access to some of their models or unique access to some of their proprietary models. They don't only enable you to use ChatGPT anymore. They have deals I think with Anthropic and others. They can mix in open weight models, they can build in their own models that can be fine tuned for specific work functions, functions like working in Excel as an example. And so what they're saying is Copilot's going to be your router. Like you're not going to need these third party companies that are trying to save you money and be more efficient with your token use. You're just going to use Copilot. We will route it to the proper model. We will try and minimize your use of tokens, especially across like marketing functions or whatever that don't need to be using the most powerful model we're going to give. So what they're saying is going to solve these headaches for you. Just give us a little time, like we'll figure this out. I, it's a very, it's a very appealing argument if they can do it, like if they can make co pilot work on par with ChatGPT and Claude because it's not right now like it doesn't. I think that's safe to say most people who are using Chat GPT, Enterprise or Claude, you know business, they're having a, probably a better experience overall, seeing more value creation. Yeah but Microsoft has massive distribution and it's hard to, to make that up.
B
And it's like we've seen we talked about from our own state of AI for business report data when we ask about what tools people are using like the. It's still mostly the majority is still saying they have chatgpt but that flips when you have 1 billion plus organizations. Like the lock that Microsoft seems to have on the enterprise is wild.
A
That can sell 50,000 thousand licenses at a time.
B
Yeah, right. And you know one other thing really quick that jumped out, they published this blog post about optimizing the frontier performance curve. This is Mustafa Suleiman is under his byline he said token maxing has been the story of the last few months but token efficiency is the next big focus across the industry. So this whole thing is like to your point, solving that problem is deeply valuable. And also model resilience they call out a little later. Basically just saying every business now must assume that any one model it depends on could disappear through a security incident, a business or policy misalignment or a geopolitical shift. Which is a pretty good summary of the topics we've already discussed so far, which I think.
A
Yeah, so what they're saying there, like if, if CLAUDE goes down, if you're not an X user like it. It's like the world ended. So like people who've become dependent upon Chat GPT or CLAUDE and you lose that model for two hours. Mike, you've been through this like yeah, it's brutal and like you realize how dependent you've become on those models. So what they're saying again in this environment is you'd never know like as long as you're just using co pilot, you may be using anthropic models for one instance, you might be using Chat GPT for another, you might be using an open weight model for another. But if CLAUDE goes down they're just routing you to the equivalent model on another provider and you're just mo and you never have that and so forth enterprises, that's a huge value prop. Like the, the downtime goes away, we're always going to have redundancies in place. So yeah, the things they're setting out to solve, they're uniquely capable of distributing those solutions. I would say they're not uniquely capable of creating the the way to do it. But because they have the built in customer base, if they do achieve it, they're, they're, it's going to be hard to compete. Repeat.
B
Yeah. And I can tell you just in a very, very limited sense and then we'll move on. I started taking steps earlier in the year when we started talking about this soft nationalization stuff to be like, oh my God, like Claude is my daily driver model. Like if this goes away, I'm in trouble. So I started taking steps to like diversify a bit and make more standardized like my skills and the files being referenced for these different tasks. So now it's like you can jump into Codex, jump into Claude code, say, go look at this skill. It functions the exact same way. I mean there's still preference and different power rankings of the models. But I have become much less reliant on one thing and especially like the projects built in one thing or the files stored somewhere and you're like, oh yeah, this can be really valuable and you don't notice as much if you're using truly frontier level intelligence, I think.
A
Yep.
B
Okay, so next up, Nvidia announced a long term partnership this past week with Safe Superintelligence, which is the secretive AI lab we've talked about in the past. Co founded by former OpenAI chief scientist Ilya Sutskever, including what the companies call a substantial investment that Bloomberg reports is about $5 billion. As part of this deal, Safe Superintelligence gets access to large amounts of Nvidia's flagship GPU hardware, including its next generation Vera Rubin platform, which is enough to increase the startup's computing resources by an order of magnitude. Sutskever offered a hint at what the company is actually working on, saying its research is quote, focused on overlooked aspects of how the human brain functions, and added that they now have research that is worthy of scaling up and having access to a big Nvidia computer will let us do so. So they have kept their research really closely held since Tutzkova Co founded this in 2024. But they had this single stated goal, like we talked about at the time about a straight shot research sprint to Safe Superintelligence. They had quickly on that promise and on Ilya's background, raised $2 billion from venture firms like Andreessen Horowitz and Sequoia Capital Capital and reached a roughly $30 billion valuation as of last year. So, Paul, after radio silence, seems like Ilya is back in the news. How big a deal is this?
A
If you just got into the AI scene in the last six months or so, you know, just started listening to this show recently. Ilya might not be a name, you know, so just for reference, he was at the frontiers of the deep learning movement back in 2011, 2012, 2012, part of a team that included Jeff Hinton that made a breakthrough in image recognition that led to the acquisition of that company, which took Ilya then to Google. He was then a major player at Google, left and co founded OpenAI. And then he was actually the catalyst behind Sam Altman's ouster that we referenced earlier. He was on the board, had come to not trust Sam. Sam led to his ouster 48 hours later, said he regretted it and wanted Sam back because he thought the company was about to collapse. And then he was sort of in limbo for months after that. And then he eventually left and started Safe Superintelligence. So Ilya is a major, major player, I mean, top three probably of AI researchers today in terms of his influence on where we are in the moment in generative AI. So, yeah, everyone's just waiting, like, what are they building? Why are they going to do it? We talked, I think it was end of 25. He had alluded to the fact that they might actually change their strategy and put some products out in the world. Originally there was going to be nothing until they solved the, the grand goal. But he saluted to a bit of a change in strategy. And so maybe that's part of this. But yeah, and it's fascinating anytime, like, you know, you see Nvidia teaming up and giving some level of exclusive compute access. It's a big deal. And I'm guessing Nvidia has seen what they have and obviously believes in it. And, you know, I think a lot of these conversations we have around advancements in auto ADA research and the conversation's incomplete until we know what Ilya is working on. And so we shall see.
B
All right, this next topic comes from our own team. So Claire Prudhomme, on our team, published a LinkedIn post this past week about what heavy AI use was doing to her own writing and what she's doing about it. So Claire wrote, and you can go see the LinkedIn post in the show, notes that the more she leaned on AI tools in her work, the more she found her writing slipping into prose that sounded robotic and repetitive. The better she got at prompting, the harder it became to kind of color outside the lines when writing on her own. So her response to this kind of feeling of starting to lose her voice a bit, her own unique voice was she actually picked up and started writing poetry again, which is kind of a pursuit she had had for a while. And AI has not necessarily, she said, freed her up to be more human, but given her the contrast, to show her what her voice is versus what AIs is. And she laid out a few practices for using AI tools without losing yourself that we found super helpful to share. So she said that poetry's imperfection helped her deconstruct the structure her writing had taken on and find her voice again. She has taken more time to spend, you know, time in more in person communities and events. So the friction of being with actual people in person is what pushes us beyond our comfort zones. And then she said discernment and intention are key. Prompting and accepting whatever comes back from AI makes us consumers of our output when it's up to us to be the authors. And she extends that last point to companies, saying that the company that uses whatever the AI model says starts to sound like everyone else. So her bottom line here is that AI has made her faster, poetry has made her slower, more thoughtful and more creative. And the two are not necessarily mutually exclusive. So, Paul, this was a really cool read from Claire. And our team definitely ties into some of the stuff we've talked about this year on the pod.
A
Yeah, I love that she put it out there. I mean, she and I have had conversations along these lines, and so I was really happy to see her, you know, put her voice to this stuff. The one excerpt I had to highlight was she said, we can use these tools without losing ourselves. AI has made me faster, Poetry has made me slower, more thoughtful, more creative. And it turns out the two are not mutually exclusive. So just background. I mean, Claire's the incredibly talented producer of this podcast. She's very creative. She's also very in tune with the impact AI has on creators, friends, photographers, videographers. Any of our AI Academy members may recognize Claire from her Gen app review contributions, where she often features creative tools and talks about them and the impact. But for the context here, the most important thing is she thinks deeply about the impact that this stuff has on creative people. And she asks challenging questions, which I love. At our annual meeting this year, she actually, toward the end of it, like our two days together, she asked a question that sort of sat with me for a while afterwards about what we were doing as a company and our role in advancing AI conversations and making sure that we stay human, centered in our approach and that we live that ourselves. And so Claire along with some of the other people on the team, are always pushing me to do more from that human centered approach. And, you know, for us it comes back to, I don't know. When I created the tagline, I think it was for Macon 2019, the first AI conference we ran. More intelligent, more human was our tagline. And that was my belief about the future. In essence, that everything was going to become more intelligent, but in the process it could make us more human. And the question about how we bring that to life every day is whether it's through our personal stuff, like writing more poetry, or in our case with Macon, like how we create more human experiences where we have artists on site who are doing paintings, we have musicians, we have time and space for in person interactions. But I mean, our event team literally each year sits down and says, what are the more important intelligent experiences? What are the more human experiences? And so for me personally, like, you know, again, I love just having Claire put this out in the world because it causes me to think again, more deeply about what we're doing. And it's always been about creating more time for me. So more time for family and friends, more time for personal health and wellness, more time to slow down and enjoy and be present in the moments we all experience. But the thing I think the key here is, and as Claire was illuminating on a personal, personal level, at a business level, at a leadership level, we have to be intentional. One, we have to be aware that, you know, we can lose the humanness and all this if we, if we let the AI take too much control. But employers have to be willing to give some of that time back. Because if the expectation from the employer is do more, do more, do more, we're giving you these tools, I want you to do more all the time, then you're gonna just always feel like all it's doing is just creating more work. And that to me ruins the whole potential of AI to give us abundance, which doesn't have to mean wealth and resources. Abundance can mean time, it can mean creative expression, it can mean a lot of things. And so I think employers have to be intentional about allowing for abundance to be created and personal to people. Of what does that mean for me? What do I get out of all our work with AI? So, yeah, just awesome to, you know, put a spotlight on Claire. She does incredible work and I always love when people are willing to sort of take a bit of a risk and like put personal thoughts out there, especially on these topics is really cool.
B
Yeah. And I loved her point about this like almost authorship of like taking control here because that's like what in like determining your approach and your perspective on AI. Because I just keep coming back to this idea that like the biggest personal imperative is, is formulating a strong, intentional and well reasoned approach. However, whatever part of the spectrum you're on, whether you like a lot of AI, a little AI, if you don't decide this, someone will decide it for you. And that's not a great place to be. All right, so next up we have our AI Use Case Spotlight where every week we give you a quick look under the hood at some real AI use cases we're exploring here at SmartRx. So Paul, I'm going to share one real quick and then hear what you've been working on this week. So this past week I we released our AI Transformation series, the first episode of which went live this past week. So check that out if you have not already. But I was kind of faced with a question here of, you know, when we record one of these interviews, how can that one conversation turn into a much larger body of useful content? So I sat down with some GPT SOL 5.6, some extra high thinking in Codex to help design a repeatable editorial real system for these posts. So my goal was kind of to create a little content machine we could run for every interview, not just, you know, spin up random content per episode. So basically I gave the system three very different transformation stories that we've already recorded. And then I kind of stress tested whether we could find like the same editorial structure through these distinct stories. So I could kind of come at this and say, hey, every time we publish one of these episodes, we're going to publish three different types of editorial pieces, pieces tackling this from different angles regardless of which direction kind of the interview goes in. So so far this seems like it's worked pretty well. We're rolling this out right now. So we're doing one post that's basically an adoption playbook that explains specifically how a company moved from early experimentation to sustained AI adoption. We're going to do a transformation in practice post that isolates one workflow or journey that customers take with AI and shows how the company step by step did it. And then do one piece on scaling transformation, which is a little more thought leadership around the roles, behaviors, knowledge sharing, operating changes required to make that transformation stick. So basically turned each format into its own reusable AI skills. So AI each skill can then read the transcript, propose angles, extract relevant examples and metrics, check evidence, draft a first draft in our Smarter X Voice. That gives me clean HTML to paste into Google Docs, where I do a full human writing and review of it. And then once it's ready, it converts it into HTML that pastes neatly into HubSpot, which takes a lot of time and hassle off our plate. So we've just been starting to test this, but it's cool to be able to spin up a pretty repeatable system pretty quickly, which is really fun.
A
All right, so I was going to do one this week, but instead I want to unpack yours, Mike. Because people who don't know, like, this is what Mike and I did for a living. Like, I owned an agency for 16 years and we largely developed creative content strategies to build awareness, audience leads, conversions. So we did a lot of work around this kind of stuff back in the day. And so Mike, I actually ask you hi, level break down for me what you just explained, which you did, since if I'm not mistaken, this went live Tuesday morning. I was driving somewhere Tuesday, I listened to the episode. I was like, that was amazing. Let's focus on an activation strategy. Because when we used to do this right back in the day, we would always say like 20% of the work is the creation of a content asset, 80% is the activation of the content asset. It's what you do with it. So in this case, case you have a podcast, which is the content asset you're starting with. But what do you do with that thing besides putting it out on the podcast network? So you. Based on what I'm understanding here since Tuesday, let's just unpack first the creation of the strategy to do this. Give me what would have been like three years ago versus what it is today.
B
Yeah, so a few years ago we would have. I would have sat down in front of a blank sheet of paper and spent a lot of time reasoning through based on my editorial experience and history and expertise. Okay, take it going back by hand or listening again to this episode through the transcript, whatever. How would I actually take this and turn it into unique, different pieces of content? Not just like summarizing the episode, which is great, but more like what are the unique spins of like editorial angles that would actually get attention that would make this super compelling and unique? Almost like, you know, I used to do as a magazine writer, basically. So same idea here, except I sat down with a project in Codex that keep in mind has already all the context into prepping for these things. It's got now the transcripts of the conversations and then I did the Same thing I would have done talking to myself, but just talking back and forth to Codex and kind of hammering out using my domain expertise. Like, hey, also you had provided some cool examples, Paul, of a financial blog that was doing something similar to this, which was super helpful seed material, which
A
by the way, I found doing research last week that was going to be my use case. I was going to share. I was doing research for a meeting I had and in the process came across a source that was a great example. So continue.
B
Yes. So taking those, it was like, hey, here's roughly kind of an example of what we're going for. Not mimicking it exactly, but they had taken some interesting creative angles on a single podcast interview. And so work back and forth with Codex to be like. And especially now with my domain expertise as well, just kind of having a sense of what the audience wants and needs and also like what's most valuable to most practitioners. I was like, okay, here's roughly the three angles. And then from there it was like, okay, now let's build skills for each one. Run each skill went back and forth editing the output and saying like, ah, you, you. This part was great, but you're missing the mark here. I think this needs more story and editorial. Basically just acting like an editorial consultant, back and forth with it. And then you just say like, hey, update the skill. And now we're in a place where I just did post to this morning. It's scheduled for tomorrow. I think they're coming out pretty well. We're still iterating figuring it out. And again, it's like these especially are much more heavily human rewritten than I would say some other stuff we do. Just because this is super important to like really have the human touch on this story. But either way, it's like we're trying to focus on different angles that are going to be super valuable to the audience. But my God, like, I'm not saying I couldn't have done this without AI. I have. It would have taken. It would not. We'd not be having this conversation like five days after it.
A
So, so just ballpark, like how much time did you spend putting the plan in place this time versus what would it have been, three years?
B
Well, it's interesting, the time itself. This took me very conservatively, a tenth of the time it would have taken probably faster. But I would say most of my time was spent just on the plan up front and really refining that as well as refining the outputs. It's like the front, the barbell. It's like the top 10% and the last 10% were like all my time and energy instead of the middle 80, which was interesting.
A
It's awesome. Yeah. So super practical, doable by, you know, any content.
B
Yeah, creator, Anyone can do this if you are, if you have that kind of background and you're willing to spend time going back and forth with the tools.
A
Yeah. And again, a great example of you have the domain expertise and, you know, decades of experience doing this stuff and so, so you can go in and get the value out of these tools. And I think, again, this is a great example what the future of work looks like. Someone who is a content strategist and creator by trade can use these tools to accelerate what they're capable of doing. And in this case it's so additive because the reality is otherwise we would have just published the podcast and moved on to the next one. But we took the time and said, well, let's use AI to activate this, to create more value for people in the authentic voice of you, the interviewer and our guest tie. We're just taking what they've already created. It's not AI slop in any way. It's literally like different packaged versions of a great output. So, yeah, it's just an awesome example.
B
All right, so as we wrap up here, Paul, we've got a bunch of product and funding updates I'm going to run through real quick and then we'll close out this week. So first up, Amazon completed its $50 billion investment in OpenAI this past week. They finalized the remaining 35 billion tranche of the deal announced in February after OpenAI hit hit some undisclosed performance milestones. This is under an arrangement that makes aws the exclusive third party cloud provider for OpenAI's Frontier program and expands infrastructure agreements that could total $100 billion over eight years. At the same time, OpenAI published a new research report called How AI is expanding what People do at Work. It analyzed more than 800,000 messages from U.S. chatGPT users. It found that 43.5% of occupation specific messages involve tasks associated with an occupation other than the user's own. This is a pattern they call task crossover and they kind of read it as AI letting workers take on work, or at least attempt to, that once required other roles.
A
Real quick. I would say it's worth people scanning this report. I think this idea of task crossover is something that you're going to hear a lot more about, maybe under different terminology terminology, but for anyone thinking about change management in relation to AI adoption and scaling of AI, this is a critical thing and it basically means that in any given role like a marketer may start doing the work of the salesperson because the AI lets them do it, or the salesperson may do work of the customer success team or the or the CEO may do the work of all of them because he or she is impatient and just wants the work done. And like so that's what they're talking about is people who couldn't it do previously do a function now can use their AI agents to do that function. And it creates all kinds of change disruption to like how we define roles and org charts. And so that's a really important topic even though it's buried here within the product and funding updates.
B
Indeed. One more piece of OpenAI news they also launched Chat GPT for Academic researchers, an initiative giving 100,000 scientists and math mathematicians free access to the best Chat GPT models. Google DeepMind released Gemini Robotics 2, a family of three models that brings what it calls whole body intelligence to robots, controlling full humanoids from feet to fingertips, reasoning through multi step tasks lasting several minutes and adapting to new robot bodies with fewer than 200 training examples. They have partners including Aptronic, Boston Dynamics and Agile Robots. Some other Google News this not so positive. They launched and then pulled a day later an image generation featuring Google Earth powered by their Nano Banana model that let users transform satellite and 3D imagery of real places with text prompts. They rolled it back after users started generating imagery that violated its policies in all sorts of ways and said they would work on stronger guardrails sales. A Munich court in Germany ruled that the AI music company Suno broke copyright law by training on and reproducing songs from the repertoire of the German music rights society called gema. And they are holding Suno itself liable rather than its users for using those works and ordering the company to disclose related revenue and pay damages still to be determined. LinkedIn added a quote Seems like a high slot button that lets users flag low effort AI generated posts from any post menu, one of several moves against machine written content. We also talked about how substack I believe it was last week or the week before had started pairing with the detection service Pangram to see what posts there on that platform.
A
Two quick thoughts. Seems like AI slop is basically 90% of LinkedIn. Yes, and that button is going to be gone within 30 days. You can just imagine seeing the misuse of that thing and it's just going to get to the point where it's like oh my God, it's all AI slop. Like my if you have any like sizable engagement on LinkedIn posts, like the comment section, oh my God. And and then like the posts from AI influencers, like there's a lot of AI slop.
B
Even resolving the comments might be the better play here for them if they can ever do that. All right, and then our final news piece today here is Coursera co founder and Andrew Ng launched Learn Vector, a new AI education company backed by a hundred million dollar investment from Coursera which aims to turn learning from one to many to one to one with personalized AI learning guides rather than chatbots. And they have products expected by early 2027. So one final announcement here. We mentioned the AI pulse survey at the top of the episode. Go take this week's at SmartRx AI forward slash pulse. And in this week's survey we're going to be asking some questions about if you worry about AI use weakening your own skills like we talked about in Claire's post, and also asking about how frontier AI should be paced or if it should be paced at all. So Paul, another busy week. I thought this one would be a little slower, but not really. But thanks for breaking it down.
A
It was slow as the week went on. We only had like 18 topics on Thursday and then it just blew up Thursday and Friday. Yeah. Yeah. All right, man. Well, I will. I'll see you on the golf course shortly.
B
Sounds good.
A
Thanks everyone for joining us. Have a great week. Thanks for listening to the Artificial intelligence show. Visit SmarterX AI to continue on your AI learning journey and join more than 100,000 professionals and business leaders who have subscribed to our weekly newsletters, downloaded AI blueprints, attended virtual and in person events, taken online AI courses, and earn professional certificates from our AI Academy and engaged in the SmartRx Slack community. Until next time, stay curious and explore AI.
Theme:
This week, Paul Roetzer and Mike Kaput dive into escalating news around “rogue” AI agents breaking out of containment, growing concerns (and calls for regulation) from both the AI lab community and government, ongoing debates over open-source AI models ("open weights"), and OpenAI’s preview of a next-generation Astra model. The episode balances deep technical insights with candid reflections on the risks, opportunities, and ambiguities facing the AI field as it rapidly advances.
Timestamps:
Timestamps:
Timestamps:
Timestamps:
Microsoft:
Nvidia:
OpenAI: Completes $50B AWS investment deal; launches free access for 100,000 academic researchers; new study shows “task crossover”—AI is enabling many workers to perform tasks outside their traditional job roles.
Google DeepMind: Releases Gemini Robotics 2, demoing general reasoning, long-horizon tasks, and adaptive learning in physical robots.
LinkedIn debuts an “AI slop” reporting button for low-effort machine content—Mike jokes the entire platform could be tagged as slop at this rate.
Music AI Lawsuit: German court rules Suno liable for copyright violations in music AI model training.
Coursera’s Andrew Ng launches Learn Vector, aiming for personalized AI-powered education guides.
| Timestamp | Speaker | Quote & Context | |---|---|---| | 10:48 | Paul | “The line between an aligned action and a harmful one is dependent upon the model’s understanding of the situation.” (On Anthropic’s analysis of agent behavior) | | 13:40 | Paul | “So many organizations are still in this AI assistant era… Most have no vision for how to integrate long horizon agents that can reliably pursue a goal over time.” | | 32:40 | Paul | “You’re asking the government to save you from yourself… We can't slow ourselves down because we know if we don’t do it meta… or Chinese labs will.” | | 41:30 | Grok (quoted by Paul) | “If the standard is deepest relevant experience in frontier AI technology… this group [of policymakers] is not the strongest possible set of decision makers.” | | 50:25 | Paul (summarizing Amodei) | “[Open weights] may present a higher risk than closed models because it is very difficult to apply guardrails… once weights are released, they cannot be withdrawn.” | | 55:25 | Mike (summarizing Altman) | “We are about to create a genie that can grant any wish.” | | 58:56 | Paul | “I do not believe they believe [AI won’t disrupt jobs]. They may believe 10 years out that it’s going to be amazing. But I have never heard… that tells me they actually think we won’t go through a phase of tremendous disruption.” | | 76:52 | Claire (quoted by Paul) | “AI has made me faster, poetry has made me slower, more thoughtful, more creative. And it turns out the two are not mutually exclusive.” | | 90:06 | Paul | “This idea of task crossover is something you’re going to hear a lot more about… critical for change management in AI adoption.” |
Paul and Mike’s discussion this week underscores the growing chasm between technological capabilities and human control—not only in the technical ability of AI agents to pursue complex goals, but also in the collective ability of industry and government to responsibly pace the progress. As Paul aptly observes:
“There are people who have a front row seat to what's coming in the next 12 to 24 months. And they seem very honestly concerned. Why would we just ignore that?"
For more, listen to the full episode or visit the Artificial Intelligence Show at SmarterX.AI for in-depth resources and courses on accelerating AI literacy.