Loading summary
A
Some in the government apparently concerned that very powerful open source models could pose a threat to the US Frontier Labs.
B
I actually think that there's two races that are going on right now and China has already won one of them.
C
American companies will use Chinese open source models hosted in America, which will hurt Anthropic and OpenAI. It's the same thing. If you said, hey, you have to use gasoline that costs $800 a gallon. Everyone else, you use gasoline that costs $8 a gallon, they're going to just be able to have way more economic activity.
D
There are risks of potentially having like back doors might have significant cybersecurity risk for come for these companies that self host those models.
C
Once they've kind of taken the lead then they'll okay, well now these models are ours now and they won't be open source anymore. They just want American companies to fail. Thanks to our friends at PayPal, the exclusive sponsor for this Week in AI try the payment and growth platform that's
B
trusted by millions of customers worldwide.
C
PayPal open source start growing today at paypalopen. Com.
A
Welcome to another episode, episode 23, I believe of this Week in AI. We're very excited. You may be looking at your phone in confusion. That's not Jason Calacanis. That's not our usual guest host, Alex Wilhelm. No, it's me, Lon Harris. You've seen me co host sometimes. I'm here in the captain's seat for this week. Jason is in Japan. We are still doing this week in AI and we have an incredible lineup of guests for you today. Joining me from Patronus AI we have their CEO and co founder Anand Kinopan. Thanks for being here, Anand.
B
Great to be here.
A
All the way from A21 Labs. He's calling in from Tel Aviv, the heart of Tel Aviv, Israel right now. Their co CEO and co founder Ori Goshen. Ori, so glad you could make it.
D
Great to be here. Thank you for having me.
A
Thank you. And joining me as co host and to make sure if I say something extremely dumb and outside the box on AI there's there's somebody here to be like lon, please, come on, bring it back to earth. He's the CEO and founder of Henry Intelligent Machines. He's one of my favorite AI influencers on YouTube. Give it up for Alex Finn.
C
Good to be here, Lon. Very exciting and I'm sure you won't say one stupid thing. We'll be fine.
A
Not what if you're watching.
C
Not a single thing.
A
If you're watching at home right now. Grab the transcript, put it into Claude and see if I said anything stupid. I think I'm going to get away with it. I think I'm going to pull this off. So this is. There is a ton going on this week we gotta talk about with this panel. We're gonna go. I definitely wanna dig into a little bit about Petronas, about AI21, what all of you guys are working on. But I think first, before we jump into anything else, we gotta talk about, are they going to ban the Chinese open source models that everyone loves? Axios, of course, reporting that the Commerce Department, White House and NSA this week have looked into potential ways to limit access to these foreign open source AI models. Of course. Kimik 3 making huge headlines, is it? Overtaking Fable 5 on some of the benchmarks. Trump at an AI summit this year said, and I quote, our children will not live in a planet controlled by the algorithms of the adversaries advancing values and interests contrary to our own. But the concern, as I said, may not just be about cyber security, but also AI supremacy. Some in the government apparently concerned that very powerful open source models like Moonshot's, Kibi K3, GLM 5.2 from Zai and others could pose a threat to the US frontier labs. Of course, there are a few options on the table for what the government could actually do. But I want to take this to the panel Anand start us off. Do you think this is actually likely to happen or is this just diplomacy being conducted in the pages of Axios and Politico and other blogs?
B
I think that's a great question. I'd say that I think in some ways it's both, and here's why. So I actually think that there's two races that are going on right now and China has already won one of them. And the first race is the layer of intelligence that can be measured cleanly and copied very cheaply. And so whether that's coding or math or anything that can be verified correctly, in many cases, I think that China has basically taken that. And although in many cases we could say that China has been maybe three to six months behind in terms of how the frontiers evolved, but they're catching up very quickly, especially if you just look at the charts in terms of performance across a lot of these benchmarks, I think that the race that's still in play right now is the layer that's harder to measure and harder to copy. And so that includes, can agents actually solve real world problems? Can it solve problems over longer time horizons? Can it solve tasks that are super reasoning heavy and where there's a lot of real expertise that's required to actually evaluate performance. And I think that a lot of the closed source model companies have done a lot to be able to get there. And of course that's why they've had a ton of traction, especially for example, anthropic and enterprise. And so it's going to be really exciting to see where both those races ultimately end up over time.
A
It's interesting. I mean, Chamath. In a memorable tweet. It's here the docket, Jacob, if you can find it, pull it up. He, he pointed out that there are a lot of jobs that AI companies from the US could do. Even if we take out building the next big frontier model from the equation. Sort of like the, you might think of it as the sort of picks and shovels argument that there's all these like side, side jobs or side hustles we could do. Here's that Chamath tweet. Right now he's talking about things like making chips and hyperscalers, Neo clouds, rack manufacturers. So what do you guys think of that? I mean, is America ready to say we're sort of out of the chase to make the most powerful, most efficient, cheapest model and now we're going to focus on these other things, or is that too premature? Alex, what do you think about that?
C
Yeah, I mean, here's where I come from. First of all, I'm very pro open source. I talk about it a lot. But I also operate from the lens that I believe America has to win the AI model war. Right. I think that in the future all wars will be determined by who has the best AI models because it'll be robots controlled by these AI models. And so I'm operating from the lens. America needs to win that. I don't know how America wins that. If our best labs go bankrupt. How does America win the AI model war? If the people making the best AI models in America go bankrupt, the best AI labs in America will go bankrupt. If, if China puts out models that are 95% as good but 1% of the price and they're operating on a different battlefield than us, they're operating according to different rules. They are willing to steal from us and they are willing to subsidize all of their labs. I don't know how American labs win if the companies operating on different rules are able to make much cheaper models because the companies here, if it's between really expensive models and cheaper models, choose the cheaper ones. And if no one's using anthropic and OpenAI in the enterprise, they can't survive. They simply cannot survive. You know, and I don't know if the answer is government regulation. That feels weird to me. I don't know if the answer is government banning models. That also feels very weird to me. I don't know what the answer is. But I also don't know how America wins the AI model war, a war in which my opinion we must win if our best labs go out of business.
A
There was an interesting tweet from Dean Ball that you sent along. He's an advisor at OpenAI and, and he was suggesting you could do something like a soft ban, like Give Give, put enough rules in place to where companies feel like insecurity and doubt about using open source models and that drives them to using the open AIs and the anthropic releases. Would you be in, would you be in favor of something like that?
C
The issue with regulation, now I know this kind of sounds like I'm playing both sides of this, but the issue with regulation is if you say, hey, American companies, you have to use models that are very expensive. Everyone else in the world, you can use models that are very cheap. If our intelligence is way more expensive, we're going to lose. Like we're, our companies are going to lose the economic war, the global economic war. It's the same thing. If you said, hey, Americans, you have to use gasoline that costs $800 a gallon. Everyone else, you use gasoline that costs $8 a gallon, they're going to just be able to have way more economic activity. And, and so at the same time, regulation, I don't see how that works either. So this is a, this is very challenging. I do. I haven't heard a single good solution to this problem yet.
A
Orey, you're kind of looking at this. You've got an international view that you get to take. I mean, looking at this now, what do you think America should do? Is it free market, let the market decide? If OpenAI loses? Oh well, another company will rise from its ashes. Or should we be taking a little bit more measures to make sure, you know, we're not just getting our models distilled and then defeated by foreign adversaries?
D
Yeah, I'm with Chamath on this, to be honest. I think the free market, we should let the free market decide, you know, and also encourage innovation. I mean, if we overregulate, I think we may decelerate or essentially decentivize this race Right. So I think there's actually plenty of room for innovation where models are being vertically integrated with specific applications or are vertically integrated on the hardware side. So I think the opportunity to innovate and create new value, the pie, is just going to increase and be bigger and bigger and bigger. That's why I think the angle that America needs to win on a global view, not just domestically. And it is so important that the rest of the stack, not just the models themselves, but the hardware and the applications, will also be dominated by US companies. So to Alex's point, I think it's super important that, you know, be careful with that to over regulate and create some dynamics in the market where the rest of the world can get access to this intelligence much more cheaply because that would put America in a disadvantage in the long run.
A
You'll all have to humor me. This is something I've been curious about. I keep trying to get Jason to ask this on the show and he never does. Let's go around the horn. When you say America needs to win the AI race, like we, what does that actually mean to you? Like what do you think of as being, this is what it would mean for America to win versus what it would mean for China to win versus or nobody wins or everybody wins. I mean, how are you sort of defining that category?
D
It starts with the infrastructure, it starts with the hardware layer, right? I mean that's the basis like in telecommunications, right? We have, you know, Huawei now is basically dominating the global landscape in terms of telecom infrastructure. And I think of course now Nvidia is dominating that space in the AI world. And I think you want to have a world where. The US based technology is basically everywhere and is not being taken by other alternative that offers it for a much more cost effective way. So I think it starts with the infrastructure layer, but it goes up the stack. I mean if you have applications that are using other techniques, other models that are based on more efficient infrastructure, that would obviously create an advantage for foreign based companies and activities. So I think we're currently at a point where this sort of war is actually healthy. It creates, it incentivizes people to create more innovations to make this technology more efficient. And we need to keep that balance. Otherwise we'll see some, we may see actually even backlashes. Like I can imagine a world where, and this has been rumored a few weeks back, where China is making export controls on their models and that may have serious implications for other countries in the global landscape. So I think we should be very careful with Regulation at the this very early point in time of the industry.
A
Anan, let's go to you. When you hear America has to win, we got to win the AI race. The Chinese, they're almost catching us. Look how fast the robots are. Like, what are you thinking? How do you define AI supremacy or winning the AI race?
B
I think it's a great question and I like Ori's point around infrastructure. And I think it depends on if we're talking about what does winning the race actually mean. I think you can look at it from a few different lenses, look at it from the lens of, let's say, economic and market dominance. And where's the concentration of most of the intelligence in terms of the power of intelligence actually happening? And where is the diffusion actually happening? I think the second, maybe from the context of the US is largely around geopolitical and strategic power. And so I think that is also an important category. So my take on a lot of this, especially with regards to your previous question earlier as well, is that I don't think that. I think there's this new talking point now that China has some kind of principled thinking or principled approach to why open source models are extremely important to them. I think that's just flat out false. I think that they only care about open source models because they are extremely bottlenecked by compute. And if they were not in that situation, then they would certainly move extremely quickly to launch closed source models. And so I actually think that their tune on all of this may change in the coming years, especially if the bottleneck to compute starts to reduce. Of course, I think we've all seen examples around how some of the labs are working with subsidiaries in Singapore to be able to get more compute. Of course, they're making a ton of progress inside the country to resolve some of these bottlenecks. And so I think once that happens, they're certainly going to march a lot faster towards closed source. I think this is actually very important to them because they care about the concentration of power. And so if we think about the concentration of power, I actually think that that closed source will likely be the strategy that China may take in the long run.
A
Yeah, I mean, if I'm running an open source model, I can ask it all kinds of questions about Tiananmen Square and stuff that the CCP doesn't want me to be able to ask about. So I think, I think, I think that's a great point. Alex, I'll wrap it up with you. We're. When you, when I mean, you you were one of the big people saying this like we need anthropic, we need open AI because America's got to win this war. Like what, what is, what is winning the AI war look like in your world?
C
Well, I think everything is downstream of military, right? Because if your military is the best and has the best robots and has the best AI, which I think everyone probably agrees the future of warfare is robots and AI and that's it. Probably, no, not many human beings involved. If you have the best military robot, AI, military, you have complete economic control. Then China can take Taiwan, which gives them a tremendous amount of economic control over the entire world because their control of the chip production. So I think everything is downstream of military and so robotics. And I think the most important part of that stack is the model itself, the intelligence, because the intelligence can quickly improve the robotics and the hardware.
B
Right?
C
Everything's kind of downstream of the intelligence as well. And so I think winning the AI war is having the best model which gives you the best military now. So that's number one when it comes to closed source or in open source. I believe China only is doing open source to beat America and tank our companies, right. If they can put out models and they're open source, even if we ban the use of Chinese models which they see coming, well, the fact that they're putting it on the Internet for anyone to download, companies over here will still find a way to take it, put it on their servers, host it, which basically kills our companies from the inside, right. American companies will use Chinese open source models hosted in America, which will hurt anthropic and open AI. I believe when those companies fall behind anthropic and open AI and China is able to get ahead because they're not getting the same investment as they used to because of open source models, then they'll be closed source. Once they've kind of taken the lead, then they'll be okay. Well now these models are ours now and they won't be open source anymore. So I believe they're waiting to take the lead before they close source models. You know, I don't believe there's any sort of morals. They just want American companies to fail. And so that's why I believe it's critical that our Frontier Labs are building the absolute best models, need to some way or another be protected.
D
I just wanted to reiterate what Alex said and also highlight another risk which I think is coming from a cybersecurity background. I think once you people think that when you self host these models These Chinese models, you're completely safe because you control the environment and so on. But I think it's important to remember that these models are now becoming more and more agentic. They call external tools, they actually generate code that may actually call for another action and so on. And they're long running and so there are risks of potentially having backdoors and issues there that might have significant cybersecurity risk for these companies that self host those models. So I think we should very, we should be very, as an industry, be very prudent about how we deploy these models and also take that risk profile into consideration. And I think this is probably one of the angles that the government is looking at as well.
A
Yeah, I mean, I think we've seen, you know, OpenAI even recently they took their kind of coding assistant Codex. It kind of baked it into this more agentic overall platform ChatGPT work and it upsets some engineers. But do you see like that maybe is a way for OpenAI and anthropic to continue having some kind of edge to like move away from the chatbot structure and make it more like an all in one agentic workspace that enterprises and people who are maybe a little less technical feel very comfortable just jumping right into and running.
C
I don't think there's a moat there. Is there a moat there in the application layer? You know, ChatGPT work? You see, ChatGPT works just a copy of Claude. Coworkers, right. And now Cursor's coming out with their own work. Those any advantage any of these AI companies get at the application layer within two weeks is gone because all the other labs do it. The advantage is the intelligence. That's the only reason why Anthropic has been dominating the last two years is because Claude point blank has been the best model out there. It just has been. Right. And it's more expensive, but it's been number one at the enterprise because it is the smartest intelligence there is. And so I think the only moat you have in this space is how smart is your models.
A
I do want to move on and talk about agents because we've got two a Gentech CEOs working on this exact problem with us right now. Just as Alex said, all of these agents are being powered by incredibly smart models. We even hear people talk about, you know, AGI is here, so. And yet both of your companies, Petronas and AI21, they're, they're kind of both focused on like agents still aren't, aren't that reliable. We Got to test them, we got to grill them harder, we got to watch everything that they're doing like a hawk. So explain this imbalance to me. Like our models are getting so much more powerful. Why do our agents still make so many errors and we need to sort of watch over them so closely.
B
Yeah. So I'd say that a lot of it comes down to the long, long time horizon nature of a lot of the tasks that we want agents to be able to solve. And so there's a chart I was just looking at yesterday. There's AI research focused benchmark in terms of can agents can actually solve research problems. That came out in November 2024. And one of the experiments that was run was can you run an agent alongside a human to solve the same kinds of tasks? And at what point does the human or the agent actually overtake on performance? And so at the time it was actually that humans would overtake or surpass agent performance on the same kinds of tasks at the three hour mark. And if you look at it now, over a year and a half later, that is now at the 24 hour mark. So essentially in the first 24 hours, agents are able to solve problems more accurately than humans can. But at some point, humans continue to surpass in terms of performance. And if you actually look at the chart in terms of what that performance looks like, it actually looks like a log scale. So essentially it looks like the agent is better than the human for the first 24 hours, but after 24 hours it continues to taper off in terms of how good it can be. And so I actually think that one of the ways in which the labs can actually solve this problem is through things like continual learning and recursive self improvement, where instead of that chart looking like a log scale chart for agent performance, it can actually start to look a little bit potentially closer to exponential. If for example, agents can actually improve as they see and experience new things. And that's actually what we're missing right now, which is why agents continue to
A
be error prompt orey anything, anything to add. And I mean, I'd also ask why is the learning curve for AI agents proving to be so demanding? Is there a lack of flexibility sort of ingrained in there?
D
Yeah, I think Ananda made a really great point. One of the missing ingredients in current AI architecture is the continual learning. And there are ways and techniques to bypass that and try to emulate type of learning. But the reality is that you get a model, the model is frozen. And when you're running these long horizon Tasks, you just change the context and you add and add and add more and more and more context throughout the long horizon task. And in that sense the model is kind of attentive to the new context and the new inferences that were made during the, during the task, but is not explicitly incorporating new knowledge and new learnings through that process. So it is a missing, it is one of the, I think most interesting research areas where and when we spoke about innovation and areas potentially to innovate and create advantage and better intelligence, I think that's probably the most interesting area. The other thing I would say is there is some sense of brittleness. Right. We need to kind of understand that these large pre trained models and then post trained models are trained on huge amount of information, not necessarily of the uniformly level of quality of data. And also some of data points are subjective. It's not necessarily objective data that is being inserted during these training phases. So this may skew some of the inferences we get. We're still working with statistical machines, very powerful statistical machines, but they are statistical machines. And to be, you know, they may be 95% of the time brilliant but still, you know, 5% of the time dumb as nail. And I think that's the challenge and I think that's the challenge that we always, when we think about intelligence, we have this abstract notion of human intelligence and that's what we're trying to mimic here. We have a very powerful machine, but it's kind of different. It's very capable, very powerful, but still very different. So we need to, I think, I always think we need to make sure we have programmatically or when we're dealing this with human interaction. We need to realize both the capabilities which are great, but the limitations of these AI agentic systems.
A
I'm interested in Ananda, sort of Patronus. They're trying to catch agents from failing prior to deployment by sort of running them through world models and testing them rigorously. And Ori, you're kind of focused on stopping agents who are failing at the moment they're actually running. Are both sides of these pipelines sort of permanently necessary? Is this going to be what agentic security looks like moving forward? Or is there a time where maybe if we get the advanced simulation good enough, we won't need sort of runtime validation? What do you both think about that?
B
Yeah, I can start. So I'd say a large part of what we're really focused on is ultimately being able to provide the necessary evaluation and simulation that ultimately can help you measure and ultimately improve models that you're developing. And so the simulation piece is extremely important to your point, because if we think about why agents are error prone right now, especially as we try to get them to solve longer and longer tasks, it's because a lot of the data was not brought into in distribution and it doesn't exist right now. And so, for example, if, let's say an agent can solve a problem, a given problem, over 10 hours, but we wanted to solve a similar problem that takes 40 hours, it just hasn't been proven yet. And so in order for us to be able to do that, we may need to simulate it. And so I think theoretically, if we want to be able to solve for the N plus 1th model capability, we have to simulate it first and then train on it and then understand can it actually now solve it, and then of course, we can roll it out to the real world. And so by definition, in order to solve the capability that does not exist today, simulation is extremely important.
A
Ori, what's your take? Are we always going to need the simulation upfront and then a system like Maestro at the orchestration layer, making sure things are happening the way they're supposed to?
D
Yeah. First of all, our focus these days are on scaling agents. Yeah, I mean, if you think about like six months ago, enterprises in the world was basically on kind of a build and experiment phase. And just in the last six minutes, six months, we started seeing this huge takeoff of these agentic deployments at scale. And now when companies are starting deploying these systems, they start facing new types of challenges because you're facing a scale like we haven't faced before. And one of them is the economics. Right. It becomes very costly and it's become very token inefficient. Like there's a lot of token waste. So we actually focus on the agent optimization piece, where we basically build tool and systems that help people to deploy these agent systems, but doing it very cost efficiently. And remember that as agent builders, you always try to find the right sweet spot between this triangle of quality, cost and latency. These are the three basic parameters that you're tweaking with. And that's a very cumbersome and manual process these days, especially for AI engineers. So our goal right now is basically to equip these AI engineers with tools that will help them find the right sweet spot and kind of operate on the operational frontier.
A
Got it, got it. And Alex, you're over at Henry, you're sort of assembling and scaling fleets of compounding micro businesses I'm sure a lot of this is being done agentically with agents. Where do you find the most problems in terms of setting up and then running agents? Are they flawless at this point? Where are they sort of running into trouble? Where do they need the most human in the loop intervention still?
C
Yeah, for me, it's. It's all around managing context. When you have a lot of really complex tasks these agents are doing based on a tremendous amount of data about the user, it's about, okay, what context do we include for the agent? What about the user should we include that's relevant? What about their past actions they've done in the app? Should we include that's relevant. So all I'm doing most of the time is trying to figure out what's the right context I need to include, because if you include way too much, it could slow it down a ton. It can make it a significantly more expensive. And so context management for me and for what I'm building in Henry, is 99% of the game of what I'm tweaking and trying to make more efficient, as Orey said, cheaper and more performant.
A
Are you is this trial and error? You're just like, oh, I left way too much context in there that time and I dropped $200. I got to like, tweak this moving forward.
C
Yeah. Tremendous amount of trial and error. You know, obviously that's something I'm trying to do agentically. Building different tests and harnesses so it can test itself over and over again to judge its own self and its own quality and its own cost. Right. But a tremendous amount of trial and error and putting money into the machine to figure it out.
A
Sure. Thank you for that wonderful segue back to Anant. I do want to talk a little bit about what Patronus is, is building. Then obviously we'll jump into AI21 as well. So over Patronus, you're building digital world models that, you know, sort of help agents explore a simulated version of the web before we just throw them out there into the, into the wild west of the Internet. I'm curious, like, how those are built. Where does the data come from? I mean, we know real, you know, world models where they're looking at dash cam footage. Like, what are you looking at? And how are you training your. Your digital world models?
B
Yeah, that's a great question. So a digital world model is essentially the analog to the world models that we've all heard about. And those world models are focused on the physical world, so focused on being able to do things like spatial reasoning, 3D applications like robotics. And the analog here is that a digital world model is focused on being able to predict and simulate latent dynamics of the digital world. So essentially the world in which agents operate, whether that's agents operating over web apps, mobile apps, desktop apps, corpus of documents or code bases, anything that we might do in the digital world. And these digital role models, architecturally speaking, they're diffusion models. And what makes them unique is that they are really great at being able to produce highly diverse data, which LLMs are notoriously bad at. Of course, diffusion models are also great for speed. And so if you want to be able to simulate at scale and in real time settings, that's also really, it's an architecture that lends itself really well for that kind of use case as well. And so what we actually do is we actually train digital world models to be able to scalably generate all things related to agent simulations, evaluations that then we can use to do evaluation and post training. And so you can imagine that we, for example, we may simulate synthetic agent trajectories. So for example, all the things that an agent might do in a given domain or a given use case, or a given task even, and then we take those directories which we generate at scale across a lot of different kinds of parallel rollouts, and then we use that to be able to do post training, or we use that to do some reverse thinking, to do data mining, to develop the kinds of tasks that we can then use for training as well. And so it's essentially a very important piece of how we develop simulation data, which then helps all of our customers who are training or fine tuning models accelerate performance across lots of different kinds of use cases.
A
I do have one more question for you, unless anybody else has one before we move on. Petronas used to build products for LLM evaluations. So do you think of this as sort of a pivot from that to digital world models, or is this sort of the ultimate form of evaluating LLMs now we get to see how they interact with the rest of the Internet.
B
Yeah, so it's more of an evolution than a pivot. And the reason it's an evolution is that it's predominantly focused on agent evaluation as opposed to element evaluation and element evaluation referencing. What you mentioned earlier was largely focused on single turn or multi turn chatbots. That's those are really most of the use cases back in 2024, for example. But in the past year, most of the value that we're seeing, a lot of research teams and enterprises get from AI is by being able to solve a lot of different kinds of workflows with Aegis and Aegis that can solve longer and longer tasks. And so the way in which we use the world model is to be able to simulate what agent evals can actually look like across different use cases at scale. And so that's exactly why, why we do what we do and how that evolution happened in the past year.
A
All right, I want to jump over to Ori now in AI21. Talk a little bit about what you guys are working on. You have some slides. Let's forget my question, let's just go in and look at the slides you brought. I want to hear about this presentation.
D
Cool. Now this is just a few data points I wanted to share about what we're focusing these days. Okay, so I wanted to start with this slide. So Brian Armstrong from Coinbase, he just posted a few weeks ago, which I found this very interesting. This is internal usage of AI in their company and I think a lot of companies are aspiring inspiring to see that chart. I mean the black line basically represent the cost of AI, the AI spend. The black line is representing the usage which goes up and up, you know, you see it's growing exponentially. And the bars represent the cost, the actual AI spend. And where you see here in the middle of the chart is where they start applying an AI gateway that they've built in the company itself, which basically is doing several things like switching to default models. So speaking of our conversation earlier, they switched the default to a GLM 5.2 code generation model. They've done some routing, they've done some other optimizations that is described there. And this is actually pretty non trivial. And I think this post by Trey saying only 0.1% of the companies can actually do it at scale is right. It's very hard to do if you want to keep the same frontier quality while controlling the cost. And I think that's the kind of winning picture that companies are aspiring to achieve these days.
A
Can I, can I just pause here and ask you why is this so we have so many numbers. I feel like there's so many benchmark, there's so many tests, there's so many evaluation strategies for these models. Why is it so difficult for people to figure out which model to assign which task?
D
It needs to be battle tested. I mean we've seen, you know, we've seen models that on the face of this are performant, but in reality they provide subpar performance on various tasks. And it's Very nuanced. It's not, let's say a very coarse grain. Let's give these sepsis of models, these types of tasks and those other models, the other types of tasks and also the cost aspects of it. You may think that weaker or open weight models that are cheaper on the face of it, like they have a lower cost per token, that they may be cheaper, but in some cases they may be cheaper on a cost token basis, but basically because they may run longer, they may generate more and more and more tokens on an aggregate level on a kind of dollar per task or price per task, they'll be actually more expensive. So estimating the actual cost of using a specific model and actual what would be the actual performance is a non trivial exercise. So doing this and applying this systematically across different agents is non trivial.
C
Ori how about from an implementation perspective, right? It sounds like one challenge is understanding the cost per task and which tasks each model is best at. But also how about from implementation, how complex is it to make it so you have some sort of routing system that makes the right decisions.
D
So here you need to be very diligent about providing performance guarantees because you are okay with routing, but you want to make sure that you don't hurt performance, right? So you need to have a mechanism that basically predicts the ability of specific model to complete the tasks successfully. You also don't want to. You also need to consider other factors like cash, right? You have models that hit the cache so they have lower cost. If you switch to another model, you're basically losing that cash advantage. So you need to have a full picture of the environment before you make those routing decisions. That's, that's part of why it's non trivial from an implementation perspective.
A
I didn't mean to interrupt. Keep going. I know there's more slides.
C
Yeah, yeah.
D
And just to give a sense, this is a collection of data points about why model routing is so important these days. Why now? And there are basically four factors. One, if you look at the Pareto frontier, these are the points that are dominant on like the points on the chart price performance chart that dominate every any other models, right? So if you go a year back, There were basically 2.6 on average across tasks, 2.6 points on, on that Pareto frontier. Now There are about 5.2 points on average, which means there's more choice, there are more interesting points, more interesting models to choose from. The second thing is the spread. What's the difference between the cheapest sort of model and the highest performing model? The spread is pretty high 60 times. If you look at the recent, if the recent example, it may be in two orders of magnitude, so it's pretty substantial. And then if you look at the frequency, and I think this is very intuitive for us, if you look at the frequency in which models are getting into that new Pareto frontier, it's about six to eight weeks. So it's pretty frequent. You need to keep up with new and updating models. And in terms of opportunities to route differently because of these agentic systems that Anand described earlier, long horizon and they tend to have sub agents, which means you now have the opportunity to route to a different model. So you have a lot of a typical agentic flow actually has a lot of opportunity to route to more performant or cheaper models. So I think this gives you a sense for why model routing is becoming such an important, I believe will become a very important building block in AI infrastructure. My last slide here is about smart division of labor. Showing when you use a portfolio of models in a smart way, and this is something that can be learned, it's not something that has to be dictated but can be learned by a system, a system can actually perform much better and much cheaper than the state of the art. And this is a work we published just a few days ago, showing if you use an ensemble of models like a minimax and GPT 5.2 and Fable together, you're actually better to achieve much better performance in terms of quality, like a new state of the art, much cheaper, like 3 times cheaper than Opus for example. And the cool thing is that it's very human, like what happened here. And this is a Swebench Pro, it's a coding eval. And what you see here is pretty cool, where the system actually allocated the, you could say junior type of work to the weaker open weight model to the Minimax. So it explores the space, it generates many, many ideas for solution to solve that specific problem. And then a better model comes and extracts relevant context to enrich those proposed candidate solution and proposed solution. And then finally you have like an expert, like a fable, like a super powerful model that actually decides and creates the final patch. And this combination is really at the frontier both in terms of quality and cost. And I think it's very human nature. If you think of a legal firm, you have all these interns, right? They do all the busy work and then it rolls up to the associates that tries to synthesize and kind of bring more context. And then you have the partner that makes the final call. So I think it's basically the same thing only applied in AI.
A
Is this one of those sort of like model council type setups where they're sort of all consulting with one another or is this more of just like you do the busy work, I'll sit over here and do the higher levels stuff.
D
I think it's more the latter but it's a more optimal allocation of compute in terms of types of models that are doing different types of jobs and that gives you the result. So I think the intelligence will see and we spoke about innovation and how to increase the level of intelligence is actually going to be by smartly assembling this portfolio of models and smartly orchestrating these different types of models to achieve the best outcomes.
A
I was going to ask my one last question here for you. There are a lot of other products that are kind of in this sort of routing help you choose the best model space stuff like Unity, AI Gateway from Databricks or router from Ramp. How does your sort of the AI21 take on it? Like what are you doing differently? What's your kind of unique spin on here's how we're going to help you route the right work to the right model.
D
Yeah, so our kind of secret sauce, we do this on a per agent basis. So we learn the traffic per agent. Each agent has its own distribution and then we know how to construct online these routing policies. So basically get a much more fine grained and more efficient routing policies per agent. And we've seen many of these products, when you apply them into different agents, you see very different types of savings. What we've realized is when you have a continually learned router, you actually can get these savings consistently over time.
A
Interesting. Awesome. All right, before we wrap up what everyone's working on corner, I think we'll probably have time for one more newster. I want to go to you Alex. You recently tweeted Open source has officially caught up to the frontier in your opinion. So I want to know now that you have all of this extra intelligence for so much cheaper, what are you most excited about? What are you using it to build right now in your own workshop?
C
So I'm using it to build a lot of things. Obviously my main focus is Henry intelligent machines but I, you know I'm building my own home AI lab. I have a bunch of Mac studios connected together. I have three 512 gigabytes. I have a couple sparks, an AMD computer that just came out. I'm working on building an ambient AI in my own home. AI lab right now that kind of watches what I'm doing and adds intelligence to everything I'm doing. So every piece of content I put out, it automatically takes, repurposes it for my different channels, automatically edits different videos I put out, automatically is able to keep an eye on my email and alert me when certain things happen. So you know, I think one of the big positives of open source AI and being able to run these models locally is this idea of ambient intelligence that just constantly watching everything you're doing and helping you along the way proactively. So that's my big focus at the moment. Now that we have like actually really good intelligence we can run locally.
A
And is it like I have agents as well that I try to use for stuff like this and I find that it's great when it's task oriented, when I'm like, hey, figure this out. When it's proactive, it's suggestions still not, still not that great. Like not as good as if I had a junior human. How are you training them over time to get smarter about what kinds of things to proactively suggest to you from versus what stuff is not going to interest you as much?
C
Yeah, I find the proactive suggestions usually aren't good when you don't have good kind of context and guardrails on it. And so for instance, I, I, I was trying to get to proactively repurpose my content in newsletters and then send it out to my audience because newsletters would take up a tremendous amount of time. I have 40,000 subscribers. It's just taking up a lot of my time. I didn't want to spend as much time on anymore. And so what I actually did was I hired a newsletter consultant who wrote, who repurposed my content to newsletters for about two months. I then took their repurposed newsletters, fed it to my agent, my local AI as contacts, that this is what the, this is the newsletters they repurposed, this is the sources they use, the different content they repurpose from and the locally I actually reverse engineered what they were doing and now the output's a lot better. Now it proactively creates these newsletters and sends them out at a significantly higher quality. And so I find when it's just like, hey, take my content, repurpose it, yeah, the output's really bad. But when I was able to get like kind of expertly done repurposing and have my AI reverse engineer that, now it's actually really powerful. I have GLM5 2 running on a Mac Studio right now. And it's able to do it really, really well.
A
You Merkored it. You created your own, like mini Mercour where you put a human expert in the loop.
C
Yeah. If you can just get a human expert in the loop for kind of different parts of your tasks, get as much data on what they did as possible and hand it to an AI, it can recreate whatever they're doing really well.
A
That's awesome. We do have. I have one more video that you posted that Jacob has ready to show here. He swears to me, your LCD typewriter system. I want to know what was the thought behind this? Like what? What's the advantage of having a system you could run your agents on, but it doesn't have like a browser or the usual desktop stuff?
C
So I have this here. This is like my favorite device ever. So I bought this. This is a Pamera DM250, I think it's called. It's just a digital typewriter. Like all it is is a keyboard with a black and white LCD screen with a word processor on it. That's it. That's all this is. And it's built for like writers and authors to write things on a focused level. Right. Just a typewriter. I have a tremendous amount of add. The hardest thing in the world for me to do is focus on literally anything at one time. And so, you know, I vibe code a lot and use Claude and Codex and all that, but I have like this two monitor set and I'm just constantly distracted by Twitter and scrolling. It's just incredibly distract. Email. And so I wanted a device I can just lock in on, open it up, put everything else away and use AI, talk to an agent or write in codecs or have cloud code, build something. And so I figured out I could actually reverse engineer this digital typewriter, get Linux installed on it, and then have it SSH into my Mac studio. So it connects my terminal sessions here and then I can have a terminal up on here and just use Claude code and codecs. Zero distractions, zero anything. All it is is an AI device where I can talk to Claude and ChatGPT and it can build things out on my Mac studio. And so for me, someone who is just constantly distracted at all times, this has been amazing. I can lock in my time from like 12pm to 4pm every day. I open it up, all I got is Claude and codecs. And all I can do is sit there and build things out, not get distracted by anything. And so it's the point being is like, it's really amazing what you can do when you kind of point Claude code at something and say, hey, hack this. Figure out different ways we can use this. And it's like, okay, here's what we're going to do. Plug a SD card into the computer I'm on right now. I'm going to get a version of Linux on it, then take that SD card, put it into your device, do all this, and it'll be set up on there. And it like walked me. It like prompted me to hack the device, which was amazing.
A
If. If Sam Altman and Johnny I've are watching this right now, I think you're screwed. I think they're gonna. I think they're gonna borrow this idea. It's very visionary.
C
I mean, it's just a cool kind of weekend project. Find devices you're not using anymore and then see what new ways you can use them with claw.
A
Repurpose them so you don't get distracted by tweets.
D
I should, I should try it on my Sinclair from the 80s.
C
I bet you could. As long as you can, like have it ssh into another device, which doesn't take really any compute at all, it can do pretty much anything you want.
A
All right, one more news story before I let you all go. It's almost wrap time here at this Week in AI, but I do want to talk about the Hugging Face breach. This is the first, they're saying, publicly known case of a AI on AI cybercrime involving a major platform, of course, Hugging Face, popular platform for sharing and hosting AI models and data sets. They published this blog post recounting a major attack on their production infrastructure that was, and I quote, driven end to end by an autonomous AI agent system. The company's own AI then detected and responded to the attack. Here's the interesting wrinkle on this one. VentureBeat then reported that Frontier AI models like Fable 5 declined to help hugging Face's team during the attack. They cited safety concerns that violated their guardrails. So the Hugging Face team ended up running their entire investigation into the hack on GLM 5.2. This was echoed by our former aizar and friend of the pod, David Sacks. He tweeted that Kimmy K3 fixed 15 critical security bugs that Codex and Fable refused because of cyber guardrails. There's no reason to limit American models on tasks that Chinese models handle without issue. We're only making ourselves less competitive. So, Anand, I'll go to you first. Do you think AI on AI Attacks are more common than we've maybe heard about. And, and this is just the first one that sort of Hugging Face decided to admit to openly. And do you think that these frontier models should be able to help us out without putting up guardrails that prevent people from using them for cybersecurity?
B
Yeah, definitely. Well, I'd say yes to both. I think there are a lot of AI attacks that have already happened that have certainly been swept under the rug, and it's certainly the case for companies that don't have as large of a public profile yet as hugging Face does. But I also expect that the number of these kinds of attacks are only going to continue to go up, especially because in the same way that we think about humans, if agents can solve longer and longer problems, and there will certainly be a ton more surface area for these kinds of things to happen. And to the second question, I also certainly believe that there will be a regression to the mean. I think we've gone quite far ahead in terms of deploying a lot of guardrails to the point where, of course, now it is hurting capabilities. And so I think we're starting to see this tug of war happen between capabilities and safety, which hasn't happened before. And I think this is the first time. And so given where things are now, I expect there will be aggression to the mean in terms of us loosening some of the guardrails or at least being a lot more thoughtful about how we deploy them. And I expect that that's going to continue to happen over time, too.
A
Orey, what do you think? Have you run into an issue where the AI, you know, frontier AI models won't let you ask a question? You run into the guardrails. Has this been part of your experience? And what do you think about the way that we're sort of handling AI security at the moment?
D
Yeah, I think it's. We're kind of cuffling our hands here. We've these guardrails on one hand, on the other end. It totally makes sense to, to put them together. Right. Because I always said when, even back when, you know, in 2021, when the first GPT came out, that, you know, people spoke, spoke about doomsday and, you know, what, what, what will these machines discover and what will be the impact on society? I always say that cybersecurity is probably the most risk area for these types of models as they become more and more capable. And what I think the solution, or at least a path to be very practical, is actually kyc, you need to know, and you need to give access. If the person or the organization that is using this technology is clearly
C
a
D
non malicious actor and can have more permissive rights to use more capabilities, then at the very least we should loosen the guardrails for these types of organizations that need this technology to face these cybersecurity risks. Right. I think that's a very natural way. I think that's what we're going to see happening in the next year or so. So the guardrails will not be like. One size fits all. It will be also according to your KYC and the risk profile associated with it. So that's where I think we're heading with this. At least I hope so organizations will be equipped to deal with these very challenging cyber attacks.
A
Yeah, Alex, we'll, we'll close on you. You're a big Fable 5 fan. Did you ever run into guardrail issues? And what do you think about, are we kneecapping these models unnecessarily to try to make them safe?
C
I had a funny situation yesterday where I'm like, I'm trying to build my own benchmark for models and one of them is Claude's building this huge app with like 500 different bugs in it. And I'm going to use it as a benchmark with other models to see how they, how many of those bugs they can solve. And, and during the building this benchmark, Claude actually went. I actually foresee Claude safety guardrails stopping models from using this benchmark. So I'm going to build into the benchmark. I'm going to say it's like a play, it's like a pretend situation rather than a real situation so we get past the safety guardrails. So I was just in this funny situation where Claude was building workarounds to its own safety guardrails in something I was building yesterday. I think that open source is the big kind of bomb in this situation that's going to force the hands of a lot of different of these labs because open source models are not going to have a lot of these guardrails. There's going to be models out there that are completely uncensored in the future, especially as it becomes a lot more efficient and optimized to both host these models and fine tune these models and train these models. People are going to take these models and make them so there's no safety guardrails. And so if your Frontier lab puts out a model that's way too restrictive on guardrails, but there's an open source one that's really cheap to run, that I can run a Mac Mini that allows me to do whatever I want. I'm not going to use your Frontier model. And so I think right now, at least, yeah, they need to be looser. It just goes back to Opus way too much with Fable. But in the future, I mean, it's going to have to be much looser because it's going to be too easy to use alternatives that have no guardrails whatsoever.
A
By the time there's GLM 6.2, we won't be worried about it.
C
Right.
A
All right. That is our show, folks. We made it. I feel like 10 times smarter. I feel a real sense of accomplishment. Before we go, I'm going to finish it off with Jason's patented question. We'll go to you and on first, are you hiring at your company right now? And do you want to make a case to the viewers for why they should come work for Patronus?
B
Yeah, we are hiring a ton of we're hiring a lot of AI researchers and engineers to continue to help us build what we believe is the most important infrastructure to unlock the next frontier of intelligence. And if I had one reason for why you should join our company, I think that we're working on some of the most intellectually stimulating problems out there. And so if you care about solving some of the most interesting problems, interesting technical problems, then we're the right place for you.
A
Alex, thank you for being here. Thank you for making me feel so much more comfortable, comfortable in my own skin during hosting this episode of this Week in AI. Always appreciate you being here. Thanks so much for joining us, everybody. We'll be back next week probably with a different host for episode 24. Thanks for watching this Week in AI.
America’s position in the global AI race is under threat, with a specific focus on open-source AI models originating from China and their competitive, economic, and security implications for US "frontier" AI labs. The expert panel debates if America can keep its AI lead, what “winning” the AI race means, the right regulatory approach, how agentic AI systems remain error-prone, and emerging best practices for deploying and optimizing AI at scale.
“There is a ton going on this week we gotta talk about with this panel.” — Lon Harris (02:13)
Main Topic: US government exploring potential restrictions on Chinese open source AI models—are these real threats or just policy noise?
Each panelist offers a different vision:
Anand Kinopan, on China’s gains:
“China has already won one of them. The first race is the layer of intelligence that can be measured cleanly and copied very cheaply ... they’re catching up very quickly, especially if you just look at the charts in terms of performance across a lot of these benchmarks.” (03:38)
Alex Finn, on US labs going out of business:
“If China puts out models that are 95% as good but 1% of the price ... I don’t know how American labs win if the companies operating on different rules are able to make much cheaper models ... I also don’t know how America wins the AI model war, a war in which my opinion we must win, if our best labs go out of business.” (05:41)
Ori Goshen, on hardware and innovation:
“It is so important that the rest of the stack, not just the models themselves, but the hardware and the applications, will also be dominated by US companies.” (08:50)
Anand:
“I think there’s this new talking point now that China has some kind of principled thinking or principled approach to why open source models are extremely important to them. I think that’s just flat out false ... They only care about open source because they are extremely bottlenecked by compute.” (13:10)
Alex Finn:
“I believe China is only doing open source to beat America and tank our companies, right. ... When those companies fall behind Anthropic and OpenAI and China is able to get ahead ... then they'll be closed source. Once they've kind of taken the lead.” (16:01)
Anand Kinopan:
“There’s a chart ... benchmarking if agents can actually solve research problems ... At the time ... humans would overtake or surpass agents at the three hour mark. Over a year and a half later, that is now at the 24 hour mark. ... But after 24 hours [human] performance continues to surpass.” (20:25)
Ori Goshen:
“We’re still working with statistical machines, very powerful statistical machines, but they are statistical machines. ... They may be 95% brilliant but still, 5% of the time, dumb as nail.” (22:22)
Anand Kinopan:
“We've gone quite far ahead in terms of deploying a lot of guardrails to the point where ... now it is hurting capabilities ... expect regression to the mean ... loosening some of the guardrails.” (54:10)
Ori Goshen:
“You need to give access. If the person or the organization ... is clearly a non-malicious actor and can have more permissive rights ... then at the very least we should loosen the guardrails for those types of organizations.” (55:39)
Alex Finn:
“If your Frontier lab puts out a model that’s way too restrictive on guardrails, but there’s an open source one that's really cheap to run ... that allows me to do whatever I want, I’m not going to use your Frontier model.” (57:46)
| Topic / Quote | Timestamp (MM:SS) | |-------------------------------------------------------------|-------------------------------| | US considering ban/restriction on Chinese AI models | 00:00–03:38 | | Panel intros | 00:59–02:13 | | “Two AI races” (Anand: China already won one) | 03:38–04:55 | | The “picks and shovels” (Chamath tweet, US jobs up the stack)| 04:55–05:41 | | “America must win the AI model war” (Alex) | 05:41–07:15 | | Regulation vs. free market (Alex, Ori) | 07:38–10:20 | | What does “winning” mean? | 10:20–15:17 | | On military and economic dependency | 15:17–16:01 | | Security, backdoors in open-source models | 17:19–18:32 | | Application moats & intelligence as only advantage | 19:01–19:44 | | Why agents still fail (long horizon tasks) | 20:25–22:22 | | Simulation & model evaluation (Patronus, AI21) | 25:17–27:12 | | Context management (Alex: Henry AI) | 29:25–30:34 | | Patronus: Digital world models | 31:07–34:31 | | AI21: Model routing at scale (Coinbase example) | 34:44–39:42 | | Ensemble models & division of labor analogy | 43:42–44:11 | | AI21’s agent-based router | 45:25–46:16 | | Open source equals ambient intelligence (Alex’s home lab) | 46:39–47:36 | | Practical mentorship for AI agents (Alex’s approach) | 48:02–49:18 | | Repurposing old devices for AI, digital typewriter segment | 49:55–52:37 | | Hugging Face AI-on-AI cyberattack & guardrails debate | 52:37–57:46 | | Closing remarks, hiring asks | 59:35–59:52 |
This episode offers a highly engaged, expert-level view of how the global AI model race is playing out in real time, what it means for practitioners and enterprises, and where both the threats and opportunities lie as America faces the reality of open-source disruption.