Loading summary
A
The best open source models are coming from China. It's basically going to hurt the AI strategies of tons and tons of companies that are building with an open source first mindset.
B
Mathematically untenable if one model GPT4O level from OpenAI is $6 an hour. So you can't use 31 because the math blows up on you unless you use open source. If you run our Constellation all on OpenAI models, it'll cost you $105 an hour. That's more than the human. The real power of all of this AI is not replacing work you do today and making it cheaper. That's actually super. Well, thanks to our friends at PayPal, the exclusive sponsor for this Week in AI. Try the payment and growth platform that's trusted by millions of customers worldwide. PayPal open start growing today@paypalopen.com hey everybody.
C
Welcome back to this Week in AI. My name is Alex and today we are talking about the possible end of open rate models from China. AI live bearishness, benchmarks, the great AI data race and more. To help us understand it all. We've brought two leading AI startup founders to the show today. And in one corner we have Anastasios Angelopoulos from the company arena, now previously known as L and Marina. It's a brilliant data set for tracking comparative AI model performance. It's now a big darn startup having raised $150 million in a blockbuster series A. Anastasios, welcome back to the show.
A
Happy to be here.
C
Glad to have you back. We also have Moonjal Shah from the hot shores of vertical AI. His company Hippocratic AI is working on all things clinical voice agents in the healthcare space. With more than $400 million raised including 126 million at a $3.5 billion valuation recently and version 5 of its Polaris err, model out the company is shooting for megascale. Moon Jael, welcome to the show.
B
Thanks for having me. I'm excited to be here.
C
I'm so glad to have you both here because AI news is dropping like you wouldn't believe. The latest thing that's going on that we all have to talk about is the possibility that Reuters reported this morning that China may end the release of its AI models to the world. Now, we all know that Chinese open weight models have been incredibly super powerful, super cheap. A lot of startups build on them and if that ends, it does seem to kind of reorient or reorder the broader AI landscape. Anastasio, starting with you, first, reactions to this news because I think we're all still digesting a bit.
A
I mean, listen, a lot of people are going to be made really upset by this because open source models, the best open source models are coming from China. And so it's going to do a couple of things. The first thing it's going to do is it's basically going to, you know, hurt the AI strategies. Tons and tons of companies that are building with an open source first mindset. The GLM is like basically a frontier quality model. It's comparable to a lot of the frontier offerings from OpenAI.
C
GLM 5.2.
A
Yep, that's right. GLM 5.2. And so, you know, if we can't get 5.3, 5.4 or GLM 6 or whatever, you know, what have you, what's going to happen is that people are going to have to reconsider what models to use, whether they're going to go with a closed source strategy or, you know, and this is the second thing that's going to happen. It's of course going to promote open source development inside the US and Europe. Models that, you know, we can actually rely on to, for to be released within the confines of our actually sovereign nation.
C
So do you think that this is a net benefit then to American and European AI labs or also American and European AI startups?
A
Well, yeah, there's, that's exactly the right question because there's a bunch of game theory you need to play out. If you're China, if you're trying to, you should be thinking, okay, how do I make China win? And so if you shut off open source AI now, let's be honest, China is, is doing well there, but they're not exactly winning the AI race yet. So by shutting it off and not giving Chinese companies the ability to compete on the global stage, not giving them the ability to sort of access global data, not giving them the ability to access global customers, what they might actually be doing is sort of, you know, cutting the budding trees that they have in China related to AI and then promoting the development of Western open source AI. So to some extent it might actually be good for our national security.
C
Mandal, where do you stand on the winners and losers from this? I know your company makes its own models, so I presume you have a foot in all open and closed source camps.
B
Yeah, actually, you know, so we only use open source as our base. We do massive post training to those models. We put them in a configuration of 31 models together and that's how we ensure clinical safety. We Actually found one model kind of wasn't enough and one model open source without significant modifications wasn't enough. Like we actually modify 100% of the parameters in the open source open wave models that we have. And so we started originally by building on the meta of open source products, and those were state of the art at one point. But you know, we have migrated to actually working and leveraging the Chinese open source models at this point because they are the best. And so one, I think one of the big implications is open source today allows startups to compete with the big model companies. And so this will actually destroy their ability to do that, including ours. Like, it's not a great thing if this happens. So second, people don't realize the cost implications. So, you know, we sell our AI voice agents for about $9, $10 an hour. Okay. We're using 31 models in parallel because for voice you need very low latency. Right. So I gotta fire them all at once. Now, just one model GPT4O level from OpenAI is $6 an hour. So you can't use 31 because the math blows up on you unless you use open source. But actually using those 31 creates a level of safety we've not been able to do with one model, no matter how smart it is, because you have a model checking a model checking a model kind of thing, and models staying very focused on just doing one thing, like looking for overdoses in a call. And so while the models can scale the context windows, you can't scale the attention span nearly as well. Right. You put in a million tokens. Great. Which tends to focus on the last 10,000 or the first 10,000 or a random 10,000. You don't know.
C
No one knows that. One million context window. It's on a billboard.
B
Context window being good? No, not always. Right. So, but I think that, I think that now you can't run multiple models if you don't have open source. And so there's an entire class of use cases that actually need a multimodal architecture to ensure safety, accuracy at a high precision that become mathematically untenable. If you run our Constellation all on OpenAI models, it'll cost you $105 an hour. That's more than the human.
C
Yeah. If you read it all on Fable 5, how much would that cost you? $10 billion an hour.
B
$10 billion an hour. But so look, I think there's a cost implication. There's that the other part is, you know what Alex Garp was saying last week, which is That a lot of the businesses don't want to share all the queries back with some of the frontier model companies. And open source hosted in your own environment on your own servers is which is what we do by the way. We take the open source and run it in our own place. We don't run it because we modify it so much. We have a different set of weights by the end. But that is also. So there's a security implication too which in health care one does care a lot about the security, the privacy of the data.
C
I just want to double click on something. When you say you use 30 plus models in Harmony with one another, does that mean you have post trained each one of those from an open weight base or is there one model that you trained heavily and then sharded into smaller components?
B
Largely most of them we've taken an open weight base of different sizes because we don't need all of them to do the same thing. Like one is checking and ensuring that HIPAA authentication is always done correctly. One is an escalation engine that actually transfers the call to a human if you're having chest pains and runs something called the Schmidt Thompson Protocol, which is what nurse triage call centers run to ensure safety. I mean we're one of the few LLMs that will literally hand off properly. We'll run a escalation, we'll hand off the call to human if needed. But we do not, not all of the 31 are post trained and sign or fine tuned. They're based, you know, there's I think a few of them that are just a good set of prompts on an open source model. But that's, that's a very small number out of the total.
C
Do any of them qualify as an slm, a small language model? I'm just curious if there's, if there's any place in your multimodal setup. No.
B
Okay, now this is the fallacy. I mean it's like I'm a small language model. I'm like that's not a thing. This thing only worked because it was a large language model, folks. And so what, what we, what we found in our system. So for example our entire Constellation is now 5 trillion parameters. The main and so what we found was every time you use a big model it catches more out of distribution things better. Which in health care matters we had the other day a scheduling call where everybody's like scheduling, that's not hard. How difficult is that? Well, this guy called up, he said I need an appointment with my neurologist The AI was like, tell me why you need appointments. I get your right appointment type. He's like, I was struck by lightning yesterday at work.
A
Huh?
B
Like, like, you know, and the AI had to reason because the answer isn't send him to the ER911, it's the next day. He's obviously okay.
C
Not, not an imminent risk, but also
B
not schedule him six weeks from now. It's like, you better get him in today or maybe tomorrow, you know, a more high priority appointment. The AI that's not in the rule set, if struck by lightning yesterday and still standing, do this.
C
That's the power of generative AI kind
B
of big models to do that and everything else. They won't handle those. And so I think the smallest one we use is 7 billion parameters. The biggest one we're using is over a trillion in an MOE structure mixture
C
of experts, which is a way to reduce the overall processing load of a larger model by looking for one slice of it that has a specific expertise. Anyways, Anastasios. So on the point of startups using a lot of open weight models to build their own thing, there's been a lot of conversation about how open models from China have been getting closer to what I'll just call the American frontier performance. I'm curious from the arena perspective how true that is, because on one hand it does seem that there has been a closing gap, if you will, but also it seems like a bit like that truck gift where the truck never actually hits the pole and they're never actually going to quite get there. So I'm just curious about how much performance we might lose overnight if these models were kind of taken off to market.
A
Well, it's absolutely true that China has continued catching up to the US frontier. And so a lot of the reason why people are saying that is because of Arena. They're looking at arena and saying, hey, GLM 5.2 is. Yeah, exactly, is, you know, near the top of the leaderboard right now. If you look at agentic performance, we have agent area, an agent. Arena is measuring the ability of models to do general purpose agentic tasks, like, you know, the cloud code type task or cloud code work type tasks in your browser. So we have millions and millions of traces that we collect every week that allows us to assess these capabilities. And what you'll see if you do that is that GLM 5.2 is about a GPT 5.5 level model, which is
C
impressive because that just came out.
A
Super impressive. Yeah. And you know, it's, it's probably like tied around. Tied with 5.5 high, which is, which is pretty incredible performance. The X High version of GPT is probably, is a little bit higher and what you'll see is that it's, it doesn't quite have the same level of tax success rate. And yeah, you pulled it up on, on the screen here. It doesn't quite have the same level of steerability, but it's quite close and especially with a little bit of fine tuning, can actually exceed the performance of these models on an enterprise's workload. So it's definitely true that China's catching up. And what happens if you lose these models? Well, the next best open source model, I mean, let's look down the leaderboard. If you exclude China, it's like you're losing vai, you're losing Kimmy, you're losing Deepseek, you're losing Minimax, you're losing Quinn. And then where are you like Gemma?
C
Well, no, well, kind of. I think you're more at Nvidia's Nematron models which have been pretty. Oh, Moonjal is nodding his head back and forth. Moonjal, what's your opinion on the Nvidia Nematron model family? Because I think you have one.
B
No, no, I mean, look, I think it's awesome that they're doing it. I really wish that they continue to progress and get to the same level that we're seeing with the Chinese models. Right. In terms of performance, because we need it. But you know, I think they have yet to release Nematron Big or whatever their kind of version of the largest one is, if I remember right, unless it came out already, Anastasia, you know, I guess better than me. But like, you know, we need these open source and actually there's a different reason we need them. I love these leaderboards because they allow us to rank. But I also believe the coverage of all the tasks you would need the model to do that's covered by the benchmarks is very small. And so, you know, there's this artificial jagged intelligence concept that kind of says, hey, you know, these models are good at some things, really bad at other things as they've been improving. What we notice on the tasks that we need for healthcare activities and automatic conversations with patients is that the peaks get better but the troughs don't move much. And so we have to do a lot of post training in RL to move the trough up. And so, you know, because this, I mean if you think about all the possible outcomes of a model that are tested in, in Terms of the benchmarks, it's still a very small coverage area and we try to get benchmarks that are, you know, indicative of other areas. Right. But it's, it's not perfect. And so we have actually built a whole set of new benchmarks that we're using in the healthcare setting and then actually starting to benchmark all the different models on them, including the voice to voice models by the way, which are super, which are really dumb. There are no voice to voice non cascaded models that are good. Everybody's like they sound amazing. I'm like, they do, but they're dumb as heck. And, and so I mean, I think that we just have to realize that the other value of open source is that it allows you to improve these troughs and in, in kind of the blind spots where the current benchmarks are not kind of measured.
C
We're going to get to benchmarks in just a second. So I have questions about that, but I want to make sure that we're getting this right. If the Chinese government does decide to preclude global access to open wave models from its myriad high performing, sometimes public AI labs, there isn't an obvious immediate replacement from broadly the West. And so we would be in essentially a open weight model drought. Is that fair, Anastasios?
A
Yeah, I mean, listen, who are the losers? The losers are all of the enterprises that are building on those open source models now and don't have a good alternative. And the winners are the current American open source ecosystem which is not, you know, it's not caught up with China, but likely will if all of the traffic moves to them. So those players would be players like nemotron, it'd be rc, it'd be Mistral, it'd be Google with Gemma, Reflection, AI and Reflection and Reflection.
C
Mistral is French, but I mean we'll count it in our bucket. They're part of NATO, so it's all the same. I mean, one big happy family. Right? There's no attentions whatsoever in that, in that domain.
A
The 51st state.
C
Yes. I thought that was Canada, but you know, we'll take them now. If there isn't a replacement for these and no company currently can match, we're talking about seeing once again a lot of American businesses being overly dependent on China for a key piece of their operations. And I feel like we kind of ended up in the same place we were with manufacturing, with AI and it's just disappointing. But the reason why Anastasia is I'm a little bit skeptical of your claim that American OpenAI open weight AI, just let's say it that way. To be clear, we'll catch up as quickly as. Even in China they're having a hard time monetizing these open weight models. There's a story in the Times, I think today about how Alibaba is really struggling to monetize its Quinn Group. And if you look at the Minimax and Z AI's earnings as public companies, very small compared to American closed source. So I wonder if there's even a financial incentive to do this or if we're just going to see companies like Hippocratic kind of stuck without having the same, you know, ingredients to keep making improved soup with with these open weights.
A
Yeah, I think it's a great question. It's a question that I brought up actually many times in, you know, various podcasts, which is like how, what, how do you think about the, the monetization strategy and the business model behind open source? So let's just to, you know, get deeper into the question, what is the reason why we're asking this with software? The reason why? Oh, you can create an open source business model. You can say, okay, I'm going to have this open source software and then I'm going to become the best place to run this open source software. I'm going to for example, have Spark and then I'm going to build databricks on top. Databricks is not just Spark. Databricks has like a huge thick value layer on top of Spark. They're the best place to use Spark and the best place to store your data. Blah, blah, blah, blah, blah.
C
Apache Sparks, the open source project.
A
Exactly. And then that's how you build your value. Now with open weights AI it's a little different because you can just download the weights. Nobody else can contribute to the open source project really because nobody has access to the data and the compute and the blah, blah, blah to be able to train a big model and then you just can run it on any, basically any cloud.
B
Yeah.
A
And so there's really, you know, it's unclear what the business model will be. Now what I have heard that people are doing in order to try and create an actual business model around open weights AI is to do these licensing agreements where in the license of the model they have a. The basically if you spend enough or if your company is large enough or earning enough revenue that you, that the fact that you're using this open weights model will entitle you entitled the company to a rev share. So they'll start saying, okay, if I'M if you're like a $10 billion company with, you know, you know, over $1 billion in revenue and blah, blah, blah, then 20% of your revenue goes to us because you're using an open weights model.
C
Moon job. If you had to cough up a chunk of your revenue to a open weight American or Western AI lab, would that dramatically change your economics? Or could you afford a 20% margin, hit on your inference costs and still offer your product at an attractive rate?
B
It wouldn't be a big deal. I mean I, I think that it's so much cheaper than running it on the closed source frontier models when you like. It's just so much cheaper because I mean it's very simple. You just can't. You as a company cannot sell on top of a frontier model because you can't double stack software margins. That's it. It's that simple. You can't take an 8% margin, add another 8% margin and then sell it. And you certainly can't put together lots of models around it. And so if there is a, I mean, I think Anastasia is right. There needs to be a business model there because there isn't the same network effect that you get in traditional open source software where everybody's contributing to it and it's all getting better and everybody's gaining from it. But at the same time this is critical to us not having a bipolar, unipolar world of frontier models and everybody having to pay the tax and everybody being worried like Figma, that they're going to be stealing them. And now in the, we serve pharma as well. And as you can see, Anthropic recently came out and said, oh, we're going to build drugs too. I'm like, oh great. I'm sure the pharma guys are super excited to keep using Anthropic after that.
A
Right?
C
We're going to subsidize our own execution. I think I'll pass. Thank you.
B
I think I'll pass. And so this is very critical and we. There's a structural. I actually had this conversation a long time ago actually even with Dave Sachs when we were at the White House together. And I was like, David, the most important thing we got to do is ensure that US Open source continues to exist and is a, is on the leading edge of open source. And I, so I think, you know, maybe this gives us more impetus to get that initiative going and maybe Nvidia, with obviously an infinite supply of GPUs, is the guy to help pioneer that with Nematron. That would be awesome.
C
I just don't want Nvidia to become the next open source AI giant because I don't trust them to not bias that in their own favor. They are a enormously profitable company. They're the most viable company in the world. I don't think they need more power. But I'm curious why, and I mean this with nothing but love and respect, your aspirations aren't higher here because you're already doing all this work to collect the data. You need to test it out in practice to fill in the troughs of the, you know, artificial jagged intelligence line. And you're doing all this work already. Why not just do it all yourself? Soup the nuts. Or form a consortium with some friends to share the cost if you had to.
B
So two things. One is we have found the safest models are the biggest models. Biggest models are really expensive to train. And then I got to do it every year.
C
Hey look, my phone's ringing.
B
David.
C
David Sacks is calling with a check from his enormous pile of money.
B
Yeah, I mean unless David, I mean it's probably like $400 million and you have to build the team to do all of the pre training and to gather all the data and you have to figure out some of these tricks that are hard to figure out that somehow China's getting, figuring out as well. Like not an easy place to be. And honestly, that asset you're making is valuable enough to other people who don't compete with you that it really should be a shared cost. Now maybe there's some consortium of it, but I do think that the person best positioned, you know, also I had conversations there and I Nvidia is at least open to not only open sourcing the weights, but open sourcing the training data, open sourcing everything. So all of us can do continuous pre training and other neat things that we want to do on it. So but, but net net, it's a very expensive lift to have to keep making your own open source because you got to do it like every single year. Right. It's not a one time thing. No, it's continuous and you have to stay on that. I don't think one startup by itself could ever be kind of capitalized properly to do that unless it just becomes a frontier lap, you know, effectively. And, and, and you know, there's not an easy way to have a consortium around that. I think this is a critical piece of infrastructure that we need. And luckily we had it in the beginning. It then the dominant players became Chinese in IT and now Nvidia maybe can lead us back, but I don't see how startups continue to fight this war even and verticalize and build unique products without it. And actually the companies that have built on the frontier models, what's happening to every single one of them?
C
They're getting eaten by the frontier models.
B
Yeah, well, not only that, but right now their unit economics are all horrible or upside down even. Right. So they're all like, oh, don't worry. I mean, recently I heard one of them raised, I won't say which one, a vertical AI company that's in many verticals. And it was like, yeah, we know we're totally upside down, but by the way, we're going to build on open source later and swap out these models and you know, Cursor did it and these guys did it, it'll work. And I'm like, first of all, I mean, if coding is probably very specifically trained on it, but net, they won't be able to do that swap if there is an open source to exist. So I think this is super critical infrastructure.
C
I mean, look, I'm a, I'm a capitalist and I'm not a big fan of state capitalism, but if we're going to have a national AI project, maybe this is the place we could be putting some more of our work.
B
But I also don't see myself right. I'm not sure China's advantage right now is that their products being used as open source. I think they will. They need to find a different answer than we're just going to block this. Because as soon as they block this, like what influence do they have on the worldwide AI game? I mean, they're just going to build a frontier model company and say, hey, that generated a lot of money, so we're going to generate one, but ours isn't quite as good. So then what is it? Why am I buying it? Is it cheaper? It's got to be something otherwise, like, why am I using. It's still an inferior model, technically speaking, to the frontier models.
C
All right, let's go back to benchmarks, which we touched on just a little bit ago. I was going to bring this up in the context of the latest from our dear friends over at Doordash, clearly a leading company in the broader AI conversation. That's sarcasm if you didn't catch it. They recently dropped a thing called Dash Bench, which allowed them to train, essentially show pairs of AI models working together to find code to fix. And they found some interesting things about this. To me, this kind of felt like We've reached the possible apex of the number of benchmarks out there in the world. So Anastasios, from your end, clearly you have your own approach to this, but have we reached kind of like peak benchmark saturation today and have they lost a lot of their bite? Because I don't care as much anymore about benchmarks versus actually using something and seeing if I like it myself.
A
You know what, the benchmarks are only going to be accelerating. And here's the reason. It's because you cannot tell whether something is good until after you put it in production. It's the post deployment evaluation of models that actually matters. And it's not like whether or not it's good at a multiple choice test. That's exactly the philosophy that we have at arena, by the way, is that all of our benchmarks are based on the continuous usage of tens of millions of people of like actual AI products and reality. We just place them in people's hands, see what they do with them. It gives us the most diverse benchmark in the world because we can collect millions and millions of agentic traces every week from people and then measure how it's doing, not just, you know, on history questions, but how well it's doing for, you know, everybody in every country around the world in math and coding and instruction following and multi turn tasks and legal and medical verticals and so on. And, and it needs to go beyond that. You know, our benchmark, you know, we, that that's our philosophy, that it needs to be at the, about the post deployment utility to real people. But in order to really, you know, continue to fulfill that mission, we really need to be inside every company in the world. We need to be helping every company define what good means, define outcomes, help them measure these outcomes and ensure that they are getting the reliability, the performance that they need at the cost that they want within their particular vertical. And I think that Moon Ja was talking about this earlier, that means that we need as many benchmarks as there are businesses.
C
So then they don't become benchmarks in a general sense, but they're more just company specific quality checks.
B
So I'll draw an analogy for you. I think we've all gotten too obsessed thinking everybody's building the same type of vehicle, but some people need pickup trucks and some people need station wagons and Some people need SUVs and some people need convertibles and some people need Caterpillar tractors that are super rugged. And so I'll give you an example of where the benchmarks are actually coming in now. That we're seeing. So if you do a two by two grid of latency fast enough for voice, we call that about 500, 600 milliseconds in the LLM because you need voice in, voice out, and the whole thing's got to be under about 1.8 seconds or less, ideally, or latency above that. And then high intelligence, low intelligence. You notice all the innovations come up here in the top left, which is high latency, high intelligence. So if you have all the time to wait in the world, the models have gotten really smart in the last year. If you go to the bottom right, which is you need it in 500 milliseconds, 600. Have they gotten much better in the last year? Because how do you make a model smarter without deep thinking and planning and all the other things? We basically saturated the amount of tokens. We had to kind of hyper train them with more and more data because we ran out of data quite a while ago. And oh, by the way, all this synthetic data stuff turned out to be false, which I'm not surprised by because it was basically inbreeding. I'm like, I knew inbreeding wasn't going to work, but okay, let's ignore that.
C
But in the top icebergs, just saying.
B
So in the top, right. You know, I think that very few people need high intelligence, low latency. If I'm taking orders at Taco Bell, I don't need it. Fine, I mess up a little and give you two tacos instead of three. I mean, that's what happened anyways when I go through the Taco Bell, drive through with the human anyways, on the other hand, but in healthcare, we needed it. We need very high intelligence, very low latency, and we needed a way to build that. So, you know, Hippocratic, we ended up, we've actually changed kernel implementations. We forked VLLM and created a VLLM that runs faster for moes, where you hit a few experts more than the others, by the way, which we do. And. But it works best for large ones. And we've actually sped that up almost 10x. And we take all this latency surplus, use it to deploy bigger models because the bigger models have better reasoning and the better reasoning creates more clinical safety. And then we find the next tech innovation and we do it again and again. We've done flywheel after flywheel at the very low level to be able to do this. And I think that when you come back to this benchmark question of it, I would say there's not even Just a question of all the benchmarks. There's also a question of all the different use cases drive it. We're in this low latency, high intelligence bucket that very few people need. But in healthcare we need, and there will be a whole new set of benchmarks to even talk about that low latency, low intelligence. There's a whole bunch of very simple, you know, if I'm giving you your bank balance, like I don't need it to be that smart. Like it's fine and so and then. But most of the world right now is innovating only in the top left quadrant, which is infinite latency. Like right now I'm like, write me this report, Claude, I'm going to go get a sandwich. You can take 10 minutes. I could care less. It's better than me writing it. And so I think we've gotten too obsessed on kind of thinking everybody wants the same car. They don't. They need different types of vehicles and those will drive even different benchmarks.
C
Yeah, no, I think this is a really interesting point, but Anastasia, when I look at for example the text arena leaderboard, it's super helpful for me to understand probably to Manjal's point, the most intelligent but maybe not the fastest models. So how do you work in the latency point and the other things that Manjal just brought up into creating a reasonable public facing metric for model quality?
A
Yeah, absolutely. Well, I think there's three, there's the trifecta, there's performance, there's cost, there's latency. And right now the hottest topics, performance, performance, cost. Because what people are finding is that the token maxing era happened and people are just spending, spending, spending tokens and you'll get crazy things. You'll, because of the way the context works, you'll say, you know, thank you to Claude and you'll basically tip it $3 because it had like all this big context that it's using to like read and say, what do I need to look at for this? You know, for this? Thank you. And it's like, you're welcome. $3. That'll be $3. And that's happening, you know, millions and millions of times around the world. You know, it's happening today, right now, somebody's tipping cloth. And so you have to start asking the question, am I getting the outcome or am I just spending money for no reason? And so Arena's definitely trying, you know, trying to help with this. If you look at the text arena leaderboard and you go to the Pareto frontier plot, we are trying to help developers and businesses around the world who are making choices about their models. And so what you'll see exactly is that Fable 5 is on the extreme end where it's the best model on the text leaderboard, but it's also by far the most expensive. And the scale, by the way, is logarithmic. So as you go to the right, you know, every line is an order of magnitude cheaper and cheaper and cheaper models. And so what you'll see is that a lot of that Pareto frontier is dominated by Google.
B
If you look at it, one of the things we did in looking at it and creating a novel set of benchmarks is look, here's some very specific things that are not in the normal benchmark. Drug name disambiguation. Can patients say drug names right? No, they leave entire syllables. Right. They're like, what are you on? I'm like, something statin. I'm like, that's what I'm on. I'm on something statin. Okay, well which statin is that? Super statin? Simba statin? Like, which one is it? And so, but like, look at the things we compared to. First, did we Compare to Gemini 3.3? And at that time we read this, it was GPT5? No, we just said too slow for voice. Too slow for voice. Too slow for voice. Okay, we did. Now then we looked at the speech to speech models that everybody's excited about, but it turns out they're all very low parameter count. They're just not very smart. Now then we went to the Cascaded model where you have an ASR that takes the voice and then converts it into text and put it in the LLM and then you give it to a tts, a text to speech engine and speak it back out. And you notice there we compared to all of those and it got better. But now if you look. Then we looked at the Polaris model and said, all right, how good is that at it? And then we looked at the Polaris model where we actually, we look at our 4.0. We looked at our 5.0 with just the main model and then the whole constellation of the 31 working together. And we basically showed step by step by step that you know, hey, these are the things you need to know. Do models know toxicity limits of of OTC meds? Oh, I Can I take ibuprofen? Sure you can. Don't take more than 2,000 milligrams. Does it nail it? Every single time it needs to. Does it get. There's different Mental health questions and muscle skeletal questions. There's wound and skin questions. Like we went really deep and this is the, there's payer questions on that. There's compliance questions on HIPAA and anti kickback things for pharma calls. Like these are the detailed things and you can benchmark them and you can show how much of the contributions coming from the model, how much is it's coming from the post training we're doing, how much of it's coming from the constellation of the post trained models altogether. This is why I think he's right there. Literally. We kind of proved his point. Every business is not only one new rubric or one new eval, it's hundreds.
C
So you think this type of rundown is going to become the norm for pretty much any company that puts AI to use inside their product or service, regardless of the industry or vertical?
B
Yes. And now. But let's, let's talk about our friends at what's the legal one?
C
Lagora. Harvey.
B
Harvey. Right, Harvey. Guys, that oh my God, we got beat on our own benchmark by the open by the Frontier model files. Actually part of that I think is and why didn't that happen to Hippocratic is again the latency. Their use case has infinite latency.
A
Right.
B
Meaning you can wait however long to generate the document and the big guys are gone after that in a big way and have improved everything they can to do that. So A, I do think this is exactly where it's going. But B, part of what's interesting is like I said, they were the big guys are making a convertible and they were making a convertible and the big guys beat them making a convertible. I'm making a Caterpillar trail, you know, Caterpillar tractor that needs reliability over everything else, even if it costs more.
C
Okay. Now on these benchmarks, I'm really curious because I looked through some of these before we jumped on but I just scrolled through all of them went on for quite a long ways there quite a lot of your scores for the Polaris 5.0 constellation, which is your newest and best model fired on full power, quite a lot of your scores are like 99.97. First of all, 10 points for getting yourself high marks. But how much work is there left to do? Like what's Polaris 6 going to bring if Polaris 5 is already so damn close to perfect for these tasks that you have set as the critical work that it needs to do?
B
Look, in healthcare safety there's still patients that are hurt in the point three Right. It actually, like, you don't stop there. You stop at five nines. And so that's part of it. Luckily, right now we've done 200 million patient interactions without one significant safety incident. So we're, we're, you know, feeling like we caught everything. But I still think there's a need to get even better. The second part is actually, it's the blind spot of what benchmarks are not there that's the improvement. Because actually, how did we even find those? In fact, Anastasia and I, we were talking at that football game last year, remember?
A
Yeah, we met at the Super Bowl.
B
Yeah, Yeah. I tried not to say which one, but now you just said it. But like the.
C
But were you in the same box or just like, did you range other at the bathrooms? Like what?
B
We had the same friend who took us to the same box.
C
Oh, okay. I, I understand, I understand.
A
Yeah, we just had a friend. He took us to the Super Bowl.
B
Us entrepreneurs don't, don't go on our own. We.
C
He runs a billion dollar company. I run a billion dollar company. He is a friend. I have a friend. No, we went for free.
B
Just officially, we went for free thanks to the generosity of a good friend. But the. Because us poor entrepreneurs, you know, we, we just try to get into any game we can. But the. But we were talking and he said, hey, put your benchmarks on our site, you know, and we can run them all. And in some ways I'm excited to do that, but in some ways I wasn't because actually a lot of the IP of the company is knowing is now seeing so many real world examples where we had failure cases and we had to address those failure cases and deal with that. I'll give you another example we found on the TTS side recently where, yeah, 11 labs is not drug name stable. It not only says the drug name wrong, it says the drug name differently each time. It's okay. Kind of to say it wrong because patients say it wrong. If you're consistent, you're actually okay. Although that's still not great because if they do, like, people don't know how to say it, but they know when you said it wrong. And your clinician, your AI clinician loses a lot of credibility if you say it wrong. What kind of clinician are you? You can't even say the drug right.
C
And So I think 100%, if I was on the phone with somebody and they mispronounced ibuprofen, I would be like, okay, click. Like, I'm not gonna.
B
So. And 11 labs recently released a new version that I think improves that. So good for them. They figured this out. But, you know, we found it was drug name stable 95% of the time. It just wasn't drug name stable 99.9% of the time. And so we ended up building our own TTS to fix that because it was such an important criteria. And by the way, our pharma customers had, were having none of it because it's their drug with their name. And so they're like, no, you can't say Manjaro. You got to say Mounjaro the right way. And you know, and so, I mean, I think that these are the different elements that, that come out of this. But a lot of the, a lot of the improvement is in the blind spot of what benchmarks have we not created and what are the queries in the benchmarks? And there is actually a tremendous amount of IP in what we're testing on that we had to learn the hard way.
A
Yeah, I think that's exactly right. I would take it even one step further to say that the benchmarking, the valuation is actually the fundamental and core IP of any business in the future. And the reason is because it encodes what you know about your domain and encodes everything about what it means to achieve success. And so you should know that and nobody else should, because if they have a verifiable method for hill climbing on what it means to do a great job in your domain, they can just replicate your business.
C
So, Anastasio, I want to double click on this. Your point is that if you as a company can come up with the correct benchmark for what you're serving through the medium of AI, you know, what matters to your customers, what your system can do. And essentially it becomes kind of like a functional DNA of your in market performance.
A
Exactly. I mean, listen, like moonjal runs Hippocratic AI. It's true. Imagine that Moon Jael releases all of the benchmarks, all of the data that tells exactly, that gives an exact roadmap to all of the AI companies on how to build the best voice assistant in medicine.
B
Medicine.
A
And then they just can climb that and build the perfect voice assistant for medicine. You know, that's a pretty bad outcome for you. And that's, that's not just true for your vertical. It's true in all verticals, in all AI native companies. Because. And the model labs are going after these. You know, the model labs are now competing with Harvey, and so Harvey better not release the benchmarks that they use. And the definition that they have of you know, what it means to be good because it's not that hard to climb once you have the North Star.
C
Okay, but then if I'm a customer of Hippocratic AI and they say we have a 99.59 for I'm going to pick a thing here, Child Protective services for these agentic voice calls and I say cool, how'd you measure that? And then Moon Jal says, well I'm not going to tell you, that's my secret sauce. Does that break customer trust or is that simply just protection of in house IP in the same way we've seen IP protected in like patents and so forth?
B
Actually there's a very easy way. Go ahead. We would just show them, I mean that customer under NDA we'd be like hey, here's the queries we did and here's Al because I'm not worried the customer's gonna go Hill climate, right? It's highly unlikely and so happy to be very transparent and show them. But like in that case, what that actual feature is, is it's actually if somebody, a 14 year old being called by the AI or a 15 year old right over the thing about their health care says it's not safe here by law in 50 states you have to notify Child protective services within a certain time frame in a certain way and you have to flag those and you have to do that. And if you could release an AI that is an AI clinician that calls patients, calls teenagers and if it doesn't have this feature, you can't go live, you really can't. So there are some very detailed things. But but net net, there is IP in this, there is learning in this. I'll give you another thing. We learned that that is on the periphery, not even in the core model when we used to fail on 25% of all calls because of background TVs. Because when you're in healthcare, who you calling? Typically older folks because they use healthcare more than younger folks. And what are they doing at home when you're calling them? They're watching tv. Is the TV loud? Oh, it's loud. They just go to my parents house. My dad definitely struggles with his hearing and the TV is blasting and getting rid of background noise is easy. Background speech though is the thing you're listening for. And so we had to build an entire layer of this algorithm we call MRX that sits in front of all of our processing and listens for background TV and cleans it up. And now we still fail on about 1% of all the calls, but it's better than 25. But we didn't even know that existed as a problem. Same thing for handling post stroke victims with slurred speech. Same thing for, you know, there's numerous features both in the core clinical model, outside the core clinical model, in the tts, on both ends, on the input, on the output in the middle. And these are all things you learn from just failure. But you, you can't, you know, you could replicate everything I've done and you'd still probably fail on 25% of calls, which is unacceptable if you didn't get this background TV thing right. Wow.
C
Well, you guys also just introduced cough detection, which I presume was easier than background TV removal, but still, because it goes to the point of how many things you have to deal that are edge cases that aren't really, but instead come up quite often, or at least often enough to take your Back to the five nines. Point five nines down to 1, 9, 9 9.
B
Yeah. I mean, the next thing we're actually working on a future feature we haven't launched yet is, is SOB detection. Because think about this case. I'm talking to you and your, your words are saying something different than the rest of you. So you're like, yeah, I'm totally fine. I'm totally, totally fine. Could you imagine if the LLM responses, I'm so glad you're having a good day, Manchal, you'd be horrified, right? You'd be absolutely horrified. You'd be like, that is horrible that the LLM did that because you and I know they're not having a good day and they're misleading you with their words. But right now the ASR just takes the words and here's the words, please, LLM, give me an answer. And so these are critical features that, that no truly empathetic responsible clinician would ever want. Like, they'd be like, no, I can't deploy this. Same thing with the, you know, post stroke people slur their speech. Well, that's not that often in the normal world. It's very often post discharge calls in healthcare.
C
Well, you've convinced me I'm not going to build a competing startup to Hippocratic AI because it sounds very tricky. Now, we've talked about data a couple times and Mujal, you've mentioned how many data points you have that you've put into the training of Polaris 5 in preceding versions of it. Anastasia, your company launched a product that lets people collect a certain type of data from your user base and turn that into a commercial product last year and it's grown to 100 million run rate pretty quickly. How does that information you can bring impact the conversation about benchmarks and getting AI to this kind of five, nine levels of repeated performance?
A
Well, arena is all about the post deployment utility of the AI to people model. It's exactly the conversation that we've been having, which is if you actually take a model and you put it in the hands of individuals in the wild and they're using it for all of their super diverse use cases, how is it going to perform? It's data that we have access to that almost nobody has because we have one of the largest consumer apps in the entire industry. We have tens of millions of users on arena that are coming all the time to use AI for whether it's their, you know, work for their legal, for their medical also, or for their software engineering or for their personal tasks, or, you know, asking for research, for planning, for advice, for lookup. And so we're able to take that post deployment data and turn it into very careful descriptions of a model's strengths and weaknesses and analysis that helps businesses build better models and understand which models to choose. I think that, you know, the future of the company is really around helping the rest of the world understand how to get the best out of their own data. Because it's not enough to just look at an external data set, no matter how diverse. You know, I think arena has the best coverage of any data set in the world simply by virtue of the way that it's collected. It has all of those things that you didn't know you need to measure because of the fact that it's in the wild. We don't determine the distribution. The distribution is determined by this like super broad set of users. And so it inverts the problem from a like. Let me describe what my problems are to let me look at the data and see where the problems arise, which is a fundamentally more scalable and organic way of doing benchmarking. So can we take that template and help every business in America do better benchmarking? That is something that I would be excited to work with someone like Moon
C
Jaal on Moon Jal. Is that something that you could use at Hippocratic Ax? I'm trying to figure out a little bit through the weeds here what arena is selling and how it grew from 0 to 100 so quickly. Because I was very impressed, Anastasia, by that rapid commercialization.
B
Yeah, I mean, look, I think that there is this. We are actually always seeing this as an issue as well, which is that we've built all these detailed features. We know the product works. Works better than anybody else's product on all these corner cases. It's not always easy to get that across to customers efficiently and fast. And, you know, in the early days of the company, it was like, you have voice AI. I've never heard voice AI. Then it's like, oh, my God, you guys have a clinical voice AI, and you've ensured it's safe, and here's how you've tested it, and here's how you've done it, and the protocols. And so that was also differentiating direct. We're still the only ones that will escalate and transfer the call to a human, which is an important kind of safety element. And so you have unique safety. And we continue to have lots of people looking at different parts, but nobody wants to do clinical. They're all scared of it. And we're the only ones that have kind of gone into that. But at some point, somebody will. And there is a need to articulate some of these deep differences where otherwise it's just. And why they're important. Like, if I just told you background tv, I figured out how to do it, you're like, that's great. But it was. When I told you I failed on 25% of all the calls because of it, you're like, oh, shit. That's a really important feature. And the same things on all the different clinical elements that we have, it's like, well, how often do you have to do child or adult protective services? How often do you have a patient that has a case that you got to handle better on the clinical side? How often are. And. And so I think that's. That's where it's important to be able to articulate that. You do build a reputation over time of, hey, these are the safe guys. These are the ones taking the most things. But I think the other part that's interesting is showing why the model architecture is actually bringing us many of these advantages. Why is the constellation better than a single model? Why is big models better than small models? You know, by the way, one thing I forgot to mention when we were talking about the open source. I don't know if you realize this, but one of the reasons Kimmy and all these guys have not gotten as much adoption of their open source is that the code isn't there to train them all. Like, it's amazing how much code my team has found is missing to do some of the core trainings you want to do on them to optimize them. You run them on the inference engines. They're super unoptimized, so they actually don't run very fast. And especially the big ones, people have done a lot of optimizations on the small versions of them because they like to run them on their laptops. But the big ones have actually been largely ignored. We've had to do a lot of core work that we shouldn't have had to do, frankly. But, you know, now that this is. It would be great to create, you know, I think over time you'll end up with a benchmark for every vertical, for every type, or a series of benchmarks that will be in a compounded benchmark and people will do it. But today we're still exploring, meaning we don't even know all the things that would need to go into that benchmark that actually matter and how to weight the benchmark. Like, let's say they're. They're an aggregation of 500 different, smaller evals. Okay. But your aggregate benchmark needs to be a weighting on those, and it's not an equal weighting because some problems occur.
A
Well, you know, the other aspect is severity. 500 may not be enough.
B
It's probably not. I mean, it's probably 5,000 or 10,000 or 100,000, but.
A
Yeah.
B
Well, the thing is, you need that. And the weighting, you would have to.
A
You have to ask the question, do you think that sort of ad ridiculum. Do you think that every person needs their own benchmark
B
is the limit of X as we go to infinity? Basically every single person has a set of benchmarks around it.
C
I mean, I have that with my own Codex instance, which I've trained to talk to me in a certain way. And I view its performance through the lens of what we can call, Alex bench, if you want.
A
Yeah, because you have this. Your own set of tasks that you're doing, and you have the model that you like to talk to you in a certain way because you are Alex and I'm on a stasis. And we have a different way of communicating, and we might like different types of people. We might, you know, want different types of information. We might have a different task distribution that we care about. We might have our own different failings that we need the model to cover or help us understand.
B
And to your point, each language we've also found. So, for example, we have made Mandarin safe enough to use for clinical actions, as long as you're not Talking about drugs. No, we have not figured out how to get the drug part right.
A
Exactly.
B
Yeah, we actually couldn't figure it out in Arabic either. And what's the reason? Well, the speech recognition systems error rates are so high in these low resource languages. Now, hopefully China can fix the Mandarin thing, but we literally have not found an ASR that's good enough to be able to use for drug name identification for anything relating drug names. And so the error rate's so high. So we won't let our customers do Mandarin for anything related to drug names.
A
It makes sense. And I think that the sort of conclusion that we're arriving to is that we do actually need that level of fine grained detail to understand much more than we could ever possibly write down. And the implication of this is that benchmarking as an area, like building benchmarks, is actually not a scalable solution to this problem because you will never be able to write down all of the things that you care about. So you need to do something different. You actually need to invert the problem and you need to be looking at the data first. It needs to be completely data driven. And we probably need to be training models to do it too. Training models to help us understand on an individual basis, you know, even more finer grain than a business because you have customers. It's not like your business is just one business serving one individual. You have, you know, thousands, hundreds of thousands, millions of customers that you need to serve. And they, they probably have their own customers that they need to deal with. And so, you know, when you're getting to that level of granularity, what needs to happen is you just need to be monitoring, you need to be looking at what is happening. And you need to have automatic systems for classifying where the errors are happening. What are the sort of the principal components that are going into the failure rates, you know, for who it's working, for who it's not working. And those types of systems are likely to be the future of what we now call benchmarking and evaluation.
C
Well, in that case then I want to bring my own context from my own AI world to that conversation. And I wanted to be able to plug into Polaris 5 from our dear friends at Hippocratic AI. So it knows. Oh, Alex tends to talk this way. He's hyperbolic and likes to kind of make some jokes. Don't take him too seriously. This is a medical call, you know, drive to the point. And for you, Anastasia, it could be he's a complete stoic or whatever, but I would Want to bring that. And I don't think that a top down benchmark to your point, would work. So then how do we create a system by which not only can I aggregate and collect my own personal or personal business AI context and bring that to bear in a setting like what Hippocratic AI pulls together? Do we need a new NCP for people?
A
Very, very good question. And I think you're getting to something where there's a human performance and then there's a superhuman performance. If you were a human on the other end, the intelligent human would be talking to Alex on the phone. You know, Alex, you know, has, is bleeding out or something like this. And he's funny, he's making jokes and blah, blah, blah. But you're kind of listening to who this is and you're trying to understand them as you are, you know, working with them from a medical setting. And then you kind of learn that, oh no, this is serious because you, because there's something going wrong and he's telling you, no, I'm getting lightheaded, I'm dizzy, I'm about to faint and you know, things are going wrong. But also, lol, you know, I'm fine, I guess, like, it's okay. Like dog with the fire meme. Okay, so maybe that's you. But the other, the person on the other end, an intelligent, emotionally intelligent and intellectual individual on the other end will learn that. But where AI can go farther is it could potentially learn that across all of your interactions in every surface that you've ever seen.
C
We don't have the way to collect that yet. But Moon Jal, if I was a patient of a hospital group that uses your technology and I had to talk to this voice agent, let's say twice a month, is it able to learn from me as Alex, and therefore how I communicate and so forth, and bring that to bear the next time? Or does that just. Does my interactions flow into a kind of a shared bucket that shapes overall performance in a less individualized way?
B
So two things or three things. One, we actually do allow, we do store memories from the prior calls and bring them up. By the way, super hard to do. You think it's easy? I'll just take the whole prior call transcript and put it in this context window. Yeah, until you blow the context window. Because when you exceed 10,000 tokens in the context window, you slow down latency, even if you have space. So now you got to compress the memories a bit, right? You got to say, I'm in jaw has two kids. These are their ages, blah, blah, blah. You can't just store everything. And so that's one issue. And it's a latency thing more than it's a context window size thing. The second element. Yes, by the way, when you remember things about prior calls and putting it up, the superhuman ability of that. Patients love it. Like, they just absolutely love it. Second, we've now built in a dynamic personality that we had to build because people were annoyed. I'm in a rush. It's not in a rush. If I'm in a rush, you're in a rush. If I'm joking, you're joking. So the third thing is I. I want to argue a little bit the other direction for this kind of everything's got to be specifically personalized. Look, there's a thing, a really good thing we did in healthcare in America. It's called a standard of care. We standardized a lot of things, and as long as you do the standard of care, you're fine. Second, do I need to personalize to every patient and be there? Well, my first priority is to get the things that have a safety risk. So should I personalize everything for you? I mean, it'd be nice. It's not that critical. What is critical is, for example, figuring out that you're one of those stoic guys that even when your pain level's 10, you're not going to tell me it's a 10 that I should do. I should say, hey, in my prior calls with Alex, Alex is kind of an understater versus, you know, you've seen patients on the other end. You, like, touch them with one little needle. They're like, ah, like, okay. You're like, okay, there's. The scale of your perceived pain is not an absolute scale. Right. So now where do you have a risk? The people who overstayed it? Not nearly as much, maybe. I mean, a little bit of an opioid risk, but we have good controls for that these days. In medicine, on the other hand, it's really the understated ones that are super in pain and not telling you. And if you notice that there's a pattern with Alex and that you definitely want to get. So one of the things we do at Hippocratics is we don't just try to get every last thing and every land. We're like, which ones matter? And create a medical safety issue. Let's prioritize the heck out of those. Let's get those right. Let's make sure we don't mess up on those. Let's Build separate models to double, triple check those, you know, and there's something called condition specific disallowed OTCs, for example, which is the. It's a personalization if you think about it. It says, hey, can I have ibuprofen? Sure you can. Don't take more than 2,000 milligrams. Oh, shoot. You have chronic kidney disease stage three or four. Oh, now you can't have any right? You got to get that right. You got to get that right every single time. And it's personalized to the fact that you have chronic kidney disease stage three or four. And so like you can't say, oh yeah, my model gets that right. 95% of the time it's good. 98% of the time it's good. No, I will kill you if I tell you to take ibuprofen and you have CKD4. And so these are the different elements. And the last part is, this is what I love. I love finding verticals where incrementalism matters. So let's say I'm doing voice AI for Taco Bell. Let's say you have a product, I have a product. Your product is 95% as accurate and you. But you're half the price or 110 the price. And I'm 99% accurate at getting the order right. But you know, I'm clearly cost more. I will buy you every day of the week, but in health care safety, I don't buy you at all. How can I tell my boss, yeah, I saved half the money, but I knew I was going to hurt people.
C
And then you get sued for 10 times the money and you end up backwards.
B
That's right. Yeah.
A
Listen, I. All I wanted was a creme brulee. Taco Bell Crunchwrap.
C
Wait, did they, did they. They haven't actually made a creme brulee Crunch.
A
They have a creme Brulee Crunchwrap.
C
Time to go back to my stoner roots. All right, a couple of other topics before we move on. Anastasia, I mentioned that your company recently announced a revenue milestone. Well done. Putting you on track to be public company size whenever you'd like to be. Moon Jal. Your company, as far as I can tell, has been very close to the chest regarding its financial performance. I'm just kind of curious why and if the recent success of SpaceX's IPO is pushing founders like yourself towards maybe being a little bit more IPO friendly than they were before.
B
We actually did recently put out an announcement. We announced that we were at a 60 million run rate in basically we only started selling the product January of last year. So basically 18 months, which you know, healthcare takes a little longer. We got to do integrations, we got to roll things out.
C
That's very impressive. Don't, don't talk down that number. That's great.
B
Well, you know, in this world where everybody and their grandma is going to 100, I, I definitely happy. But the, but I will tell you that we're seeing tremendous adoption. I think we've now gotten 50 health systems over 2 billion. We've got five of the top seven payers, we've got seven of the top 20 pharma all in 18 months. It's been an absolute terror. And, but the, but as for the ipl, I don't know. I have two words. I mean on one hand you're right, the market's there. On the other hand, what a pain in the neck it is to be a public company. Like, I mean I, I would rather, I mean there's like a hundred painful things I would rather do in my life than be a public company CEO. But I mean it may be though, but I don't let my personal predilections drive the right thing for the business. Like if that's the right thing for the business, that's the right thing for the business. You know, we'll find a way to make that work. But I mean I, I think that there is an advantage. I don't think SpaceX could have raised as much money had it not stayed private as long and invested in its product correctly. I think that you're going to see a massive focus on margins. Everybody's saying, oh, you know, they've been lowering the cost of tokens continuously. Yeah, you watch once they're public, if they keep lowering the cost of tokens and now how it works when everybody can see your whole entire P and L, it turns out that there's a lot more scrutiny on your margins. And I think the days of costs coming down dramatically are not, we're just not going to see them the same. And so I think when you go public, you kind of freeze a bunch of long term investment ability or lose it and you eventually sow the seeds of the next generation of companies that'll come get you. And so I would like to invest long term and you can always do that better private. And today the private markets are so deep you used to need it because you couldn't raise the hundred million dollars we both raised privately now. I mean there's a zillion people you know, I mean, we both of us probably get emails every single day from somebody wanting to put in money, so.
C
Well, that's an incredibly depressing answer to my question. I was hoping that you were going to say everyone's so excited now about going public. Elon did it twice. We're going to join the ranks. But basically
A
I think that's the other way around that we should be really thinking about, which is how do we give ordinary people access to private markets? Because it is really not fair. Yeah, the only people that have access to the highest growth companies in the world are quote unquote accredited investors and not even really them because they don't have access to the deal flow. So I think the sorts of things that Robin Hood is doing around this in terms of just like giving people the ability to invest in private companies more directly through these like, you know, aggregate vehicles is really, really good. And I think it's an issue of financial fairness and freedom.
C
Yeah, but I thought Moon Jal was going to say that, look, you know, my investors are really encouraging me to, you know, just stay focused, stay private and grow. Because to me that would be, that would allow for more venture capital or private investor, if you will, broadly take rate of the company's future success and growth and valuation. But instead when he said that, it's just, it's too annoying and I just don't have to, I mean, to me, man, it just. Is that the common perspective amongst CEOs of AI companies your size and age? Manjal, is that just the kind of the de facto view?
B
I think this is a function of, you know, I'm a, this is my fourth company. I built a bunch of companies before. I've suffered through some of them, some of them gotten really great. You know, I sold a company, Google, that went great. I had a company didn't work out so well, that didn't go so great. I had another company that, you know, we did well on. But when you've suffered, I think, you know that it's just never something for nothing. And so, you know, right now, I mean, it is partly what you said. I mean, we are investing, we are able to invest, we are able to do smart things, but our burn rate's very reasonable. We've kept it pretty tight despite raising so much money. So I have a ton of money sitting on the balance sheet because I have seen what happened. I mean, the wind is blowing at our backs, right? And it's a gale force wind at the moment I've never seen ever before. But I've also seen what happens when the wind stops blowing or when the wind blows at your face. And I just, you know, so we're running as fast as we can while the going's good. But I, I don't, I think you need to plan for a day when margins are going to matter, when cash won't be so easy to raise. And, and so, I mean, I think, I think the public markets are interesting, I think they have opportunity. But I, I think one should just think very carefully about whether you're.
A
That.
B
The other part is I remember I invested in a startup that became a unicorn and I won't say which one it was and it went public, but in going public, it was one in the advertising space, online advertising. So like, you know, it was buying ad inventory and selling it on the other end as actions. And what happened was all on both sides of its marketplace. Everybody saw its margins.
C
Oh, no.
B
And right, because it went public. And so then every quarter actually when every renegotiation came up, they actually lost margin, lost margin, lost margin, lost margin, loss margin. And so the public company process is one that degrades competitive advantage by definition.
C
This is what Google advantage, it used
B
to give you was access to capital, which gave you a reverse competitive. It increased your credit. But now you can get almost as much capital, maybe even, I mean, maybe you can't raise 70 billion in the private markets. Like, and so he needed it, but he's also got a very capital intensive business that we don't have.
C
OpenAI's last round was what, 122 earlier
B
this year, Also a very capital intensive business.
C
For sure. For sure. But I mean, just you said 70 might be the cap. I think the cap is really uncapped currently.
B
Oh, you're right. They didn't do it and they didn't do it private. So you're right. Your points made. Well made.
C
So Anastasia, before we jump to the show, you wanted to talk about Trump accounts and I was kind of curious how we'd fit that in. But you mentioned, you know, the fairness factor of this and if companies don't want to go public, and that does preclude people like, you know, me, Mr. Index funds from getting access to their upside as an American taking part in this economy. What about private companies donating some of their shares to the Trump accounts of the kids?
A
You know, I absolutely think that's a great idea. It's something that I've even considered myself because I think that it's just such a universal positive to get young people in this country, financially educated, interested in investing, interested in owning businesses of the future. I think there's a lot of questions about what's going to happen in terms of our socioeconomic structure in the age of AI, wealth inequality, blah, blah, blah. There's some things I feel good about. There's some things I don't feel so good about. I don't feel that great about the way the government is using our money, but I would feel great about giving part of my money to the children of America.
C
I know you're really glad we ended up hearing the conversation, but you were at the White House. You know, a lot of people in the policy spaces that are making the sausage here. What do you think the appetite would be for something akin to a let's all do the 1% pledge, but instead of giving it to charity, put it into Trump accounts for kids.
B
I'd probably say two different things on it. Not a bad idea. Actually giving it to the kids feels better than giving it to the government by 100x. So I'm not a big fan of nationalization of assets. I think we tried that in the past. Is it the greatest way to build an economic model?
C
What the is Trump doing buying shares in all these American companies? Leave them alone. Sorry. Back to you.
B
Anyways. Well, at least he's buying them instead of just saying I take over all of your oil company producing. I mean totally. He's. The other way was, was worse. I where you're buying is still it's fine. But I think of buying it, you know, is not any different than a, a GIC or, or some of these national sovereign funds, you know, buying assets and startups, which they do so and buying is fine, buying is clean. But you know, nationalizing. When does 5% given to the government become. Well, you know, now you have to give us 5% more next year. Now you got to give us 10%. Now you got to give us the whole, I mean, like it's a slippery slope. But I would probably say if I were to take Hippocratic shares and have a foundational element, which we actually did put some shares into a foundation that is designed to benefit health systems and patients. When we first started the company on day one, by the way, which has now actually turned into quite a bit of money, three and a half billion dollar valuation. But I would probably find something that's more akin to, you know, is there a Trump account for your healthcare costs for the people who can't afford it, who get in a pickle? And I would rather Put some of our shares into some sort of structure like that if I could. I don't know there is one that exists, but I would try to find something that aligns better with the mission of the company and with really the goal we're at of kind of like an hsa. Yeah. If I could say everybody's HSA gets a little bit awesome. It's just, you know, it's, it's at least mission aligned if we're going to, you know, kind of add that dilution to the company and, and do that.
A
But you know, it's interesting. I don't barely. Yeah, I don't necessarily see it exactly the same way. I understand why you're saying that it needs to build a mission and of course I respect your mission. But you know, I think that it's, it's also very good for both sides, both for the company and for the US population for people to be educated about these businesses. Businesses like literally open them so like understand them. I have, I have a portion of Hippocratic AI, let's say I have a, you know, one unit of this. I'm like, I'm gonna want to understand it if I'm like, you know, 15 years old. I don't have that much else to do with my time. Maybe I'm doing my homework. But I have like a thousand dollars and here's my stocks. I'm gonna go look up these businesses. I'm gonna go.
B
It's just a matter of who you give. I mean, anyways, like I said, it's not a bad idea to give it to you kind of the, the Trump accounts, you know, basically for future college funds. But if there's a mission aligned way to give it, then I might want to look for that as well. But I think, I think in general this is not again, there's difference between giving it and buying it. And we should just be, you know, the, the nationalization of our assets has never in the history of economics shown to create better, stronger companies, better stronger economies or more efficiency. I don't know why we've forgotten all that. I feel like there are many lessons from my upbringing on capitalism being not the perfect system, but still better that are being lost.
C
Well, yep, I think. Well, I have a lot of thoughts about that. But I mean just kicking this can one bit further down the road. What if, if buying is better than taking, which we all agree, why not allow the Trump account Universe to purchase 5% of each startup's round in the proximate investment opportunity and therefore Buy shares that way and then appreciate on the upside and therefore no one's getting socialized and the kids get the get a chance.
B
That would be fine. But remember, these are very volatile assets. I even tell people, I'm like, do not exercise your options with your kids college funds when you work here at Hip Hop. That is not the right answer, my friends. Yeah, that money needs to be there. When it needs to be there, you put in your money you were saving for pure fun, your vacation fund for that trip, that if you lose it, you're not going to. I mean, you have to remember these are still very high volatility assets. They're not the place to put money in the you 100% need. Yes, you're missing out on tremendous crazy growth. But there's a reason there are different asset classes with different risk profiles. This is possibly the highest risk profile possible. But I think there's also one thing we're all missing. We're jumping on this because we think there's this runaway train and we want to catch the train and everybody should catch the train. But I think that we're missing something, that that's the biggest idea in it. And I see this in healthcare the most, which is. And everybody says it, but they don't have good examples of it. But I do. Like we're actually doing it. Which is the real power of all of this. AI is not replacing work you do today and making it cheaper. That's actually not working super well. What is working super well is things you never thought to do until you had an infinite supply of clinicians at almost at a much lower cost. And I'll give you an example. There's been a heat wave in the east coast, right.
C
I know.
B
Yesterday, I think it was, we called 50,000 people and did a heat stroke assessment. And if they're having issues, educated them on where they could go or even called them an Uber to get onto a cooling center. You could have never done that without AI. You'd have to find how many clinicians at a moment's notice. And they're busy in the hospitals because the hospitals are overflowing, because people are overheating. Like, this is the power of AI and healthcare is one of the few places that can absorb that abundance. If I give you an infinite supply of accountants, is your business going to really get better? Probably not. An infinite supply of lawyers? Probably not. But an infinite supply of clinicians? Yes. Because the ideal staffing is one clinician to one person. Education has the same property. The ideal staffing is one teacher to one student. There is a power here. We should focus on this power's impact on something, society instead of just trying to divide up the, the chips. At the moment, I think it's very shortsighted.
A
I actually love your example because I think it's also a great example of how the use of AI can increase demand for human workers. By virtue of the fact that you called, you know, 50,000 people identified heat stroke or whatever, you're actually increasing load on the hospital system because those people are going to the hospital now.
B
You know, what we do is we send them to the cooling center because the idea is to not get them to the hospital because A, it costs a ton of money, B, by the time you're having that level of an issue, you're in a bad space. And if we could have just sent you to the air conditioned cooling center, right, we could have avoided pain for you. I mean, it's good for everybody.
A
That's even better. Even better. Yeah, absolutely.
C
The thing, the thing that I'm concerned about because we have to close here in a second is that the AI abundant future that we're all describing here is far enough away that people who aren't equity holders at the family level are going to be pretty mad. And I'm concerned that when I talk to people outside of our world that the anti AI sentiment is sufficiently high at the moment in many places that it could lead to regulations or otherwise changes to the economy that could preclude us reaching that relatively more abundant future. So that's why I care.
B
Yeah, I mean, look, like I said, it's not off the table in my mind. I think we just construct it properly think through it and just understand the implications of it. Just be thoughtful about it. I don't think it's such a simple answer. We just give 5%, we're done. I mean it's fine, we sell 5% of VCs left, right and center. It's not a big thing on our front. But I mean, I think if we sold that same to the government, like fine, money's money, but I think we should just think through the implications. But the real power is I don't think the dividend for society is the capitalization of these companies. The real dividend is the abundance of healthcare workers, at least in our vertical context, that will be there. I mean my mother had high blood pressure and they gave her meds and she didn't take them. And that was five years ago, six years ago and a year ago she didn't take them. They made her dizzy and they made her mouth dry and she didn't like that. And then a year ago, you know what happened? We ended up in the ER for 220 blood pressure 1 night and the beginnings of congestive heart failure. And I was like okay. But her doctor knew she wasn't refilling for five years because she didn't ask for a refill. Somebody knew. Now they do not staff to call out to every patient that doesn't refill and say or they call once and my mom blew them off and didn't thanks to hipaa. I didn't know she was even on the meds which you know, makes me, I feel so bad as a child I'm like I should have probably snooped in her medicine. I mean I should have done something. I, you know, she would have taken care of me when I was young. And so there's an abundance here to call Mrs. Shah and say Mrs. Shah you didn't do it. Can you take your blood pressure right now? Can we get you a refill? Can we have a courier to your house? Like this is what we can do with abundance. That is the real dividend that AI is going to bring and it's not far off. We are doing it today. We just did it in New York yesterday. Yeah.
C
And I just hope that people are able to see that and that it will reflect how they approach voting and who they elect in the future because I do think that really matters and I think this is not a one party point, it's more of a multi party point. Moon, Jal Anastasia, such a pleasure to talk with you both. I'm optimistic about the future. I'm not educated on benchmarks and I'm really worried about the lack of American open source AI but we'd love to have you back. In the meantime where can people find you on the Internet and is there a job you're looking to fill that you want to shout out to the audience we have here today? And Moonjal start with you.
B
We're looking to fill every job. If you're interested in our company please come and apply. I think we are a 210 person company with 130 open recs so in every last thing you can imagine. So please come because we're growing pretty fast and then and you can find us at, @hippocratic AI.com so thank you.
A
You can find us at Arena AI and if you want to apply for a job, Arena AI jobs. We're looking for all types especially if you are an incredible machine learning researcher, scientist, engineer. We would love to work with you on the future of benchmarking.
C
Amazing. Thank you both so much. This has been this week in AI. We'll bring both these guys back as soon as we can. In the meantime, see you next time.
A
Cheers.
B
Cheers.
Title: What happens if China pulls the plug on open-source AI?
Host: Jason Calacanis
Guests:
This episode tackles a pivotal and emerging issue in the global AI ecosystem: the possibility that China may halt the release of its leading open-source AI models to overseas users, as recently reported by Reuters. With open-weight models from China currently representing the best value-performance option for many startups and enterprises, the hosts and guests explore the ramifications at a technical, economic, and geopolitical level. Discussion dives deeply into the current state and future of open-source AI, the role of benchmarks and evaluation, business models, and societal impacts.
[00:00–04:02]
[04:11–10:04]
[10:04–15:32]
[16:45–21:15]
[21:15–23:11]
[24:38–40:12]
[52:00–57:07]
[40:29–54:49]
[48:18]
[64:17–79:36]
This episode elucidates just how foundational Chinese open-source AI models have become—both as a technical lever for startup innovation and as a political bargaining chip. Their withdrawal would disrupt not only technical strategies but also startup economics, national security calculations, and even AI’s rate of improvement through open benchmarking. The conversation reveals the complexity and dynamism of AI benchmarking, the persistent need for genuine open-source infrastructure, and the dangers and opportunities posed by both competition and collaboration at every layer of the stack.
Closing Note:
If China were to stop releasing its models, startups may find themselves squeezed between unaffordable closed models and a lagging Western open-source ecosystem—unless radical reinvestment and collaboration close the gap. Meanwhile, ever-finer-grained, vertical, and even personal benchmarking—and the post-deployment data that underpins it—will define the winners and losers in the era of AI abundance.
[End of Content-Based Summary]