
Hugging Face CEO Clem Delangue joins Equity to talk about why the open vs closed source fight matters in the wake of Anthropic’s halted Fable release, and why he's worried about the possibility that a handful of big companies could end up controlling everything.
Loading summary
A
When you need to build up your team to handle the growing chaos at work, use Indeed Sponsored Jobs. It gives your job post the boost it needs to be seen and helps reach people with the right skills, certifications and more. Spend less time searching and more time actually interviewing candidates who check all your boxes. Listeners of this show will get a $75 sponsored job credit@ Indeed.com podcast. That's Indeed.com podcast. Terms and conditions apply. Need a hiring hero? This is a job for Indeed Sponsored jobs.
B
Open source AI is booming, according to Hugging Face CEO Clem DeLonge. He's seen the same story play out again and again. Companies start out on Frontier APIs, but as they scale the costs, push them towards open source models. So today, instead of our usual Friday news roundup, we're turning things over to Rebecca Belan, who talked to Clem about why the open versus closed fight matters so much to the AI industry and why he's worried about the possibility that a handful of big companies could end up controlling everything. Stick around.
C
Welcome Back to Equity TechCrunch's flagship podcast about the business of startups. I'm Rebecca Balon and today we're joined by Clem Delong, the co founder and CEO of Hugging Face. Clem, welcome to the show.
D
Thanks for having me. Super happy to be here.
C
Yeah, I'm really excited to talk to you. I mean you've had a pretty incredible career. You're almost like you're, you've got a little bad boy status, right? You're like turned on opportunities at Google and ebay earlier in your career, you know, starting to trying to stick to startups and open tech and then, you know, founded Hugging Face and it's been, it's taken off. Right? You've raised nearly $400 million. And I'm kind of curious, you know, I brought you on today because there's been so much talk about open, open source again, right? Like it's just feels like there's these peaks where open source comes into everyone in the mainstreams attention. So when you look at Hugging Face today, I'm curious, like what is a data point that best captures how much open source AI has changed in the last year?
D
Yeah, there are a couple of data points. First, the volume and the number of models and data sets that are being shared on the platform is quite astonishing I think. Now there's a new repository created every seven seconds on the platform. So that's almost 3 million public models, 1 million public data sets that have been shared on the platform. So it shows A little bit a different picture than the one model to rules them all and everyone talking just about one model. I think in reality we're seeing more and more companies using a lot of different models with a lot of specialized customized models for their need for their specific use cases. I think we're seeing also in terms of adoption by enterprises now we have half the Fortune 500 using hugging face and so using either their own private models or open source models. So we've really been seeing kind of like a trend in the past few months as more people get interested in open source models.
C
So you're saying enterprises are using, are they actually deploying open models in production or are they still experimenting, like how far away from the mainstream?
D
Yeah, a lot of them are deploying them in production. I would even argue that the typical kind of like flow is that companies starting by using frontier APIs maybe at the beginning to experiment to launch the new feature and then when they really hit production and they hit scale, the cost is starting to be too big really with frontier models. And so that's usually when they switch to open source models to power their workload. So, and I suspect this trend will continue maybe in a few years, kind of like frontier models will be for experimenting and some really kind of like high value tasks. And then most of the production workloads will actually be powered either by private models within companies or by open source models.
C
Yeah, that would be pretty devastating for companies like OpenAI Anthropic which are desperately trying to increase the amount of like usership that they have, right?
D
Yes and no. Because I mean I, I, I still think they can build very, very valuable companies by you know, being kind of like at the frontier level, for example, for, for reasoning or, or for some, some tasks ultimately kind of like AI will be so big that I think there's a world where both an OpenAI and Entropic are companies and also most of the workloads are powered by other types of models. So I'm not too worried for them. And they're probably going to be the most valuable companies like somewhere by the end of this year or next year. So they're going to be fine. I'm not worried about them.
C
Yeah, there seems to be this kind of perfect storm of factors that's bringing open source back into the public attention. Right. I mean, you mentioned tokens, right. We recently saw Alex Karp, Palantir's Alex Karp, ranting about how token infrastructure being used by big lives like OpenAI Anthropic is, you know, Criminal and he's pushing for more open models. And then so we've got the token issue. And then on the other side, we've got again the conversation around the Chinese models catching up. And then within all of that, there was this absolute cluster of the Trump administration limiting the release of private AI models. What stands out to you as kind of the thing that's pushing it the most?
D
Well, I think the point of Alex Karp that, you know, companies want more control basically, and transparency in the systems they use is the one that resonates the most with me because obviously it's something that we've been saying for a while, which kind of really makes sense. Like, you know, if you're an AI company or technology company, you don't want to outsource your core capabilities AI to another company. They are kind of like a black box API that you don't control, that you don't have any visibility on, that you don't really kind of have any sort of ownership. So this kind of like, idea that companies need to own AI and own models instead of renting them and outsourcing them to someone else is kind of like the thing that I hear the most from companies and customers and community members these days, which makes a lot of sense in my opinion.
C
I mean, not to mention, if you own your own model, you are not at risk of if the government decides we're going to shut this down for safety reasons, that your kind of left in the lurch there, right?
D
Yeah, yeah, yeah. It sounds like a more sustainable way of building the technology, by the way, much closer to what we've seen with software. Right. Like, I mean, that's how software has always been built with like everyone being able to write their own code and build their own software stack instead of kind of like delegating that or outsourcing that to other companies. But if you consider that AI is kind of like the next generation of Software or software 2.0, I think this approach makes much more sense.
C
Does that require, is there enough talent of people that are able to implement these models into businesses? I mean, larger enterprises is sure. But I mean, across the board, I think so.
D
Because I mean, what we're seeing at Tweeting face is that we had a few hundred thousand users like, like three years, three, four years ago, and now we have 16, 17 million AI builders using, using the platform. What we've seen is that a lot of the software engineers, especially with agents, are now able and it's becoming easy for them to train models, run models themselves. Optimize models themselves. So I think with kind of like agents, we're seeing that it's becoming easier and easier for software engineers to run their own models, optimize their own models, train their own models. And we're seeing that across the board, not only from smaller startups or very kind of AI native startups, but also in enterprises.
C
Interesting. And so are they, are they just kind of using, like how would a startup, let's say, if they want to work on their own model, like how would they use Hugging Face to do that? Because I know that you're, I feel like Hugging Face is in an interesting place right now. It's like you're part GitHub for AI, but you're also kind of like becoming a little bit AWS in terms of services. Right. So how might they use Hugging Face?
D
Yeah, it's a very kind of like a modular platform. So it depends a lot on your skills, on the structure of your team, on your goals and things like that. But most users really start from kind of like an open source based model, right? So maybe they're going to use like a GLM 5.2 for genetic workloads, OpenAI, open GPT for some other task, and Nvidia, Nemotron, any sort of model, and kind of like deploy them directly on their own infrastructure and basically start running workloads. That's usually kind of like the starting point. And then progressively you see teams kind of like wanting to do some optimization, particularly for example if they have compute constraints or if they have cost constraints or speed constraints. So they can start kind of like optimizing these models, optimizing these weights for their specific use cases. And then they'll at some point start to post train these models to basically be more accurate for their specific use cases. And that's kind of like usually the process you start from really off the shelf solutions and then you end up by really controlling a lot of the workloads yourself and building a lot of the systems yourself, which creates actually your differentiation from other companies and other organizations. Right. Like you want to build these skills of building AI systems better than your competitor and that's what's going to differentiate you in the long run.
C
Yeah, that's a good point. Now my brain is stuck on one of the models you mentioned. You mentioned GLM 5.2, which, to take it back to kind of the geopolitics of it all. Um, so this is one of the Chinese models that's been getting a lot of attention for its amazing agentic capabilities. Recently and Hugging face's own spring 2026 report says that Chinese models accounted for most of the downloads. Right, like 41%. So China's surpassing the US monthly and in overall downloads. What are some other findings about how Chinese models are doing on the platform versus US models? And what do you think this says about the state of open source right now?
D
In my opinion, this is a very big challenge because in an ideal world, I think we would want more of the open source models, especially that are used in the US to be shared by American companies instead of Chinese companies. So I know that a lot of organizations in the US are working toward this goal, right? Like you have Nvidia that has. I've told. I've become like the king of American open source AI lately by sharing a lot of very interesting data sets, a lot of very powerful models themselves like Nemotron, and there are a lot of startups like rc.
A
Reflection.
C
Reflection, yeah.
D
There are more and more organizations, I feel like, in the US that are sharing in open source, but we need much more. Because if you think of it kind of like open source is kind of like the foundational stack for the rest of the AI stack. And I think you'd want kind of like every country to have kind of like some sort of sovereignty on each parts of the stack. So I think that would be much, much better to have a world where a lot of the open source used in the US is actually created by American organizations.
C
Yeah. Are you seeing like a lot of American organizations using Chinese open source models? Like, is there not, you know, some kind of a stigma against that?
D
No, no, we're seeing a lot of them using Chinese open source. Right. Some of them famously shared about it. Right. Cursor talked about how their models was built on top of Chinese open source. You have Brian Chesky from Airbnb that has been very vocal about open source AI. The majority of the scale ups in the US that are using open source are now using open weights from China also. We don't talk about them a lot anymore, but all academia, so if you go to Stanford, if you go to Harvard, because the only way to really learn, study, do research on AI is to have open source and open weights. Right? You can't really study an API because it's a complete black box.
C
Right?
D
And so all the, all academia, all the research community is basically using Chinese, Chinese, open source.
C
What's the risk of that though? Like, why is it, why does it matter? Maybe they're just better at open source. Right? Like, what's the problem.
D
So there are a couple of challenges. The main one, I would say, is that open source is kind of like both the foundation and a very strong accelerating factor for AI in general. In my opinion, the reason why the US is ahead now is because from 2016 to 2022, 2023, the US was super open with open research, open source AI, everyone collaborating and sharing with each other. Right. The famous example is The T in ChatGPT came out from Google. Sharing in open source. Transformers. Right? And so open source creates, in a way, the conditions for your AI leadership. And so almost automatically, if China keeps leading in open source, keeps sharing all this research openly in China, it creates kind of like this accelerating development of the field. And I wouldn't be surprised if as a result, China starts to lead AI in general probably next year or the year after.
C
Well, I mean, what would you say to people, what would you say to people who argue that China's open source is only improving so well because they're really good at distillation attacks and, you know, copying the homework of closed frontier models?
D
I would say it's very reductive and very simplistic because China has some of the best developers and researchers in AI. Now, we know distillation to be a very, very small factor in the ability to create good models. It's a practice that everyone is doing, including companies in the us so if it was as easy just to do distillation, to get good at building AI models, there would be many other countries, including in the us where we would be much, much better in open source AI. The reality is that they have really, really good research teams in China doing really well and taking a much more open and collaborative approach to AI than in the US and that's why they're successful.
C
What about the risks? Is it riskier? Are open models riskier because they're harder to control? Right. Obviously, Trump's. The Trump administration had limited the release of Claude's sorry, Anthropics, Mythos and Fable, and then also OpenAI's GPT 5.6 due to cybersecurity concerns, et cetera. With open models, you know, you're seeing open models catch up and as a result, there's a lot more cybersecurity attacks because it kind of, you know, while the models aren't mythos level or fable level necessarily, they are good enough and they have fewer guardrails to stop them from carrying out these kinds of attacks. So how do you balance safety with access?
B
This episode is brought to you by Google Chrome. You Think you know a browser, but Gemini and Chrome, that's new. It can help you with practically anything on the web, like restoring a vintage motorcycle from a 50 page restoration block. Or finally break down that long article you've had open for weeks. Gemini and Chrome is here for it, ready to make anything online make sense. There's no place like Chrome. Check responses set up, required compatibility and
D
availability various 18 yeah, so historically, open source has always been less dangerous than kind of like some closed source, secret, kind of like behind closed doors initiatives. And the reason why is because it's more transparent and so it's easier to understand the capabilities of it and to create kind of like mitigations for them. For example, for defenders to patch the cybersecurity risks that they know open source models can do. The reality is that these risks already exist by the fact that models just are here in the frontier labs, right, and already distributed to organizations through APIs. The reality also is that guardrails or APIs are very shallow and quite ineffective. I think that's something that we've seen in the past few weeks. Like you can put some guardrails and have a feeling that it's safe, but the reality is that it's very, very easy to jailbreak them. It's possible to steal the weights. And if these kinds of capabilities start to happen and to be possible in one lab, it's very probable that other labs will be able to also replicate the same thing, not only in the US but in China.
C
So is the argument less that there should be guardrails or.
D
I think the argument is that you don't really make it safe by keeping it big behind closed door for just a few players. You actually make it more dangerous because you create asymmetry of power and asymmetry of capabilities between some actors and that have access and can't use, can steal, can, can of like use in a malicious way these weights and other people who can't defend themselves. The way you make the world safer, in my opinion, is by by leveling up the playing fields and really kind of like creating transparency on these models and giving them both for the attackers and the defenders and obviously making it harder and illegal to do the attacks, right? It's the same as if you take other pieces of technology. You don't really kind of prevent risks by making it illegal to share some materials to create kind of like dangerous things, you, you make it illegal to kind of like use this kind of like tools, right? With this approach you enable kind of like innovation. You enable Competition, you enable job creation, you don't create monopolies. Right. In, in my opinion, the biggest risk in AI is concentration of power.
C
Right.
D
We, we mentioned, you know, some of the AI companies becoming the most valuable companies in the world very, very soon. In my opinion, they're also becoming the most powerful companies in the world. You've seen that, for example, with their interaction with the Department of War. If you would have told me a few months ago that an AI company could be in a situation of power compared to the American Department of War, I've told you that's crazy, but it's actually what's happening. And so one of the, in my opinion, one of the biggest risks in AI is that you end up in a world where a few companies are completely dominating AI, getting to an amount of power and an amount of wealth that you've never seen before. It's basically similar to if there was just one or two companies being able to do software. And in my opinion, this is the real dangerous, scary scenario. If you don't do anything to fight that, if you actually just enable a few companies to build frontier AI and if you kind of like create the regulatory environment that allows them to do it and no one else, you end up, in my opinion, in a very, very scary and dangerous world.
C
Yeah, so, so do you think that there's moves that the US government should be making to. You know, I think that the Trump's AI plan kind of hand waved at open source, but I don't really know that much has been done. Is there anything like in particular that you would point out as something that you'd like to see the US Government do to support open, open source development?
D
Yep. Public support is kind of like very important right now because. Yeah, for, for some reason, uh, in the US every time we talk about open source now, we, we talk about how it's unsafe, which is kind of weird. And it creates this weird counter incentive for anyone to actually do open source versus what it used to be. Where open source was an open research was really celebrated as something that, you know, companies doing this were actually putting collective contributions ahead of profit maximizing. Like I remember 10 years ago, if you had a company that was sharing their research, sharing open source, people would be cheering for it, which in my opinion is the right thing to do versus right now when the company is doing that, people are questioning it and be like, oh, but is it safe? Should the, should they be doing that? So I think if the, if the U.S. administration can continue to really show support for open Source contribute also like we have, interestingly, we've had a lot of American organizations, American public organizations contributing to open source AI. For example, a few days ago the, I think it's called the National Design Agency released an open source model for PII detection on hugging face. So the public organizations can contribute a lot to open data for open source, which is still kind of like a big bottleneck. These are some of the things that I think could be interesting for them to do to support and foster open science and open source AI more in the US Yes.
C
Yeah, I think open source maybe needs a little bit of a rebrand or makeover certain people. But you know, there is a risk when you have open data sets and sharing open data sets. Right. Like hugging face yourself. You're embroiled in a recent lawsuit, right? Evox Productions don't know who they are. They're suing hugging face alongside stability and Runway, claiming that for hugging face specifically. I think the claim is about you guys hosting data sets that have copyrighted images. I'm curious, you know that that lawsuit's ongoing, so you probably can't comment on it too much. But how do you think about legal risk as a platform that's hosting other people's other companies models and data sets? Does this change how you vet or moderate content at all?
D
So of course, you know, we, we follow all regulation and, and kind of like for all the, the rules, all the legal kind of like things that, that we need to, to follow as a, as a platform. It's been important for us right from the start and we actually did a lot of initiatives to kind of like give more legal clarity to the field. For example, a few years ago we introduced a new type of license that allows open source models, open weight models to have kind of like more clarity on the kind of use cases they can be used for. You know, again, I think the, the challenge and the trade off is between kind of like doing it in, in private and doing it in, in public and how, how different this is. The reality is that we know now that a lot of the closed source labs have been actually, you know, basically using the whole web without any sort of kind of like copyright.
C
Yeah, I guess with open source you don't need a subpoena, right? You can, you can get it. Well, I mean with open source, not open weight, a lot of the time open weight models, you can't actually see the data sets. So.
D
Yeah, so making them public instead of like, you know, keeping it private behind closed doors gives more attribution. It Allows it allows people to actually know what is used and what is not being used. And also it's kind of like used in a different way. Like if you look at how fair use has been built, has been designed for copyright, there's always this balance between obviously giving attribution and protecting creation, but also allowing innovation to continue giving tools for people to learn and to get education. Right. So for example, you can use copyrighted material if you're a teacher to teach kids, which makes a lot of sense, right? You don't want to make teachers having to pay for everything they use or they teach if it's for the public good. And you see the same thing for open source, right? If some data and some data sets are shared in open source for everyone to use for free, not for profit. This is very different than lab using that to make billions of dollars of revenue without any public contributions.
C
This is America, Clem. What are you talking about? All right, well look, okay, just to quickly pivot because I think we can go into the benefits of open source all day.
D
I can talk about all of that for hours.
C
Yeah, and I could listen for hours, honestly. So one thing I want to ask you about, since you're clearly. And I made a joke. This is America, Clem. But speaking of the way American companies run things versus the way you're doing things at Hugging Face, one thing that I thought was really interesting was yeah, you've raised 400 million, um, but not like for three years, right? You haven't done around in three years. And I think you also turned down a huge investment from Nvidia last year. Right? So I'm like, how are you thinking about fundraising in this environment? Like you have become such a huge part of the AI infrastructure at the moment, yet you're not following the Silicon Valley, you know, fundraise at all cost rules.
D
Yeah, we've always taken a little bit of original unique approach to things. We feel like we're building kind of like a platform for the community and they're trusting us with kind of like sharing their data and their models on the platform. So we have some sort of kind of like a long term responsibility to them. So we've always taken kind of like a bit of a conservative approach to things and not necessarily kind of like maximizing short term revenue, but instead kind of like focusing on long term sustainability. Compared to most AI startups and companies, we quite capital efficient in the sense that we don't need billions of dollars of compute to run. We kind of like close to profitability, we Just recently started to touch the money that we raised three years ago. So we, we're in like a, in a position where we optimize more for kind of like long term sustainability of the company than kind of like a short term, you know, profits or fundraising maximization. And we're pretty happy to be in this disposition and I think it's quite aligned with what we're building. Right. We become kind of like the storage and collaboration platform for AI builders. Right. We have kind of like strong network effects on the platform, but also we are platform so we have to create kind of like, you know, like 100 times more value than we would be creating if we weren't a platform and capture like one person, 2% of this value. So these things also take time and take a long term, long term approach to it. Similar to kind of like a social network if you think of it. But also I think in the long run there quite unique and quite, quite interesting. I think we ending up with this approach on a quite strategic and interesting position in the field. Right. Like we're not in a very highly competitive position. We're more kind of like in the unique position where we can keep creating value for the community and for AI builders.
C
Now when I'm thinking about value and capital and where it's flowing, right, you, you have kind of a bird's eye view over I guess like where capital's flowing, but also what are some underrepresented opportunities like where, where on hugging face, like what kinds of data sets are forming in certain industries that you're not seeing capital going to at the same rate that they're, you know, joining. Hugging face.
D
Yeah. This is a very, very big disparity. That's why a few, few months ago someone asked me if they, if we were in a AI bubble and I answered that we were probably in LLM API bubble, but definitely not in an AI bubble because there are a lot of domains topics that are under invested, for example local AI. Right. Like the ability to actually run AI on your phone, on your laptop, on your own data center rather than kind of like running it on the cloud. You know, there aren't a lot of like companies investments there.
C
Is that a hardware issue?
D
Well, I think it's a lot of, it is kind of like an investors kind of like mimetic behavior issue where you know, like we do this on the cloud. Yeah. A lot of the investments kind of like focus on the few of the very hot and kind of like common topics that everyone is talking about. Another one is Obviously, you know, biology, chemistry, all these domains have seen very, very little investment compared to text LLM APIs in the past few years. And so there are a lot to do that.
C
And then of course there's robotics. Right. Like you have hugging face has Richie, I think I see one behind you. Is that.
D
Yeah, I have a couple of them behind me.
C
Is that Riichi mini?
D
Yeah, it's a couple of like the first, first iterations of, of Riichi mini.
C
So okay, so then is robotics where open source has a big advantage because like no one company can collect all the physical data?
D
There are a couple, yeah, differences between kind of like robotics and the rest of AI. As, as you mentioned, I think data is not going to be only just more important, but also more difficult. For example, when we look at the robotics data sets on hugging face, they are huge. They're really massive. Just because, you know, video image data sets are really much, much harder to work on than text data sets. We starting to talk to people who are posting on hugging face, you know, petabytes sized, you know, data sets for robotics. The second important thing also is I think the trust issue with robotics. You know, I have a couple of like rich minis at home and I have like babies, right. And so when I think about kind of like having a robot at home that interacts with my environment, interacts with my family, interacts with my privacy, I think it's even scarier than for the rest of AI to have like a black box system just controlled by a few organizations. Especially if these organizations CEO is kind of like not the most stable.
C
Wait, who are you talking in the world?
D
And so I think for robotics, even more for the rest of AI, you need more transparency, you need more open source to have a lot of different companies competing. You need to understand how it's working, why it's working like that.
C
Now that's a really good point. I hadn't thought of that. Right, like you're gonna have a robot in your house. You're like, okay, well I'd like to know what's going on underneath the hood. I like to know what you're collecting about me. Like when you're, are you actually off when I say you're off, et cetera. So open source is probably a less scary version of.
D
And you want choices. It's the same for AI in general. Like, like a world where you have one or two choices is a very scary world because you give up some of your agency, you gave up some of your ability to kind of like decide and reward different things. So you want open source for competition to kind of like really empower not just one or two companies, but hundreds, thousands, tens of thousands of companies to be able to build different things and give people choices.
C
I feel like I want to have a whole second episode with you about robotics. So maybe you can join us again sometime. But in the meantime, we have gone over. We are out of time. Clem. Where can our listeners connect with you online?
D
Twitter or LinkedIn? You know, it's like usually the best, the best way to follow a little bit what, what we're doing and people can reach out to me there too.
C
Okay, great. Well, thank you so much for joining. This has been great to our listeners. You can find me on Twitter and LinkedIn as well. You can find Equity at Equity Pod on X and Threads. Talk to you next time. Equity is hosted by TechCrunch senior reporters and produced by Teresa Loconsolo with editing by Cal. Subscribe on YouTube or wherever you get your podcasts and find out what's next@techcrunch.com events. Thanks so much for listening and we'll talk to you next time.
Equity Podcast Summary: "Open source AI matters more than ever, according to Hugging Face's Clem Delangue"
Date: July 10, 2026
Host: Rebecca Bellan (TechCrunch)
Guest: Clem Delangue (CEO and Co-founder, Hugging Face)
This episode of Equity dives deep into the evolving landscape of open source AI with Clem Delangue, CEO and co-founder of Hugging Face. As open source models rise in prominence, the discussion covers why enterprises are gravitating towards open models, the competitive dynamics between US and Chinese AI innovation, the risks and opportunities of open source, what’s being overlooked by investors, and why concentration of AI power could be the greatest danger of all.
“You don't really make it safe by keeping it behind closed doors for a few players.... you actually make it more dangerous because you create asymmetry of power.” — Clem ([19:56])
“If you actually just enable a few companies to build frontier AI... you end up in a very, very scary and dangerous world.” — Clem ([22:45])
On the open source imperative:
"Open source is both the foundation and an accelerator for AI in general." — Clem ([14:36])
On the future of the AI industry:
"AI will be so big that both open and closed approaches will have their place, but the bulk of workloads will move to open and private models." — Clem ([03:16])
On legal and copyright risks:
"Making [datasets and models] public... allows attribution and lets people know what's used and what isn't. And it's used differently when shared for free, for the public good." — Clem ([27:08])
On robotics and black boxes:
"It's even scarier than for the rest of AI to have a black box system just controlled by a few organizations—especially if these organizations' CEO is not the most stable." — Clem ([34:27])
Takeaways:
Open source is not only thriving but reshaping how AI is built, deployed, and governed. Transparency, community, and diversity of models drive innovation and mitigate risk. Clem Delangue articulates both the promise and the perils: open source is foundational for broad, democratic AI progress, but only if power is distributed and public support endures.