
General Intuition CEO Pim de Witte joins Equity to dig into why world models trained on gaming data might be the next big leap in physical AI, how the company spun out of gaming platform Medal TV, and where the ethical red lines are when your models could end up being used for defense applications.
Loading summary
A
Hello and welcome Back to Equity, TechCrunch's flagship podcast about the business of startups. I'm Rebecca Bellon and this is the episode where we bring on an industry expert to help us explore a trend in the tech world and dive deep. For those of you who've been watching, we've been talking about embodied AI, physical AI and world models for some time now. So today we've got a show that brings that all together with video games, robotics, and a $2.3 billion valuation. Today we're joined by Pim DeWitt, CEO of General Intuition, a New York based startup found that just closed a $320 million round led by Khosla Ventures, with backing from Jeff Bezos, Eric Schmidt and researchers at MIT and Google DeepMind. And we're really excited to talk to him. Welcome to the show. So one of the things that got my attention about your deal is the heavy hitters behind it. You have some serious backers. Talk to me a little bit about how that came to be, because General Intuition is only a few months old,
B
really, as a baby, yeah, we really sought out people that we thought could make a difference. Bezos is obviously heavily focused on physical AI, so naturally there's a lot of applications downstream of solving this problem. And then same thing for Schmidt. So we focused on people who are sort of directionally aligned in the research. And then we obviously had the luxury of coastline betting a very, very large portion of their fund, which I think them obviously already sitting on the board and having all the information to sort of make such a bet. I think made the rest a bit easier. And then the other thing is, as the round progressed, the models just kept getting better and better and better, which just made closing easier. And yeah, I'm glad it all came together. I think it's one of the benefits of working with like large funds that have a reputation for not really being concerned about just backing their founders all the way. And I think it's one of those times where I think also that paid off.
A
Definitely. And I think that so much of what your company is building and your proprietary data set is, is a big part of that. So to take a step back for our customers who don't know General Intuition yet, which I don't know how they don't, because you guys have been a big deal in the last couple of weeks. Talk to us about high level, what the company does, how you spun out from metal.
B
Okay. So when you're training large language models, you're taking initially a large sort of clone of all the text on the Internet. And you're just asking a model to predict what the next text token would be in order for the model to develop an understanding of text, in order to predict text sequences, which is how you get what sort of the perceived intelligence coming out of LLMs Today, this works incredibly well. Text is a very, very good way of compressing information. Entire fields of science are built on top of texts, right? Math is largely a subset. You get this kind of jagged intelligence that feels really good in some areas, but really bad in some areas. And it's very applicable for some and not applicable for others. And the reason for that is text fundamentally removes a lot of the information that the real world needs, particularly information around space and time. So if you look at games, on the other hand, games sort of perfectly marry the information density of the Internet. So lots of interfaces, lots of information on interfaces that need to get reused when you're solving tasks. The same amount of discourse around different topics that you get on the Internet are largely present in video games or people interacting in video games. But it marries that with also the spatial and temporal dynamics of, of the real world. So these worlds are active, they evolve over time. You're interacting with other people. And so a lot of the things that text lacked, these models don't lack. We were able to raise such a large round in part because we have pretty much the only data set that has this diversity and scale represented at the level where you could take an Internet skill pre training, bet in the trillions of tokens. In the same way that the early LLMs required trillions of text tokens in order to really get their breakthrough results. It is a bet on what is already somewhat validated to be the next scale of pre trading. The question is still the next scale of pre trading, to what extent? So for example, to what extent can these models predict text? Eventually we, we don't know to what extent can these models do science? We don't know. We know that they're really, really good at controlling behavior in simulation. We know that they transfer over spatial temporal reasoning.
A
Right. They're good at understanding space and time, how they move throughout space and time.
B
Right, exactly. And one other kind of philosophical way of looking at the models, I think, is that text is always in the perception of the author. So if you write something down, you always bias it in some way with some other information that you might be thinking at that time. And you write through the lens of another way of looking at it is sex is a one dimensional sequence of outputting multidimensional thoughts and emotions and all these things that you have going on when you're reasoning. And so you have what I describe as described, reality. Whereas our models are trained in perceived reality. So they are forced to be trained in a very neutral, unbiased way to your own personal preferences. And so as a result, you get just a very different type of intelligence that feels very unlike LLMs. And in some ways that's good, in some ways that's bad because it also makes it harder to interpret in some cases because sometimes those stated preferences are incredibly useful for safety and things like that. So we look at it as the next stage of pre training and there are these fundamental issues that I think LLMs had that these address, but not without then also introducing new challenges.
A
So I want to dig into the pre training bit, right? So just to take a few steps back, general intuition spun out of Metal tv. Metal TV is another company that you own. This has hundreds of millions of hours worth of video game clips. And so not only do you have the data of. This is so cool to me. Not only do you have like the video data of, you know, recording people's clips, but also the action data. Right? Like how people, what buttons they pressed, when they pressed them, at what time. So that kind of, that kind of data, that's your unique set. Right. And so when you talk about doing pre training in a different way, when I came to your office, I was so struck by the fact that, you know, the first thing they showed me was an agent navigating through a video game on its. On its own. They've been doing so for, you know, 100 hours. I'm like, okay, yeah, well this is your bread and butter, video games. And then you had a little, well, not a little, kind of a giant quadruped walking at me and I was watching it navigate your office and you said, I almost got this wrong. In my, in my draft you said it took only eight hours. Sorry, I did it again. Eight minutes. Eight minutes of real world robotics post training data. So the model, the underlying model itself was pre trained on all of the proprietary data that you have. I'm sure you threw some other stuff in there just for good measure as well. And then you only had to give it just a little bit of fine tuning, which kind of like turns a lot of what has happened with LLMs on its head.
B
Yeah, that's a great way of looking at it. And that's right. So it was eight minutes of data. And not only was it eight minutes of data it was eight minutes of real world data from the streets. So the model had never seen yofus before. And you could argue that games have a much larger representation of outdoor navigation than indoor navigation. So the fact that it was actually able to zero shot on just the front camera, no sensors, the office, which with dynamic objects being introduced, people walking by, was a very big surprise to us and I think is a sign of what's to come. I'll caveat that by saying that some actions will work better in others. So you could for instance, make the argument that while the office is not in distribution, the act of navigation itself is. Right. And so there's levels of sort of generalization and in distribution. But it's the same is true for LLMs. Before LLMs came around there was this class of models called BERTS B E R T. They were quite general, people trained them for specialized tasks, moderation, safety, data filtering, that is things like that. They were never used in sort of this aspect of an intelligent chatbot. You obviously need a lot more data for that. But that's how I view it. There's a lot of companies right now doing lots of specialized work focused on individual embodiments, individual environments, individual robots. Right. And what's going to happen here is there's a general pre trained base that's going to come out where the generalization of the model itself is the product. Right. The fact that it has a base level of reasoning about space and time, that is going to be the reason why people stop collecting hundreds of thousands or millions of hours of real world data. Because the reality is you only need a few minutes. And I'll say also that. So a few minutes was specifically related to navigation. Right. That means that if you have a robotic arm, for example, which also ships as a game controller, which is great, but is not very present in video games like navigation, you may need a lot more. Right. So I'll caveat this by just saying that it's.
A
Yeah, I guess it depends on what you know, what's your. So yeah, it's interesting though that what your company's named after general intuition. Right. And so when I was talking to Vinod Khosla about this, who led the round, he said that the thing that, you know, with LLMs, the next big leap was reasoning. Right. And that completely changed the game. He's saying with world models, the next big leap is intuition. Right. And, and understanding this. And so I found it really interesting to talk to Vinod about your company because it seems in going back to what we Started with having backers that really support you. General intuition was born out of rejecting an acquisition offer. Right. And that you have confirmed for me. You've not confirmed who the acquisition offer came from. The information reported it was OpenAI, which you can confirm or deny right now if you like. But, you know, I'm sure there's been others who have also tried to reach out to get this data because, like, you know, it is a big data gap. You mentioned Vanilla would kill you if you sold and none of you seem to be wanting to be acquired. Like what? You know, this is an incredible data set, but why do you think they see you as more of a generational company rather than an M and a target?
B
I think it's a team. It takes a lot of special things to line up in order for things like this to become possible. If it was only one thing, that is the data, then we should have just sold. But it's, I think the data, it's the fact that we have an amazing team who can actually do this.
A
Yeah. Tell us about your co founders a little bit. One of them came from Wave, I believe, and worked on some kind of a world model there. A lot of your team is European as well, which you're based in New York, which I love. Me too. And so you've got kind of a different DNA than a lot of Silicon Valley companies.
B
So my co founders are the authors of a paper called diamond, which was a 2024 Neurip Spotlight paper, which was really one of the first diffusion based open source world models. And Vincent and Eloise specifically also are the authors and creators of Delta IRIS and iris, which are some of the foundational world model architecture papers, on top of which a lot of the other world models are then built. So look at it kind of as if this is a very abstract comparison, but as if my co founders invented kind of the Linux equivalent of this space. And again, it's very abstract, but they created something that was sort of demonstrably great. And a lot of the things that came after were modeled after those inventions, but they did that while they were still in their PhDs and so they made the choice to finish their PhDs. World models also weren't really that understood yet. And then when they did diamond, the beauty of it was that they got a real time world model running in counter strike on a 4080. So very basic GPU, real time 10fps, most importantly on 100 hours of data of which only 87 were actually used for training. And I believe it was an 8 were used for evaluations and holdouts and things like that. So they delivered a state of the art paper in one of the most constrained environments in the world. Right. They didn't have a lot of compute available to them. And when you're that good, so how
A
did you get them? You know, you're just coming from coming out with a startup like, you know, a lot of founders have a great idea, maybe they've got, you know, some other kind of data set. How are they going to get the talent they need when you know, you have, you know, OpenAI poaching from Google and you know, in our case, does it help? Does it help? Being like Europe based, you know, is like overwhelming.
B
It does. I think the most thing that helps is that you have just a good understanding of the world versus any specific location. I think you need to understand that there are geopolitical implications to the work that you're doing and that you work with people who are sort of value and mission aligned in what that is going to play out like versus where you already preemptively are opposed to in your ethics and your values. But I think other than that, it doesn't really matter where you do the work other than obviously having now expert laws to consider. That said, the thing that really did it, I think was the data. Every researcher in the world wants to do their best work and their best work and the frontier innovations is always going to come from. Where's the data? Think about it. Look at the reporting coming out of the other labs right now. A lot of people, what are they doing? They're doing data work, right? Why are you doing data work is because everything is just downstream from the data. So if you want to, if you want to be in the frontier on world models, if you want to publish frontier papers, if you want to publish the best models, you go where the data is. And in world models that's nowhere because it doesn't exist, right?
A
Well, yeah, but it does exist with you, right? Like no one was looking at gaming really. Like I don't think VC was looking at gaming. But I'm curious what other kinds of consumer products might suddenly become AI infrastructure, right?
B
The indicators I think are the number of environments the product is being used in need to be very diverse. That is in our case it was 2D interfaces, 3D games, lots of different types of games. Then we figured out that games had more diversity, interface to interface than like a SaaS application to a SaaS application, for example, because most SaaS applications are designed for the same type of group of users. And so you get less learning sort of per teaching. There were all these kind of things that I think a lot of those show by just talking to the labs about the data as well and having a good understanding of what they want to do with it, then you start putting things together, I think. And it was very rare that we were able to do this. I don't think that there are many as general categories where this can be done. I do think that there are maybe subfields. So I think there are likely very deep buckets of data in biology, in healthcare, in physics, specific types of research. Look at for instance, what ASML is doing with their 2nm. But there's an enormous amount of IP, I'm assuming underneath that that allows you to go really deep in a specific type of model class built on top of the generalization capabilities of the frontier models. So the question is, is the data set general or is it specialized? If it's specialized, then you are likely looking at a harness approach or a fine tuning approach, or if there is a specific economic problem that you can do pre training for, that's fine, but it's unlikely. And then if it's a broad data set, one thing you can do is just fine tune the multimodal models and see if you can beat some of the benchmarks. Right. Computer use benchmarks or specific benchmarks in your field. Once you can do that, you've essentially learned how to process the data, you've learned how to apply some of it, and then now you actually have a model you can kind of start seeing. Right. Who's interested in this type of stuff. And so this is one of the things that we did. And then right later on we realized, oh, we have so much of it that we can actually just do from scratch, pre training. But all of this is just a journey, it's just right.
A
Yeah. You don't come into thinking that this would happen necessarily.
B
Yeah. And the worst case outcome is you get good at this stuff and you can get a job at one of the labs. Right?
A
Yeah. Okay, well just like, for like quickly, if you can answer for me. So you're saying like a lot of this general data is hard to get, but like There's a million YouTube videos out there. Right. Like what is different about world models learning off of video versus learning off of the data set that you have.
B
Yeah. So the big difference with YouTube is that. So a lot of labs will make the argument that, oh, you can actually get action information from pixels, from the video but the reality is, if you're landing a plane and you're moving the rudder, that's not going to be in the pixel stream. Right. It's not going to be in the frames. And the reason why labs say this is because the benchmarks test for a fairly general set of things. But what customers actually want is they want the models to do incredibly well when you hit those edge cases. And that's when the fully accurate long horizon data actually starts mattering and the errors accumulate. If you use inferred data. Right. Imagine if the labs trained on all the tech state and Internet, but wrote a randomizer where out of every 98 words it would switch a few letters. Think about the downstream impact that something like that would have. Right. It allows the models to separate out you from the environment, because how do you interact with the environment is through your actions. So if you don't have the action labels, there's a completely different scaling law that needs to be considered where a model actually learns to separate out the actions from the environment and then you just get completely downstream different applications, which models that are trained from scratch on ground truth data will blow out of the water.
A
Yeah, that. Having it be separate is important, I think. And something I noticed while playing around with your world models and your agents versus like when I played around with genie, for example, I think that the way that GENIE does it, when you have a little agent there when you're doing a demo, the same world model that's generating itself is also generating the agent versus yours, the world model is generating and the agent is deciding how to move through the space separately, right?
B
Yeah. So GENIE does this reasonably well. The thing about genie, so I wouldn't necessarily ding GENIE for this one in particular, although my understanding is they do have a large amount of video data in there. GENIE purposely went really broad. Right. So they purposely went on any picture. So you could view world models as just predictors of unfolding dynamics over time in pixel space. So you start with an image. Right. And then you need to correctly predict the unfolding dynamics. Because Gini went so broad in terms of the capabilities you can feed in any picture. As a result of that, you're most likely going to encounter environments where the model hasn't internalized the physics and the collisions of that particular type of environment. So apples to apples, broadly, yes, is correct. But I will also say that that is more true for the world models other than Gini than for Gini. Gini has actually done a reasonably good job at internalizing physics because Demos has Said this publicly. They also utilize a lot of physics engines, for example, to ensure that these models have internalized these types of things. Yeah, but the difference is still that you see, indeed is ground truth versus inferred. One gives you intuitive, predictable environments and the other does not. And when you're training in them, you want predictable environments.
A
Now, you know, I could nerd out on the tech all day, but I do have a couple other things I want to ask you about. So you mentioned ethics and aligning with ethics. One of the things that we spent a lot of time talking about is, you know, where your red lines are for defense. You know, you've said that you won't build for lethal autonomy, you won't let your API be used for anything like that. You're okay with search and rescue. I'm curious, can you realistically build embodied AI without eventually serving defense in that way? Like, Anthropic had a deal with the DoD, but the US wanted to change that deal. Then over the course of those negotiations, and the US ended up being like, all right, you're a supply chain risk because you won't do what we want you to do with your technology. So does that make you think differently maybe about positioning your work with the US Government one day?
B
Yeah, I think at the end of the day, no good researcher wants to build models that harm humans. Our stance is not, oh, we're not going to do defense at all. We're not against defense. Our stance is that we don't want to be part of harming humans whatsoever. And so that is important all the way. So, for example, difference there is, do you use a violent clip to reward a model positively or negatively? Right. These are architectural, technical decisions that you
A
need to make with gaming. It's a lot of, there's a lot of violence in gaming. Right. It must be hard to keep that out of the base model.
B
Yeah. And the way you solve that is to use those examples to negatively reward the models. By the way, if you're going to train a model to try to prevent those things, the exact thing you would want is lots of examples of that negative behavior happening. Because you need to score how the model does in those situations. And the best way to do that is to run it through a scenario where you know the outcome is negative.
A
So if you do defense work, would you kind of pitch it as like, hey, just FYI, our base model, like, doesn't know how to kill people?
B
Yeah, exactly. The way that we're going about things is we train models where the actions going into the models. Everything we should be proud to talk about publicly, and we should be happy to talk about publicly. If there is an escalatory event that happens in the world. Let's say Putin invades NATO or let's say that Putin, Taiwan gets invaded, we may revisit those things. Right. The way that we service those beliefs is by hiring good people who are very, very good at evaluating those types of decisions. Right. I was in the humanitarian field. Brianna was in the humanitarian field. So there is an aspect. So, yes, of course, the models that you're building are going to have some impact on national security, and you want to be able to use them to advance democratic values in the world. Now, at the same time, I think the most important part is that you do not become part of an escalatory system.
A
Yeah, well, it's just hard to do today because the Silicon Valley is very gung ho for war.
B
Exactly. So I think you want to be proud that you can build incredibly general capabilities. And you want to also be proud that as a country, look at, for instance, During World War II, how many factories were turned into making bullets or whatever. Exactly, exactly. So these things will happen. Right. And you're part of a democracy, which means that these things, they might be inevitable. Right. You're not more important than the democracies you serve. I believe this very, very, very firmly. Now, you can make your senses clear. And the most important thing is I think you can choose the actions that you take leading into development of those models and making sure that those are in favor of not harming humans. Everything comes down from that.
A
Yeah, I think, I mean, you know, best of luck if, you know, the anthropic thing, I think, just has a lot of people shook. But we don't have to harp on about it too much because I want to talk to you about one more thing before we have to go. You know, another way that you kind of are showing ethics through your company is recognizing that with AI comes job loss. And, you know, you're a big part of the gamer community. You have been since you were a kid. I think it's really interesting that you launched Nerve recently. So Nerve is a marketplace for gamers. And just talk a little bit about that. What kinds of opportunities? I think it's anything from like data labeling eventually to teleops.
B
So we looked very seriously at, okay, these models are going to have a lot of impact on the economy. You can do two things. One, you can hire an economist. An economist. And you'll have an answer in a year. Or maybe six months or something like that. And I think it will largely be a spiritual, philosophical answer as opposed to actually doing something or directional answer. Or you can just put as much effort that you do into frontier research into creating jobs. And I think that's the answer. I think the answer is just if every single lab were to just put as much effort into research to create tools that create jobs in this new economy, which as labs we have a front row seat to, a lot of this shockwave that could happen can just be absorbed. And I think if that's the case, then you'll see some of the doomerism of the lab leaders go down as some of the largest self owns in history. The reality is that I don't want to hear how you think that models are going to affect jobs unless you're doing something about it. Right.
A
And so something more than here are workshops so you can use our tools more.
B
Exactly. And so what we're doing about it is we're getting really, really good obviously at understanding which actions people are really good at. And so we can use that to put really, really great jobs in front of them. General intuition is the first customer for this. Indeed. We're using, we're creating data labeling jobs, we're creating teleoperations jobs for our partners that want to have really skilled people teleoperate robots.
A
And that creates a feedback loop. Right. Because then you can collect the data on how these teleoperated machines are or like what actions humans are taking as they're doing that. Right?
B
That's right. Nobody has all the answers. That's the reality. And I think pretending that we do is absurd. So the best thing we can do is just on the input side is just do something. Just create the jobs and build flywheels that create the jobs. And I think that's the simple answer. And that's just what we got to do.
A
Amazing. Well, I know that you got to jump right now. So before I let you go, is there anywhere that our listeners can connect with you online?
B
Yeah, I'm mostly on Twitter or X imdewit. And yeah, look forward to hearing from you.
A
Okay, thanks so much for coming on the show to our listeners. You can find me on X and LinkedIn, et cetera, et cetera. You can find Equity at Equity podcast on XN Threads. Talk to you next time. Equity is hosted by TechCrunch senior reporters and produced by Teresa Loconsolo with editing by Cal. Subscribe on YouTube or wherever you get your podcasts and find out what's next at techcrunch. Com events. Thanks so much for listening, and we'll talk to you next time.
Equity Podcast Episode Summary
Episode Title: Your gaming data could be the secret to AGI, according to this Bezos-backed startup
Release Date: July 8, 2026
Guests: Pim DeWitt (CEO, General Intuition)
Hosts: Rebecca Bellan (TechCrunch), Kirsten Korosec, Anthony Ha, Sean O'Kane, Theresa Loconsolo
This episode explores the intersection of video games, embodied AI, and world models, focusing on General Intuition—a startup that recently raised a $320 million round at a $2.3 billion valuation, with backing from high-profile names like Jeff Bezos, Eric Schmidt, and institutions such as MIT and Google DeepMind. Host Rebecca Bellan interviews CEO Pim DeWitt to unravel why gaming data is being touted as a cornerstone for the next generation of artificial general intelligence (AGI), how their unique dataset is shaping the conversation, and what this means for the future of robotics, jobs, and AI ethics.
LLMs vs. World Models:
Breakthrough in Robotics Transfer:
YouTube vs. Gaming Datasets:
Comparisons to GENIE and Other World Models:
On Gaming Data vs. Text Data:
“Games … marry the information density of the internet… with the spatial and temporal dynamics of the real world.” —Pim DeWitt [02:23]
On Training with Video:
“If you’re landing a plane and you’re moving the rudder, that’s not going to be in the pixel stream... Customers want models to do incredibly well when you hit those edge cases.” —Pim DeWitt [16:51]
On Team DNA:
“As if my co-founders invented the Linux equivalent of this space.” —Pim DeWitt [11:20]
On Ethics and Defense:
“No good researcher wants to build models that harm humans… everything we should be proud to talk about publicly.” —Pim DeWitt [20:27, 21:35]
On Tech Industry Attitude:
“It’s hard to do today because Silicon Valley is very gung-ho for war.” —Rebecca Bellan [22:24]
On Jobs and AI:
“If every single lab were to just put as much effort into research to create tools that create jobs… a lot of this shockwave could be absorbed.” —Pim DeWitt [23:45]
| Topic | Start Time | |---------------------------------------------|------------| | Heavyweight investing & company origins | 00:55 | | Gaming data vs. LLM text data | 02:23 | | Robotics transfer experiment | 06:02 | | Deep dive on co-founders and talent | 11:05 | | Consumer products as future infrastructure | 14:11 | | Video data vs. action data | 16:33 | | World models beyond GENIE | 18:03 | | Defense and ethical boundaries | 19:44 | | Launch of Nerve, jobs for gamers | 23:12 |