
Last week, World Labs announced its acquisition of SceniX, bringing together two teams working on one of AI's biggest unsolved problems: how to give machines a true understanding of the physical world. Martin Casado sits down with Fei-Fei Li, co-founder and CEO of World Labs, creator of ImageNet, and pioneer of spatial intelligence, alongside Yunzhu Li, co-founder of SceniX and assistant professor at Columbia University. They discuss why World Labs acquired SceniX, how simulation can unlock the next generation of robotics, and why training robots may require a fundamentally different approach than training language models. The conversation explores real-to-sim-to-real pipelines, world models, robotics foundation models, evaluation, synthetic data, and why the future of AI depends not just on understanding language—but on understanding and interacting with the physical world.
Loading summary
Fei-Fei Li
We are building the next frontier of AI, which is what we call spatial intelligence.
Yun Zhu Li
At Synix, we are developing what we call a real to sim to real pipeline. We can replace all the data, all the evaluation we need in the real environment by using the data that can generate at a scalable way in our digital world.
Fei-Fei Li
Think about human intelligence. We do a lot of simulation in our head. You know why? There's a very important role simulation plays that real world data doesn't play, which is counterfactual reasoning.
Yun Zhu Li
What we are building is a consistent world. Consistent both over space, over time, over different viewpoints and over different types of interactions. My North Star is I want the robot to work.
Fei-Fei Li
The world we live in can be multiverse, that we create technology to allow people, builders, developers to act within different spaces.
Martin Casado
Do you believe we'll ever be able to build robots that have the power, efficiency of a human being? How far away are we from this? Is this like five years or this is like never?
Podcast Host / Narrator
The TLDR is language models transformed how AI understands words. The next frontier is teaching AI to understand and act within the physical world. Following World Lab's acquisition of Sinex, Martin Casado sits down with Fei Fei Li and Yun Zhu Li to unpack the vision behind the deal. They discuss spatial intelligence, world models, simulations and why solving robotics will require a new generation of AI built for three dimensional reasoning, not just language.
Martin Casado
All right, well, it's great to have you both here. So Fei Fei, for the listeners that may not have the background, maybe you can give an overview of what World Labs does.
Fei-Fei Li
Yeah, well, World Labs is a two year old startup. I think we should just recognize it's a frontier model lab. We are building the next frontier of AI, which is what we call spatial intelligence. And spatial intelligence is about creating AI that has the ability to generate, understand, reason with and interact with spaces, whether it's physical or virtual. And of course a means to an end towards spatial intelligence is building large world models. And that's what World Labs is mostly focused on.
Martin Casado
Yeah, so you've been saying this since the very beginning, which is the machine's ability to perceive and reason about spaces and act on spaces. But I always had the assumption that the acting on spaces was some long distance future thing. But now you're acquiring a robotics company and so maybe talk a little bit about the timeliness of this and the intentions.
Fei-Fei Li
Yeah, so first of all, it doesn't just take robotics to act within spaces or to interact. Right. I mean, look at the creative field, whether it's VFX or gaming or design many use cases you can create and act within virtual spaces. World Labs thesis has always been that the world we live in can be multiverse. That we create technology to allow people, builders, developers to act within different spaces. Having said that, the ability to act within the physical space is one of the most exciting and most profoundly important capability of the future AI World. So robotics is very much that. So worldlab has always believed that robotics is an important application as well as use case of spatial intelligence and world modeling. So by joining force with inviting Sinix and Scenics team to World Labs is part of our long term vision and mission. We've always committed to that.
Martin Casado
Amazing. So Yunchu, you're the co founder of Scenics, so maybe provide everyone with a quick overview of your background and what Scenics does.
Yun Zhu Li
Yeah, so I'm Yunchu. So I'm currently co founder of Sinix and also assistant professor at Columbia University. So my research started from my PhD at MIT and then postdoc with Fei. Fei.
Martin Casado
Really? That's great.
Fei-Fei Li
The world is small.
Yun Zhu Li
The world is small. It is. Throughout my career my goal has been very simple. Trying to help the robots better perceive and interact with the physical world. So I'm a very practical person. I want my robot to work in the real physical environments. So for Synix, the unique opportunity we see is that there has been a lot of bottlenecks. Right now we see faced by the development of general purpose robots, especially around training and also around evaluations. So at Synix, we are developing what we call a real to sim to real pipeline. We're about to map the real environments into the digital world that has the best alignments with the real environments. By alignments we mean that whatever happens in the digital world is also going to happen in the real environment. Such that we can replace all the data, all the evaluation we need in the real environment by using the data that can generate at a scalable way in our digital world. So that is how everything started in Synix. We put together a very, very strong and best teams around robotics, robot learning and also simulation and rendering trying to build this realtosity real stack to solve some of the key bottlenecks.
Martin Casado
It's amazing that you two work together. Yeah.
Fei-Fei Li
And there is a funny story here because you would think because we work together, he was my amazing postdoc. We've been talking about this Synix and World Lab integration for a long time. It's actually not true. They came into World Labs as a customer really. When we released the first version of Our generative model called Marble. Last winter, around November, December, Sinix just signed up.
Martin Casado
No kidding. As a customer?
Fei-Fei Li
Yes. And I didn't even know what it was. And then I realized this is Yunzhou's company. I called Yunzhou. I'm like, wow, this is your company? And then we realized there's so much synergy.
Martin Casado
Maybe Faye, if you just quickly describe what Marble is.
Fei-Fei Li
Yeah. Marble is the codename for the base model that WorldLab has been training and iterating on. The fundamental capability right now of Marble that is publicly released is to take a prompt. It can be an image, it can be a few images or a text and turn that into a geometrically consistent world that can be represented in 3D geometry, whether it's Gaussian splat or mesh. Really what Sinic's team is doing is trying to solve this extremely difficult problem in robotics, which is the lack of data, the lack of data in training, the lack of data in evaluation. This is very, very different from language models where data is abundant on the Internet. And we know that in order for robotics to work, we have to somehow unlock the power of scaling law. But where does that come from? This is something that there's a profound problem that everybody's battling with in robotics.
Martin Casado
It'd actually be great to talk about this energy. You have put together a very, very talented team. You have put together a very talented team. And so to what extent is there overlap? To what extent is this an extension? Maybe talk a little bit about that. Yeah, that's how complementary it is.
Fei-Fei Li
It's actually the TLDR is very complementary and with a shared mission. So Yun Drew is one of the three technical co founders. The other two are Chang Xi Zheng, another Columbia professor who has been a world technologist in simulation. And Cheng Xi has his background in also vfx. He worked at Weta, he worked at Tencent, he's being an entrepreneur. Then there's Sunny Hu, who is a phenomenal engineering leader who was also in a startup that was acquired by Amazon many years ago. So he worked in many different tech stacks in the computer vision field in Amazon. So when we started talking more seriously, I recognized that a couple of things that Sinix has from a talent point of view, is extremely complementary to World Labs. One is obviously Yunzhu's incredible thought leadership and just technical prowess in robotics. Right. So from really from hardware, full stack robotics. And even when he was my postdoc at Stanford, at that time you already had your faculty offer, so you were there only for one year. I Wanted you for more than one year. But he had to go become have the real job. So he was a full stack researcher in robotics from modeling to hardware. And of course Yunzhou and his students as scenics was that pool of talent World Lab hasn't had yet. Then on the Changxi side is just incredible simulation capability. Right. He's such a senior researcher and technologist in simulation and World Labs is doing is very much interfacing the world of simulation. So I think what they don't have obviously is on the generative model side as well as the computer vision 3D reconstruction side. We're also very strong at World Labs. So that's a technology that CNX needs. So together these two sides come together and make it much more complete.
Martin Casado
Jifei's motivation in this is like this is an extension and a complement to get into robotics. You know, having been in your situation, which is deciding when to sell a company, it would be great to hear from you on like how you think about joining World Labs and kind of the fit there and like why you made the decision to do it.
Yun Zhu Li
Yeah. So at the very beginning we were deciding, okay, do we want to just keep going? But after chatting with Fei Fei, after seeing all the synergies that happen in the middle, it just makes perfect sense for the forces to join each other. So in essence, as cynics, what we've been doing is real. To seem to real is to do dense reconstruction of the environment. So we capture the appearance of the environment, geometry of the environment, and also the dynamics of the environment, meaning how the environment is going to change when you apply actions. So this dense reconstruction right now is still a little bit on the heavier side. And what World Labs right now has been doing involves a lot of profound capabilities around sparse reconstruction and generations. So we see a lot of opportunities of leveraging marble and other capabilities as World Labs and in order to do very efficient reconstructions and modeling of the environment.
Martin Casado
So can we expect a foundation model for robotics from World Labs?
Fei-Fei Li
World Lab is building a foundation model. As you know, Martin, we're building a base model and as the technology has been evolving, some of the most exciting base models are omnimodels. Right. They take multimodal input, they have multimodal outputs. And what is a foundation model for robotics? It's very likely going to involve actions. It's very likely going to involve the output of actions in addition to the state of the world. And we're definitely not ruling this out.
Yun Zhu Li
Yeah, great. So for example, for the foundation models, it Essentially needs to be a multimodal model. So it has to take into account frame, text, image depth and different kind of modalities. And action is a very, very important part of that modality. So if you think about frame actions as inputs that essentially a forward simulator that is going to predict how the environment is going to change when you apply a specific action, when the action is output. This is essentially a policy model that is trying to predict given a specific goal, like what should be the action you take in the real environment to get you closer to that goal. So this kind of omni models actually can benefit a lot and actually provide huge amount of values for the robot robotics communities in trying to understand how to model the environments and at the same time how to act in the environments. And this can also act as a backbone for you to fine tune into specific robotic applications to making sure it's really live up to the reliability and efficiency that expected by the clients.
Martin Casado
You know, if Yunchi, if you don't mind kind of a lay investor question, I see a lot of robotics companies and a very popular approach right now for the robotics companies that come in is like we'll use a video model, you know, and like, you know, that's the predominant method where this is, you know, 3D and simulation. It's a very different approach. And so maybe you could contrast, you know, this popular approach of just using video only versus kind of what the ambition here is.
Yun Zhu Li
Yeah, so in order to create words with the robot kennel, the words as I mentioned, need to capture the essential structure of the problem. And one of the very important necessary like requirements for those words will be consistency. So that is where I actually see there's very, very strong synergies with marble because what we are building is a consistent world, consistent both over space, over time, over different viewpoints and over different type of interactions. And marble, the generated world from marble is also provide an infrastructure, a component of that entire world that we believe is necessary for the robot tuner. Imagine if a robot's pushing an object forward, the object just magically disappear, which has been a problem or many of the existing video prediction models. This one provides good enough signal for the robot to know what is the right thing to do. But obviously right now there has been a lot of investigation on building better and better and stronger and stronger video models. So we actually see a way where some of the infrastructure we build can provide as initial momentums and to going through this data flywheel of going from this more simulation driven models into like robot policy models. Which going to do the execution in the real environment, collecting new data. The data will come back in where the model doesn't necessarily have to be physics only or learning only, but somewhere in the middle, which be able to capture the essential structure of the problem, but at the same time be able to scale and become better and better as you accumulate more data.
Martin Casado
You know, I've worked now feifei, very closely for a while and you've always had this North Star which has driven this and you've articulated variously as kind of 3D and in a number of other ways. And I'm just wondering for you, is there also a similar philosophical North Star or you're more the pragmatic like I am. Build the system, like do the thing.
Yun Zhu Li
My North Star is to make robots work.
Martin Casado
Amazing.
Yun Zhu Li
Yeah. In the real environment, I'm a very practical person. I want the robot to work. One interesting thing that's actually coming from my collaborations with Fei Fei during my postdoctoral. We are building this kind of benchmark. We actually send out surveys asking the general public what they want their robots to do for them. Among the thousand tasks we collected, one third of the tasks are about cleaning. People just don't like to do those dull and dirty tasks. And those are the scenarios that we really want to making sure we have robotic solutions to deal with.
Fei-Fei Li
One thing I really like about Sinix Martin, especially continuing your question, there's a lot of robotics companies building models and all that. One thing I truly like about Cynics is Rindu and his co founders have such an incredibly pragmatic approach to robotics. Especially they come from academia. Right. Sunny doesn't, but Yunzhu and Xi come from academia. But their first instinct is work with design partners and customers in real industry, whether it is labset industry labs or warehouses or electronics, electronics assembly. That is such a refreshing, actually a refreshing way of approaching robotics. And that really true made me very excited to work with it.
Martin Casado
This is for Yunchu, but I just feel this is personal curiosity, which is. It seems to me that for robotics you have to be pretty exact. I mean, not perfect, but pretty close. But for the creative use cases, which World Labs has done, a lot of you kind of don't need to because, you know, I mean, even sometimes being wrong is stylistic or intentional or whatever. And so from a technical perspective, what is the challenge here for reconciling these two things? Or do they never get reconciled? There'll always be two points in the
Yun Zhu Li
design space, so they will be reconciled in the long terms, of course. And modeling of the environments doesn't have to be perfect, the model doesn't have to be perfect in robotics.
Martin Casado
And by the way, again, this is pure curiosity, but is there like a bit more formal way to say that? What does that mean not to be perfect? It has to be pretty close.
Yun Zhu Li
So let me put it this way. For example, models over the developments of all different kind of robotic applications has been a very important cornerstone. If you look at all the existing robotic applications, like plane, Jones, Roomba, or even for quadruped robots, bipedal robots, model has been the way for them to actually work and be able to transfer from simulation to the real environment. But if you look at those locomotion robots, like quadruped robots, bipedal robots, they can walking on snows, they can walking on bushes, but you don't need to have a simulator. You can simulate all the bushes and snows like very precisely. You need to have a simulation that captures the essential structure of the problem and do whole different kind of randomizations inside the digital environments. So that is what we're aiming for. So basically with cynics and together with world labs, we're trying to investigate what is the level of fidelity. We need to model the massive, massive worlds besides the robots, such that we'll be able to transfer the robotic systems, training, the simulated environment, digital worlds back into the real scenarios.
Martin Casado
As an investor, I've heard other researchers say, like Sergey Levine say simulation will always eventually deviate from the physical world. And real world data collection is absolutely critical. And so maybe talk a little, little bit about like the viability of this approach where simulation is a cornerstone as opposed to some other approach so they
Yun Zhu Li
don't contradict with each other. So if you're thinking about the simulation, simulation essentially trying to predict how the environment is going to change when you apply the actions. And this is essentially a model of the world that doesn't necessarily have to be pure physics. It can be a combination between both physics and also learning. We are collecting real world data, we will be using those real world data. It's just at different stages of this, like a data flywheel. Maybe at the very beginning we have stronger emphasize on we have more physics to making sure we have the right consistency and right structure for us to learn the world, for us to train the robot policies. But as we accumulate more and more data, both through data collection and also through the collaboration with our clients, we'll have the data that will be moving towards more, towards more learning based like modeling of the environments. So this kind of transition and Also, this kind of data fly is really an enabling factor of both getting the best of both physics and the geometry and consistency, as well as all the power and magics from the data and compute.
Fei-Fei Li
I want to add to this and be slightly philosophical here is there isn't a binary choice between simulation or no simulation. All this comes in together to make robotics work. Think about human intelligence. We do a lot of simulation in our head. Why? There's a very important role simulation plays that real world data doesn't play, which is counterfactual reasoning, is that you play out events that hasn't happened or cannot happen, or you don't have enough data to make it happen in real world. And while, while you play it out, you learn how to act in it. Humans do this all the time. We probably don't, you know, we just. I know you were at the World Cups.
Martin Casado
I was at the World Cups.
Fei-Fei Li
Congratulations to Spain winning. I'm sure in the planning of every game there is simulation, whether it's digital or on the whiteboard or whatever, that simulation, the role simulation plays is counterfactual reasoning. And that's really important in robotics because we just do not have, cannot possibly have enough real world data for that. Here's a real life example. The industry of self driving cars. Waymo has officially said they use billions of hours of simulation. And actually Waymo is more simulation heavy than just real world data heavy. So these are real examples. And as you know, Martin, Andrew too. Cars are the simplest kind of robots.
Martin Casado
Yeah.
Yun Zhu Li
2D.
Fei-Fei Li
Yeah, yeah, they sort of. So clearly simulation plays a huge role in robotic learning.
Yun Zhu Li
I also want to add to that. So if you put things more specific, simulation can provide two levels of benefits. The first one is reliability and the second one is efficiency. So for reliability, if you're thinking about a robotic system working reliable in the real environment, you need data to provide systematic coverage of all the state space and variations that robots might encounter. That's how you can learn of how it gets robust. So with simulation you can do systematic randomizations and control and variations of lighting, frictions, geometries, object types, and also all different kind of physical parameters to making sure you have sufficient coverage of the state space. So this is what can give the robotic systems reliability. And second is about efficiency. So right now many people are doing teleoperation and if you look at many of the teleoperation device, imagining all the actual skeletons you are using, you are actually collecting, tracking the data at a speed that is actually slower than human actually doing the task. But for many of our clients, human speed to them is not good enough. They want faster than human speeds. So for the robot to move faster, it's not as simple as just drive the robot faster because the gravity doesn't change. But in simulation you can do systematic speed up of the robot's behaviors to train the robots, such that it considers all the dynamics, changes of the environment. So this is what can give our clients for them efficiency. So both for the reliability and efficiency, there are some kind of very unique values where simulation can provide.
Martin Casado
You've talked about the technology and the platform, what it does, maybe talk about the specific use cases people use it for.
Yun Zhu Li
These are essential. Like two specific use cases, especially around those training and also around evaluations. Starting from the evaluations. So evaluation is something like people often overlooked in the robotics. But if you are training like robotic models, you have to know how well it works and that is the only source of information for you to iterate.
Martin Casado
By the way, a lot of every AI person really understands what evals are and uses it all the time. Non AI people, it often means something a little different. So maybe it's even worth just describing specifically what you mean by evaluation.
Yun Zhu Li
Okay, so what I mean by evaluation is you'll be able to understand for this specific checkpoints, how well does it perform? Does it perform for example, 95% of the time or 99.9% of the time? And the key criteria people use in industry is how long does it take? How long in time does it take for you to distinguish between a checkpoint that is 90% from a checkpoint that is 92 points. And if you only do that in the real environment, that's just takes so long for you to do the distinguishments. And if you really think about also the robotic evaluations right now, people are doing in the real environments, the iteration speeds is multiple orders of magnitude slower than iterations of those language models. So not only is like the robotic tasks very varied, very diverse.
Martin Casado
Oh yeah, because like you actually have to do the thing,
Yun Zhu Li
break the glass
Martin Casado
atoms, have to move through space.
Yun Zhu Li
Exactly.
Martin Casado
The laws of physics have to be.
Fei-Fei Li
Have you watched those robotics videos? Every video has like 10x8x because it moves so slowly.
Yun Zhu Li
Exactly. So not only is slow, it's dangerous, it's costly. But at the same time the speed is also like multiple orders of magnitude like slower. So some of our clients actually need this digital environment that can be used to evaluate their robotic systems. And because our digital environment has proven alignments with the real world, so meaning whatever happens in the sim is also likely to happen in the real environment. If a checkpoint is working better in the simulation, it's also highly likely to also work better in the real environments, as we have also been discussed in the blog post. So that actually give our clients very strong confidence in actually using the data, using the signal from the digital environment to do scalable, safe and much faster evaluations of their robotic systems. Great, so that is on the evaluation then on the training. So on the training side. So basically, like I also mentioned, it's about controllability. If you want to control all the different possible variations of states, parameters, lighting, frictions, physical parameters, like even object geometry, object types. So you want to making sure you have sufficient coverage of all different kind of scenarios such that you'll be able to generate like informative data for your robots to be robust. And this is just going to be so hard to do just in the real environments, like we discussed. If you do teleoperation, the speed at which you are collecting data is slow. You're also limited by how many robots you have, how many teleoperation devices you have. There's a whole different kind of challenges around all the data operations around it. But in simulation, everything can be controllable, everything can be systematic and everything can be understand at a level where you know exactly and making claims about exactly what distribution you have covered to develop confidence about within the distribution. We know the robot will work. So those kind of confidence and efficiency and scalability is something that our clients also value, to use our digital words, for the training of robotic systems.
Fei-Fei Li
Here's the crazy thing. Even before Sinix and we are talking, our inbound customers for marble were already seeing this kind of demand. We just cannot serve these customers. But we are already getting a lot of phone calls from robotics, early stage robotics companies who are developing their models all the way to downstream very pragmatic use cases. And we're seeing these needs.
Martin Casado
When people hear you're going into robotics, what they're going to envision is you're pulling out a 3D printer and you're going to be making hardware and then you're going to be programming the brain of a robot and sticking it in the robot. And then you've got a robot. And I don't think that's what you guys are talking about here. So maybe talk about where this fits in the life cycle of creating a robot and where you will end and where the rest of the ecosystem will begin.
Yun Zhu Li
So what we have been building, you can imagine, is an infrastructure with the softwares around these infrastructures for people to for example build worlds such that robot can learn and evaluate. And this infrastructure is naturally model agnostic and embodiment agnostic.
Martin Casado
So I just want to be very clear, just because this is actually a very subtle, I mean for you it's obvious, but it's a very subtle point which is from what you said, that's not building a robot. It's building an environment which another company can place their robot brain to navigate and to learn.
Yun Zhu Li
Yeah. So for our customers right now they have all different kind of robots. Some are using for example single robot arm, some are using bimanual, some are using a fixed arm, some are using like mobile manipulators, some using grippers, some are using some more elaborate versions of the undefactors. So our platform right now is just naturally embodiment agnostic. We can very easily integrate different kind of robotic embodiments, be able to put them into the worlds we generated, we digitalized such that we will be able to give those individual robots capabilities of, of doing the right tasks and at the right levels of reliability and efficiency in the real environments. And we are also for example model agnostic. So we can just using the data generated by our worlds to train different models either from scratch or doing post training of existing foundation models like vision language action models or word action models. So to us it doesn't matter. We just want to making sure we have the infrastructure, we have all the words such that the robot can work reliably in the real environment.
Martin Casado
You have told me that you think a lot of the predictions around humanoids are a little bit aggressive and we're likely to see more constrained rollouts like warehouses or whatever. Can you talk a little bit about that and how that impacts what you're going to be tackling here at World Labs?
Yun Zhu Li
So that's a very good question. So if you look at for example all the progressions of robotic applications in the real environments, it has always followed the trend from going from fully structured environments into semi structured environments and then into unstructured environments. For fully structured environments, what do we mean? That you have knowledge and control over all the configurations within the environments.
Fei-Fei Li
Like factories.
Yun Zhu Li
Like factories or for example car manufacturing lines. Those has been automated for decades. And then you have for example, semi structured environments which you have certain controls over the environments. For example like the Amazon warehouses or for example restaurants, hotels, where you have certain control over the environment to just make the task easier for your robots. But there are obviously many other objects or for example clothes. Those are the objects you don't have control. And then for the unstructured environment, it's like your home and my home. Those is, I would say, the grand
Martin Casado
challenge, especially my house. Trust me. Three dogs, five year old dogs.
Yun Zhu Li
Exactly. If you're thinking about where does the robustness coming from? Robustness coming from a sufficient coverage of the scenarios that robots might encounter. So it's so much easier and more approachable, at least like right now to focus more on the semi structured environments before we move on to fully unstructured environments. So we will move into that direction. It's just we want to take a more sustainable and more realistic approach towards it.
Fei-Fei Li
I think your point here is that humanoids mimics human body and evolution has optimized human body for unstructured environment. And so our fingers, our legs are not the best apparatus to do one thing. For example, if our only goal as a species is to climb trees, we will not have this body necessarily. So we'll have different kind of fingers. But what humans end up having are evolved into is this body shape that can be very general, but not necessarily best at everything. And that is for the survival of unstructured environment. But from a business point of view, from a pragmatic technology point of view, that and this unstructured environment and a generalized body is actually the hardest problem to solve. It's not necessarily even the right way to solve the problem. So we take more specialized body to solve a narrower problem. But the challenge for cynics is that to be more body agnostic so that their infrastructure can serve different bodies and different semi structured environments.
Martin Casado
A common lens to look at exactly this question is an economic lens, right? Which is you compare it to generative LLMs. They can create prose or code 10,000 times faster than a human being, a bunch cheaper than a human being. So the economic case makes sense because our brains aren't very efficient at that. However, our brains and our bodies are very efficient at 3D navigation, like movies of the world or picking things up. And so this is just a prediction question, but do you believe we'll ever be able to build robots, at least in the foreseeable future, that have the power efficiency of a human being when it comes to menial tasks. So let's say just basically, you know, minimum wage or something like that, like how far away are we from this? Is this like five years or this is like never?
Yun Zhu Li
I think it's going to take a very long time. So if you're really thinking about like robots in the real environment, in the end it will always Be a system. So every working robot in the real environment is a system work. You need to be very mindful and thoughtful about how the systems are coming together. The hardware, the software, the brain, even to the details of for example, what's the friction coefficients of your fingers. So there's a lot of things you have to consider to make these things a reality. And it will take iterations. But what I am excited about is that I have always been at the state of the art of robot learning and also trying to push the state of the art forward, but the state of the art always moving faster than I expected. So what I'm focusing on and trying to investigate right now is very different from when I started my PhD. So this is a speak to how fast the whole ecosystem has been evolving and all the moving pieces started coming together or building this robotic system. But we also have to be calibrated about our predictions. So we will see a lot of progress. But to achieve for example human level efficiency and capabilities, it will take longer.
Fei-Fei Li
MARTIN the hardest thing in today's AI is to have the right measured optimism.
Martin Casado
It's totally true.
Fei-Fei Li
I mean even LLMs does not have human brain efficiency. Human brain operates on 30 watts.
Martin Casado
Yeah, that's true.
Fei-Fei Li
That's so we are far from that.
Martin Casado
So but that, but I mean performance to power, it may be close, right?
Fei-Fei Li
In narrow tasks like software engineering, like
Martin Casado
generating an image or software engineering than it is, right? Yeah, I think so. I don't think we're anywhere close when it comes to robotics. Does this change how you think about your like strategically the level of ambition that your team can go after? I mean, does it is it changed that or is it still very much in line with what you expected to do when you started?
Yun Zhu Li
It definitely changed the trajectories in a very profound manner. So we see a lot of unlock in be able to do this whole process through the modeling of the environments in a much more efficient and much more scalable manner, especially in partner together with World Labs. And I also want to add to fei fei if you're thinking about for example the current states of the language models. So those are models that with incredible capabilities but still you don't just just blind trust it to book your flight tickets or make your hotel reservations. You still hopefully there's still a person proofreading the output from those language models. But that is very different from how people and will be using for example robotic models. Because for robotic models out of the box, the robots has to work reliably in the real environment. And we don't even have the data, we don't even have all the necessary infrastructures around those for the robots to just out of the box work reliable in the real environment. So for that reasons be able to create these digital worlds, these scalable digital worlds where the robot can learn and evaluate within. Yes. It's going to unlock so much more potentials for being able to replace all the costly and unsafe data in the real environments with the data generated from the worlds for the robots to be able to do scalable learning and evaluations.
Martin Casado
I've seen many of these integrations. They actually work very well at this stage when they have this much alignment, which is great. But there's always this question, question of do you integrate now into what's happening now or do you keep things quite separate and provide kind of like a long term trajectory that will be realized in the year timeframe? How are you thinking about this Fei Fei, Is this something that integrates right away or is this kind of a separate longer term?
Fei-Fei Li
This is a great question. I think at this point, Yunzhu, Changxi, Sunny, Justin, Ben have been talking about this. At this point we are going to take it thoughtfully. We're not rushing to integrate everything from code base to teams because I think Syndex does have a very well thought and I wouldn't call it standalone completely but fairly contained tech stack as well as their customers as well as the kind of products they're building. We're going to take time. We definitely will. We already on the simulation side as well as the potential base model, action condition model side, we already are starting to talk and also they are using Marvel as a internal customer. So we will be integrating but we're not rushing to blend the team as like a full salad bowl.
Martin Casado
How are you thinking about geographies with this? Will scenics move? Is it going to stay in the
Fei-Fei Li
same place we're going to? Yunduro is going to move.
Martin Casado
Oh well, welcome here.
Yun Zhu Li
I'll be moving to San Francisco.
Martin Casado
Yeah, Florence during the Renaissance. Perfect.
Fei-Fei Li
I think Aurora Labs is officially becoming a bicoastal company where the headquarter is in San Francisco. I live in Palo Alto. I feel like I'm in a different state but. But I'm actually excited that we're going to have an office in New York that can help us to attract talent on East Coast. And also we have been talking about making sure that in both offices we set up the robots so that we get to basically test out and mature our engineering stack so that we can Work with robots remotely because we have to do that for our customers anyway.
Martin Casado
So maybe just to be very concrete faith say maybe let's just pencil out like what is the perfect success case in two years? Like what product do you have? Who's engaging with it? How do they use it? Just the crisp like.
Fei-Fei Li
I would be very happy that Cynics Team World Apps Team will have validated customers in in a small number of important vertical use cases where our system, our infrastructure has proven to be truly beneficial to their automation needs. And these customers became our lighthouse examples to scale our business.
Martin Casado
How early, let's say someone listening to this is running a robotics company. How, at what stage do they engage with World Labs? Is it really early on? Is it somewhere in the middle?
Yun Zhu Li
So right now for our customers, because we are building this kind of real to sim to real pipelines where the simulation is essentially the words we're gonna provide the training and evaluation grants. Some customers, they need only the real to sim part. They want to digitalize the task they care about and be able to do the evaluations of their robotic system. Some custom customers need this real to sim to this entire pipeline such that they will be able to have policies running on their hardwares. So our platform is also designed in a way that is flexible, depending on what our clients needs. And at the same time the clients we are working with are actually pretty close to the deployments like a stage. So basically they are working on very, very practical tasks. Those tasks when replaced when we have a robotic solutions that are there can just create value value immediately and they have at least tens or hundreds of this kind of situations they are thinking about to do the automations for. So as together with wordlabs we'll be able to develop reliable solutions for those scenarios. As we have already showed, we have a number of scenarios already instantiated in our blog post and we'll be able to further our investigation to see how they can actually solve the key like requirements and also constraints faced by the real world deployments. Great.
Martin Casado
I want to be very specific about this. Is it ever too late or too early to call World Labs if you're a robotics company?
Yun Zhu Li
No, we want everybody to call us. We want to learn about your use case. Wonderful.
Martin Casado
If you're listening to this and you're anywhere close to a robotics project or robotics company, please track World Labs.
Fei-Fei Li
Yes, thank you. Definitely. Open for business.
Yun Zhu Li
Open for business business.
Fei-Fei Li
Not too early.
Martin Casado
All right. If you're doing robotics, call World Labs. Thank you both very much for coming.
Fei-Fei Li
Thank you.
Podcast Host / Narrator
Thanks for listening to this episode of the A16Z podcast. If you like this episode, be sure to like, comment, subscribe, leave us a rating, or review, and share it with your friends and family. For more episodes, go to YouTube, Apple Podcasts, and Spotify. Follow us on X16Z and subscribe. Subscribe to our substack@A16Z.substack.com thanks again for listening and I'll see you in the next episode. As a reminder, the content here is for informational purposes only, should not be taken as legal, business, tax, or investment advice, or be used to evaluate any investment or security, and is not directed at any investors or potential investors in any A16Z fund. Please note that A16Z and its affiliates may also maintain investments in the company's discussion discussed in this podcast. For more details, including a link to our investments, please see a16z.com disclosures.
Episode Date: July 28, 2026
Host: Martin Casado (Andreessen Horowitz)
Guests: Fei-Fei Li (Co-founder, World Labs) & Yun Zhu Li (Co-founder, Synix & Assoc. Professor, Columbia University)
This episode dives into the state-of-the-art intersection of spatial intelligence, simulations, and robotics, focusing on the rationale and synergies behind World Labs’ acquisition of Synix. Fei-Fei Li and Yun Zhu Li discuss the current bottlenecks in robotics–particularly the shortage of high-quality data and the limitations of real-world training and evaluation–while mapping out their shared vision for building robust, generalizable AI systems that can perceive, reason about, and act within both physical and virtual environments. The conversation balances deep technical insights with strategic outlooks on the path toward real-world, impactful robotics.
| Time | Topic | |-----------|-------------------------------------------------------| | 00:00 | Defining spatial intelligence | | 01:06 | The next frontier: physical world reasoning | | 01:49 | World Labs' mission and world models | | 04:09 | Synix: “Real to sim to real” & robotics bottlenecks | | 06:21 | Marble explained | | 11:05 | Foundation/omnimodels: multimodal input & actions | | 13:07 | Why not just video? The need for 3D world models | | 14:58 | Philosophical vs. pragmatic North Stars | | 18:16 | Simulation vs. real-world data in robotics | | 21:30 | Simulation's roles: reliability & efficiency | | 23:10 | Evaluation and training in simulation | | 27:46 | Simulation platforms vs. physical robots | | 29:41 | Structured to unstructured: deployment priorities | | 33:25 | Timeline for human-level power & efficiency | | 37:06 | Integration strategy post-acquisition | | 38:20 | Geographic expansion | | 39:07 | Defining a two-year success case | | 40:13 | When robotics companies should engage World Labs | | 41:57 | Open call for robotics partnerships |
If you're anywhere close to a robotics project or company, they're open for business—call World Labs!