After nearly two years, Fei-Fei Li finally sits down with a16z again, diving deep into Spatial Intelligence and World Labs' comprehensive strategy in the robotics field.
The entire session was packed with high-density information—one couldn't afford to blink!
Our goal is to establish Spatial Intelligence as the next frontier of AI.
The ability to act within physical space remains one of the most exciting and profoundly meaningful capabilities in the future AI world, and robots are at its core.
Robotics is the first proving ground for spatial intelligence and a crucial application scenario for world modeling.
SceniX tackles the most challenging problem in robotics: both training and evaluation data are extremely scarce.
To make robots truly functional, we must unleash the power of scaling laws.
Real-world evaluation is not only slow but also dangerous and expensive. In contrast, simulation environments offer complete controllability and analyzability; they provide unique value in both reliability and efficiency.
The specific model used isn't that important because models will always be upgraded, but infrastructure won't.
As soon as the video was released, netizens in the comments section let their imaginations run wild. Some even brainstormed a Robinson Crusoe-style adventure with robots, laying out all the necessary equipment! (Is that...appropriate?)

One has to admit, this generation of netizens possesses such wonderful mental states that it makes you want to hand them a clapperboard and yell 'Action!' right away.

Recently, World Labs announced the acquisition of SceniX, planning to collaborate with them to build a scalable digital training ground for robots.
Shortly after, World Labs CEO Fei-Fei Li and SceniX Co-founder Yunzhu Li appeared together on the latest a16z podcast "Fei-Fei Li is Solving the Hardest Problem in Robotics", publicly dissecting for the first time the core problem they are jointly solving: enabling AI to not only see the world but also physically change it.
Interestingly, Yunzhu was once a postdoctoral researcher personally supervised by Fei-Fei Li. What story led to this collaboration?

Without altering the original meaning, we've had an AI meticulously proofread the transcript of the interview.
Enjoy~
Why Are Robots the First Proving Ground for Spatial Intelligence?
Martin Casado (Host): Fei-Fei, could you first introduce World Labs' work for listeners who might not be familiar?
Fei-Fei Li: Sure. World Labs is a two-year-old startup, a cutting-edge model lab.
Our goal is to establish Spatial Intelligence as the next frontier of AI.
Simply put, it's about enabling AI to freely generate, understand, reason about, and interact within real or virtual spaces.
Building large world models is one path towards spatial intelligence, and it's the primary focus for World Labs currently.

Host: From the beginning, you've emphasized that machines need to perceive space, reason about it, and act within it. I initially thought acting in space was a distant goal, but you've now acquired a robotics company. Why is the timing right for acquisition? What's the intention behind this?
Fei-Fei Li: Robots aren't the only way to act or interact within space. In creative fields like VFX, gaming, and design, people have long been building worlds and engaging in immersive interactions within virtual spaces.
World Labs has always held a belief: the world we live in can become a multiverse, and we can create technology that allows users, creators, and developers to operate across different spaces.
However, the ability to act within physical space remains one of the most exciting and profoundly meaningful capabilities in the future AI world, and robots are at its core.
World Labs has always believed that robotics is the first proving ground for spatial intelligence and a crucial application scenario for world modeling. Inviting SceniX and its team to join World Labs is a natural part of our long-term vision and mission.

Host: Yunzhu, you're a co-founder of SceniX. Could you briefly introduce your background and what SceniX does?
Yunzhu Li: Hello everyone, I'm Yunzhu Li, co-founder of SceniX and an Assistant Professor at Columbia University.
My research journey started during my PhD at MIT, followed by a postdoc at Stanford under Fei-Fei. My professional goal has always been simple: to help robots better perceive and interact with the physical world.
I'm a very pragmatic person; I want robots to work in real physical environments.
We've observed that general-purpose robots face many bottlenecks, mostly centered on training and evaluation. Therefore, we're developing a real-to-sim-to-real pipeline that maps real environments to a highly aligned digital world.
Alignment here means that what happens in the digital world should also be possible in the real world.
This way, data collection and evaluation tasks that originally had to be done in the real world can be partially replaced by scalably generated data in the digital world. That's where SceniX started.
We've assembled a team covering robotics, robot learning, simulation, and rendering, and we're trying to build a complete real-to-sim-to-real tech stack to solve these key bottlenecks.

Fei-Fei Li: There's also a funny story here (laughs). You might think that because Yunzhu was my postdoc, we were discussing merging SceniX with World Labs early on.
Actually, no. SceniX initially came to World Labs as a customer.
It's a small world. Last winter, around November or December, we released the first version of our generative model Marble. SceniX was among the first to register and use it.
At first, I didn't even know it was Yunzhu's company. I found out later and called him, realizing there were many opportunities for collaboration between World Labs and SceniX.

Host: Could you briefly introduce Marble?
Fei-Fei Li: Ok. Marble is the foundational model that World Labs continuously trains and iterates on. The core capability of the current public version is to take one or more images, text, or combinations of these modalities and transform them into a geometrically consistent world.
This world appears very realistic and artifact-free.

And SceniX tackles an extremely difficult problem in robotics: both training and evaluation data are extremely scarce.
This is different from language models, which can access abundant data from the internet. Robots cannot.
To make robots truly functional, we must unleash the power of scaling laws.
But where does this data come from? This is a deep question the entire robotics community keeps asking.
How Do World Labs and SceniX Complement Each Other?
Host: You've both built strong teams. How much do your capabilities overlap, and to what extent do they expand World Labs' abilities?
Fei-Fei Li: I believe the two teams are highly complementary and share a common mission.
Yunzhu is one of SceniX's technical co-founders. The other is Changxi Zheng, a professor at Columbia and a world-class simulation expert with a VFX background. He worked at Tencent and has startup experience.
Sonny Hu is an excellent engineering lead who worked at a startup later acquired by Amazon and participated in various computer vision tech stacks at Amazon.
Yunzhu brings thought leadership and technical strength in robotics, covering hardware to full-stack robotics. Even during his postdoc at Stanford, he was already a full-stack robotics researcher spanning modeling and hardware. Yunzhu's students and the SceniX team add talent that World Labs previously didn't have.
Changxi brings strong simulation capabilities. World Labs' work intersects heavily with simulation worlds. What SceniX relatively lacked was generative models and 3D reconstruction capabilities based on computer vision, which are World Labs' strengths and the technology SceniX needed.
Combined, our capabilities will be more complete.

Host: Yunzhu, when deciding whether to stay independent or join World Labs, how did you assess the fit? Why did you ultimately make this decision?
Yunzhu Li: Initially, we considered staying independent. But after talking with Fei-Fei and seeing the synergy, we felt moving forward together made a lot of sense.
SceniX's real-to-sim-to-real work involves dense reconstruction of environments, including capturing appearance, geometry, and dynamics—how the environment changes after actions are applied. Currently, dense reconstruction tasks are still relatively labor-intensive.
World Labs has deep capabilities in sparse reconstruction and generation. We see many opportunities to leverage Marble and other related capabilities to perform environment reconstruction and modeling more efficiently.

Host: Can we expect World Labs to introduce robotics foundation models?
Fei-Fei Li: World Labs is building foundation models, and we certainly don't rule out this direction.
Yunzhu Li: Yes. Robotics foundation models are essentially multimodal models.
They must handle frames, text, images, depth, and other modalities, with action being a crucial one. If frames and actions are inputs, the model can act as a simulator, predicting how the environment will change after specific actions.
If action is the output, it's essentially a policy model, predicting what actions a robot should take in the real world to achieve a specific goal.
Such Omni-Models can not only help embodied intelligence researchers and developers better understand environmental modeling and decision-making but also serve as foundational backbone models that can be fine-tuned for specific robotics applications to achieve the reliability and efficiency desired by customers.

Why Emphasize 3D and Simulation?
Host: Currently, many robotics companies adopt a pure video model approach, while you emphasize 3D and simulation. What's the difference between the two approaches?
Yunzhu Li: To create worlds where robots can work, these worlds must capture the essential structure of the problem. One necessary condition is consistency—spatial, temporal, across different viewpoints, and different modes of interaction.
This is where we see strong synergy with Marble. The worlds generated by Marble can provide part of the infrastructure robots need.
Imagine a robot trying to push an object forward, but the object mysteriously disappears. Many existing video prediction models can have such issues. This kind of result doesn't provide the correct signal for the robot to act.
Video models are advancing rapidly, but the infrastructure we're building can provide the initial momentum for a data flywheel, moving from more simulation-driven models towards robot policy models that can be deployed in the real world, then collecting new data and feeding it back.
The model doesn't have to rely solely on physics or solely on learning. It can reside between the two, capturing the essential problem structure while continuously scaling and improving with data accumulation.
Host: Fei-Fei has long had a "North Star" concept around 3D and spatial intelligence. Do you have a similar concept, or are you more pragmatic, focusing on building the system first and making it work?
Yunzhu Li: My North Star is to make robots work in real environments. I'm a very pragmatic person.
During my postdoc with Fei-Fei, we collaborated on building a benchmark and surveyed the public about what they wanted robots to do for them. Among the thousand-plus tasks we collected, one-third were related to cleaning. People dislike these tedious chores, and these are exactly the scenarios we want robots to solve.
Fei-Fei Li: What I truly admire is the extremely pragmatic working methods of the SceniX co-founders.
Despite most of them coming from academia, they chose early on to collaborate with designers and customers in real industries, entering labs, warehouses, electronics assembly, and other real-world settings.
This approach to robotics R&D is refreshing and makes me very excited to collaborate with them.

Host: Robotic systems need to be quite precise, but creative applications can tolerate and even embrace errors, sometimes errors even become muses for creativity. Technically, Yunzhu, how do you reconcile these two requirements?
Yunzhu Li: I think they can be reconciled. The world model doesn't have to be perfect, and the robot model doesn't have to be perfect either.
Host: What do you mean by not perfect? It still needs to be very close to reality, right?
Yunzhu Li: Models have long been an important cornerstone for various robotics applications. Drones, robotic vacuums, quadrupedal and bipedal robots all rely on models to work and to transfer from simulation to the real world.
Locomotion robots can walk on snow or through bushes, but the simulator doesn't need to replicate every bush or patch of snow exactly.
Simulation needs to capture the essential structure of the problem and support different forms of randomization in the digital environment.
We're researching what fidelity is needed to model the world around a robot so that robotic systems trained in simulated or digital environments can transfer back to real-world scenarios.

Is a Simulation-Heavy Approach Feasible?
Host: Some researchers believe that simulation will always diverge from the physical world, making real-world data collection crucial. Is a simulation-heavy approach feasible?
Yunzhu Li: I think the two aren't contradictory. Simulation is essentially about predicting how the environment will change after applying actions—it's modeling the world.
This model doesn't have to be purely physics-based; it can combine physics and learning. We are collecting real-world data and will use it at different stages of the data flywheel.
In the early stages, we might emphasize physics more. We'll let the robot learn about the world first to ensure it has an appropriate learning pace and knowledge structure to facilitate effective learning.
As we accumulate more data through collection and customer collaborations, world modeling can gradually shift towards being more learning-driven. This transition and data flow can combine the advantages of physics, geometry, and consistency with the power of data and compute.

Fei-Fei Li: Yes, it's not a binary choice of "simulation yes" or "simulation no." All elements must combine to make robots truly functional.
Just like humans conduct extensive mental simulation before acting. Simulation plays an important role that real-world data cannot—counterfactual reasoning.
People reason about events that haven't happened, cannot happen, or lack sufficient real data, learning how to act through that reasoning. People always do this, and robots can too.
I believe planning for every soccer match involves some form of simulation, whether digital, on a whiteboard, or otherwise. The role of simulation is to support counterfactual reasoning. This is crucial for robots because we cannot collect enough real-world data for all situations.
Autonomous vehicles are a real-world example. Waymo mentioned using billions of hours of simulation data, and their approach heavily relies on simulation, not just real-world data.
Cars represent a relatively simple robotics problem. Therefore, simulation clearly plays a significant role in robot learning.

Yunzhu Li: Simulation can bring benefits on two levels: reliability and efficiency.
Regarding reliability: For a robotic system to work stably in the real world, data must systematically cover the state space and variations the robot might encounter.
In simulation, we can systematically randomize and control lighting, geometry types, physical parameters, etc., to ensure sufficient coverage of the robot's state space, thereby improving system reliability.
The second benefit is efficiency. Many teams collect data via teleoperation. Using current teleoperation devices and exoskeletons, data collection speed can be even slower than a human directly performing the task.
For many customers, human collection speed isn't enough; they want robots to be faster than humans. But making robots faster isn't simply about increasing actuation speed, because gravity is constant, and simply speeding up doesn't solve the resulting dynamic imbalance issues.
In simulation, we can systematically speed up robot behavior while training the robot to account for changes in environmental dynamics.
So, simulation offers unique value in both reliability and efficiency.

How Does SceniX Use Its Platform to Train and Evaluate Robots?
Host: How does SceniX use its platform to train and evaluate robots?
Yunzhu Li: There are two core uses: training and evaluation.
Let's talk about evaluation first; it's often overlooked in robotics. When training a robot model, we must know how it performs, and evaluation is the source of information that drives iteration.
Host: AI practitioners understand evaluation, but everyone's interpretation might differ. What specifically do you mean by evaluation here?
Yunzhu Li: I think evaluation refers to understanding the performance of a specific checkpoint, for example, whether the success rate is 95% or 99.9%.
A key criterion in the industry is how long it takes to distinguish between a checkpoint performing at 90% and one at 92% using a given metric. If you have to deploy the model in the real physical world each time to verify this 2% improvement, the verification cycle would become unbearably long.
Currently, robot evaluation in the real world is orders of magnitude slower than language model evaluation. Robot tasks are also highly diverse. Robots must physically execute actions, move in space, and obey physical laws.
Robot demonstration videos are often played at 8x or 10x speed precisely because robot actions are very slow.
Real-world evaluation is not only slow but also dangerous and expensive. So, some customers need digital environments for evaluating robotic systems.
Since our digital environments have been shown to align with the real world, if a checkpoint performs better in simulation, it's likely to perform better in the real environment too. We can leverage signals from the digital environment to perform evaluation that is safer, more scalable, and much faster.
Furthermore, for training, controllability is paramount.
We want to control variations in states, parameters, lighting, friction and other physical parameters, even geometry types of objects. Training must sufficiently cover a diverse enough set of scenarios to generate data with enough information, enabling the robot to achieve robustness. This is extremely difficult to achieve in the real world.
Collecting data via teleoperation is not only slow but also constrained by the number of robots, remote operation devices, and a host of data operations issues.
In contrast, simulation environments offer complete controllability and analyzability.
In the digital world, you can explicitly define the coverage of the data distribution and be confident in the robot's reliability within that distribution. This certainty, efficiency, and scalability are the core motivations for customers choosing digital world training for robotic systems.

Fei-Fei Li: Yes. Even before SceniX approached World Labs, customers proactively sought out Marble, proving this demand is real. Although we weren't equipped to serve them at the time, we received numerous inquiries from robotics companies, including early-stage model R&D teams and companies focused on deploying practical applications. We're paying attention to these needs.

Host: When people hear World Labs is entering robotics, they might assume you're building hardware, developing robot brains. But you don't seem to be doing that. Where do you sit in the robot development lifecycle? What are the boundaries with ecosystem partners?
Yunzhu Li: We are building a set of infrastructure and supporting software that allows developers to create worlds where robots can learn and be evaluated. This infrastructure naturally isn't tied to specific models or specific robot morphologies.
Host: So, you're not building robots, but building environments where other companies can place their robot brains for navigation and learning?
Yunzhu Li: Our platform is agnostic to robots and models. Whether single-arm, dual-arm, fixed-base, or mobile-base, even various grippers and complex end effectors, they can all be directly placed in our digital world for training.
The goal is singular: to make them work reliably and quickly in the real environment.
The same goes for models. The data we generate can be used to train new models from scratch or to fine-tune popular VLAs (Vision-Language-Action models) or World Action Models.
The specific model used isn't that important because models will always be upgraded, but infrastructure won't.
Simply put, we don't take sides on hardware or bet on algorithms. We do one thing: build a universal digital world where any robot can train effectively and perform reliably when deployed in the real world.

What's the Right Path for Robotics Deployment?
Host: You mentioned that expectations for humanoid robots might be too aggressive in the short term, and more realistic deployment points are constrained environments like warehouses. How does this judgment affect World Labs' focus?
Yunzhu Li: That's a great question. Real-world robotics deployment typically follows an evolution path: from fully structured environments, to semi-structured environments, and finally to unstructured environments.
Fully structured environments are like factories or automotive assembly lines, where all variables are known and controlled. We've been automating these for decades.
Semi-structured environments are like Amazon warehouses, restaurants, or hotels—environments are largely controlled, and tasks are adapted for robots, but unexpected items like clothing can appear.
And unstructured environments, like your home or mine, that's the ultimate holy grail in robotics.
Robustness stems from sufficient coverage of scenarios. At least at this stage, tackling semi-structured environments first is more feasible than directly charging into fully unstructured ones. We will eventually get there, but we want to take a more sustainable, pragmatic path.
Fei-Fei Li: Humanoid robots adapt to unstructured environments by imitating the human body, and humans have evolved to adapt to unstructured environments.
Our fingers and legs aren't necessarily the optimal tools for any single task. If the sole goal of the human species was to climb trees, we'd likely have differently shaped fingers.
What humans ultimately evolved is a body morphology that is very general-purpose but not necessarily optimal for any single task. Because this morphology aids survival in unstructured environments.
But from a commercial and pragmatic technical standpoint, unstructured environments plus a general-purpose body is the hardest problem to solve.
For specific problems, this might not even be the right path. More specialized body morphologies might be better suited for solving narrower tasks.
Therefore, the challenge for SceniX is to remain agnostic to robot morphology, allowing the infrastructure to serve different body forms and different semi-structured environments.

Host: Generative language models can generate text or code far faster and cheaper than humans. But the human brain and body are highly efficient at 3D navigation and object manipulation. In the foreseeable future, can robots reach human-level energy efficiency in everyday physical tasks? Is that five years away, or nearly impossible?
Yunzhu Li: I think it will take a long time. For robots to work in real environments, it's ultimately a systems problem.
We must seriously consider how hardware, software, and the "brain" integrate, even down to details like finger friction coefficients. To truly deploy the system, many elements need to be in place simultaneously, and continuous iteration is needed.
But what excites me is that the cutting edge of robot learning has been advancing faster than I anticipated. The problems I research now are completely different from when I started my PhD. This shows the entire ecosystem is evolving rapidly, and the various components needed to build robotic systems are starting to converge.
But we must calibrate our predictions. There will be significant progress, but achieving human-level energy efficiency and overall capability will still require a longer time.
Fei-Fei Li: The hardest thing for AI right now is to maintain a measured, calibrated optimism.

Even large language models don't match the human brain's energy efficiency. The human brain operates on roughly 30W, so the gap is still huge.
Host: But I feel that for narrow tasks like image generation or software engineering, energy efficiency might be approaching human levels. It's just that robotics might still be far off.
What Are the Goals Two Years After the Merger?
Host: Does joining World Labs change the team's strategic trajectory, or rather, raise the ceiling of what you can pursue?
Yunzhu Li: It has changed our trajectory in a very profound way. We see many capabilities being unlocked, especially after collaborating with World Labs. We can accomplish the entire environment modeling process in a more efficient and scalable way.
Current language model capabilities are amazing, but you probably still wouldn't let it book a flight or hotel for you without checking, you usually have someone review the model's output.
Robots are different. An out-of-the-box robot model must work reliably in the real environment.
Currently, we neither have enough data nor possess the complete infrastructure for a robot to work reliably out-of-the-box.
Therefore, creating scalable digital worlds where robots can learn and be evaluated will unleash tremendous potential.
We can use data generated in the digital world to replace expensive or unsafe data collection in the real environment, enabling scalable learning and evaluation for robots.
Host: So, will you integrate immediately, or will SceniX remain relatively independent initially, proceeding at a more long-term pace?
Fei-Fei Li: We will proceed thoughtfully; we're not in a rush to merge all codebases and teams immediately. SceniX already has a well-thought-out, relatively complete tech stack, its own customers, and product. We'll give it time.
However, integration will definitely happen. We've already started discussions on simulation, potential foundation models, and action-conditioned models. SceniX is also using Marble as an internal customer.
We'll gradually advance integration, but we won't rush to throw everyone into one "salad bowl." (laughs)

Host: Will the team stay in place?
Fei-Fei Li: Yunzhu will move with some team members to San Francisco. Once settled, World Labs will officially become a company spanning both US coasts. Our headquarters are in San Francisco, and I live in Palo Alto. We will establish an office in New York, which I'm excited about because it helps us attract talent from the East Coast.
We've also discussed deploying robots in both offices to test and refine the engineering stack for remote operation of robots. After all, we'll eventually need to provide this capability to customers.
Host: To be more concrete, what would be the ideal success case in two years? What kind of products will you have, who will be using them, and how?
Fei-Fei Li: In two years, if the SceniX team and World Labs can secure validated customers in a few important vertical use cases and demonstrate that our systems and infrastructure genuinely help meet their automation needs, I would be very satisfied.
These customers would become lighthouse cases, helping us further scale our business.

Full interview:
https://www.youtube.com/watch?v=-tabaM5l3s0
Reference links:
[1]https://marble.worldlabs.ai/
[2]https://www.worldlabs.ai/blog/scenix
This article comes from the WeChat public account "QbitAI", author: Wenting






