What Exactly is a World Model? How Do Fei-Fei Li, Jun Zhu, and Others Interpret It?
The article discusses the conceptual ambiguity surrounding "World Models" in AI and how leading researchers like Fei-Fei Li, Yann LeCun, and Jun Zhu are defining and advancing the field.
It notes that the term "World Model" is loosely applied to diverse systems like video generators, interactive environments, latent space predictors, and robot controllers, creating confusion. To address this, the piece outlines three key interpretive frameworks:
1. **Fei-Fei Li and World Labs** propose a functional taxonomy: Renderers (output pixels), Simulators (output world states), and Planners (output actions).
2. **Yann LeCun** advocates for models that learn predictable world structures in an abstract latent space to aid reasoning and planning.
3. **Jun Zhu and his team** define a "General World Model" from first principles, centered on three core capabilities forming a closed loop: **Understanding** (inferring current state), **Imagination** (predicting possible futures), and **Action** (intervening and learning from feedback).
Zhu's team further proposes a five-level evolutionary roadmap (L1-L5) to measure progress, from generating consistent worlds (L1) to enabling multi-agent collaboration in open environments (L5). Current systems are seen as reaching L3 (acting in the world).
The article uses the example of Shengshu AI's dual-track approach to illustrate this vision: their **Vidu** video models provide broad world knowledge from internet-scale data, while their **Motus/Motus2** embodied AI systems ground this knowledge in physical action and learning. The recently introduced **Motus2** architecture integrates policy generation, action-conditioned simulation, and outcome evaluation into a single self-improving loop, representing a step toward L4 (autonomous agents).
Finally, the piece positions this visual-dynamic path to AGI, exemplified by Shengshu AI, alongside the language-model-centric path, suggesting future convergence. The conclusion argues that true progress in World Models will shift from impressive demos to a system's ability to understand, imagine, act, and crucially, correct itself based on real-world feedback.
marsbit1h ago