Fei-Fei Li's Manifesto for World Models

marsbitОпубликовано 2026-06-09Обновлено 2026-06-09

Введение

"Feifei Li's World Model Manifesto" draws a crucial distinction between current AI's linguistic prowess and its lack of understanding of the physical world. Citing Wittgenstein, Li argues that true intelligence requires moving beyond text statistics to comprehend physical laws like optics, inertia, and collision. The article diagnoses the current confusion around "world models" and proposes a clear taxonomy based on the Partially Observable Markov Decision Process (POMDP) framework. Li identifies three core, interdependent pillars for building such models: 1) The **Renderer**, which masters visual plausibility and pixel generation (e.g., Sora, image models) but lacks structural integrity. 2) The **Simulator**, which prioritizes strict adherence to physical laws (mass, friction, collision) and is essential for robotics and real-world application, though it is computationally demanding and data-hungry. 3) The **Planner**, which connects perception to action, enabling decision-making in complex, unstructured environments. Li posits the **Simulator as the critical nexus** linking rendering and planning, highlighting NVIDIA's Omniverse as a leading example. Mastering physical simulation is key to industrial AI applications. Despite challenges like scarce annotated 3D data and "physics-unrealistic" generative outputs, a convergent trend is emerging. The future lies in a **unified foundational model** that seamlessly integrates rendering, simulation, and planning into a dynamic, i...

"The world is everything that is the case."

In 1921, Ludwig Wittgenstein wrote this famous sentence in *Tractatus Logico-Philosophicus*. A century later, it is quoted by AI pioneer Fei-Fei Li as the opening of her latest technical blog post.

In the landscape of deep learning, people have become accustomed over the past three years to AI's disruptive impact on language, starting with ChatGPT which endowed machines with expression, programming, and reasoning abilities far surpassing humans.

However, behind this digital miracle lies a blind spot that is often overlooked: machines can talk about the world, yet remain ignorant of its physical essence. The blog post released by Fei-Fei Li serves as a sobering reality check.

Today, as generative AI has become an indispensable tool globally, the industry's internal definition of "world models" is becoming increasingly chaotic. Whether in video generation or embodied intelligence, various companies are vying for the interpretive authority of this concept.

After Fei-Fei Li published this blog post, many believed she was attempting to reclaim the definition of "world models." But on the contrary, I think what Fei-Fei Li truly aims to do is to issue a declaration: The world is not constituted by language, but by the rigorous laws of physical space and time.

For machines to truly step into the human physical world, they must break free from the comfort zone of text statistics and instead understand the refraction of light, the inertia of objects, and the logic of collisions. This is not only a paradigm shift in technology but also a necessary path for AI's advancement toward embodied intelligence.

01 We Need a Taxonomy

It must be admitted that in the AI lexicon, "world model" has devolved into a catch-all pronoun; any project involving image generation or environment simulation seems capable of being linked to it. This ambiguity stems precisely from the multi-dimensional human need to define the "world."

When a technology is just starting out, there naturally won't be unified doctrines to confine it within clear boundaries. This chaos in defining "world models" is not uncommon in history. When ancient Greek philosophers debated whether the essence of the world was water, fire, or indivisible atoms, they were essentially searching for a cornerstone for their reasoning.

The AI field now faces a similar problem: When a video generation model produces visuals that are extremely realistic yet physically impossible, how should we define it? Fei-Fei Li's blog mentions an ancient and robust foundational definition: the Partially Observable Markov Decision Process (POMDP).

This is also the core axiom of reinforcement learning mechanisms, revealing the eternal closed loop of interaction between an agent and the physical world: The agent takes an Action, leading to a change in the world's State. However, the agent lacks a god's-eye view and can only construct a partial perception of reality through Observation.

Essentially, a world model is the abstract model of the world that a machine builds in its "brain" to survive within this closed loop. If any part of this loop is not clearly defined, then the so-called world model remains merely a blind stacking of pixels.

02 The Three Pillars of Building Intelligence

This loop sounds simple, with each component's function easily understood. However, upon careful analysis, each contains countless details with blurred definitions. To explain the chaos within, Fei-Fei Li deconstructs world models into three core components. They serve both as a technical taxonomy and as the three pillars for AI's journey toward embodied intelligence.

1. Renderer

The core logic of the renderer is visual plausibility. Its output is pixels, striving to make the imagery appear natural, coherent, and aesthetically pleasing to the human eye.

This is currently the most mature field commercially. Models we are familiar with, such as OpenAI's Sora and ByteDance's Seedance 2.0 for video generation, and OpenAI's GPT-image-2 and Google's Nano Banana 2 for image generation, are essentially the most sophisticated visual probability machines available. By learning from billions of internet images and videos, they have ultimately mastered the distribution patterns of light, shadow, and form.

This seemingly beautiful reality comes at a cost, as Fei-Fei Li points out. While these top models can generate magnificent architecture, attempting to interact within their generated physical structures would likely cause the building to collapse instantly due to a lack of support structure. In other words, they don't understand what "support" is; they generate only what the viewer "sees," not what the world "is."

2. Simulator

What the simulator pursues is precisely the structural fidelity that the renderer lacks. It doesn't care at all whether a video looks good; its sole concern is whether the world follows physical laws. When a simulator outputs a mundane cup, it must include the cup's mass distribution, material friction coefficient, gravity response, and physical boundaries during collisions.

With a simulator, the content in videos gains a claim to realism. However, simulators are not only severely underestimated but often outright ignored in the current AI wave.

From the case of the cup above, the existence of a simulator transforms "discussing art" into "studying physics." Constructing a simulator that strictly adheres to physical laws requires unimaginable computational resources and annotation costs. But for robots, visual aesthetics are almost a useless attribute; physical precision determines everything.

If a simulator isn't accurate enough, robots trained within it can never enter the real world. The Sim-to-Real challenge is objectively real. Test actions that pass 100% in the lab can be completely paralyzed by minute friction in the real world—this is what we often call the "Moravec's paradox."

3. Planner

The planner is responsible for action output. As the connection point between perception and feedback, it needs to solve the core question with no standard answer: "What should be done next?" In Fei-Fei Li's framework, this is also the final component of the entire "perception-action" closed loop and simultaneously the most frontier-challenging domain.

All current Vision-Language-Action (VLA) models are attempting to enable systems to make decisions in unstructured, complex worlds. The planner doesn't merely predict the future; it chooses, from countless possibilities, the path most likely to achieve the goal. It is the key for machines to evolve from "observers" into "practitioners."

03 The Hundred-Billion-Dollar Hub

Among the three categories Fei-Fei Li outlines, models corresponding to the renderer and planner are relatively common; the remaining simulator has logically become the most difficult component to realize. Fei-Fei Li also offers an insightful judgment: The simulator is the link connecting rendering and planning, and the core hub of the entire system.

The company performing most excellently in the field of simulators is not OpenAI, Anthropic, or Google, but Jensen Huang's NVIDIA.

NVIDIA's Omniverse claims to support trillion-dollar digital twin dreams precisely because it grasps the essence of the simulator. On NVIDIA's platform, the operations of factories, supply chains, and warehouses have all become complete digital mirrors. For the industrial world, this is no longer a visual demo but a core infrastructure for productivity.

This is not an exaggeration but a trillion-dollar market opportunity visible to all.

From virtual visualization in architectural engineering to molecular dynamics simulations in the pharmaceutical industry, and scenario testing for autonomous driving. What these industries lack is not vivid image or video generation models, but a high-fidelity simulator. It's no exaggeration to say that mastering the ability to simulate the physical world equates to holding a priority ticket for AI industrialization.

But the difficulties in reality leave this field with almost no technological optimists. Fei-Fei Li also admits that a huge gap persists.

First is the issue of embodied intelligence data, which we have repeatedly mentioned before. Video data on the internet is abundant, but 3D data with explicit geometric structure, material properties, and physical feedback annotations is extremely scarce.

Second, the application of generative AI will always be accompanied by hidden risks. AI-generated geometric models can at best achieve visual perfection but are often physically unreasonable—like cups intersecting with tabletops, or objects colliding and losing volume. In human terms, the brief phrase "clipping through" can summarize these bizarre phenomena, but in real industrial applications, this spells disaster.

04 Toward a Unified World Model

Despite the immense difficulties, Fei-Fei Li offers a positive prediction of industry trends: The boundaries between rendering, simulation, and planning are becoming increasingly blurred.

This is not a distant vision but a reality already unfolding. After exploration, Fei-Fei Li's World Labs team believes humanity is already moving towards a unified foundation model. In this architecture, imagination and logic can merge into one.

The models of the future will no longer be a patchwork of single-function add-ons, but a unified neural network foundation. It can simultaneously render realistic scenes via Gaussian splatting and generate the collision meshes required by physics engines in real time. Simply put, a unified foundation model will achieve seamless switching between the visual patterns humans need and the state patterns physics engines require.

From another perspective, traditional models are static, while future world models will possess stronger interactivity. Renderers will no longer be passive video generators but will gradually begin to accept action instructions; simulators will become more editable and controllable; planners will also be capable of logical reasoning, automatically adjusting strategies based on environmental changes.

05 The Long Arc of Spatial Intelligence

Finally, returning to the macro level, why is all this about "world models" important?

In Fei-Fei Li's view, decades of AI research have been searching for that key to allow machines to enter the physical world. Today, we already possess language models adept at handling logic; what we need next are models that handle space. The core of spatial intelligence lies in how machines interact with the physical world they inhabit.

This battle is not about who possesses more computing power, but about who can define the digital standard for the physical world.

World models are by no means a simple algorithmic optimization, but a grand feat of AI evolution.

"Language gives machines the ability to talk about this world, while world models are the way machines ultimately understand, imagine, reason, and interact with the physical world."

Every person in this era is transitioning from the stage of talking about the world toward a new epoch of truly understanding and reconstructing it.

Nonetheless, world models are merely an intermediate node on the path to AGI, and the AI created by humans still has a long way to go before reaching a truly meaningful "world model." Here, the somewhat extreme view of another world model luminary, Yann LeCun, is worth sharing:

Optimistically, it will take at least another five to ten years for machine intelligence to barely approach that of a puppy.

This article is from the WeChat public account "Silicon-Based Spark," author: Siqi

Связанные с этим вопросы

QWhat is the core problem with current AI models highlighted by Li Fei-Fei in the context of 'world models'?

ACurrent AI models, particularly generative AI, are proficient at processing and generating language but have a fundamental blind spot: they can talk about the world but lack an understanding of its physical essence—the laws of physics, space, and time. They operate in a 'text statistics comfort zone' without grasping concepts like light refraction, object inertia, or collision logic.

QAccording to Li Fei-Fei's framework, what are the three core components (or pillars) of a world model?

ALi Fei-Fei's framework identifies three core components: 1. The Renderer, which focuses on visual plausibility and aesthetic output (pixels). 2. The Simulator, which prioritizes structural fidelity and adherence to physical laws. 3. The Planner, which is responsible for action output and decision-making, connecting perception to action.

QWhy is the Simulator considered the crucial 'hub' in Li Fei-Fei's analysis of world models?

AThe Simulator is the crucial hub because it connects rendering (visual plausibility) with planning (action). It provides the essential understanding of physical laws that allows models to generate content that is not just visually appealing but also structurally sound and interactive. This makes it foundational for applications in embodied AI, robotics, and industrial digital twins, representing a massive market opportunity.

QWhat major challenges does the development of effective world models (particularly simulators) currently face?

AKey challenges include: 1. A severe scarcity of high-quality, well-annotated 3D data that includes geometric structures, material properties, and physical feedback, unlike the abundance of internet video data. 2. The risk of 'physics-unrealistic' outputs from generative AI (e.g., object interpenetration or 'clipping'), which are catastrophic for industrial applications. 3. The immense computational resources and labeling costs required to build high-fidelity simulators.

QWhat is the predicted future trend for world models as mentioned in the article?

AThe trend is toward a unified foundational model where the boundaries between rendering, simulation, and planning become blurred. This model would be a single, interactive neural network capable of seamlessly switching between generating visually realistic scenes and producing the physical state representations needed for simulation and planning, thereby combining imagination with logic.

Похожее

When LPs Teach Me Investment with Doubao: A Self-Narrative of a Private Equity GP Switching Careers

When LPs Use Doubao to Teach Investing: A Transition Story of a Private Equity GP AI is making life increasingly difficult for small private equity fund managers, as a former GP of an offshore dollar fund reveals. The fund, managing tens of millions in US stocks, outperformed the Nasdaq but struggled with fundraising. Its traditional Cayman SPC/BVI structure failed to attract major Asian LPs, who now prefer Hong Kong LPF or Singapore VCC frameworks. The rise of AI-powered quantitative strategies has further squeezed the space for funds like his, which relied on subjective, discretionary investing. AI tools have leveled the information playing field, empowering LPs—often high-net-worth individuals, entrepreneurs, or family offices—to analyze investments themselves using chatbots like Doubao. This has eroded trust in GPs' expertise, leading to more frequent challenges over investment decisions and even withdrawals, especially during market rallies when retail investors sometimes outperform funds. Friction arises not necessarily from AI's capabilities but from how LPs use it. Many rely on conversational AI for validation rather than rigorous analysis, sometimes receiving misleading or hallucinated advice. While AI democratizes research, effective investing still requires discerning real insight from plausible-sounding output. Ultimately, AI is unlikely to fully replace GPs. Asset management remains a trust-based service. However, the industry must adapt. The future may see "human私募" (private equity) learning from AI and focusing more on providing value beyond pure analysis—perhaps by mastering the emotional intelligence and trust-building that machines cannot replicate.

Odaily星球日报6 мин. назад

When LPs Teach Me Investment with Doubao: A Self-Narrative of a Private Equity GP Switching Careers

Odaily星球日报6 мин. назад

AI Leads to Layoffs? Research Shows AI Is More Expensive Than the People It Replaces

Title: AI Layoffs? Research Shows AI is More Expensive Than the Workers It Replaces. This year, nearly 50,000 employees have been laid off due to AI, with companies initially believing AI could replace human jobs. However, recent findings indicate that the actual operational costs of AI often exceed the expense of the human labor it was meant to replace. Examples include Uber exhausting its annual AI budget in just four months, Microsoft cutting Claude Code licenses due to high costs, and an Anthropic employee incurring $150,000 in API usage in a single month. A CloudZero survey reveals that 45% of enterprises spend over $100,000 monthly on AI, yet only 8% of S&P 500 companies report any AI-related revenue, and half struggle to measure ROI. Analyst Scott Galloway predicts a shift toward cheaper Chinese AI models, which are 10 to 30 times more affordable than American counterparts. Data shows Chinese models' usage among developers surged from 1% in 2024 to over 60% by mid-2026, with 80% of U.S. AI startups adopting them. This trend may prompt regulatory responses, such as potential restrictions from the Trump administration.

marsbit20 мин. назад

AI Leads to Layoffs? Research Shows AI Is More Expensive Than the People It Replaces

marsbit20 мин. назад

Wang Chuan: After Investing in Storage Stocks and Seeing a Thirty-Fold Return, How to Remain Unanxious (Part 7) - A Quarter-Century Cycle

Wang Chuan: Reflections on Investment Anxiety and Market Cycles After Observing a 30x Gain in a Storage Stock (Part 7) – A Quarter-Century Cycle This article examines the cyclical nature and inherent risks in technology hardware investments, using the storage and semiconductor sectors as examples. It criticizes the misleading practice of "annualized" Net Dollar Retention (NDR) rates, where short-term growth is extrapolated unrealistically. A key concept explored is "reflexivity" – demand driven by panic, exploration, and liquidity during market booms, which can vanish just as quickly when conditions reverse. This reflexivity exists both in product demand and among speculative stock buyers, creating powerful feedback loops that inflate prices during upturns and exacerbate crashes during downturns. The author highlights a major risk for hardware sectors: unlike assets with defined cycles (e.g., Bitcoin's halving), there's no guarantee of a swift recovery post-crash. Companies like Micron, Intel, and Cisco took roughly a quarter-century to surpass their 2000 highs, enduring drawdowns exceeding 80%. This is attributed to the "bullwhip effect" in supply chains, where demand collapses instantly but过剩产能 persists, and a migration of narrative-driven capital. High-valuation stories吸引 speculative funds during growth phases, but these funds quickly depart for the next hot narrative once growth slows, leaving behind stronger companies with much lower valuations. The piece warns of dangerous mental models formed during bull markets: 1) equating current strong demand with perpetual high growth, and 2) believing that making fast, large profits is easy. Citing巴菲特, the author notes that easy money undermines rationality, likening speculators to Cinderella at a ball with a clock that has no hands. The current phase presents an asymmetric risk-reward scenario: potential for further gains exists, but the downside risk is an 80%+ drawdown and a multi-decade wait for breakeven, which reflexive speculators cannot tolerate. The hypothetical investor "老王" (Lao Wang), who achieved a 30x return, is used to illustrate potential pitfalls. Leverage could lead to a wipeout during a sharp correction. Even without leverage, ingrained beliefs in easy money would likely lead him to double down after losses, expecting a quick rebound. Instead, he might face a protracted decline, depleting his resources through frantic trading as the high-growth narrative fades. The conclusion references Schopenhauer, comparing those who have seen multiple market cycles to an audience seeing the same magic trick repeatedly—once the illusion is understood, its power is gone.

marsbit28 мин. назад

Wang Chuan: After Investing in Storage Stocks and Seeing a Thirty-Fold Return, How to Remain Unanxious (Part 7) - A Quarter-Century Cycle

marsbit28 мин. назад

US Stocks Too Expensive? This Top CIO Scoured the Globe and Found 5 Stocks More Attractive Than NVIDIA

Summary: Main Street Research CIO James Demmert maintains his bullish 8,100 target for the S&P 500 but argues that greater opportunities now lie overseas. He identifies five international stocks with superior valuations poised to benefit from the AI revolution, suggesting international markets will outperform the US for years. Key Recommendations: 1. **ASML (Netherlands):** A foundational chip manufacturing technology provider, offering crucial AI exposure and geographic diversification. Demmert's top long-term pick. 2. **HSBC (UK/Asia):** A global bank with a 9x P/E ratio, better growth prospects than US peers like JPMorgan, and strong Asian presence. 3. **Siemens Energy (Germany):** A direct play on global power grid expansion driven by AI, crypto, and EV electricity demand. 4. **BHP Group (Australia):** A "hidden AI play" and "second derivative" of the trend due to massive copper demand for data centers. Trades at a 16x P/E. 5. **AstraZeneca (UK):** An undervalued healthcare stock with a strong pipeline (18x P/E, >20% growth), expected to benefit from AI's impact on medicine. Core Thesis: International outperformance is driven by both attractive valuations and a major policy shift. While the US tightens fiscal policy, Europe and Japan are launching unprecedented stimulus, reigniting growth. Demmert recommends allocating 45% of a portfolio internationally, citing excessive US investor conservatism as a key mistake.

marsbit33 мин. назад

US Stocks Too Expensive? This Top CIO Scoured the Globe and Found 5 Stocks More Attractive Than NVIDIA

marsbit33 мин. назад

a16z Partner: Three Paths for Crypto Projects to Find PMF

Author: Jason Rosenthal. Compiler: Shenchao TechFlow. Finding Product-Market Fit (PMF) is the most critical variable for a company's survival. In the crypto space, misaligned growth hacking and airdrops often mask the absence of true PMF. However, leading teams are now finding PMF faster. Here are three proven paths for crypto projects to achieve PMF: 1. **Co-build with Anchor Clients:** Partner with the most sophisticated potential clients in your field and develop the product based on their specific needs. Their adoption serves as the strongest validation, more valuable than media coverage or TVL metrics. This approach is shaping current product roadmaps, as seen in collaborations between crypto startups and traditional finance. 2. **Position Ahead of an Exponential Curve:** Identify and position yourself ahead of a major emerging trend before the market fully realizes it. The most evident current curve is the rise of AI Agents as autonomous economic actors. Projects like AgentCash by Merit Systems, which enables AI Agents to pay for API access with crypto, are building foundational payment rails for the impending Agent economy. 3. **Be Your Own First and Best Customer:** The most enduring infrastructure companies don't wait for external validation. They first build and prove their technology by using it to power their own applications at scale before offering it to others. Matter Labs exemplifies this by anchoring its ZKsync technology in a concrete application, Cari Network, which enables U.S. regional banks to conduct real-time, on-chain interbank transfers of tokenized deposits. The underlying logic is consistent: the fastest path to PMF involves choosing the right battlefield and executing with conviction—by co-building with clients whose validation compounds, positioning ahead of the curve before consensus forms, or becoming your own best case study.

marsbit34 мин. назад

a16z Partner: Three Paths for Crypto Projects to Find PMF

marsbit34 мин. назад

Торговля

Спот

Фьючерсы

Обсуждения

Добро пожаловать в Сообщество HTX. Здесь вы сможете быть в курсе последних новостей о развитии платформы и получить доступ к профессиональной аналитической информации о рынке. Мнения пользователей о цене на S (S) представлены ниже.

Fei-Fei Li's Manifesto for World Models

Введение

01

We Need a Taxonomy

02

The Three Pillars of Building Intelligence

03

The Hundred-Billion-Dollar Hub

04

Toward a Unified World Model

05

The Long Arc of Spatial Intelligence

Связанные с этим вопросы

Похожее

When LPs Teach Me Investment with Doubao: A Self-Narrative of a Private Equity GP Switching Careers

AI Leads to Layoffs? Research Shows AI Is More Expensive Than the People It Replaces

Wang Chuan: After Investing in Storage Stocks and Seeing a Thirty-Fold Return, How to Remain Unanxious (Part 7) - A Quarter-Century Cycle

US Stocks Too Expensive? This Top CIO Scoured the Globe and Found 5 Stocks More Attractive Than NVIDIA

a16z Partner: Three Paths for Crypto Projects to Find PMF

Торговля

Популярные статьи

Как купить S

Sonic: Обновления под руководством Андре Кронье – новая звезда Layer-1 на фоне спада рынка

HTX Learn: Пройдите обучение по "Sonic" и разделите 1000 USDT

Обсуждения

Топ вопросы

Популярные категории

Популярные теги