# World Model Related Articles

HTX News Center provides the latest articles and in-depth analysis on "World Model", covering market trends, project updates, tech developments, and regulatory policies in the crypto industry.

Just Now, The World's First Ultra-High-Frame World Model Was Born, Nvidia Content 0, Racing to 50 FPS

Just Now, Global First Ultra-High-Frame World Model Born, 0% NVIDIA, Speeds to 50 FPS A Chinese team has developed MoWorld, the world's first Flash World Model, achieving real-time interactive inference exceeding 50 FPS. Crucially, it is entirely built on domestic NPUs (National Processing Units), bypassing NVIDIA GPUs. Developed by Moxin Technology in collaboration with Zhejiang University's Pan Yunhe academician team, MoWorld represents a complete, closed-loop system from training and distillation to deployment on domestic computing power. The model tackles the critical industry bottleneck of real-time performance, essential for applications like robotics, gaming, and digital worlds. MoWorld achieves this through a full-stack redesign for NPUs, including a proprietary 3D-annotated data pipeline, system-level optimizations for long-sequence training (up to 2000 frames), and inference optimizations like dynamic mixed-precision quantization. On a Huawei Ascend 910C platform, a 14B MoE parameter model achieves over 50 FPS, reducing typical inference costs by 70% compared to equivalent GPU solutions. This breakthrough lowers the deployment barrier, potentially accelerating the industrialization of world models. Key application areas include gaming/entertainment (offering 6-DoF camera control for immersive exploration), embodied AI/autonomous driving (providing a high-fidelity digital training ground), film pre-visualization, and 3D reconstruction/digital twins due to its strong geometric consistency. MoWorld demonstrates that a full-stack domestic compute ecosystem can support cutting-edge, real-time world models, positioning China at a competitive starting line in defining next-generation spatial intelligence standards. The project underscores a shift in competition from model scale to real-world usability and cost-effective deployment.

marsbit07/08 04:15

Just Now, The World's First Ultra-High-Frame World Model Was Born, Nvidia Content 0, Racing to 50 FPS

marsbit07/08 04:15

HKEX Welcomes Its Largest IPO of the Day

Today (July 8th), Momenta successfully listed on the Hong Kong Stock Exchange, becoming the "first Physical AI stock." The company, founded in 2016 by Tsinghua University alumnus Cao Xudong, focuses on autonomous driving as an entry point into Physical AI research. Momenta's IPO price was HK$295.6 per share. With a market cap exceeding HK$70 billion post-listing, it was the largest among the five companies debuting that day. The offering raised approximately HK$6.8 billion and attracted a "star-studded" lineup of 14 cornerstone investors, including top-tier international funds, leading strategic industrial investors like Mercedes-Benz and BYD, and major Chinese financial institutions. The company has pioneered a "flywheel" strategy, integrating mass-produced advanced driver-assistance systems (ADAS) with its full-self-driving (L4) development. Data from over 1 million vehicles equipped with its systems fuels its AI models, enabling continuous improvement. This massive real-world data scale is a core competitive advantage. In April, Momenta launched its self-developed R7 World Model for mass production, a foundational model designed to understand and predict physical world dynamics. The company positions itself not just as an automotive tech supplier, but as a platform-level Physical AI company. Its technology platform has the potential to expand beyond autonomous vehicles into areas like logistics and embodied AI. Financially, Momenta's revenue grew from RMB 743 million in 2023 to RMB 2.413 billion in 2025, with licensing income surging 42-fold during this period. While still reporting adjusted losses, it is nearing breakeven. The company boasts partnerships with 24 global automakers, including 9 of the world's top 10, and holds a 65% market share in China's third-party urban NOA segment. The listing marks a significant moment for Physical AI in global capital markets, reflecting strong investor confidence in Momenta's unique technology path and commercial execution.

marsbit07/08 03:10

HKEX Welcomes Its Largest IPO of the Day

marsbit07/08 03:10

StarDynamics Secures 2.5 Billion in Two Months, State-Owned Capital Consortium Joins In

Star Era Raises 25 Billion Yuan in Two Months with State Capital Leading the Charge. Chinese humanoid robotics leader Star Era has secured a new 10-billion-yuan funding round led by state-owned capital, including funds like Chengtong Fund under the SASAC, marking 25 billion yuan raised within two months. The company, a spin-off from Tsinghua University, has built a comprehensive capital matrix combining state guidance, top-tier financial backers, and industrial partners. Founded in 2023 by Dr. Chen Jianyu, one of Tsinghua's youngest doctoral supervisors, Star Era stands out for its early and pioneering work on "world models" for embodied AI, notably releasing its PAD world action model ahead of major global players. The company follows an AI-native, full-stack R&D strategy from data and AI brain to control, dexterous hands (XHAND series), and robot bodies (bipedal L7, wheeled Q5). A core innovation is its fully direct-drive dexterous hands, which act as high-fidelity data collectors for training its AI models like the ERA-42 and VLAW, creating a virtuous cycle of data and intelligence. Star Era claims to possess one of the world's largest real-world dexterous hand datasets. Commercially, Star Era has achieved product-market fit, most notably in logistics, with robots operating 24/7 in distribution centers for partners like SF Express and China Post, handling over 1,200 parcels per hour. It is also expanding into high-end manufacturing (Samsung, Geely) and commercial services. Its hardware components are used by nine of the global top ten tech firms and leading research institutions. The article positions 2026 as an inflection point where success shifts from model capabilities to proven, scalable commercial deployment. Star Era's rapid funding and industrial traction highlight its position in this competitive race.

marsbit07/06 01:35

StarDynamics Secures 2.5 Billion in Two Months, State-Owned Capital Consortium Joins In

marsbit07/06 01:35

Li Fei-Fei's Latest Long-Form Article: When Video Generation, Robotics, and NVIDIA All Call Themselves World Models, We Need a Taxonomy

In a new article, Dr. Fei-Fei Li addresses the widespread and often inconsistent use of the term "world model" in AI. She proposes a clear, functional taxonomy rooted in the classic Partially Observable Markov Decision Process (POMDP) loop (agent → action → state → observation → agent). According to this framework, current systems called "world models" are different projections of this loop, categorized by their primary output: 1. **Renderers**: Output observations (pixels). Their goal is visual fidelity for human consumption (e.g., video generation models like Sora). They are the most commercially mature but are limited by a focus on appearance over physical accuracy. 2. **Simulators**: Output states (geometric, physical, dynamic representations). They provide a structurally accurate world for both human professionals (e.g., architects) and computational agents (e.g., robots for training). Li argues simulators are the crucial, underappreciated bridge, as they can underpin both rendering and planning. 3. **Planners**: Output actions. Given an observation and a goal, they decide what an agent should do next (e.g., robotic action models). This area is highly promising but remains the least mature for real-world deployment. Li highlights a key trend: the boundaries between these three categories are beginning to blur, as they all rely on a shared underlying understanding of geometry, physics, and dynamics. The logical endpoint is a unified world foundation model capable of switching between rendering, simulation, and planning based on downstream needs. This convergence, she concludes, is central to advancing spatial intelligence—enabling machines not just to talk about the world, but to truly understand, imagine, and interact with it.

marsbit07/05 09:24

Li Fei-Fei's Latest Long-Form Article: When Video Generation, Robotics, and NVIDIA All Call Themselves World Models, We Need a Taxonomy

marsbit07/05 09:24

Introduction to the Concept of World Models: A Story from Psychology to the Main Battlefield of AI

**World Models: From Psychology to AI's Core Concept** "World model" is a trending but often confusing term in AI, describing a system that allows machines to internally simulate, predict, and rehearse potential outcomes before taking real-world action—like a mental "sandbox." While definitions vary—Yann LeCun emphasizes physical understanding, OpenAI's Sora is a video-based "world simulator," Google DeepMind's Genie 3 creates interactive 3D environments, and companies like Alibaba and Tesla focus on practical applications—the core goal is consistent: reduce reliance on vast real-world data by creating an internal, predictive model for safer and more efficient AI. The concept has deep roots, tracing back to psychologist Kenneth Craik (1943). In AI, it was revitalized by researchers like David Ha and Jürgen Schmidhuber (2018). Major technical approaches include: 1) generative video models (e.g., Sora) for visual realism; 2) abstract predictive models (e.g., LeCun's JEPA) for efficiency and physical reasoning; and 3) explicit 3D simulators (e.g., NVIDIA Omniverse) for precision. Fei-Fei Li proposes a classification based on the AI action loop: renderers (output observations), simulators (output world states), and planners (output actions). The emerging "World Action Model" (WAM) paradigm aims to unify future prediction and action generation. An industry framework is forming: upstream (data, compute, sensors), midstream (general and vertical platforms), and downstream applications (autonomous driving, robotics, gaming, etc.). Autonomous driving is currently the most mature use case. The current lack of a unified definition reflects the field's early, dynamic stage, similar to past tech revolutions. Different approaches—focusing on pixels, physics, or behavior—represent parallel explorations of how best to compress and understand the world. This diversity, while seemingly chaotic, signals that world models have moved from an academic idea to a critical industrial battleground, ultimately aiming to give machines the ability to understand, imagine, and reason about the world.

marsbit06/29 05:09

Introduction to the Concept of World Models: A Story from Psychology to the Main Battlefield of AI

marsbit06/29 05:09

Domestic First Explosion-Proof Certification, World's First Fueling Brain Solution: How Did They Secure Two 'Firsts'?

China's embodied AI sector is booming, with over ¥37 billion in funding this year. The focus has shifted decisively to real-world application, particularly in hazardous, repetitive tasks humans should avoid. A key, often prohibitive, barrier to entry for robots in environments like gas stations and oil fields is obtaining explosion-proof certification, requiring meticulous hardware and circuit design from the ground up. The article explores three main application areas. At gas stations, the challenge lies in executing a long, precise sequence of actions (opening caps, handling the fuel nozzle) with millimeter accuracy across diverse car models. For facility inspections, robots need sustained autonomous patrols combined with real-time anomaly detection and response. Port scenarios introduce the complexity of multi-robot coordination. Addressing the core challenge of long-horizon tasks, the piece highlights a technical breakthrough: a "world model"-driven approach. This enables predictive planning, allowing the AI to visualize the desired end-state (e.g., nozzle returned, cap closed) and work backward to synthesize intermediate visual frames. This "imagination" of the task trajectory, as implemented in the H-GAR architecture, guides action generation, significantly reducing cumulative error in multi-step operations. The three-step H-GAR process involves generating a coarse action draft, synthesizing target-conditioned observation frames, and then refining actions based on visual context and a memory of past successful motions. The conclusion emphasizes that success in specialized, safety-critical fields requires long-term commitment and deep integration of the "embodied brain" (AI) with a purpose-built, certified physical "body." Mastering this brain-body-data闭环 (closed-loop) is positioned as a crucial competitive advantage for commercialization.

marsbit06/26 03:49

Domestic First Explosion-Proof Certification, World's First Fueling Brain Solution: How Did They Secure Two 'Firsts'?

marsbit06/26 03:49

The War Without a Unified Name: The Domestic Tech Giants' World Model Landscape

The article outlines the diverse and fragmented landscape of "World Models" in China's tech industry, where major players are pursuing similar goals under different names like world foundational models, physical AI, or integrated within autonomous driving and embodied intelligence systems. The core aim is to enable AI to create an internal, dynamic environment for simulation, reasoning, and learning, reducing reliance on infinite real-world data. This "data engine" allows for unlimited generation, experimentation, and iteration. The report categorizes the approaches of different companies: * **Internet Giants:** Alibaba is developing models for linguistic, virtual, and physical worlds (Qwen-AgentWorld, HappyOyster, Qwen-RobotWorld). Tencent's HY-World focuses on 3D, game, and social scenarios. ByteDance leverages its vast video data for a potential "digital twin" model. Huawei integrates its model into industrial applications like smart cars and robotics without separately branding it. Baidu embeds world model capabilities within its Apollo autonomous driving and Ernie systems. * **Automakers:** Companies like NIO, Li Auto, XPeng, and Geely are using world models as virtual "driving schools" and "testing grounds." They generate complex scenarios (e.g., rain, snow) to train and validate autonomous driving systems in simulation, aiming for more capable and safer AI drivers. * **Autonomous Driving Suppliers:** Firms such as Momenta, Horizon Robotics, Haomo.ai, and DeepRoute.ai are building the underlying "world engines." They focus on large-scale video generation for simulation, reinforcement learning, and enhancing end-to-end autonomous driving models, often integrating these capabilities into commercial products. While startups bring focus and innovation, they face challenges like limited data, compute resources, and deployment channels. Large companies possess these advantages and are rapidly transitioning world models from research projects into core business infrastructure powering products in vehicles, games, and industry. The conclusion is that world models represent an evolution and convergence of existing AI fields into crucial industrial infrastructure, moving the competition from simply building a model to effectively deploying it to understand and interact with the physical world.

marsbit06/25 06:52

The War Without a Unified Name: The Domestic Tech Giants' World Model Landscape

marsbit06/25 06:52

GPT-5.6 Countdown: Abandon the Illusion of a Single API, Computational Iteration Can't Outpace a Single Page of Compliance

In mid-June, three seemingly independent industry events—the compliance-driven throttling of Fable 5, the open-sourcing of GLM-5.2, and the leaked release timeline for GPT-5.6—are pushing the global AI industry toward a watershed moment. These shifts signal a fundamental restructuring of the industry's underlying logic. First, **"usability" has substantially overtaken "advanced capabilities"** as the primary weight, pushing the global large language model (LLM) supply chain into a "dual-track" phase of controlled closed-source and local open-source coexistence. Second, **the competitive moats of closed-source giants are shifting**. Their technical focus is moving from "language intelligence" toward "spatial intelligence (world models)"—a domain heavily reliant on computing power. Third, faced with常态化 transnational compliance risks, **a "model-agnostic" decoupled design has become a survival necessity for application-layer developers to maintain business continuity.** The article details how Anthropic's Fable 5, despite its advanced engineering feats, was restricted for non-U.S. citizens within 72 hours of launch, highlighting how geopolitical compliance can instantly limit even the most advanced models. In response, the open-source camp, exemplified by Zhipu AI's MIT-licensed GLM-5.2, is gaining market share by offering stable performance improvements and significant cost advantages (up to 70% savings for enterprises), while achieving full adaptation with domestic semiconductor platforms. Meanwhile, closed-source leaders like OpenAI are pivoting. The anticipated GPT-5.6 reportedly shifts focus from language to spatial intelligence and world models, aiming to rebuild a generational gap in areas like 3D understanding, simulation, and industrial design that demand immense compute. The core conclusion is that the LLM supply chain's logic has changed. Enterprises must now evaluate infrastructure based on a composite of technical performance and policy compliance. For developers, complete reliance on a single closed-source API poses unacceptable risk. Implementing a truly model-agnostic architecture—enabling swift switches to compliant, locally deployable open-source alternatives—is no longer just good practice but a fundamental baseline for business continuity.

marsbit06/21 04:40

GPT-5.6 Countdown: Abandon the Illusion of a Single API, Computational Iteration Can't Outpace a Single Page of Compliance

marsbit06/21 04:40

He Just Raised 2.7 Billion, and Li Fei-Fei Also Invested

Pete Florence, a former senior research scientist at Google DeepMind and a key contributor to the Vision-Language-Action (VLA) model architecture, is deliberately distancing his startup, Generalist AI, from the trendy "world model" label. He argues that the industry should prioritize concrete goals over buzzwords. His goal is to create robots that can perform a vast range of unseen tasks with high speed and success rates, without needing task-specific training data. Recently, his company raised $400 million (¥2.7 billion) at a $2 billion valuation. Notable investors include NVIDIA's NVentures, Bezos Expeditions, NFDG, as well as Xiaomi co-founder Lin Bin, Zoom founder Eric Yuan, and renowned AI scientist Fei-Fei Li. Florence's approach stems from his academic background at MIT under Professor Russ Tedrake, focusing on understanding the physical world. After joining DeepMind, he developed models like Transporter Network and co-created the VLA framework. He left in 2025 to found Generalist AI. The company has launched two models: GEN-0, which demonstrated that scaling laws apply to physical motion, and GEN-1. GEN-1 was trained on over 500,000 hours of physical interaction data collected via a specialized wearable device. It achieves a 99% success rate on precise mechanical tasks like folding boxes and maintains performance three times faster than its predecessor. Florence believes GEN-1 is reaching a commercial utility threshold similar to the GPT-3 inflection point. The substantial funding round, following GEN-1's release, signifies strong investor confidence in Generalist AI's practical, goal-driven path to creating versatile, useful robots, regardless of the "world model" terminology.

marsbit06/20 06:06

He Just Raised 2.7 Billion, and Li Fei-Fei Also Invested

marsbit06/20 06:06

活动图片