# World Model Related Articles

HTX News Center provides the latest articles and in-depth analysis on "World Model", covering market trends, project updates, tech developments, and regulatory policies in the crypto industry.

10000 Hours of Human Data, Trains the World's First Whole-Body Mobile Manipulation Implicit World-Action Model

"Being-M0.7" is the world's first latent world-action model for whole-body mobile manipulation in humanoid robots, developed by Zhi Zai Wu Jie (Beyond Being). Trained on over 10,000 hours of human-centric multimodal data, it aims to overcome key industry challenges: the high cost and scarcity of real robot demonstration data, the computational inefficiency of pixel-level video prediction models, and the lack of full-body coordination in existing approaches. The model is based on a Vision-Motion Mixture-of-Transformers (MoT) architecture, which allows training on a mixture of paired video-motion data, pure video data, and pure motion sequences. A key design is a unified motion representation that bridges human and robot morphology, enabling knowledge transfer from vast human behavioral data to specific robot control. After pre-training on human data, the model is adapted to a real robot (Unitree G1) using a small amount of teleoperated demonstration data via a lightweight "Action Expert" module. This process decouples low-frequency world planning from high-frequency motion control. The model was tested in four challenging real-world demos: fishing a toy fish from water (liquid interaction), retrieving an object using a mirror (visual reasoning), a multi-step pick-and-place task, and obstacle avoidance while carrying a box. In comparative tests against other models, Being-M0.7 showed stronger performance in tasks requiring indirect reasoning and full-body coordination. This work represents a shift in humanoid robotics competition from hardware spectacle to model capabilities rooted in scalable data and training paradigms, using human experience as a foundation for physical world understanding and action.

marsbit07/15 01:48

10000 Hours of Human Data, Trains the World's First Whole-Body Mobile Manipulation Implicit World-Action Model

marsbit07/15 01:48

Is the iPhone Moment for Embodied AI Coming Soon?

Is the "iPhone moment" for embodied AI approaching? This article, based on a roundtable discussion, presents expert insights on the current state and future of embodied AI. The consensus is that the pivotal "iPhone moment" is still distant. The field is likened to the "brick phone" era, with technology paths—such as VLA and world models—not yet converging. While robotic "motor skills" (e.g., walking) have matured, the "brain" (decision-making, generalization) remains far from commercial readiness. A major bottleneck is data: an estimated tens of millions of data points are needed for a breakthrough, but only around 500,000 currently exist globally. Currently, cost remains prohibitive for widespread labor replacement, making the economic case challenging. However, experts see a three-tiered market potential: a billion-level market for emotional companionship (e.g., entertainment, basic care), a trillion-level market for commercial services (e.g., guides, receptionists), and a massive, long-term opportunity for physical labor in factories and homes. The discussion suggests that while humanoid robots face hurdles, non-humanoid embodied AI applications (like existing service robots) can be deployed sooner. The ultimate vision is for AI to operate seamlessly in the physical world, not just behind screens. Regarding AI tools, participants noted their widespread use for boosting efficiency in coding, research, and teaching. However, they warned against over-reliance due to risks of AI "deception" and the erosion of critical thinking, emphasizing that core judgment must remain with humans. In summary, embodied AI holds immense promise but requires significant progress in brain models, data collection, and cost reduction before achieving its transformative potential. Its development is expected to be gradual, advancing through specific use cases rather than a single explosive moment.

marsbit07/14 05:07

Is the iPhone Moment for Embodied AI Coming Soon?

marsbit07/14 05:07

Just Now, The World's First Ultra-High-Frame World Model Was Born, Nvidia Content 0, Racing to 50 FPS

Just Now, Global First Ultra-High-Frame World Model Born, 0% NVIDIA, Speeds to 50 FPS A Chinese team has developed MoWorld, the world's first Flash World Model, achieving real-time interactive inference exceeding 50 FPS. Crucially, it is entirely built on domestic NPUs (National Processing Units), bypassing NVIDIA GPUs. Developed by Moxin Technology in collaboration with Zhejiang University's Pan Yunhe academician team, MoWorld represents a complete, closed-loop system from training and distillation to deployment on domestic computing power. The model tackles the critical industry bottleneck of real-time performance, essential for applications like robotics, gaming, and digital worlds. MoWorld achieves this through a full-stack redesign for NPUs, including a proprietary 3D-annotated data pipeline, system-level optimizations for long-sequence training (up to 2000 frames), and inference optimizations like dynamic mixed-precision quantization. On a Huawei Ascend 910C platform, a 14B MoE parameter model achieves over 50 FPS, reducing typical inference costs by 70% compared to equivalent GPU solutions. This breakthrough lowers the deployment barrier, potentially accelerating the industrialization of world models. Key application areas include gaming/entertainment (offering 6-DoF camera control for immersive exploration), embodied AI/autonomous driving (providing a high-fidelity digital training ground), film pre-visualization, and 3D reconstruction/digital twins due to its strong geometric consistency. MoWorld demonstrates that a full-stack domestic compute ecosystem can support cutting-edge, real-time world models, positioning China at a competitive starting line in defining next-generation spatial intelligence standards. The project underscores a shift in competition from model scale to real-world usability and cost-effective deployment.

marsbit07/08 04:15

Just Now, The World's First Ultra-High-Frame World Model Was Born, Nvidia Content 0, Racing to 50 FPS

marsbit07/08 04:15

HKEX Welcomes Its Largest IPO of the Day

Today (July 8th), Momenta successfully listed on the Hong Kong Stock Exchange, becoming the "first Physical AI stock." The company, founded in 2016 by Tsinghua University alumnus Cao Xudong, focuses on autonomous driving as an entry point into Physical AI research. Momenta's IPO price was HK$295.6 per share. With a market cap exceeding HK$70 billion post-listing, it was the largest among the five companies debuting that day. The offering raised approximately HK$6.8 billion and attracted a "star-studded" lineup of 14 cornerstone investors, including top-tier international funds, leading strategic industrial investors like Mercedes-Benz and BYD, and major Chinese financial institutions. The company has pioneered a "flywheel" strategy, integrating mass-produced advanced driver-assistance systems (ADAS) with its full-self-driving (L4) development. Data from over 1 million vehicles equipped with its systems fuels its AI models, enabling continuous improvement. This massive real-world data scale is a core competitive advantage. In April, Momenta launched its self-developed R7 World Model for mass production, a foundational model designed to understand and predict physical world dynamics. The company positions itself not just as an automotive tech supplier, but as a platform-level Physical AI company. Its technology platform has the potential to expand beyond autonomous vehicles into areas like logistics and embodied AI. Financially, Momenta's revenue grew from RMB 743 million in 2023 to RMB 2.413 billion in 2025, with licensing income surging 42-fold during this period. While still reporting adjusted losses, it is nearing breakeven. The company boasts partnerships with 24 global automakers, including 9 of the world's top 10, and holds a 65% market share in China's third-party urban NOA segment. The listing marks a significant moment for Physical AI in global capital markets, reflecting strong investor confidence in Momenta's unique technology path and commercial execution.

marsbit07/08 03:10

HKEX Welcomes Its Largest IPO of the Day

marsbit07/08 03:10

StarDynamics Secures 2.5 Billion in Two Months, State-Owned Capital Consortium Joins In

Star Era Raises 25 Billion Yuan in Two Months with State Capital Leading the Charge. Chinese humanoid robotics leader Star Era has secured a new 10-billion-yuan funding round led by state-owned capital, including funds like Chengtong Fund under the SASAC, marking 25 billion yuan raised within two months. The company, a spin-off from Tsinghua University, has built a comprehensive capital matrix combining state guidance, top-tier financial backers, and industrial partners. Founded in 2023 by Dr. Chen Jianyu, one of Tsinghua's youngest doctoral supervisors, Star Era stands out for its early and pioneering work on "world models" for embodied AI, notably releasing its PAD world action model ahead of major global players. The company follows an AI-native, full-stack R&D strategy from data and AI brain to control, dexterous hands (XHAND series), and robot bodies (bipedal L7, wheeled Q5). A core innovation is its fully direct-drive dexterous hands, which act as high-fidelity data collectors for training its AI models like the ERA-42 and VLAW, creating a virtuous cycle of data and intelligence. Star Era claims to possess one of the world's largest real-world dexterous hand datasets. Commercially, Star Era has achieved product-market fit, most notably in logistics, with robots operating 24/7 in distribution centers for partners like SF Express and China Post, handling over 1,200 parcels per hour. It is also expanding into high-end manufacturing (Samsung, Geely) and commercial services. Its hardware components are used by nine of the global top ten tech firms and leading research institutions. The article positions 2026 as an inflection point where success shifts from model capabilities to proven, scalable commercial deployment. Star Era's rapid funding and industrial traction highlight its position in this competitive race.

marsbit07/06 01:35

StarDynamics Secures 2.5 Billion in Two Months, State-Owned Capital Consortium Joins In

marsbit07/06 01:35

Li Fei-Fei's Latest Long-Form Article: When Video Generation, Robotics, and NVIDIA All Call Themselves World Models, We Need a Taxonomy

In a new article, Dr. Fei-Fei Li addresses the widespread and often inconsistent use of the term "world model" in AI. She proposes a clear, functional taxonomy rooted in the classic Partially Observable Markov Decision Process (POMDP) loop (agent → action → state → observation → agent). According to this framework, current systems called "world models" are different projections of this loop, categorized by their primary output: 1. **Renderers**: Output observations (pixels). Their goal is visual fidelity for human consumption (e.g., video generation models like Sora). They are the most commercially mature but are limited by a focus on appearance over physical accuracy. 2. **Simulators**: Output states (geometric, physical, dynamic representations). They provide a structurally accurate world for both human professionals (e.g., architects) and computational agents (e.g., robots for training). Li argues simulators are the crucial, underappreciated bridge, as they can underpin both rendering and planning. 3. **Planners**: Output actions. Given an observation and a goal, they decide what an agent should do next (e.g., robotic action models). This area is highly promising but remains the least mature for real-world deployment. Li highlights a key trend: the boundaries between these three categories are beginning to blur, as they all rely on a shared underlying understanding of geometry, physics, and dynamics. The logical endpoint is a unified world foundation model capable of switching between rendering, simulation, and planning based on downstream needs. This convergence, she concludes, is central to advancing spatial intelligence—enabling machines not just to talk about the world, but to truly understand, imagine, and interact with it.

marsbit07/05 09:24

Li Fei-Fei's Latest Long-Form Article: When Video Generation, Robotics, and NVIDIA All Call Themselves World Models, We Need a Taxonomy

marsbit07/05 09:24

Introduction to the Concept of World Models: A Story from Psychology to the Main Battlefield of AI

**World Models: From Psychology to AI's Core Concept** "World model" is a trending but often confusing term in AI, describing a system that allows machines to internally simulate, predict, and rehearse potential outcomes before taking real-world action—like a mental "sandbox." While definitions vary—Yann LeCun emphasizes physical understanding, OpenAI's Sora is a video-based "world simulator," Google DeepMind's Genie 3 creates interactive 3D environments, and companies like Alibaba and Tesla focus on practical applications—the core goal is consistent: reduce reliance on vast real-world data by creating an internal, predictive model for safer and more efficient AI. The concept has deep roots, tracing back to psychologist Kenneth Craik (1943). In AI, it was revitalized by researchers like David Ha and Jürgen Schmidhuber (2018). Major technical approaches include: 1) generative video models (e.g., Sora) for visual realism; 2) abstract predictive models (e.g., LeCun's JEPA) for efficiency and physical reasoning; and 3) explicit 3D simulators (e.g., NVIDIA Omniverse) for precision. Fei-Fei Li proposes a classification based on the AI action loop: renderers (output observations), simulators (output world states), and planners (output actions). The emerging "World Action Model" (WAM) paradigm aims to unify future prediction and action generation. An industry framework is forming: upstream (data, compute, sensors), midstream (general and vertical platforms), and downstream applications (autonomous driving, robotics, gaming, etc.). Autonomous driving is currently the most mature use case. The current lack of a unified definition reflects the field's early, dynamic stage, similar to past tech revolutions. Different approaches—focusing on pixels, physics, or behavior—represent parallel explorations of how best to compress and understand the world. This diversity, while seemingly chaotic, signals that world models have moved from an academic idea to a critical industrial battleground, ultimately aiming to give machines the ability to understand, imagine, and reason about the world.

marsbit06/29 05:09

Introduction to the Concept of World Models: A Story from Psychology to the Main Battlefield of AI

marsbit06/29 05:09

Domestic First Explosion-Proof Certification, World's First Fueling Brain Solution: How Did They Secure Two 'Firsts'?

China's embodied AI sector is booming, with over ¥37 billion in funding this year. The focus has shifted decisively to real-world application, particularly in hazardous, repetitive tasks humans should avoid. A key, often prohibitive, barrier to entry for robots in environments like gas stations and oil fields is obtaining explosion-proof certification, requiring meticulous hardware and circuit design from the ground up. The article explores three main application areas. At gas stations, the challenge lies in executing a long, precise sequence of actions (opening caps, handling the fuel nozzle) with millimeter accuracy across diverse car models. For facility inspections, robots need sustained autonomous patrols combined with real-time anomaly detection and response. Port scenarios introduce the complexity of multi-robot coordination. Addressing the core challenge of long-horizon tasks, the piece highlights a technical breakthrough: a "world model"-driven approach. This enables predictive planning, allowing the AI to visualize the desired end-state (e.g., nozzle returned, cap closed) and work backward to synthesize intermediate visual frames. This "imagination" of the task trajectory, as implemented in the H-GAR architecture, guides action generation, significantly reducing cumulative error in multi-step operations. The three-step H-GAR process involves generating a coarse action draft, synthesizing target-conditioned observation frames, and then refining actions based on visual context and a memory of past successful motions. The conclusion emphasizes that success in specialized, safety-critical fields requires long-term commitment and deep integration of the "embodied brain" (AI) with a purpose-built, certified physical "body." Mastering this brain-body-data闭环 (closed-loop) is positioned as a crucial competitive advantage for commercialization.

marsbit06/26 03:49

Domestic First Explosion-Proof Certification, World's First Fueling Brain Solution: How Did They Secure Two 'Firsts'?

marsbit06/26 03:49

The War Without a Unified Name: The Domestic Tech Giants' World Model Landscape

The article outlines the diverse and fragmented landscape of "World Models" in China's tech industry, where major players are pursuing similar goals under different names like world foundational models, physical AI, or integrated within autonomous driving and embodied intelligence systems. The core aim is to enable AI to create an internal, dynamic environment for simulation, reasoning, and learning, reducing reliance on infinite real-world data. This "data engine" allows for unlimited generation, experimentation, and iteration. The report categorizes the approaches of different companies: * **Internet Giants:** Alibaba is developing models for linguistic, virtual, and physical worlds (Qwen-AgentWorld, HappyOyster, Qwen-RobotWorld). Tencent's HY-World focuses on 3D, game, and social scenarios. ByteDance leverages its vast video data for a potential "digital twin" model. Huawei integrates its model into industrial applications like smart cars and robotics without separately branding it. Baidu embeds world model capabilities within its Apollo autonomous driving and Ernie systems. * **Automakers:** Companies like NIO, Li Auto, XPeng, and Geely are using world models as virtual "driving schools" and "testing grounds." They generate complex scenarios (e.g., rain, snow) to train and validate autonomous driving systems in simulation, aiming for more capable and safer AI drivers. * **Autonomous Driving Suppliers:** Firms such as Momenta, Horizon Robotics, Haomo.ai, and DeepRoute.ai are building the underlying "world engines." They focus on large-scale video generation for simulation, reinforcement learning, and enhancing end-to-end autonomous driving models, often integrating these capabilities into commercial products. While startups bring focus and innovation, they face challenges like limited data, compute resources, and deployment channels. Large companies possess these advantages and are rapidly transitioning world models from research projects into core business infrastructure powering products in vehicles, games, and industry. The conclusion is that world models represent an evolution and convergence of existing AI fields into crucial industrial infrastructure, moving the competition from simply building a model to effectively deploying it to understand and interact with the physical world.

marsbit06/25 06:52

The War Without a Unified Name: The Domestic Tech Giants' World Model Landscape

marsbit06/25 06:52

活动图片