# World Model Related Articles

HTX News Center provides the latest articles and in-depth analysis on "World Model", covering market trends, project updates, tech developments, and regulatory policies in the crypto industry.

GPT-5.6 Countdown: Abandon the Illusion of a Single API, Computational Iteration Can't Outpace a Single Page of Compliance

In mid-June, three seemingly independent industry events—the compliance-driven throttling of Fable 5, the open-sourcing of GLM-5.2, and the leaked release timeline for GPT-5.6—are pushing the global AI industry toward a watershed moment. These shifts signal a fundamental restructuring of the industry's underlying logic. First, **"usability" has substantially overtaken "advanced capabilities"** as the primary weight, pushing the global large language model (LLM) supply chain into a "dual-track" phase of controlled closed-source and local open-source coexistence. Second, **the competitive moats of closed-source giants are shifting**. Their technical focus is moving from "language intelligence" toward "spatial intelligence (world models)"—a domain heavily reliant on computing power. Third, faced with常态化 transnational compliance risks, **a "model-agnostic" decoupled design has become a survival necessity for application-layer developers to maintain business continuity.** The article details how Anthropic's Fable 5, despite its advanced engineering feats, was restricted for non-U.S. citizens within 72 hours of launch, highlighting how geopolitical compliance can instantly limit even the most advanced models. In response, the open-source camp, exemplified by Zhipu AI's MIT-licensed GLM-5.2, is gaining market share by offering stable performance improvements and significant cost advantages (up to 70% savings for enterprises), while achieving full adaptation with domestic semiconductor platforms. Meanwhile, closed-source leaders like OpenAI are pivoting. The anticipated GPT-5.6 reportedly shifts focus from language to spatial intelligence and world models, aiming to rebuild a generational gap in areas like 3D understanding, simulation, and industrial design that demand immense compute. The core conclusion is that the LLM supply chain's logic has changed. Enterprises must now evaluate infrastructure based on a composite of technical performance and policy compliance. For developers, complete reliance on a single closed-source API poses unacceptable risk. Implementing a truly model-agnostic architecture—enabling swift switches to compliant, locally deployable open-source alternatives—is no longer just good practice but a fundamental baseline for business continuity.

marsbit06/21 04:40

GPT-5.6 Countdown: Abandon the Illusion of a Single API, Computational Iteration Can't Outpace a Single Page of Compliance

marsbit06/21 04:40

He Just Raised 2.7 Billion, and Li Fei-Fei Also Invested

Pete Florence, a former senior research scientist at Google DeepMind and a key contributor to the Vision-Language-Action (VLA) model architecture, is deliberately distancing his startup, Generalist AI, from the trendy "world model" label. He argues that the industry should prioritize concrete goals over buzzwords. His goal is to create robots that can perform a vast range of unseen tasks with high speed and success rates, without needing task-specific training data. Recently, his company raised $400 million (¥2.7 billion) at a $2 billion valuation. Notable investors include NVIDIA's NVentures, Bezos Expeditions, NFDG, as well as Xiaomi co-founder Lin Bin, Zoom founder Eric Yuan, and renowned AI scientist Fei-Fei Li. Florence's approach stems from his academic background at MIT under Professor Russ Tedrake, focusing on understanding the physical world. After joining DeepMind, he developed models like Transporter Network and co-created the VLA framework. He left in 2025 to found Generalist AI. The company has launched two models: GEN-0, which demonstrated that scaling laws apply to physical motion, and GEN-1. GEN-1 was trained on over 500,000 hours of physical interaction data collected via a specialized wearable device. It achieves a 99% success rate on precise mechanical tasks like folding boxes and maintains performance three times faster than its predecessor. Florence believes GEN-1 is reaching a commercial utility threshold similar to the GPT-3 inflection point. The substantial funding round, following GEN-1's release, signifies strong investor confidence in Generalist AI's practical, goal-driven path to creating versatile, useful robots, regardless of the "world model" terminology.

marsbit06/20 06:06

He Just Raised 2.7 Billion, and Li Fei-Fei Also Invested

marsbit06/20 06:06

The Unfinished Tale of Jueying, DaXiao Robotics Swiftly "Raises Funds"

Following a major fundraising round involving several prominent investment institutions, DaXiao Robotics, a company backed by SenseTime, has secured hundreds of millions of US dollars in financing for the first half of 2026. This move signals SenseTime's renewed and substantial bet on "Physical AI" through embodied intelligence, following the relative underperformance of its autonomous driving unit, Jueying. While Jueying achieved mass production partnerships in the smart vehicle sector, it failed to become a pivotal player in the high-level autonomous driving landscape, leading to its gradual independence from SenseTime's core financials. DaXiao Robotics now emerges as SenseTime's next major venture into the physical world. The new funding will focus on developing a "world model" and integrated hardware-software solutions for commercial applications like retail, security, and hospitality. This ambition is significantly more complex and capital-intensive than previous projects. A world model requires understanding spatial relationships, physics, and causality to guide robots in long-term tasks, demanding immense computational resources, data, and engineering. The article highlights several challenges. First, the massive funding, while substantial, may still be strained by the high costs of R&D, data collection, and commercial deployment. Second, SenseTime itself, despite narrowing losses, continues its high-investment growth model and cannot solely bankroll this new, expensive endeavor. Third, DaXiao Robotics, led by SenseTime co-founder Wang Xiaogang, carries the technical heritage and resources of its parent company but also potentially its organizational inertia. It operates in a field increasingly dominated by agile, young technical founders. Ultimately, DaXiao Robotics represents SenseTime's attempt to secure a leading industrial position in embodied intelligence—a goal its Jueying unit did not fully achieve in autonomous driving. The new venture starts with strong capital backing, but faces the critical task of rapidly transitioning from technological narrative to sustainable commercial delivery in an early-stage, costly, and highly competitive arena.

marsbit06/15 08:41

The Unfinished Tale of Jueying, DaXiao Robotics Swiftly "Raises Funds"

marsbit06/15 08:41

Three Months, 35 Billion Yuan: Investors Rush to Grab the OpenAI of the Physical World

Investors flock to a physical AI startup as the race for the "OpenAI of the physical world" heats up. Ji Jia Shi Jie (GigaWorld), a company dedicated to developing Artificial General Intelligence (AGI) for the physical world, has raised 3.5 billion RMB (approximately $490 million) in just three months, according to a report from investment media outlet Touzijie. The latest B2 funding round of 1 billion RMB attracted a wide range of top-tier investors, including sovereign wealth funds, industrial capital, and financial institutions. This brings the total funding for the young company, now valued over 10 billion RMB, to 3.5 billion RMB across three recent rounds. The company is led by Huang Guan, a post-90s Tsinghua University PhD with extensive experience in AI, autonomous driving, and entrepreneurship. Its core innovation is a "dual-pyramid" system comprising a five-layer data pyramid (from internet videos to real-world robot data) and a three-layer algorithm pyramid focused on world simulation, action alignment, and reinforcement learning. This system underpins its key models: the "World Action Model" (e.g., GigaBrain series for robot control) and the "World Generation Model" (e.g., GigaWorld series for simulating and understanding the physical world). Its models have reportedly achieved top rankings in global robotics benchmarks. Ji Jia Shi Jie argues that while current digital AGI excels in information processing, the next frontier is physical AGI—systems that can understand and interact with the real world. The company believes the field is approaching its "GPT-3 moment," a key inflection point in capability scaling. To achieve this, the company is pursuing a dual-market strategy. For the consumer (C) market, it launched the "SeeLight" brand and its S1 general-purpose humanoid robot, which has secured initial orders for deployment in real homes. For the business (B) market, it focuses on industrial automation with its Maker series robots, having signed agreements for large-scale deployment in factories, and its DriveDreamer world model for autonomous driving, which is already in use with over 30 automakers and tech companies. The report concludes that by bridging the gap between digital intelligence and physical action, Ji Jia Shi Jie aims to unlock a new wave of productivity, ultimately bringing physical AGI into everyday life.

marsbit06/15 01:30

Three Months, 35 Billion Yuan: Investors Rush to Grab the OpenAI of the Physical World

marsbit06/15 01:30

Tremble Humans, AI Continues Its Accelerated Sprint

Trembling, Humans: AI Continues Its Accelerated Sprint Yes, AI is still rapidly accelerating. While deep learning seemed to stall quickly in its early years, large models after years of development show no sign of hitting their ceiling. At the Zhiyuan Conference 2026, the focus is on enabling AI to move from the digital world into the physical world. Scaling Law remains effective, continuing to drive advancements in both large language models and multimodal models. The industry is now entering a phase of pursuing World Models, though unresolved technical paths and data issues mean this exploration may take 3-5 more years. Concurrently, breakthroughs in Agents are accelerating AI's real-world application in fields like healthcare and meetings. Making Agents truly useful requires key hardware-software co-design, evident from the strong presence of chip vendors at the conference. We stand at a new historical threshold where AI is becoming a foundational force reshaping the world. The first day of the conference highlighted AI's evolution from "knowing how to chat" to "knowing how to work." Scaling Law persists, World Models are the next key battleground, and Agents are transitioning from usable to好用 (user-friendly). Scaling Law is not ending but diversifying. New models like Anthropic's Fable 5 demonstrate scaling through parameter size, synthetic data, and reinforcement learning. Advancements in AI Coding and Agent deployment are enabling a trend of AI self-evolution, potentially allowing AI to take over digital world iterations. World Models represent the next frontier for large models extending into the physical realm, but no current model is truly impressive at solving real-world problems. Technical consensus is lacking, with debates on data sources (video, simulation, real-world). Different approaches are emerging: language-centric, pixel-centric, 3D-structure-centric, and visual-representation-centric models. Zhiyuan Institute is exploring a fifth path: unified latent space modeling fusing language and visual representations, and introduced its own under-development World Model, Physis-v0.1. On the product side, Agents are key to bringing AI into daily life. Since 2025, the "Year of the Agent," products have become more proactive and capable of complex tasks. Zhiyuan showcased four vertical Agents for cardiac diagnosis, autonomous research, meeting summarization, and protein risk discovery. However, technical challenges remain, particularly in context engineering like memory and orchestration. "Harness" – the engineering framework around an Agent – is crucial for maximizing its capabilities by clarifying intent, designing workflows, and incorporating validation and feedback. In summary, AI's breakneck pace continues on multiple fronts: foundational model scaling, the ambitious pursuit of World Models for physical understanding, and the ongoing refinement of practical Agents. The journey from capable to truly reliable and useful AI systems is well underway.

marsbit06/13 02:51

Tremble Humans, AI Continues Its Accelerated Sprint

marsbit06/13 02:51

From a Lunch Table to an Infinite Universe: Fei-Fei Li Bets on AI's Next Dimension

From a Lunch Table Conversation to an Infinite Universe: Fei-Fei Li Bets on AI's Next Frontier - Spatial Intelligence In an era dominated by large language models, AI pioneer Fei-Fei Li argues that true understanding requires spatial intelligence — the ability to perceive, reason, and interact within the physical 3D/4D world. She points to evolutionary history: spatial perception drove the Cambrian explosion 540 million years ago, while language is a far more recent, inherently "lossy" way to encode reality. Current models struggle with basic spatial tasks a child can do, like counting chairs in a video. Her company, World Labs, is pioneering this shift with "Marble," a model that generates navigable, consistent 3D worlds from text, images, or simple 3D inputs—distinct from video generators like Sora. Though smaller than models like GPT-5, due to scarce 3D data and early-stage scaling laws, Marble is already used in gaming, robot training (by NVIDIA), architectural design, and personalized therapy for conditions like OCD and acrophobia. Li envisions this technology enabling "infinite universes" for creativity, social interaction, and more. However, she cautions against utopian or dystopian extremes, advocating for a measured vision where AI enhances human dignity and prosperity, akin to how electricity transformed civilization. The journey is long — as evidenced by the 20-year path to viable autonomous vehicles — but the direction is clear: for AI to move from merely talking about the world to truly understanding and acting within it.

marsbit05/27 00:14

From a Lunch Table to an Infinite Universe: Fei-Fei Li Bets on AI's Next Dimension

marsbit05/27 00:14

Physical AI is Hot, Some New Thoughts from Me

The term "Physical AI" is gaining significant traction, marking a shift from AI that processes information to AI that understands and interacts with the physical world. Unlike traditional AI confined to screens, Physical AI involves integrating intelligence into robotic bodies to perform tasks in environments governed by gravity, friction, and inertia. The concept, formally defined in a 2020 paper, focuses on creating embodied systems that can complete perception-to-action cycles. 2026 is identified as a pivotal "deployment year," where the focus moves from demonstrations to practical utility. Companies like China's Zhiyuan Robotics have transitioned to live, unscripted factory deployments and announced mass production targets. Internationally, Figure AI, after a major funding round, shifted to its own neural system, while NVIDIA partnered with major industrial robot firms to upgrade millions of existing units with AI capabilities. A key trend is the crossover from the automotive supply chain. Companies like Aptiv and Valeo are entering the Physical AI space, leveraging their expertise in sensors, control systems, and mass production from the autonomous vehicle sector. This "technology spillover" is accelerating development, as seen with Tesla's plans to repurpose automotive production lines for its Optimus robot. The technical breakthrough enabling this progress is the engineering maturity of "world models." Previously theoretical, these AI models can now simulate physical interactions and generate vast, realistic synthetic training data for robots. Innovations from NVIDIA's Cosmos, Ant's LingBot-World, and others have made this capability more accessible, drastically reducing the cost and time needed for real-world data collection. This is driving a fundamental architectural shift in robotics: from the traditional "sense-plan-act" model, reliant on pre-programmed rules, to a "sense-reason-act" paradigm where neural networks reason and make decisions. This change represents a new paradigm where machines understand the world's physics. The competition is intense, with the landscape still forming. While the direction is clear, success will depend not just on AI algorithms but on manufacturing scalability, supply chain resilience, and efficient data strategies, with infrastructure providers potentially capturing significant value in this new era.

marsbit05/18 04:43

Physical AI is Hot, Some New Thoughts from Me

marsbit05/18 04:43

The Year of Physical AI: A Trillion-Dollar Gamble on 'How the World Works'

The year 2026 is being positioned as the dawn of the "Physical AI" era, marked by major funding rounds and technological breakthroughs. This shift signifies AI's evolution from understanding the digital world to perceiving and acting within the physical world. Key events include Yann LeCun's AMI Labs raising $1.03 billion to develop "world models," Fei-Fei Li's World Labs securing funding, and companies like Tesla deploying humanoid robots (Optimus) in factories. This transition expands the AI model competition into a broader infrastructure battle encompassing hardware, data, simulation, and real-world integration. The core debate is between two AI paths: the established LLM (Large Language Model) approach focused on text prediction and the emerging "world model" approach, which aims to understand physical states for action-oriented tasks. Hardware, particularly dexterous robotic hands, is a critical and expensive challenge. Companies are racing to build capable robotic bodies, with Tesla, Boston Dynamics, and Figure AI making significant progress. NVIDIA is positioning itself as the essential infrastructure provider for this new era, offering a full suite of development tools and platforms. A major bottleneck is the scarcity of high-quality physical world interaction data, with companies exploring solutions through real-world data collection, synthetic data generation, and human teleoperation. Substantial investments in Q1 2026, exceeding $6.4 billion, signal strong belief in Physical AI's potential, moving beyond concept validation into infrastructure building. While challenges like the sim-to-real gap, unproven business models, and safety regulations remain, the tangible engineering progress suggests this is a genuine technological inflection point, not merely a bubble. For the global Chinese community, this shift represents a significant structural opportunity to leverage their strengths in technology, engineering, hardware manufacturing, and cross-border collaboration to become key players in building the foundational layers of the Physical AI ecosystem.

marsbit04/03 09:39

The Year of Physical AI: A Trillion-Dollar Gamble on 'How the World Works'

marsbit04/03 09:39

活动图片