# Pretraining Articles associés

Le Centre d'actualités HTX fournit les derniers articles et analyses approfondies sur "Pretraining", couvrant les tendances du marché, les mises à jour des projets, les développements technologiques et les politiques réglementaires dans l'industrie crypto.

10000 Hours of Human Data, Trains the World's First Whole-Body Mobile Manipulation Implicit World-Action Model

"Being-M0.7" is the world's first latent world-action model for whole-body mobile manipulation in humanoid robots, developed by Zhi Zai Wu Jie (Beyond Being). Trained on over 10,000 hours of human-centric multimodal data, it aims to overcome key industry challenges: the high cost and scarcity of real robot demonstration data, the computational inefficiency of pixel-level video prediction models, and the lack of full-body coordination in existing approaches. The model is based on a Vision-Motion Mixture-of-Transformers (MoT) architecture, which allows training on a mixture of paired video-motion data, pure video data, and pure motion sequences. A key design is a unified motion representation that bridges human and robot morphology, enabling knowledge transfer from vast human behavioral data to specific robot control. After pre-training on human data, the model is adapted to a real robot (Unitree G1) using a small amount of teleoperated demonstration data via a lightweight "Action Expert" module. This process decouples low-frequency world planning from high-frequency motion control. The model was tested in four challenging real-world demos: fishing a toy fish from water (liquid interaction), retrieving an object using a mirror (visual reasoning), a multi-step pick-and-place task, and obstacle avoidance while carrying a box. In comparative tests against other models, Being-M0.7 showed stronger performance in tasks requiring indirect reasoning and full-body coordination. This work represents a shift in humanoid robotics competition from hardware spectacle to model capabilities rooted in scalable data and training paradigms, using human experience as a foundation for physical world understanding and action.

marsbit07/15 01:48

10000 Hours of Human Data, Trains the World's First Whole-Body Mobile Manipulation Implicit World-Action Model

marsbit07/15 01:48

Karpathy's Latest Outburst: A Single Sentence That Silenced the Entire Agent Developer Community

Andrej Karpathy, a core researcher at Anthropic, recently critiqued the current AI agent development frenzy. He argues that the biggest mistake is forcing agents to perform tasks without first thoroughly understanding the underlying large language models. Drawing from his 2016 "World of Bits" project at OpenAI—an early attempt at web-based agents that ultimately failed due to premature technology—he emphasizes that foundational model work is crucial. Karpathy offers three key pieces of advice: First, focus on getting the base models right before pushing agents. Second, recognize that creating a demo is easy, but building a real product takes a decade, akin to the journeys of autonomous driving and VR. Third, the product is the core capability, not the agent shell; a robust foundation will naturally enable advanced agents. He also suggests looking to neuroscience for inspiration, comparing agent components to brain structures like the hippocampus and thalamus. Despite his caution, Karpathy concludes that independent developers and startups, not large labs like OpenAI, are at the forefront of agent innovation. This is because the agent field is new, with no entity having a five-year head start, leveling the playing field for agile experimenters. His core message is not to abandon agent work, but to build it on a solid, deeply understood foundation.

marsbit07/06 02:33

Karpathy's Latest Outburst: A Single Sentence That Silenced the Entire Agent Developer Community

marsbit07/06 02:33

For the First Time, Pure Human Video Pretrained VLA for Dexterous Manipulation: Deployable with Minimal Fine-Tuning Data

For the first time, a purely human-video-pretrained Vision-Language-Action (VLA) model for dexterous manipulation requires only a small amount of data for fine-tuning to achieve successful real-world deployment. Achieving human-level dexterous manipulation remains a core challenge in robotics. While multi-fingered hands offer hardware potential, Visual-Language-Action (VLA) models lag behind due to the high cost of collecting diverse, high-quality robot data. A novel framework, VITRA, developed by Microsoft Research Asia and Tsinghua University, addresses this by automatically transforming massive, unlabeled real-world human activity videos into a structured V-L-A training dataset. Key innovations include precise 3D hand motion annotation from monocular video, atomic action segmentation based on hand-speed minima, and automated instruction generation using VLMs combined with 3D trajectory visualization. This process created a massive dataset of 1 million clips. Pretrained exclusively on this human video data, the VLA model (combining a VLM backbone with a Diffusion Transformer action expert) demonstrates strong zero-shot hand motion prediction in unseen environments. Crucially, it requires minimal fine-tuning (~1.2k demonstrations) on real robot data to achieve high-success-rate dexterous manipulation tasks like grasping, placing, pouring, and sweeping on hardware like the Realman robot with the XHAND1 dexterous hand. The model shows exceptional generalization to novel objects and environments. The research also observes promising scaling behavior, where performance improves with more pretraining data, paving the way for more generalized embodied intelligence.

marsbit06/08 08:54

For the First Time, Pure Human Video Pretrained VLA for Dexterous Manipulation: Deployable with Minimal Fine-Tuning Data

marsbit06/08 08:54

活动图片