They Raised 17 Billion in Six Months: Data "Tool Sellers" Became the Most Profitable Business in the Embodied Intelligence Track

marsbitОпубліковано о 2026-08-05Востаннє оновлено о 2026-08-05

Анотація

Over the past six months, data-centric companies serving the embodied AI sector (robotics) have secured over 17 billion RMB in funding in China, highlighting "data" as the most profitable niche. These "shovel sellers"—providing crucial data, models, and infrastructure for training robots—are flourishing despite robots themselves not yet being widely profitable. The surge is driven by a severe scarcity of high-quality physical interaction data needed for robot training. Companies are tackling this through five main approaches: 1) **Teleoperation/Haptic Data**: Building factories for high-precision data collection (e.g., Paxini, Noitom Robotics). 2) **Simulation/Synthetic Data**: Generating vast amounts of virtual training data (e.g., Lightwheel Intelligence, Transcend Dimension). 3) **UMI/Portable Collection**: Using wearable devices to capture human motions directly, bypassing robots (e.g., Jianzhi Robotics, Tashizhihang). 4) **Video Distillation/World Models**: Extracting actionable data from internet videos or using AI world models (e.g., Shutu Technology, Deep Genius). 5) **Data Infrastructure/Platforms**: Offering data processing, standardization, and platform services (e.g., Wuwen Zhike, Yiren Technology). Key players include LiberAI (founded by a 00-year-old PhD), which focuses on human UMI data and world models and recently raised hundreds of millions, and Lightwheel Intelligence, which became a unicorn and secured large orders. However, the industry faces a critica...

At the end of July, LiberAI completed a Pre-A+ round of financing, raising several hundred million yuan. The investor list includes 360 Group, China Development Bank Financial Leasing Co., Ltd., Sinovation Ventures, CMC Capital Partners, BinFu Capital, and Sequoia Capital China.

This company was actually founded on December 8, 2025, named "Jiangxian Technology," with the implied meaning that "wanting to be idle is the primary productive force." Its registered capital is 1.46 million yuan, and the team has fewer than 30 people.

IT桔子 found that this is already its fifth round of financing since its establishment. In the past six months or more, it has been raising funds intensively at a very fast pace, completing four rounds in one go from seed, angel, angel+ to Pre-A round. According to reports, its latest valuation has reached 5 billion yuan.

The founder, Liu Songming, is a member of the post-2000 generation, a Tsinghua University Computer Science Department special scholarship winner and top of his class, mentored by Professor Zhu Jun, who graduated with a Ph.D. at 23 and started a business. He previously led the development of RDT-1B, the first large-scale diffusion Transformer foundational model designed for dual-arm robotic manipulation, with 1.2 billion parameters.

LiberAI targets human UMI data + world models—in simple terms, it sells the data fed to robots and the brain that digests this data.

According to IT桔子 statistics, in the first half of 2026, 25 Chinese embodied intelligent data startups collectively raised over 17 billion yuan in financing. It can be said that "data" has become the most certain business in the embodied intelligence track. (Excludes individual companies simultaneously engaged in embodied data/modeling and robotics hardware, such as Xinghai Tu and Qianxun Intelligence)

To obtain data on all companies in the embodied data track and their financing data, you can view and download the data on the IT桔子 album page. https://www.itjuzi.com/album/766 (Currently includes 42 companies)

Among them, not only did Guanglun Intelligent secure significant external financing, but it also landed a 550 million yuan order in Q1.

While the gold miners haven't struck gold yet, the shovel sellers are making money first.

I. Unsolved Problem: The Data Desert of the Robotics Industry

Why did large language models explode? Because the internet has an inexhaustible supply of text corpus.

But training robots relies not on written language, but on three-dimensional physical action data like picking up, placing, walking, grasping, twisting, and wok flipping—this type of data hardly exists on the internet.

The scarcity of embodied data is a major bottleneck throughout the industry.

Grand View Research estimates that the global data collection and annotation market will reach $17.1 billion by 2030.

The industry generally recognizes that embodied data has a pyramid structure. At the top is real machine data, which is reliable in quality and can be deployed directly, but also the most expensive. The middle layer consists of simulated synthetic data + UMI data, which is moderately priced and can be produced infinitely, but suffers from the Sim-to-Real gap (approximation deviations in friction/deformation). The base is internet video and human behavior data, which has the widest source but lacks interaction details and is information-sparse.

Then, there is a batch of companies dedicated to cracking each layer—this is a coordinate system for understanding the current wave of embodied data entrepreneurship.

II. Five Schools, Over Thirty Companies Panorama

School One: Real Machine Teleoperation/Tactile—Treating Data Like an Assembly Line

This school adheres to the most fundamental logic: build factories, deploy robots, set up motion capture equipment, and mass-produce high-precision data on an assembly line.

The most aggressive is Pasini Perception.

The company ranks first globally in tactile sensor shipments. In 2025, Pasini built the world's largest embodied data collection factory, Super EID Factory, in Tianjin, with an annual output of 200 million data entries. After completing a Series B of over 1 billion yuan in March with a valuation exceeding 10 billion yuan, it announced plans to build four more super data factories in Suqian, Wuhan, Zigong, and Ganzhou.

Motion capture veterans are another branch of this line.
Noitom Robotics founder Dai Ruoli previously led Noitom Technology to capture 70% of the global professional motion capture market share, then founded the new company Noitom Robotics. In June, it completed a Pre-A++ round of several hundred million yuan with a valuation of about 3 billion yuan. Beijing, Shanghai AI Industry Funds, Shenzhen Capital Group, CICC, and Kunlun Capital collectively participated. It also released the ModalityNet full-modal data platform.

Qingtong Vision takes the open-source-for-ecosystem route. In May, it released the world's largest 1000-hour open-source dataset of human high-precision optical motion capture, with an estimated annual output of 500,000 hours. The martial arts data collection for Unitree's Spring Festival Gala robot martial arts debut "WuBOT" came from its Project Decode system.

Newcomers like Lingsheng Technology are also positioning themselves in the tactile field.

School Two: Simulation/Synthetic—"Printing" Data in Virtual Worlds

If real data is too expensive, can we create a sufficiently physically realistic virtual world in a computer to batch "print" training data?

This route has produced the world's first embodied data unicorn: Guanglun Intelligent.

In March of this year, it completed a 1 billion yuan A++ round of financing, secured another round led by Ant Group in May, and raised another 1 billion yuan in June, with a post-investment valuation exceeding $2 billion (approximately 15 billion yuan).

Guanglun Intelligent's 2025 revenue grew tenfold, and in Q1 2026, it secured new orders worth 550 million yuan, exceeding the total for the entire previous year in just one quarter.

Kuawei Intelligence is a representative of "generative simulation": its self-developed DexVerse simulation engine automatically generates massive amounts of scene data in virtual worlds. In July this year, it announced a 1 billion yuan financing round, a valuation exceeding 10 billion yuan, and initiated an IPO. Its revenue in the first half of the year was 100 million yuan, with over 1,500 models having completed commercial delivery.

Among listed companies, the "world's first spatial intelligence stock" Qunhe Technology, which went public on the Hong Kong Stock Exchange in April, is extending its massive accumulation of 3D cloud design scene data into embodied intelligence.

The simulation platform Songying Technology completed Pre-A and Pre-A+ rounds in February and July respectively, raising a cumulative total of several hundred million yuan, and is co-building synthetic datasets with the National and Local Co-built Humanoid Robot Innovation Center.

School Three: UMI Non-Body/Portable Collection—Bypassing the Body, "Replicating" Humans

Teleoperating robots to collect data has a flaw—the collected actions are not the robot's true capabilities, but compromises made "to allow the robot to keep up." So this school simply bypasses the robot body, having humans work directly wearing gloves, grippers, or wearable devices.

This is the route with the highest capital density this year.

On June 1, Jianzhi Robotics announced continuous multi-round financing totaling several hundred million yuan, led by Ant Group, Didi, and Delian Capital, with Shunwei Capital and BV Baidu Ventures participating—the largest financing in the non-body data field to date.

Founder Chen Jianxing is the former Senior Director of Algorithms at Momenta. The company, established a year ago, has already covered tens of thousands of real-world scenarios and secured a strategic partnership with Ant Lingbo.

Tashi Zhihang proposed a Human-centric data collection paradigm, launching the SenseHub wearable collection solution. In April, it completed a $455 million Pre-A round, setting a record for the largest single financing round in China's embodied intelligence sector. Hillhouse Capital, Sequoia Capital, and Meituan jointly led the round. It has also launched the "Embodied Data Spark Plan" with a target of 100 million hours.

Tsinghua-affiliated Luming Robotics's FastUMI system increased data collection efficiency threefold and reduced costs to one-fifth, securing consecutive rounds led by Mitsubishi Electric and landing orders from leading clients.

In February this year, Zhiyuan Robotics spun off its data business into an independent company, Mifeng Technology, which completed seed and angel rounds totaling several hundred million yuan in just ten days, led by Sequoia Capital China, with plans to achieve 10 million hours of annual data production capacity by 2026.

A younger crop is also emerging: Xingyi Technology, incubated from Tsinghua University's Computer Science Department, follows the Nvidia EgoScale path, focusing on high-degree-of-freedom, millimeter-precision first-person wearable collection. It completed its first round of financing and secured a Pre-A round in July. Professor Lu Zongqing's Zhizai Wujie directly uses internet videos to pre-train a general action model.

Qiongche Intelligent, incubated by Flexiv, secured nearly a hundred orders with its "production-accompanying" data collection system CoMiner, received another several hundred million yuan in financing in June, and is about to release its self-developed world model.

There's also a very new company, Yuanche Taichu (OriginFlow), which adopts the NeuroScale paradigm. It uses a self-developed electromyography (EMG) acquisition kit to capture the electrical signals of human muscle contractions, reconstructing hand posture/force/tactile feedback through the PULSE foundational model. It connects the "intention—muscle—action" native conduction pathway, compensating for the limitations of UMI visual occlusion failure and lack of force/tactile feedback.

Established for a year, the company has completed cumulative financing of over 500 million yuan across angel, strategic, and Pre-A1 rounds. Founder Qin Shentao is a post-2000 Tsinghua Ph.D. graduate with a bachelor's from Harbin Institute of Technology.

School Four: Video Distillation/World Model—Teaching AI to "Understand" the Physical World

This route solves the problem of "turning waste into treasure" at the very bottom of the data pyramid: the internet has an endless supply of human operation videos—cooking, crafting, repairs—but these videos lack the force, tactile, and 3D trajectory information robots need, remaining "visible but unusable" dead data.

The video distillation school's approach is to use algorithms to infer action trajectories and physical interaction information from 2D videos, "distilling" dead data into trainable embodied data, potentially reducing costs to a few thousandths of real machine collection.

The world model school is more direct: first, let the AI learn the laws of the physical world, then generate infinite training data "in its mind."

Shutu Technology's SynaData pipeline can batch extract multimodal embodied data like hand trajectories and object motion paths from internet videos, with extremely low comprehensive data collection costs. It has received tens of millions of yuan led by Oriental Fortune Capital, and its data is used by mainstream open-source models like Tsinghua's RDT and Zhiyuan's UniVLA.

Jijia Shijie uses the world model GigaWorld to generate training data, improving VLA model performance by nearly 300% across three generalization dimensions. It completed a 1 billion yuan B+ round in June, accumulating 3.5 billion yuan in financing within three months.

DeepWit also bets on human first-person video data, raising several hundred million yuan+ across three rounds in the first half of this year. In May, DeepWit's Z-WM world model topped the global WorldArena benchmark.

LiberAI, mentioned at the beginning of the article, is also a hybrid of this route and the world model track.

School Five: Data Infrastructure/Platform—Data Platforms as "Refineries"

This school's insight is that the bottleneck for embodied data isn't just "how to collect," but also the "how to use" link.

Different robot bodies, different sensors, diverse data formats—collected raw data is like crude oil with mixed components. Directly feeding it to models will only worsen training—who will handle cleaning, alignment, annotation, and evaluation to refine crude oil into standard gasoline?

Moreover, training grounds, collection equipment, and motion capture systems are heavy assets that small and medium-sized teams simply can't afford. Therefore, there is a need to build shared training grounds, set data standards, and sell toolchains and platform services. This school doesn't bet on which data technology route will win, but bets that whoever wins will have to pass through the road it builds.

Wuwen Zhike relies on the Deqing data collection and training ground featuring a virtual-real fusion closed loop, producing thousands of hours of data daily. It completed over 100 million yuan in financing in April and signed orders worth several billion yuan in Q1 with companies like ByteDance and Wujie Dongli.

Yiren Technology completed two 100-million-yuan-level financing rounds consecutively in April, simultaneously announcing revenue exceeding 100 million yuan in 2025 and achieving profitability, likely the first company in the track to announce profitability.

There's also Zhiyu Jishi, founded in December last year, specializing in mid-stream data cleaning, alignment, and governance. Hardware companies like Lingchu Intelligent, Qiongche, and Zhipingfang have all invested in this "upstream" player.

Listed players are also entering. Data service leader Haitian Ruisheng is collaborating with the Beijing Shijingshan Humanoid Robot Data Training Center to co-build an "Embodied Intelligent Data Training Ground."

Major tech companies are also joining in. JD.com released a full-chain embodied data infrastructure, planning to mobilize 600,000 couriers and delivery riders for crowdsourced collection, aiming to accumulate 10 million hours of real-scenario video within two years. Baidu, on the other hand, is building a "data supermarket."

III. A Dose of Reality: Is the Data Demand Side Truly Rigid?

2026 has become the year of scale for embodied data. Statistics show that by the end of April 2026, there were at least 90 data collection centers in a state of "use or planning/construction"; 64 were already operational, with the remainder under construction/planned.

But the hotter the trend, the more necessary it is to pour some cold water.

The business of embodied data has a structural problem often left unspoken: its demand side and supply side are burning through money from the same pool of capital.

Looking at the buyer list makes it clear—large model teams, startup robotics companies, traditional hardware manufacturers in transition. These data purchasers, the vast majority, are not yet profitable themselves; their procurement budgets come from freshly raised financing.

In other words, the revenue of data companies essentially constitutes a secondary distribution of downstream financing: embodied intelligence robot companies raise money, then turn around to buy data to tell a story, and if the story is good, they go raise the next round.

If the commercialization of robot hardware companies falls short of expectations, data orders and financing will recede simultaneously. This is a symbiotic relationship, where prosperity or decline are shared—a market demand dependent on the downstream.

Only when robotics companies can truly achieve sustained profitability does this business become more solid and viable.

A consensus judgment within the industry is that, in the long run, surviving embodied data companies will likely fall into two categories: one becomes an industry-standard platform, mastering ecosystem-level tools like simulation, data processing, and evaluation; the other possesses cross-manufacturer data fusion and refinement capabilities, able to continuously supply high-quality, long-tail data after robots enter real-world scenarios.

This article is from the WeChat public account "IT桔子" (ID: itjuzi521), author: Wu Meimei

Трендові криптовалюти

Пов'язані питання

QWhat is the main business model of LiberAI mentioned in the article?

ALiberAI's main business is selling data to feed robots, as well as the 'brains' (models) to process that data. It focuses on human UMI data and world models.

QAccording to the article, why is the robotics industry considered a 'data desert'?

AThe robotics industry is considered a data desert because while large language models have abundant text data from the internet, robots require three-dimensional physical action data (like grasping, moving). This type of data is almost non-existent on the internet.

QWhat are the five main approaches or 'schools' of embodied intelligence data entrepreneurship identified in the article?

AThe five main approaches are: 1. Teleoperation/Tactile Data (physical production lines), 2. Simulation/Synthetic Data (generating data in virtual worlds), 3. UMI/Portable Collection (collecting human action data directly, bypassing robot hardware), 4. Video Distillation/World Models (extracting useful data from 2D internet videos or using world models), and 5. Data Infrastructure/Platforms (building tools, standards, and platforms for data processing and sharing).

QWhat is the fundamental structural problem or risk associated with the embodied data business, as pointed out in the article?

AThe fundamental problem is that the demand for embodied data is not yet driven by a profitable, self-sustaining downstream robotics industry. The buyers (robot companies, AI teams) are largely unprofitable and use venture capital funding to purchase data. Therefore, the revenue of data companies is essentially a redistribution of downstream financing. If robot companies fail to commercialize, funding and data orders could disappear simultaneously.

QWhich company created the largest single funding record in China's embodied intelligence field according to the article?

AAccording to the article, Tashi Zhihang created the largest single funding record with a $455 million Pre-A round financing in April.

Пов'язані матеріали

Ripple Advances Full XRPL Stack Amid Expanding Tokenized Assets Market

Ripple is advancing its full-stack XRPL infrastructure amid the expanding tokenized asset market. Following strategic investments in ZILO and Licuido to broaden institutional capital market infrastructure on the XRP Ledger (XRPL), President Monica Long outlined the company's vision. She noted a shift from bank pilots to full-scale operations, with Ripple's digital asset suite now covering the entire lifecycle of tokenized assets, from issuance to utilization. The investments complement Ripple's capital markets strategy, which includes infrastructure for tokenizing funds and providing institutional liquidity. The recently launched Ripple Mint platform offers financial institutions a single interface to manage Ripple USD (RLUSD) and issue tokens across fiat and blockchain settlement systems. Additionally, Ripple's tokenization platform enables the creation of security tokens, stablecoins, fund shares, bonds, and other real-world asset (RWA) tokens. A partnership between DBS, Franklin Templeton, and Ripple demonstrates post-issuance portfolio management, allowing eligible clients to swap RLUSD for a tokenized money market fund. DBS is also exploring using these tokenized fund shares as collateral for repo agreements or lending platforms, enabling investors to switch between stable settlement assets and yield-bearing instruments. A proposed XRPL lending protocol aims to add standardized institutional lending to the network's tokenization infrastructure, managing on-chain servicing, repayments, and defaults. Separate infrastructure links XRP and RLUSD liquidity to tokenized U.S. Treasury products offering regulated yield. The Bank for International Settlements (BIS) has identified tokenization's potential for faster payments and more efficient financial intermediation, while calling for robust settlement tools and coordinated oversight. Ripple's next expansion phase hinges on validators' decisions regarding proposed lending standards, which will determine if XRPL tokenized assets can support institutional lending at the protocol level.

cryptonews.ru6 хв тому

Ripple Advances Full XRPL Stack Amid Expanding Tokenized Assets Market

cryptonews.ru6 хв тому

Торгівля

Спот

Популярні статті

Як купити DATA

Ласкаво просимо до HTX.com! Ми зробили покупку DATA Network (DATA) простою та зручною. Дотримуйтесь нашої покрокової інструкції, щоб розпочати свою криптовалютну подорож.Крок 1: Створіть обліковий запис на HTXВикористовуйте свою електронну пошту або номер телефону, щоб зареєструвати обліковий запис на HTX безплатно. Пройдіть безпроблемну реєстрацію й отримайте доступ до всіх функцій.ЗареєструватисьКрок 2: Перейдіть до розділу Купити крипту і виберіть спосіб оплатиКредитна/дебетова картка: використовуйте вашу картку Visa або Mastercard, щоб миттєво купити DATA Network (DATA).Баланс: використовуйте кошти з балансу вашого рахунку HTX для безперешкодної торгівлі.Треті особи: ми додали популярні способи оплати, такі як Google Pay та Apple Pay, щоб підвищити зручність.P2P: Торгуйте безпосередньо з іншими користувачами на HTX.Позабіржова торгівля (OTC): ми пропонуємо індивідуальні послуги та конкурентні обмінні курси для трейдерів.Крок 3: Зберігайте свої DATA Network (DATA)Після придбання DATA Network (DATA) збережіть його у своєму обліковому записі на HTX. Крім того, ви можете відправити його в інше місце за допомогою блокчейн-переказу або використовувати його для торгівлі іншими криптовалютами.Крок 4: Торгівля DATA Network (DATA)Легко торгуйте DATA Network (DATA) на спотовому ринку HTX. Просто увійдіть до свого облікового запису, виберіть торгову пару, укладайте угоди та спостерігайте за ними в режимі реального часу. Ми пропонуємо зручний досвід як для початківців, так і для досвідчених трейдерів.

355 переглядів усьогоОпубліковано 2026.07.01Оновлено 2026.07.01

Як купити DATA

Обговорення

Ласкаво просимо до спільноти HTX. Тут ви можете бути в курсі останніх подій розвитку платформи та отримати доступ до професійної ринкової інформації. Нижче представлені думки користувачів щодо ціни DATA (DATA).

活动图片