Unitree and Zhiyuan Share One Brain, Mysterious Model Demo Shocks the Scene, 10-Minute One-Take

marsbitPublicado em 2026-08-26Última atualização em 2026-08-26

Resumo

A 10-minute, unedited, single-take video has surfaced, demonstrating what appears to be a groundbreaking leap in embodied AI. The demo features two distinct robot models—Unitree and Zhiyuan—performing complex household tasks in a cluttered, real-world apartment. Crucially, both robots operate using the same underlying AI "brain," showcasing unprecedented cross-platform generalization. Key highlights include: * **Autonomous Long-Horizon Tasks:** The robots seamlessly complete multi-step chores like fetching tools, cleaning, and putting items away without error accumulation. * **Real-Time Reasoning and Tool Use:** A Unitree robot, unable to clean a window from inside, reasoned to extend its arm outside. Later, it attempted to use a box as a stepping stool to reach a high shelf, even kicking it into position when it couldn't bend down. * **Cross-Robot Collaboration:** The two robots cooperated without pre-programmed scripts. A Zhiyuan robot helped its Unitree counterpart by gently placing a scarf around its neck and later retrieving it from a high shelf. * **Interruption Recovery & Task Switching:** The robots smoothly paused their tasks upon an alarm signal, completed a new chore, and then resumed their original work. * **"Lazy" Efficiency:** The robots demonstrated intuitive efficiency, such as hanging a towel on their shoulder or lifting slippers by their straps to save effort. The demo suggests a fundamental departure from current approaches like VLA (Visual-Lang...

Today's embodied AI companies make their demo videos look as polished and exaggerated as commercials, but have you ever seen a completely unedited, genuine video?

A few days ago, a friend in the industry sent me a 10-minute demo, filmed entirely in one take, unedited, with no off-camera remote control or manual instructions. It's even a bit rough around the edges, but what it presented left me slumped in my chair, unable to calm down for a long time (doge).

In the video, a robot extends its arm out the window to clean glass; when it can't reach a high spot, it knows to move a box and stand on it; when a long-term task is interrupted, it can accurately resume...

Even more shocking, the video simultaneously features robot bodies from Unitree and Zhiyuan. These two "rival" products, with completely different hardware architectures, degrees of freedom, and sensor systems, share the same "brain," cooperating and collaborating closely, making this a globally rare sample of a cross-platform, general-purpose brain.

Not exaggerating, this might be a demo sufficient to rewrite the global embodied AI industry's understanding, overturn technical judgments in the field, and even potentially change Scaling Law...

Where on earth did this come from?

Breaking Down the Video Frame by Frame: Every Scene is Iconic

The information density in this video is simply too high. I carefully analyzed it frame by frame, and the more I watched, the more engrossed I became.

Working in Confined Spaces: Unbelievably Stable

The filming location is a roughly 15m² rental apartment, densely furnished with narrow aisles where even turning around is difficult. In this environment, teleoperation (no space for a person to stand) and preset scripts are basically impossible.

Just from the environment design, this video confronts you, saying: this can only be a video of a robot completely self-learning, self-evolving, and making autonomous decisions.

Two robots perform household chores in parallel within the same space, perceiving boundaries and planning paths in real-time.

There are no collisions, no getting lost, no stopping throughout, perfectly adapting to unfamiliar, complex real-world scenes. Even a slightly battered robot tethered by a leash showed incredible agility.

The entire video has no cuts, no retakes, and no human command intervention.

Reaching Out the Window to Clean Glass: The Pinnacle Embodiment of Dynamic Control and Autonomous Reasoning

From the very first task, the robot displayed a "delicate" and unexpectedly human-like side.

A Unitree robot uses a squeegee to clean glass. Anyone who has used a squeegee knows that if the force is too light it won't clean, if too heavy it makes strange noises—it's harder to control force and spatial awareness than using a cloth. It immediately chose the hard mode.

After wiping a few times, a spot wasn't clean. It actually judged by itself: the dirt might be on the outside.

So it performed a sequence of movements I've never seen before: turning sideways, leaning back, sticking its head out, extending its arm, and reaching the hand holding the squeegee out the window. (Steady as a rock)

Some existing embodied models excel at force control as their limit. But this robot seamlessly integrates precise spatial environment perception (the arm didn't hit the window frame) and dynamic calculations on top of accurate torque control.

Not only does it achieve top-tier performance in a single capability, but the model also shows a unified trinity of performance.

What left me unable to recover for a long time was that, when unable to clean it properly, it immediately reasoned and decided to reach out the window to clean the glass. This is the first time I've seen a model self-reason and self-evolve like a human.

This fully demonstrates the model's ability to fuse vision, touch, dynamics, etc., into unified perception and possess self-evolution capability. This time, I finally believe robots can really work, and will even do it more perfectly than humans.

Long-Sequence Task Closed-Loop: Rejecting Error Accumulation

After cleaning the glass, the Unitree robot smoothly places the squeegee back in its original spot, accurate and fluid.

Don't underestimate this finishing action. It means it completed the full closed-loop process of "fetch tool - work - store."

Meanwhile, the Zhiyuan robot nearby approaches the washing machine, takes out a cleaned cushion and places it behind the sofa, then returns to take out a cleaned doll, precisely placing it in the second layer of the wardrobe. Every item returns to where it belongs.

The Zhiyuan robot accidentally bumps into the glass, which further proves this is the robot's autonomous decision-making, because cameras sometimes can't solve glass reflections, whereas a human teleoperator could (doge).

These two scenes are a microcosm of the most basic abilities in the entire 10-minute video. In long-sequence tasks, every task is completed smoothly, accurately, and in one go, without accumulating errors and deteriorating as tasks become more numerous or longer.

This solves the traditional VLA model's biggest weakness: fear of long-sequence tasks.

Judging solely by the motion performance, the model must use a completely new architecture. Throughout the ultra-long 10-minute task chain, it continuously corrects errors, outputs stably, and every action is precise and controllable, with robustness ranking at the industry's top.

Tasks Can Be Interrupted at Will; the Model Can Resume from Checkpoints

Next, an alarm clock sound appears in the scene. (Listen carefully with the volume up, it's not obvious.)

At first, I didn't understand the purpose of the alarm. After repeated viewing, I realized the alarm is a pre-set reminder that interrupts the robots' current tasks, prompting them to start a special task of tidying the table and refrigerator.

The amazing part is, after tidying the table and refrigerator, the robots resume the tasks they were doing before the interruption. This shows the model possesses the ability to be interrupted at any time, resume from checkpoints, and recover tasks.

Compared to demos on the market that can only complete one task at a time and require careful handling, this rough video showcases robots' super-strong abilities in real work environments—not just durable, but even capable of "multitasking."

Next is a long sequence of household tidying: The Unitree robot fetches a storage bag and places it on the table. The Zhiyuan robot sees food on the refrigerator, judges it should be refrigerated, immediately acts, and incidentally takes out food that should be thawed (I guess this mysterious team might be hinting that robots will soon be able to cook a full meal).

Throughout, there are no step-by-step instructions. The robot autonomously decomposes tasks, plans actions, and step-by-step completes the full process from fetching to storing.

Currently, most robots can only execute single-task serial processes. If interrupted, the task basically collapses. But this model possesses human-level task priority judgment and dynamic scheduling capabilities, able to switch task flows at any time. The intelligence demonstrated by this flexibility and stability far exceeds any known model on the market.

Zhiyuan Tying a Scarf on Unitree: The Iconic Cross-Platform Collaboration Scene

Then, the climax of the entire video appears.

The two robots seem to have agreed: they must tidy all the clutter on the table in one go, without moving back and forth.

So the Unitree robot continuously places objects on itself until its hands are full and it doesn't know what to do next. At this moment, the Zhiyuan robot gently approaches, slowly picks up a scarf from the table, and hangs it around the Unitree robot's neck.

The Zhiyuan robot notices the scarf is long, so it holds it up with both hands. At this point, the Unitree robot seems to understand the Zhiyuan's intention, bends down, and the Zhiyuan robot drapes the scarf around the Unitree's head, hanging it on its neck.

Wait, bro?! The smoothness and the way they seem to exchange glances... is there even a hint of a CP vibe?!!

The video isn't pre-set with roles; it's a collaboration scheme autonomously explored by the model between two robots of different platforms. Two robots from different brands autonomously judge their respective capability boundaries and perfectly complement each other. This infinitely approaches human collaboration, even appearing more默契 without language.

This might be the world's first cross-platform interaction, and possibly the world's first model capable of generalizing across platforms.

I guess this mysterious team wants to use two competing platforms to tell everyone: this is the real general-purpose brain.

Hanging Towels, Carrying Slippers: Has the Robot Not Only Learned to Work, But Also Started to "Slack Off"?

After the Unitree robot, draped with various items, slowly walks away, the Zhiyuan continues tidying the remaining items.

At this point, it picks up a towel and attempts to hang it on itself. It tries once, twice, three times, and finally succeeds in hanging it over its shoulder.

Shocking triple realization!

The robot seems to know the easiest way: hanging a towel on the shoulder is less effort than holding it in hand.

The model can self-learn. After failing once or twice, it continuously optimizes until it succeeds.

The model can understand its own platform, which might be the core reason it can be applied across different platforms.

This isn't an isolated case.

Look carefully: when the robot tidies slippers, it carries them by the new slippers' hanging strap, not the slipper body itself, because carrying by the strap is easier.

Shocking again! While other robots are still carefully setting up to work, this model not only gets the job done but directly learns to "slack off"? That's just too smart.

Another Climax: The Robot Learns to Use Tools, Attempts to Stand on a Box to Reach Higher

Not finished yet. The scene that shocked me the most appears. While putting away the items hung all over its body, how to place the scarf into the third-layer cabinet stumps the 1.3m-tall Unitree G1 robot.

A god-like moment unfolds! It doesn't get stuck, nor does it give up. Instead, it reasons like a human and finds itself a box! Yes, finds a box! It attempts to stand on the box to raise itself and try again.

It finds a nearby box and pushes it to the ground, attempts to bend down to pick it up, but finds it can't bend down.

After several attempts, just when I thought it would finally give up, it astoundingly kicks the box towards the cabinet with its foot!

The whole process demonstrates:

Self-reasoning and decision-making ability. The robot understands how to reach high objects and how to move the box when it can't bend down.

The model's continuous exploration of its bodily capabilities. During the repeated attempts to bend down, the robot seems to increasingly understand its own physical boundaries and makes choices matching its bodily capabilities.

The process of self-evolution. Without any teaching, the robot actually learned to use its foot to kick the box, skillfully utilizing its own body to self-evolve in the environment.

This scene genuinely sent chills down my spine.

After all, the ability to autonomously use tools is a hallmark of humanity.

And the robot in the video actually learned to use tools. It knows a box can bear weight, knows standing on it increases height, and even knows to use its foot when it can't bend down.

Behind this set of actions lies the trinity of deep understanding of physical world rules, continuous exploration of self-capabilities, and autonomous decision-making evolution.

Cross-Platform Capability Complementarity, Cooperative Work: From Individual to Collective

The Unitree kicks the box over, but it goes askew. Just as it's thinking how to straighten it, the 1.7m-tall Zhiyuan Expedition A3 walks over. It seems to perceive the Unitree's determination and predicament, puts down the small cart it was holding (actively switching tasks), and chooses to help its companion.

Then, like a graduation ceremony, the Unitree bends down and lowers its head. The Zhiyuan slowly takes off the scarf and places it on the shelf.

Every step in the detail demonstrates capabilities that look down upon other models.

The two robots' understanding of their own and each other's hardware capabilities and cooperation transcends individual intelligence, moving towards collective intelligence.

The Zhiyuan loops the scarf around the back of the Unitree's head to gently remove it (imagine how other robots might just yank the scarf off), showcasing an ultimate understanding of force control and spatial perception.

The robot folds the scarf three times, an autonomous judgment on how to better place it in the wardrobe. This is true embodied "intelligence."

Finally, after helping the Unitree, the Zhiyuan quietly leaves, placing the last piece of clothing steadily into the washing machine (I guess it judged the clothes draped on the washer were dirty). The ten-minute one-take finally ends.

Although I don't know this model's architecture, I believe you, like me, are completely stunned.

The first cross-platform collaboration, the first time seeing a robot's autonomous evolution, the first human-like autonomous decision-making choosing optimal paths, extreme spatial perception and force control, tasks interruptible at will, learning to use tools... Every single action, every frame of this model could be a great demo on its own. But they just quietly appear in this ten-minute one-take video, presented in such a simple, direct manner.

It feels like a relaxed, effortless, almost casual display of king-like prowess.

Underlying Technology Jumps Out of Mainstream Model Frameworks

Why can the robots in the video deliver so many iconic scenes?

I learned from insiders that the most important reason is that this model jumps out of all existing frameworks of current mainstream embodied models.

Current mainstream embodied models can roughly be divided into three categories: VLA, WAM, and traditional world models. They all have bottlenecks that are temporarily difficult to break through.

△Image generated by AI

Specifically, VLA is the "data intuitionist," relying on massive visual, language, and action data for end-to-end fitting. But its essence is "guessing actions from images," lacking physical reasoning ability. Long-sequence tasks continuously accumulate errors, generalization to unfamiliar scenes is nearly zero, and it's highly tied to specific platforms.

WAM and traditional world models are the "prediction idealists," trying to model world rules first and then plan actions, but suffer from a fatal "knowledge-action gap." The model can understand scenes and predict states, but when facing practical scenarios requiring physical dynamic reasoning, it still relies on training data for deduction, unable to adapt to dynamic variables in real environments.

These mainstream embodied models share a common fatal underlying flaw: data-driven task fitting.

In other words, all actions are "probability guesses" based on massive data, not deep understanding of physical rules. This caps the model's upper limit, tightly locked by training data.

Reportedly, the model in the video was trained with only a few dozen hours of video data. Achieving such astonishing effects with such a small amount of data—does Scaling Law still hold? A big question mark is needed here. (Foreshadowing a US stock market plunge)

Judging from the details, the model in the video exhibits at least four capabilities currently unattainable by VLA and world models:

Dynamics Modeling: Trajectories satisfy dynamic feasibility, can calculate compensation torque and feedforward terms, with excellent extrapolation generalization ability;

Self-Cognition Ability: Unlike VLA's conditioned reflexes or WAM's temporal statistics, it thinks like a human: "what is the next action that brings me closest to the goal?";

Adaptive Ability: Can supplement causal reasoning, solve long-tail problems, evolving while acting;

Self-Evolution Ability: Skills self-grow, strategies self-evolve during task completion, as if capable of self-driven generation of learning objectives and actively reasoning to achieve value.

Supporting all this are three underlying technological cores.

Physically Constrained Dynamic Learning. Different from VLA and WAM's purely data-driven training modes, its core is physical-rule-first, dynamics-prediction-driven. Carrying slippers by the strap, hanging towels on the shoulder, standing on a box to reach higher—none were taught by data. They are optimal solutions the model autonomously explored through environmental interaction.

Cross-Platform Unified Modeling. Through a unified action representation space and platform-adaptive mechanisms, the same model adapts to hardware platforms of different brands and architectures. Equivalent to achieving "hardware-software decoupling" for embodied AI, ending the era of hardware-bound algorithms.

Long-Sequence Robust Closed-Loop. Relying on a global task scheduling framework, autonomously handles full-process sudden disturbances, action errors, and task switches, avoiding the traditional model's collapse from long-process error accumulation. Possesses core capabilities for real-world household and complex scene deployment.

If you understand the current state of embodied models, after reading the above, you'll probably also be jaw-dropped.

While other teams' demos are still piling up short-duration, polished, single-task videos, this mysterious video has already completed a 10-minute, unedited, cross-brand, multi-collaborative, full-process, self-evolving real-world deployment.

While other models are still "guessing actions based on data," this model seems to have achieved complete autonomous decision-making, self-learning, and self-evolution. It lets everyone see that robots can not only truly work, but their capabilities are rapidly self-evolving, giving infinite room for imagination.

More importantly, this model seems to show us that the future of robots truly capable of replacing human labor—the one everyone is heatedly debating will arrive in three, five, or ten years—appears before your eyes today.

A team from who-knows-where seems to have overturned all technical paths, overturned Scaling Law, created true embodied "intelligence," approaching the ChatGPT moment for embodied AI.

It's still unknown who this mysterious model comes from, but I believe at this moment, everyone's question is the same as mine: Who exactly is it!!!

Does any expert have insider information?!!!

This article is from the WeChat public account "QbitAI" (ID: QbitAI), author: Noah

Criptomoedas em alta

Perguntas relacionadas

QWhat was so impressive about the 10-minute unedited demo video featuring Unitree and Zhihu robots?

AThe video was a continuous, unedited, 10-minute shot showcasing two different robot models (Unitree and Zhiyuan) working collaboratively in a complex, cluttered room without any teleoperation, pre-scripted actions, or human intervention. They performed long, sequential tasks with interruption recovery, used tools autonomously, and demonstrated cross-platform coordination using a single shared 'brain' or model.

QWhat specific abilities did the robots demonstrate that are considered groundbreaking?

AThe robots demonstrated several groundbreaking abilities: autonomous tool use (e.g., using a box as a stepping stool), complex task interruption and recovery, collaborative problem-solving between different hardware platforms, self-optimization of movements (like carrying items in a more efficient way), and sophisticated physical reasoning like cleaning the outside of a window by reaching out.

QHow does the underlying technology of this model differ from mainstream embodied AI models like VLA or World Models?

AThe model reportedly differs by moving beyond pure data-driven task fitting. It incorporates physics-constrained dynamic learning for real-world reasoning, a unified modeling framework for cross-platform hardware adaptation, and a robust long-sequence task scheduler. This allows for generalization, self-evolution, and handling of dynamic real-world variables, unlike models that suffer from error accumulation in long tasks or a disconnect between prediction and physical action.

QWhat was the significance of the robots (Unitree and Zhiyuan) working together?

AIt was significant because Unitree and Zhiyuan robots are typically competitors with different hardware architectures, degrees of freedom, and sensor systems. Their seamless collaboration using a single, shared AI 'brain' demonstrated a level of cross-platform generalization and hardware-software decoupling that is rare and suggests a move towards a truly universal embodied intelligence model.

QWhy did the article suggest this demo might challenge the 'Scaling Law' in AI?

AThe article suggested it might challenge the Scaling Law because the model allegedly achieved its advanced capabilities with only 'tens of hours' of video training data. This is a minuscule amount compared to the massive datasets typically required by large models, implying its performance isn't solely dependent on scaling data size, which is a core tenet of the Scaling Law hypothesis.

Leituras Relacionadas

Why is XRP in the red now after a 50% weekly surge?

XRP fell 2.2% over 24 hours after a 49.4% surge in the previous week, moving back into negative territory. The drop coincides with major banks advancing blockchain-based payment solutions that address the need for prefunding in cross-border transfers, potentially competing with XRP's use case. JPMorgan Chase expanded its blockchain settlement system, Kinexys, to eight currencies, allowing clients to move and exchange funds around the clock without needing a separate crypto asset. Similarly, Citigroup operates a round-the-clock USD clearing network and offers Citi Token Services for tokenized deposits, aiming to speed up payments while reducing the amount of capital required upfront. While its 90-second settlement is slower than XRP Ledger's 3-5 seconds, using cash already held in a regulated bank may be more critical for companies. Furthermore, SWIFT has facilitated interoperability between different bank-issued tokenized deposits without a common cryptocurrency. In a recent pilot, HSBC and Standard Chartered completed a cross-border transaction using SWIFT's ledger to coordinate and settle obligations between their separate token systems. SWIFT reports that 17 banks across six continents are preparing for real transactions using this model. These developments in traditional finance present alternative, bank-integrated pathways for instant, cross-border value transfer, potentially impacting the demand and price trajectory for XRP.

cryptonews.ruHá 10m

Why is XRP in the red now after a 50% weekly surge?

cryptonews.ruHá 10m

The Ambition of the OEM: Dissecting Leju Robot's External Investments to Understand the 'Ecosystem Positioning' Battle in the Humanoid Robot Track

Leju Robotics, a leading humanoid robot manufacturer from China, is aggressively building an industrial ecosystem through strategic investments as it pursues a public listing. Founded in 2016 and known for its "Kuavo" robot, Leju saw its revenue surge to 258 million yuan in 2025, though it remains unprofitable. Its IPO plan, submitted in May 2026, outlines a strategy focused on cost reduction and full-stack capability development. Analysis of Leju's 13 disclosed investments reveals a three-pronged approach to securing its value chain. First, it targets core components like precision actuators (e.g., Lingxinqiaoshou), joint modules, and motors to control costs and supply. Second, it invests in software and AI layers, including companies working on foundational models and embodied intelligence datasets, to develop an independent "brain" while maintaining partnerships with giants like Huawei. Third, it forms alliances with firms in logistics and service robotics to transform from a hardware seller into a solution provider for real-world applications. While this investment network aims to create a competitive moat and support Leju's industrialization story, risks persist. The company faces ongoing losses, declining gross margins, and potential industry price wars. Many investments are early-stage with long payback periods and unproven synergistic benefits. Leju's strategy exemplifies how humanoid robot players are moving beyond pure R&D to ecosystem competition, betting that controlling the "hard work" of the supply chain will be the ultimate barrier to entry.

marsbitHá 22m

The Ambition of the OEM: Dissecting Leju Robot's External Investments to Understand the 'Ecosystem Positioning' Battle in the Humanoid Robot Track

marsbitHá 22m

Trading

Spot

Artigos em Destaque

Como comprar ONE

Bem-vindo à HTX.com!Tornámos a compra de Harmony (ONE) simples e conveniente.Segue o nosso guia passo a passo para iniciar a tua jornada no mundo das criptos.Passo 1: cria a tua conta HTXUtiliza o teu e-mail ou número de telefone para te inscreveres numa conta gratuita na HTX.Desfruta de um processo de inscrição sem complicações e desbloqueia todas as funcionalidades.Obter a minha contaPasso 2: vai para Comprar Cripto e escolhe o teu método de pagamentoCartão de crédito/débito: usa o teu visa ou mastercard para comprar Harmony (ONE) instantaneamente.Saldo: usa os fundos da tua conta HTX para transacionar sem problemas.Terceiros: adicionamos métodos de pagamento populares, como Google Pay e Apple Pay, para aumentar a conveniência.P2P: transaciona diretamente com outros utilizadores na HTX.Mercado de balcão (OTC): oferecemos serviços personalizados e taxas de câmbio competitivas para os traders.Passo 3: armazena teu Harmony (ONE)Depois de comprar o teu Harmony (ONE), armazena-o na tua conta HTX.Alternativamente, podes enviá-lo para outro lugar através de transferência blockchain ou usá-lo para transacionar outras criptomoedas.Passo 4: transaciona Harmony (ONE)Transaciona facilmente Harmony (ONE) no mercado à vista da HTX.Acede simplesmente à tua conta, seleciona o teu par de trading, executa as tuas transações e monitoriza em tempo real.Oferecemos uma experiência de fácil utilização tanto para principiantes como para traders experientes.

613 Visualizações TotaisPublicado em {updateTime}Atualizado em 2026.06.02

Como comprar ONE

Discussões

Bem-vindo à Comunidade HTX. Aqui, pode manter-se informado sobre os mais recentes desenvolvimentos da plataforma e obter acesso a análises profissionais de mercado. As opiniões dos utilizadores sobre o preço de ONE (ONE) são apresentadas abaixo.

活动图片