Not Being a Follower of Silicon Valley: A Few Young PhD Students Bet on the Integrated Brain of Bipedal Humanoid Robots

marsbit2026-08-24 tarihinde yayınlandı2026-08-24 tarihinde güncellendi

Özet

This article profiles GongShengZhiXing, a startup founded by 1996-born Ding Pengxiang with a core team of Ph.D. students, all in their twenties. They are developing an end-to-end, integrated "brain" (base model) for bipedal humanoid robots, focusing on full-body dynamic coordination rather than the industry's more common modular, layered approach. The team's stance is that for humanoids to integrate into human environments, an end-to-end model—where a single system handles perception, decision-making, and control directly—is ultimately more efficient and scalable than layered architectures that separate high-level "brain" and low-level "cerebellum" functions. They argue layered systems create information bottlenecks. A key demo shows their robot driving a go-kart at high speed, demonstrating synchronized hand-eye-foot coordination and multi-contact balance—a complex task highlighting their focus on whole-body dynamics over simple kinematics. Technically, their model outputs commands at the joint level and incorporates a "Motion Expert" module and techniques like DriftDistill to simultaneously learn task completion and physical stability from data. Acknowledging challenges like severe data scarcity for full-body mobile manipulation, they employ methods like TrajBooster to synthesize training data. Ding Pengxiang believes true "emergence" of new skills at the joint-movement level is the goal, but commercial viability likely requires scaling to millions of training hours. The...

While browsing social media at home over the weekend, a video truly impressed me.

A robot driving a go-kart sped along the track, navigating corners, accelerating, and steering—all captured in a single continuous shot.

I didn't expect our robots could already be showing off their skills on a go-kart track!

Wait a minute?

This isn't the same style I saw at the WRC exhibition!

The robots at the exhibition were walking, with sparks and lightning effects. When did the robots in the video secretly sign up for a racing training course behind our backs?

It drove like a complete pro, handling high-speed cornering and agile obstacle avoidance smoothly in one go.

Compared to previous robot demos where they slowly walked over and bent down to pick something up, the robot in this video actually got into the car by itself, achieving hand-eye-foot synchronous coordination and accomplishing multi-contact-point balancing and precise force control at full speed.

Such full-body coordination of a bipedal humanoid robot—where its eyes, hands, and feet work together to perform detailed operations during high-speed motion and complex posture changes—has rarely been seen in previous humanoid robot demonstrations.

The company that released this video is called GongShengZhiXing (Symbiotic Cognition).

What GongShengZhiXing is developing is an end-to-end sensory-motor integrated "brain" for bipedal humanoid robots—the foundational model.

The founder, Ding Pengxiang, was born in 1996. The rest of the core team members are all born after 2000, are all current PhD students, and the company was founded just two months ago.

Regarding this seemingly "unorthodox" demo, Ding Pengxiang explained:

I want robots to be more like humans, not just perform some very mechanical service functions. For example, if a human can drive a go-kart, a robot should be able to do it too.

A question naturally follows: Why would a group of young people who haven't even received their graduation certificates dare to tackle the hardest and most cutting-edge track in embodied intelligence?

Betting on the Hardest Path: Moving Beyond Hierarchical Research to Challenge End-to-End

How new is the track of "end-to-end" bipedal humanoid robot models?

So new that there are no ready-made talents available in the market.

According to Ding Pengxiang, this direction only started gaining traction in the last couple of years. Those who understand the technology are still PhD students, making it almost impossible for them to find people who meet their needs from the existing talent pool.

Those who graduated years ago can't keep up with the pace, and those who haven't reached the PhD level lack a deep enough understanding of the fundamentals. The most cutting-edge breakthroughs are precisely in the hands of this batch of top PhD students.

The team at GongShengZhiXing is exactly such a group of young and cutting-edge individuals.

They have won the most prestigious academic honors in the domestic embodied intelligence field—two Best Paper awards, with one selected as a Best Paper Candidate. They have also built the first domestic embodied foundational model whose GitHub Stars exceeded 2K.

More importantly, they are not followers who jumped on the bandwagon after the trend arrived. They are among the earliest pioneers in China exploring embodied foundational models and are the most steadfast long-termists in this track. The team has collectively published over 40 top-tier conference papers, with their technical footprint covering the full stack of perception, decision-making, control, data, and systems. A series of industry-first works, from dual systems and lightweighting to multi-configuration, all originated from this young team.

Their judgment is clear: The ultimate and most widespread form for robots to enter human life is the bipedal form.

Following the evolution paths of autonomous driving and large models, end-to-end is the most efficient technical paradigm.

This judgment stems from their firsthand research experience.

As early as December 2023, Ding Pengxiang began equipping quadruped robots with "brains."

He released QUAR-VLA as the first author, which was the first VLA (Vision-Language-Action) task paradigm for quadruped robots. It allowed the upper-level model (the brain) to understand vision, language, and task intent, then handed the decisions to the underlying motion system (the cerebellum) for execution.

This layered approach—"the brain manages decisions, the cerebellum manages motion execution"—is what the industry calls the hierarchical scheme, a method still used by companies like Figure today.

During subsequent research, he gradually discovered issues.

He describes the hierarchical approach as a "theoretically feasible but less elegant solution":

The brain and cerebellum work separately. For every new task, the upstream output needs custom fine-tuning to achieve good performance.

Deeper, there's an inherent structural defect.

The brain and cerebellum are trained separately and combined during deployment. Even if each part is optimized locally, it doesn't guarantee global optimality.

Every translation between the two layers adds cascading errors and information loss.

Ding Pengxiang concludes, "As long as it's hierarchical, there is an information bottleneck in the middle. This interface design determines that hierarchical architecture cannot achieve true Scaling."

This conclusion also comes from his long-term observation of the development of autonomous driving and large models.

Early autonomous driving relied on the coordination of multiple modules like perception, prediction, planning, and control. Later, it began exploring models that directly learn driving capabilities from data.

After Tesla's FSD switched to an end-to-end approach, its performance improved significantly.

In Ding Pengxiang's view, robots may undergo a similar transformation, allowing models to learn the complete "perception-understanding-action" process from data.

His logic is, having thoroughly explored the hierarchical path repeatedly since 2023, if it's judged that future data scaling will ultimately push systems toward end-to-end, there's no need to take another detour.

This shift is happening in 2026.

Google's Gemini Robotics 2 already uses a single policy to unify full-body actions from feet to fingertips into one model.

Yet, the consensus at the RSS 2026 top conference clearly states that a single end-to-end model cannot cover the complex full-body dynamics, and hierarchical modularization is the optimal practical solution at the current stage.

The pure end-to-end path that GongShengZhiXing is betting on is a more radical and less-trodden one.

Why bipedal specifically? Ding Pengxiang's answer is first principles: The buildings, tools, and environments of human society are all built according to the human body's configuration.

Wheeled robots can only work fixed in production lines and shopping malls. Bipedal robots can go outside, climb stairs, drive cars, and perform different tasks across scenarios.

Versatility means cost amortization. Bipedal forms also offer an anthropomorphic sense of affinity.

Timing is equally crucial. Before April 2026, there wasn't even a general teleoperation model for bipedal humanoids. Without a data source, you couldn't train a foundational model.

It wasn't until NVIDIA open-sourced Sonic that the industry gained the cornerstone for large-scale data production.

This direction is no longer in the early "want to train but have no data" stage, yet it hasn't matured enough for the technical route to fully converge. GongShengZhiXing is gambling on this exact window in between.

The Technological Watershed: Others Learn Kinematics, They Tackle Dynamics

First, let's clarify a watershed: Most current robot foundational models actually learn kinematics.

Kinematics concerns moving from point A to point B—where the hand moves and whether it grasps. This works fine for fixed-base robotic arms because they don't fall over.

When a bipedal robot reaches out to grab something, its body's center of gravity changes simultaneously.

When squatting to pick up an object, the waist must lean forward, and the ankle joints and legs must redistribute force to maintain balance.

Add friction, collisions, inertia, and contact into the mix, and the model is now facing a full-body dynamics system.

According to Ding Pengxiang, a g1 bipedal humanoid has 29 degrees of freedom, which is far more complex than the 7 degrees of freedom of a robotic arm.

In the past, robots learned "motion trajectories." Humanoid robots need to learn "how the body acts in the physical world."

Moving from kinematics to dynamics, from completing tasks to maintaining stability—this is the real technological watershed for bipedal robot models.

This explains why GongShengZhiXing pushes its model all the way to the Joint Target level.

In hierarchical schemes, the brain first outputs kinematic targets, and the downstream cerebellum then calculates them into joint movements.

GongShengZhiXing removes this intermediate step, making the model directly face the body's state with dozens of joints.

A question arises: With the cerebellum removed, can the robot still stand stable?

Relying solely on imitation learning, the model only learns standard actions. It hasn't seen various stumbling body states during execution, so encountering unfamiliar states might lead directly to a fall.

What it lacks is the ability to recover from failures and stand firm.

The core of GongShengZhiXing's approach is a dual-domain collaborative optimization mechanism called "Task Behavior Modeling — Motion Prior Distillation."

One optimization pathway continues with behavioral cloning, ensuring task accuracy and motion fitting capability. The other, via DriftDistill, transforms the stability, disturbance resistance, and Failure Recovery capabilities accumulated by the underlying controller into the endogenous motion priors of the unified model.

It's not simply adding two Losses. Instead, it integrates the two types of capabilities—"accurately completing tasks" and "stably controlling the body"—into the same model, allowing the large model to simultaneously acquire task intelligence and body intelligence.

If DriftDistill can continue to scale with the model and data volume, it attempts to solve an even bigger problem:

Beyond cognitive Scaling, can a robot also achieve Scaling in its motion control capabilities.

Ding Pengxiang refers to this motion capability aggregation module as the Motion Expert Model, currently sized close to 1B parameters, continuously absorbing capabilities from different motion control models through distillation.

Taking another step forward, the end-to-end large model must also solve the problem of the robot's output force.

Many robot systems today primarily control position. But when a robot truly enters the real world, correct position doesn't guarantee task completion.

A hand reaches the drawer handle but doesn't know how much force to apply—the door still won't open. When handing something to a person, the position is correct, but excessive force is equally unsafe.

GongShengZhiXing's technical route is to first use the Force Expert as a safe contact specialist, letting it learn force application, compliance, impedance, and contact feedback. Then, through policy distillation, it is gradually integrated into the Motion Expert and ultimately into the end-to-end model.

Position determines whether the robot can "get there," while force determines whether it can truly "get the job done well."

This is also the meaning of GongShengZhiXing's "Full-Body Physical World Model": to make the model understand how the body is subjected to forces, loses balance, and makes contact, then directly translate these constraints into actions.

Data, Emergence, and a Calm Timeline

Once model training truly begins, the first major hurdle is the data shortage.

Globally, compliant robot data amounts to only about 500,000 hours as of early 2026, less than one twenty-thousandth of that used for large language models.

The awkward part is that much of this data captures stationary operations—robots working in place within cubicles—offering limited help for full-body mobile manipulation.

GongShengZhiXing's Blog has released around 10,000+ hours of data, with several thousand hours collected by themselves.

What full-body coordination models truly lack is strongly-coupled full-body mobile operation data from real-life scenarios, such as riding a bicycle, driving a go-kart, inflating a basketball, etc.

Their solution is TrajBooster: Extract end-effector trajectories from robotic arms or wheeled robots, have the humanoid robot track this trajectory in simulation while maintaining balance, unify actions from different sources into the same space, reuse the task diversity already defined for robotic arms, and provide low-cost data for pre-training.

When discussing model capabilities, Ding Pengxiang rarely uses the term "generalization."

In his words, "That term isn't very substantive; it's as empty as saying someone is good."

What he wants is "Emergence": At the joint angle level, for instance, if the model is trained on walking and jumping data, it might combine these to produce the new skill of running during testing.

Because only at the joint level can partial motion capabilities potentially combine to form entirely new motions.

Regarding the timeline, his judgment is quite sober.

He infers that around 100,000 hours of data will likely still yield more lab-level demos. Truly meaningful large-scale commercial significance might only emerge after reaching the million-hour mark.

Right now, the industry hasn't even converged on the most suitable data form; routes like EMG, exoskeletons, motion capture, and Pico teleoperation are still competing.

In his view, the three most difficult problems for humanoid robots currently are: first, how perception and control should truly be combined; second, what type of data should be used; and third is the data scale itself.

Ding Pengxiang hopes that one day, a robot model architecture designed by a Chinese team can also be widely adopted by peers in North America.

We don't want to be followers; we want to lead technological development.

Let's return to that go-kart.

What enabled it to navigate corners at full speed was a group of young people who haven't even graduated yet.

They are betting on the hardest and most cutting-edge direction in embodied intelligence.

They are gambling that end-to-end will eventually replace hierarchical approaches.

The industry is far from being able to prove that the final route has been validated. Even Ding Pengxiang himself admits that the entire bipedal foundational model industry has not truly converged yet.

But sometimes, the cutting-edge nature of a startup team lies not in executing a mature answer faster, but in whether it can see the next truly important problem earlier than its peers and dare to start solving it when the answer is still unclear.

This article is from the WeChat public account "QbitAI" (ID: QbitAI), author: Qiao Buchi.

İlgili Sorular

QWhat is the core research focus of the company Symbio Knows (共生知行) as described in the article?

ASymbio Knows focuses on developing an end-to-end perception-control integrated 'brain' (or base model) for bipedal humanoid robots.

QAccording to Ding Pengxiang, what is the main limitation of the traditional hierarchical architecture ('brain' and 'cerebellum') for robot control?

AThe main limitation is the presence of an information bottleneck and cascading errors between the layers. This interface design makes the hierarchical architecture fundamentally incapable of achieving true scaling, as each new task requires custom fine-tuning, and the separate training of modules cannot guarantee overall optimal performance.

QWhat key technical challenge distinguishes bipedal humanoid robot models from other robotic base models, according to the article?

AThe key challenge is moving from kinematics (concerned with movement from point A to B) to dynamics. Bipedal robots must learn how the entire body acts within the physical world, managing changes in center of gravity, force distribution, balance, and physical interactions like friction and inertia during whole-body coordinated movements, which is far more complex than controlling a fixed-base robotic arm.

QWhat are the two core components of Symbio Knows's 'dual-domain collaborative optimization mechanism' for training their model?

AThe two components are: 1) Task Behavior Modeling, which ensures task accuracy and action fitting through behavioral cloning, and 2) Motion Prior Distillation (specifically via DriftDistill), which transfers stability, anti-disturbance, and failure recovery capabilities from underlying controllers into the model's innate motion priors.

QWhat does Ding Pengxiang identify as the current major bottlenecks for advancing humanoid robot models, beyond just data scale?

AHe identifies three major difficult problems, listed in order: 1) How perception and control should be combined, 2) What type of data format is most suitable (with various methods like electromyography, exoskeletons, motion capture still competing), and 3) The scale of the data itself.

İlgili Okumalar

The 'Saving U.S. Treasuries' Baton Pass: Bessent Fumbled Last Week, This Week It's Wash's Turn

"Rescuing US Treasuries" Relay: After Bessent's Miss, All Eyes Are on Walsh Last week, US Treasury Secretary Bessent's announcement to at least double long-term Treasury buybacks failed to sustainably lower yields, which quickly rebounded. The market response saw a drop in the dollar alongside surges in gold and Bitcoin, interpreted as a "pressure release valve" for anxiety. The focus now shifts to Fed Chairman Walsh's upcoming Jackson Hole speech. Markets are highly sensitive to his message, seeking clarity on the Fed's policy response to stubborn inflation and worsening fiscal conditions. Analysts warn that a lack of new guidance could disappoint markets and worsen the sell-off in long-dated bonds. Analysts question the scale of Bessent's operations, noting they are too small relative to the overall debt market and do not constitute quantitative easing. A key issue is the Fed's massive holdings of long-term bonds, which distorts the market. With the Fed holding low-yielding short-term bonds that are losing money relative to its policy rate, discussion is growing around a potential Fed-led "Operation Twist." This would involve selling short-term bonds to buy long-term ones, aiming to lower long-end yields without expanding the balance sheet. The upcoming PCE inflation data will set the stage for Walsh's speech. However, the window for action is narrowing amid political pressures. A critical threshold is the 30-year yield at 5%; holding above it could increase stress on the dollar and leveraged sectors. Overall, the article suggests that without coordinated Fed action to anchor inflation expectations, Treasury interventions may ultimately fail, with investors increasingly looking to assets like gold as hedges.

marsbit13 dk önce

The 'Saving U.S. Treasuries' Baton Pass: Bessent Fumbled Last Week, This Week It's Wash's Turn

marsbit13 dk önce

Hyperliquid's Compliance Journey: From Permissionless to Permissioned via HIP-3

Hyperliquid’s Compliance Path: From Permissionless to Permissioned HIP-3 Hyperliquid currently blocks U.S. access because its permissionless, on-chain infrastructure conflicts with U.S. market structure laws, which restrict futures trading to registered exchanges, clearinghouses, and brokers. Through its Hyperliquid Policy Center (HPC), the project is advocating for regulatory modernization, proposing that regulated entities be allowed to build products on HyperCore (its exchange and clearing layer) while fulfilling their compliance obligations. The platform’s modular stack separates roles like a traditional exchange (DCM), clearinghouse (DCO), and broker (FCM), but reconstructs them on-chain with code. This enables permissionless access, self-custody, and 24/7 global trading, but clashes with U.S. rules requiring KYC, specific margin models, and custodial arrangements. To resolve this, HPC is engaging with U.S. regulators (CFTC, SEC) to seek clarity that deploying on-chain software does not itself trigger licensing, and to establish exemptions allowing non-custodial wallets to route users to regulated derivatives. Recent political signals suggest openness to this approach. On the technical side, Hyperliquid Labs has introduced permissioned HIP-3 deployers on testnet. These allow regulated entities to launch markets, perform KYC, and whitelist compliant users. While these create separate order books, whitelisted market makers can bridge liquidity between them, ensuring deep, shared liquidity across the same L1. Features like payload-based “PA” permissions enable DEX-level account controls (e.g., reduce-only orders), mirroring traditional broker authorities. The strategy is not to open the native, permissionless front-end to U.S. users, but to position Hyperliquid as neutral infrastructure that U.S. regulated firms can use while meeting their legal duties. This paves a compliant path for U.S. investor access while preserving the protocol’s core, permissionless nature.

marsbit37 dk önce

Hyperliquid's Compliance Journey: From Permissionless to Permissioned via HIP-3

marsbit37 dk önce

Two Funding Rounds in Three Months: The Chinese Version of Palantir is on Fire

Investment Community AI has learned that Beijing Zhongshu Ruizhi Technology Co., Ltd., a domestic industrial-grade causal intelligence and high-reliability decision-making AI company, has recently completed a strategic financing round worth hundreds of millions of RMB. This round saw participation from China Internet Investment Fund, Suzhou Chuangtou National Social Security Fund, Financial Street Capital, ICBC Capital, Kunlun Capital, among others, with existing shareholders also increasing their investment. This follows a Series B funding round in the hundreds of millions completed just three months prior. The rapid succession of two major funding rounds signifies strong market recognition of the company's underlying original technology and scaled commercial implementation. Often referred to as the "Chinese version of Palantir," Zhongshu Ruizhi is entering a new phase of accelerated technological iteration, widespread scenario replication, and scaled performance release, mirroring the explosive growth of China's AI market. Founded in April 2020 by Dr. Han Han, a Tsinghua University Ph.D. and former core drafter of national AI policies, the company is mission-driven to "move AI from the digital world to the physical world." It focuses on the high-reliability, strong-decision industrial AI track and enterprise-grade AI Agent full-stack infrastructure. The team tackles the challenge of applying AI to China's vast and complex industrial and energy systems by developing a new intelligent operating system from scratch. Its core technological breakthrough lies in three proprietary底层 technologies: meta-causal cognitive theory, causal models, and a dynamic ontology engine. These address critical pain points of generative large models in industrial settings—such as AI hallucinations, insufficient reasoning, lack of temporal logic, unverifiable decisions, and multi-source rule conflicts—thereby providing trustworthy, explainable, and executable智能决策 capabilities. Commercially, Zhongshu Ruizhi has achieved scaled deployment, serving over 50 central state-owned enterprises and industrial groups in sectors like power, petroleum, and aerospace, with implementations in more than 800 highly complex production scenarios. The company reported doubled revenue in 2025, demonstrating strong self-sufficiency and a viable business model—a rarity among new-generation AI firms. The latest funds will be allocated towards advancing foundational theoretical research, replicating successful application models to expand market presence (including overseas), and attracting top-tier talent. Lead investor China Internet Investment Fund highlighted that in the current shift from general AI capability contests to deep industrial empowerment, industrial-grade causal intelligence is crucial for building China's modern digital foundation and fostering new quality productive forces. They expressed support for the company's efforts to define decision-making paradigms and trustworthy standards for industrial intelligence, aiming to secure a rule-making voice in the global physical AI arena.

marsbit47 dk önce

Two Funding Rounds in Three Months: The Chinese Version of Palantir is on Fire

marsbit47 dk önce

The Biggest Political Economy Question in the AI Era: As Robots Become More Capable, How Do Humans Share the Value?

In the AI era, the most pressing political economy question is: as machines become increasingly capable, how can humanity share in the value they create? An article originally critiquing China's tech focus has sparked a deeper debate on this global challenge. Historically, industrial progress improved efficiency but still relied on human labor for wealth creation and distribution. AI is fundamentally different—it is now replacing cognitive and knowledge work. As AI and robots take over more tasks, economic growth may continue while direct human participation in value creation shrinks, creating a core tension between productivity gains and widespread income generation. The issue is not unique to China. While leading tech companies amass enormous wealth, labor's share of income is declining globally. The core problem is a broken link: technological innovation and corporate profits are not translating into sufficient consumer income and demand. Three potential paths forward are outlined: a traditional capitalist model where profits primarily go to capital owners; a state-capitalist approach with public investment in AI; and more innovative models like digital sovereign wealth funds, universal shareholding, or AI-era basic income schemes to directly distribute AI-generated value. The future competitive advantage may lie not just in technological supremacy, but in which society can build a new, inclusive distribution system for the intelligent economy. The ultimate challenge is ensuring that as AI creates value, humans have a means to obtain income and share in the resulting widespread social benefits.

marsbit58 dk önce

The Biggest Political Economy Question in the AI Era: As Robots Become More Capable, How Do Humans Share the Value?

marsbit58 dk önce

Generating Profits for Seven Consecutive Quarters, Emerging Markets Carry Trade Outperforms Everything

For the seventh consecutive quarter, dollar-funded emerging market carry trades have delivered positive returns, marking the longest winning streak since 2008. According to Bloomberg's index, this strategy has gained approximately 22% since late 2024, outperforming U.S. Treasuries, emerging market sovereign, and corporate dollar debt. The core of the trade involves borrowing low-interest currencies like the U.S. dollar, euro, or yen to invest in high-yielding emerging market assets, such as Turkish lira bonds offering over 40% returns. Returns were amplified by favorable currency moves, with the dollar weakening against most emerging market currencies and other traditional funding currencies. For instance, the trade gained 48% on the Colombian peso in the past year. A key test came in August 2024 with a historic joint U.S.-Japan currency intervention, which caused only a modest 1% dip in the carry trade risk premium as investors shifted funding from the yen to the euro and Swiss franc. Looking ahead, the primary risk is the timing of Federal Reserve policy changes. While persistent inflation allows the Fed to hold rates, a rapid rise in long-term U.S. yields could threaten the trade. Another concern is crowding, as massive inflows increase vulnerability to a sudden reversal. High interest rates in regions like Latin America and Eastern Europe, supported by external factors like Middle East tensions and energy prices, continue to sustain the opportunity. Major investors remain engaged, favoring currencies like the Mexican peso, South African rand, and Turkish lira.

marsbit1 saat önce

Generating Profits for Seven Consecutive Quarters, Emerging Markets Carry Trade Outperforms Everything

marsbit1 saat önce

İşlemler

Spot
活动图片