Chinese Data Generation Team Makes Debut in Nature Journal

marsbitPublicado a 2026-08-12Actualizado a 2026-08-12

Resumen

A Hong Kong-based startup, Weina AI, has become China's first and the world's fourth data-generation technology company to publish in a leading Nature journal (IF>10 in recent three years) with a paper on AI-assisted kidney cancer surgery decision-making. Founded by Professor Liu Qifeng, who previously led the creation of the world's first thousand-card H800 SuperPod cluster at HKUST, the company focuses on enhancing AI's "questioning" ability to generate high-quality reasoning Q&A data, which is key for AI self-learning. The published Nature Communications paper, co-authored with medical experts, addressed a clinical challenge in renal surgery. Using an RDPM model on multi-source data from 1,621 patients, the AI achieved high predictive accuracy for long-term kidney function decline, providing a quantitative basis for surgical decisions. Professor Liu outlines AI development in three stages: from data to models, from models to token generation, and crucially, from tokens back to data—where AI actively generates questions, reasoning steps, and verified answers (cQrA: context, Question, reasoning, Answer). Weina AI's mission is to create this "data → model → token → data" feedback loop, enabling Agentic AI to autonomously evolve in professional domains. This moves beyond costly manual data annotation by using AI agents to generate scalable, chain-of-thought data. The company raised a HK$50 million seed round led by Lenovo Capital. Instead of focusing deeply on one sector, i...

A Hong Kong startup team unexpectedly came into our view.

Not long ago, a paper on AI-assisted decision-making for kidney cancer surgery was published in Nature Communications (https://www.nature.com/articles/s41467-026-73813-7). With this paper, Weina AI became the first Chinese and the fourth global data generation technology company to publish in a main Nature journal (with an Impact Factor >10 in the past three years)—prior Chinese large model companies to publish were DeepSeek and FaceWall AI.

Behind it is founder Professor Liu Qifeng, who previously built the world's first thousand-card H800 SuperPod cluster at the Hong Kong University of Science and Technology (HKUST), pre-trained China's third large model with hundred-billion parameters, and managed R&D project funds exceeding 100 million USD. Subsequently, he focused on improving AI's 'questioning' ability to generate high-quality reasoning Q&A data, which is one of the keys to the imminent explosion of AI autonomous learning. Thus, he founded Weina AI in Hong Kong.

Out of curiosity, Investment Community engaged in a nearly three-hour in-depth conversation with Liu Qifeng. The discussion started from this paper, extending to large models and embodied AI, as well as his understanding of the next phase of AI.

Starting from a Nature Communications Paper

While others go 'from papers, to papers,' Liu Qifeng goes 'from problems, to problems.'

In early 2025, a relative of Liu Qifeng suffered from kidney cancer, with the attending physician being Director Zhang Zhiling from Sun Yat-sen University Cancer Center. Like all other kidney cancer surgeries, doctors have long faced a clinical dilemma—can there be more quantified and intelligent judgment criteria between partial nephrectomy and radical nephrectomy?

The essence of this challenge is: Can AI predict complex choices in the real world?

Thus, right in the hospital ward, a collaboration spanning medicine and AI commenced: Zhang Zhiling was responsible for medical work and jointly completed data collection with multiple hospitals, while Weina AI handled AI and data processing. The co-first author of the paper, Wang Yatian, is a Ph.D. student at HKUST and an intern at Weina AI, jointly supervised by Liu Qifeng and Professor Luo Wenhan.

Addressing the challenge of multi-source heterogeneous sparse data, the team proposed the RDPM model, incorporating 3D imaging and clinical variables/indicators into the same prediction framework. It was trained and validated on a cohort of 1621 patients, achieving an AUC of 0.788 to 0.873 in external multi-center testing. The paper predicts patients' long-term kidney function decline risk, providing quantifiable support for surgery decisions highly reliant on experience.

Weina AI underwent a public test of AI prediction in a highly fault-intolerant scenario like healthcare. This also points to the other side of AI prediction that Liu Qifeng would discuss next—prediction is the underlying mechanism of large models, generating answers by predicting the next token, naturally 'skilled at answering.' However, Weina AI's focus goes a step further—making AI not only skilled at answering but also 'skilled at questioning.' To give AI 'knowledge and inquiry,' it must both 'learn' effectively and 'question' proactively.

HKUST Professor Turns Entrepreneur

Lenovo Capital Leads the First Round

'While others chase trends, he creates them.' This is how friends describe their impression of Liu Qifeng's past.

This is not an exaggeration. As early as 2001, Liu Qifeng entered the National Laboratory of Pattern Recognition at the Institute of Automation, Chinese Academy of Sciences, studying under Academician Tan Tieniu—the 2022 recipient of the King-Sun Fu Prize, the highest international award in pattern recognition. He subsequently served as a researcher at Samsung Lab, a data scientist at Yahoo! Lab, Director of the Gamma AI Lab at Ping An Group, and AI Director at the Hong Kong Institute of Innovation, CAS. In 2018, he co-founded the Hong Kong Society of Artificial Intelligence and Robotics with Academician Yang Qiang. In 2021, he foresightedly drafted the 'Hong Kong Cloud Brain' and 'Hong Kong Foundation Model' proposals for the Hong Kong government, becoming an early promoter of Hong Kong's AI supercomputing construction and large model training.

These seemingly diverse experiences all point to the same goal: enabling machines to find patterns from complex information and make judgments.

The real turning point occurred in 2023 when ChatGPT became popular. At that time, with strong support from the SAR government and university leadership, Liu Qifeng, in collaboration with six universities at HKUST, co-initiated the Hong Kong Generative AI R&D Centre with Academician Guo Yike, leading the team to build the world's first thousand-card H800 SuperPod AI supercomputing cluster. In 2024, he completed the pre-training/fine-tuning of China's third hundred-billion-parameter Mixture of Experts (MoE) large model. For Hong Kong's AI development, this was a critical juncture.

It was also during this experience that he identified the next gap: the deeper large models go, the more they rely on high-quality data—it's always 'data is king,' especially reasoning Q&A data across various industries. Therefore, enabling large models to 'question' with high quality became the primary key.

Liu Qifeng breaks down large model development into three stages: first, 'from data to model,' using massive internet data for pre-training; second, 'from model to Token,' where large models start outputting tokens to generate content or perform tasks; next is 'from Token to data'—enabling large model systems to actively ask questions, reason step-by-step, and verify answers, i.e., generating reasoning Q&A data. This forms a large feedback loop of 'data → model → Token → data,' thereby enabling AI to possess autonomous learning capabilities.

The purpose of AI autonomous learning is to acquire 'knowledge and inquiry,' and 'knowledge and inquiry' consists of 'learning' from training + 'questioning' and answering. Qing Dynasty scholar Liu Kai wrote in On Inquiry: 'The learning of a superior man necessarily involves a love for inquiry. Inquiry and learning support each other. Without learning, there is nothing to raise doubts; without inquiry, there is nothing to broaden knowledge.'

In July 2024, Hong Kong Weina AI was officially established. The company name is derived from Norbert Wiener—the founder of Cybernetics. What Liu Qifeng values is precisely the feedback loop in cybernetics. Weina AI's mission is to make AI 'question' accurately and 'answer' correctly, thereby realizing the large loop of 'data → model → Token → data,' enabling Agentic AI to autonomously evolve in professional domains.

Weina AI's task is to solve a counterintuitive problem: on one hand, large model development is advancing rapidly; on the other, large model deployment in enterprises remains very difficult. The reason is simple: low accuracy. Using student exam preparation as an analogy—having only textbooks (professional documents) but lacking exercise books (reasoning Q&A data) makes it impossible to achieve high scores (low system accuracy). Memorizing textbooks provides dead knowledge, while doing exercises practices live problem-solving abilities. What Weina AI does is help various industries supplement this 'exercise book,' enabling AI not only to 'study textbooks' but also to 'do exercises,' thereby addressing the bottlenecks of inaccuracy, difficulty in optimization, and incorrect answers currently faced by the proliferation of Agents.

There is a popular saying: large model Q&A is outdated; task execution is key. This is somewhat superficial. Execution capability depends on two pillars: the accuracy of a single agent in a professional domain and the collaborative capability among multiple agents. The reality is that current execution capabilities are far from reliable. One of the root causes is that single-agent Q&A accuracy often falls below 70%—not even crossing the threshold of 'trustworthiness,' let alone 'collaboration.'

The specific definition of an 'exercise' is cQrA: context, Question, reasoning, Answer. Context is the task scenario, Question is the generated question, reasoning is the reasoning process, and Answer is the verified answer. In other words, Weina AI enables the model to simultaneously generate questions, answers, and reasoning processes within a specific industry context.

This also distinguishes it from traditional data annotation. Traditional data annotation heavily relies on manual labor, even experts, with high costs, difficulty in scaling, providing only answers without reasoning, consuming expert experience in repetitive tasks. In contrast, Weina AI enables Agentic AI to become tireless intelligent expert teams, automatically generating cQrA data with complete chains of thought, completely breaking through the human resource bottleneck. More critical than cost-saving is that the closed-loop mechanism allows data generated in each round to feed back into the generation and evaluation models, driving continuous leaps in precision and logic for the next iteration—thus achieving a qualitative change from a 'manual workshop' to a 'self-evolving knowledge factory.'

Weina AI quickly caught the attention of the industry and investors. Shortly after its establishment, the company completed a 50 million HKD seed round of financing, led by Lenovo Capital. Lenovo Capital has consistently invested along the three key elements of AI: computing power invested in companies like MetaX and Cambricon, models invested in companies like Zhipu and StepFun, and the data element landed on Weina AI. Simultaneously, MetaX and Weina AI have deepened cooperation. In the upcoming era of the large loop 'data → model → Token → data,' one has designed the computing platform in advance for the future paradigm, while the other has defined the workload for the future paradigm in advance.

The Next Phase of AI

'Let Us Generate This World!'

Commercial validation starts with two soul-searching questions.

Question One: Will generated data be purchased by professional institutions without large-scale expert annotation?

Question Two: Can it be cross-industry and replicable?

To answer these, Weina AI, resisting pressure, broke from the traditional 'depth-first' principle of B2B tech companies—which insists on 'penetrating a specific industry' first—and instead adopted a 'breadth-first' approach. They deliberately chose four seemingly unrelated industries with high accuracy requirements: value & safety, government affairs, insurance, and horse racing, and have secured leading clients in each.

'We proved that we can achieve cross-industry replication with a small team, no industry experts, and low cost,' Liu Qifeng stated. Having now achieved validation from '0 to 4,' the next step is scaling from '1 to M x N' (M industries, each with N leading clients).

Behind this lies a long-term judgment on the value of data.

In Liu Qifeng's view, the gap between Chinese and American AI is largely due to differences in the perception of data—data has long been seen as 'dirty and tiring work,' and data engineers' salaries are generally lower than those of algorithm and model engineers.

However, the landscape is shifting. As data production moves from manual annotation to reasoning, interaction, and closed-loop feedback, large model companies are continuously increasing investment in the data side. It is now a consensus within the industry that reasoning and interactive data generation determines the upper limit of large model capabilities.

In the future, the most core element is not the model, nor even the data itself, but that 'large loop.' Just as the key to evolution is neither men nor women, but mating and natural selection—the mechanisms of chromosome replication, crossover, mutation, and survival of the fittest. Data distillation is merely one path leveraging external forces. The real moat lies in establishing an autonomous learning loop where model training and data generation drive each other, using model collaboration and feedback mechanisms to continuously generate high-quality data.

This judgment also extends to the currently hottest topic: embodied AI.

The traditional way of training embodied AI is based on imitation of humans. True intelligence should be like a baby learning to walk 'through trial and error': autonomously generating motion data through continuous falling and attempts, then iteratively optimizing decision-making models through closed-loop feedback.

Liu Qifeng says the logic of closed-loop training in the digital world has already extended to the physical world. Whether it's Agents entering industries or robots going into the field, they all rely on massive, high-quality reasoning and interactive data generated autonomously in advance. Correspondingly, cQrA evolves into cTrA—context, Task, reasoning, Action. This serves as both new fuel for training and a new benchmark for evaluation.

The journey has just begun. The answer Liu Qifeng points to leads to the not-so-distant future: 'Let us generate this world!'

This article is from the WeChat public account 'Investment Community' (ID: pedaily2012), by Wang Lu.

Criptos en tendencia

Preguntas relacionadas

QWhat is the core innovation of the paper that enabled Wiener AI to be featured in Nature Communications?

AThe core innovation is the RDPM model. It addresses the challenge of multi-source heterogeneous sparse data by integrating 3D medical images and clinical variables/indicators into a unified predictive framework. This model successfully predicted long-term renal function decline risk in patients, providing a quantifiable basis for surgical decision-making in kidney cancer, and was externally validated across multiple centers with AUC values ranging from 0.788 to 0.873.

QAccording to Liu Qifeng, what are the three stages of large model development, and which stage is Wiener AI focused on?

AAccording to Liu Qifeng, the three stages of large model development are: 1) 'From Data to Model' (pre-training on internet-scale data), 2) 'From Model to Token' (the model generates content or performs tasks by outputting tokens), and 3) 'From Token to Data' (the model actively asks questions, reasons step-by-step, and validates answers to generate reasoning Q&A data). Wiener AI is focused on the third stage, enabling AI to be 'good at asking' questions to create this high-quality data for autonomous learning.

QWhat is cQrA data, and how does it differ from traditional data annotation?

AcQrA data stands for context, Question, reasoning, and Answer. It is a structured form of data where the AI generates not just an answer but also the corresponding question and the complete reasoning chain (thought process) within a specific professional context. This differs from traditional data annotation, which is labor-intensive, relies heavily on human experts, typically provides only answers without reasoning, and is difficult to scale. Wiener AI uses Agentic AI to automatically generate cQrA data, breaking the human bottleneck and enabling a self-evolving 'knowledge factory'.

QWhy did Wiener AI adopt a 'breadth-first' strategy for its initial commercial validation, and what industries did it target?

AWiener AI adopted a 'breadth-first' strategy to challenge two core questions: 1) whether its AI-generated data (without massive expert annotation) would be purchased by professional institutions, and 2) whether its approach was replicable across different industries. To prove this, it deliberately targeted four seemingly unrelated but high-accuracy-demanding industries: values/safety, government affairs, insurance, and horse racing. Securing leading clients in each domain successfully validated the '0 to 4' proof of concept for cross-industry replication.

QWhat is the 'great closed loop' that Liu Qifeng identifies as the future core of AI, and how does it relate to embodied intelligence?

AThe 'great closed loop' refers to the self-reinforcing cycle of 'Data → Model → Token → Data.' Liu Qifeng believes the future core of AI is not the model or even the data alone, but this autonomous learning mechanism where model training and data generation drive each other. This logic extends to embodied intelligence. Instead of just imitating humans, true intelligence should generate its own action data through trial and error (like a baby learning to walk), using closed-loop feedback to optimize its decision model. The data format evolves from cQrA to cTrA (context, Task, reasoning, Action), providing new fuel for training and new benchmarks for evaluation in the physical world.

Lecturas Relacionadas

The Last Mile of Stablecoins: MoneyGram Connects Crypto Wallets to Cash Outlets in Over 170 Countries

Stablecoins have revolutionized cross-border payments with speed and low cost, but a fundamental barrier remains: converting digital money into physical cash in most parts of the world. On August 11th, global remittance giant MoneyGram announced its "MoneyGram Ramps" crypto-cash service is now live on the Solana blockchain. This allows developers, via a simple API, to connect their applications to MoneyGram's network of nearly 500,000 cash agent locations across over 170 countries. Users can now walk into a local MoneyGram agent to convert their USDC into local currency cash or use cash to purchase USDC for their wallets. Solana wallet Rift is the first to integrate the service. MoneyGram, founded in 1940 and historically second only to Western Union in global remittance volume, was privatized in 2023. Since then, it has accelerated its shift into a fintech platform, opening its compliance capabilities, cash network, and settlement infrastructure as services for developers. Its involvement with crypto is not new, having previously partnered with Ripple, Stellar, and Circle, and even issuing its own stablecoin, MGUSD, on Stellar in 2026. The Ramps API solves a critical physical obstacle to stablecoin adoption. While a USDC transfer from New York to Lagos takes seconds and costs cents on-chain, the recipient often lacks a bank account to access the funds. MoneyGram's extensive physical agent network, prevalent in regions like Sub-Saharan Africa, Southeast Asia, and Latin America, bridges this gap, connecting on-chain dollars to the offline cash-based economy. No crypto-native company could replicate this network quickly. MoneyGram chose Solana for its multi-chain expansion due to its extremely low transaction fees (often less than a cent) and fast confirmation times, making it ideal for small-value, frequent remittances. Solana's developer ecosystem is also geared toward consumer-facing applications. This partnership represents a pragmatic path for stablecoins in everyday remittances. Users don't need to understand blockchain or have a bank account; they simply visit a trusted local agent. The underlying settlement shifts from SWIFT to Solana, but the user experience remains familiar. The true path to stablecoin adoption is being paved by an 85-year-old remittance company, using its physical network to solve the "last mile" problem for crypto.

marsbitHace 13 min(s)

The Last Mile of Stablecoins: MoneyGram Connects Crypto Wallets to Cash Outlets in Over 170 Countries

marsbitHace 13 min(s)

How to Adapt to Frequent Miraculous Comebacks After the New Patch When Trading in Polymarket's LOL Section?

After recent League of Legends (LoL) game updates, major comebacks and unpredictable outcomes have become more frequent, making trading on Polymarket's LoL prediction markets challenging. This article analyzes the causes and offers adaptation strategies. The author, a longtime LoL fan turned Polymarket trader, notes that post-update, a large early-game gold lead no longer guarantees victory. The current meta reduces the decisive weight of economy, emphasizing instead "advantage conversion, resource trading, team composition strength, and teamfight execution." Key systemic changes aiding comebacks include: Bounty/Kill Gold Compensation, Strategic Objective Bounties, and reduced Baron Nashor rewards. Furthermore, the prevalence of mage picks in the bot lane, instead of traditional marksmen (ADC), can lower sustained damage and teamfight reliability in the late game, contributing to throws. For more successful trading, the article suggests: 1. **Choosing the Right Region:** Prioritize major regions (LPL, LCK, LEC, LCS) where matches matter for World Championship qualification or player contracts, leading to more serious play. Avoid or be extremely cautious with regions like CBLOL (Brazil), TCL (Turkey), and North American Challengers, citing high volatility and potential for unreliable outcomes. 2. **Trading Strategies:** * **Placing Orders:** Use limit orders below market price for strong favorites. In close matchups, consider placing equal low-ball bids (e.g., 25-40 cents) on both sides of a match; if both fill for a total under $1, you secure profit regardless of the winner. Exploit team "power spikes" by buying low before their composition peaks and selling high. * **Timely Selling:** The most crucial lesson is to take profits. If you buy the early-leading team at a low price (e.g., 50-60 cents), sell most of your position (e.g., 75%) at 85-90 cents to lock in gains and avoid a potential comeback zeroing your position. * **Position Sizing:** Practice strict risk management. Do not exceed 5-10% of your capital per match and avoid "tilting" or chasing losses after a few bad trades. In summary, adapting to the new meta requires understanding that early advantages are less stable. Success depends on analyzing team compositions, player form, and regional contexts, combined with disciplined order placement, profit-taking, and capital preservation. The next patch in August is expected to revive traditional ADCs, potentially shifting the dynamics again.

Odaily星球日报Hace 26 min(s)

How to Adapt to Frequent Miraculous Comebacks After the New Patch When Trading in Polymarket's LOL Section?

Odaily星球日报Hace 26 min(s)

Strategy CEO Announces Plan to Resume Bitcoin Purchases This Year, With Buying Volume 25 Times Selling Volume

Strategy CEO Announces Resumption of Bitcoin Purchases This Year, Buy-to-Sell Ratio at 25x In a FOX Business interview, Phong Le, CEO of Strategy (formerly MicroStrategy), stated the company plans to resume its Bitcoin acquisition strategy within the current year. This ends a pause in buying that began in May. Le revealed that since the start of the year, Strategy has purchased approximately 175,000 bitcoins while selling around 7,000, making its buy volume about 25 times its sell volume. The company remains the world's largest corporate holder of Bitcoin, with 840,447 BTC, representing roughly 4% of the circulating supply. Le explained that recent sales were used to fulfill capital obligations under a newly approved framework, including paying preferred stock dividends, funding share buybacks, and bolstering the company's US dollar reserves, which now stand at about $4.7 billion. This move marks a shift from the firm's previous "never sell" mantra, which had contributed to its market premium. The change in strategy and the associated financial pressures led to a significant drop in its stock price (MSTR) and a rating downgrade from JPMorgan. The CEO framed the company's role as the "JPMorgan of the crypto economy," emphasizing its long-term goal of increasing the amount of Bitcoin per MSTR share. The announcement signals the return of a major institutional buyer to the Bitcoin market, which could provide price support. It is also viewed as a micro-indicator of recovering institutional confidence, suggesting the company believes the most acute phase of liquidity pressure has passed.

marsbitHace 48 min(s)

Strategy CEO Announces Plan to Resume Bitcoin Purchases This Year, With Buying Volume 25 Times Selling Volume

marsbitHace 48 min(s)

ArthurHayes新文:押注日元升值,ENA未来几月或涨5至10倍

Arthur Hayes argues that the Japanese Yen is significantly undervalued and posits that its appreciation against the US Dollar is imminent. He outlines three potential mechanisms for this shift, dismissing the first two—the Bank of Japan raising interest rates and domestic institutions selling foreign assets—as politically or economically unfeasible. He identifies the third and preferred method: the Japanese Ministry of Finance (MOF) using its holdings of US Treasuries as collateral in the Fed's FIMA repo facility to borrow US dollars, then selling those dollars to buy Yen in the forex market. Hayes believes US Treasury Secretary Bessant has signaled support for this approach, which requires the Fed's Foreign Currency Subcommittee, led by Chairman Walsh, to remove lending limits on the FIMA tool. Hayes asserts that implementing this "Scheme 3" would lead to a significant expansion of US dollar liquidity. He predicts this surge in liquidity will act as a catalyst, driving up the prices of assets like Bitcoin and physical gold. Within the crypto space, he views Ethereum (ETH) as undervalued and singles out Ethena's ENA token as a speculative play with potential for 5-10x gains in the coming months, contingent on a recovery in Bitcoin basis trades that would boost demand for its USDe stablecoin. He concludes that investors should watch for the Fed's rule change as the key trigger for these market movements.

marsbitHace 1 hora(s)

ArthurHayes新文:押注日元升值,ENA未来几月或涨5至10倍

marsbitHace 1 hora(s)

Trading

Spot

Artículos destacados

Cómo comprar DATA

¡Bienvenido a HTX.com! Hemos hecho que comprar DATA Network (DATA) sea simple y conveniente. Sigue nuestra guía paso a paso para iniciar tu viaje de criptos.Paso 1: crea tu cuenta HTXUtiliza tu correo electrónico o número de teléfono para registrarte y obtener una cuenta gratuita en HTX. Experimenta un proceso de registro sin complicaciones y desbloquea todas las funciones.Obtener mi cuentaPaso 2: ve a Comprar cripto y elige tu método de pagoTarjeta de crédito/débito: usa tu Visa o Mastercard para comprar DATA Network (DATA) al instante.Saldo: utiliza fondos del saldo de tu cuenta HTX para tradear sin problemas.Terceros: hemos agregado métodos de pago populares como Google Pay y Apple Pay para mejorar la comodidad.P2P: tradear directamente con otros usuarios en HTX.Over-the-Counter (OTC): ofrecemos servicios personalizados y tipos de cambio competitivos para los traders.Paso 3: guarda tu DATA Network (DATA)Después de comprar tu DATA Network (DATA), guárdalo en tu cuenta HTX. Alternativamente, puedes enviarlo a otro lugar mediante transferencia blockchain o utilizarlo para tradear otras criptomonedas.Paso 4: tradear DATA Network (DATA)Tradear fácilmente con DATA Network (DATA) en HTX's mercado spot. Simplemente accede a tu cuenta, selecciona tu par de trading, ejecuta tus trades y monitorea en tiempo real. Ofrecemos una experiencia fácil de usar tanto para principiantes como para traders experimentados.

484 Vistas totalesPublicado en 2026.07.01Actualizado en 2026.07.01

Cómo comprar DATA

Discusiones

Bienvenido a la comunidad de HTX. Aquí puedes mantenerte informado sobre los últimos desarrollos de la plataforma y acceder a análisis profesionales del mercado. A continuación se presentan las opiniones de los usuarios sobre el precio de DATA (DATA).

活动图片