Tremble Humans, AI Continues Its Accelerated Sprint

marsbitPublished on 2026-06-13Last updated on 2026-06-13

Abstract

Trembling, Humans: AI Continues Its Accelerated Sprint Yes, AI is still rapidly accelerating. While deep learning seemed to stall quickly in its early years, large models after years of development show no sign of hitting their ceiling. At the Zhiyuan Conference 2026, the focus is on enabling AI to move from the digital world into the physical world. Scaling Law remains effective, continuing to drive advancements in both large language models and multimodal models. The industry is now entering a phase of pursuing World Models, though unresolved technical paths and data issues mean this exploration may take 3-5 more years. Concurrently, breakthroughs in Agents are accelerating AI's real-world application in fields like healthcare and meetings. Making Agents truly useful requires key hardware-software co-design, evident from the strong presence of chip vendors at the conference. We stand at a new historical threshold where AI is becoming a foundational force reshaping the world. The first day of the conference highlighted AI's evolution from "knowing how to chat" to "knowing how to work." Scaling Law persists, World Models are the next key battleground, and Agents are transitioning from usable to好用 (user-friendly). Scaling Law is not ending but diversifying. New models like Anthropic's Fable 5 demonstrate scaling through parameter size, synthetic data, and reinforcement learning. Advancements in AI Coding and Agent deployment are enabling a trend of AI self-evolution, poten...

That's right, AI is still in an accelerated sprint.


In 2016, deep learning had only been exploding for a year before it almost stagnated. In 2026, after four years of explosive growth, large models still haven't hit their ceiling.


At the 2026 BAAI Conference, Guangzhui Intelligent observed that from models to software/hardware to products, everything is striving for AI to 'run' from the digital world into the physical world.


On one hand, Scaling Law continues to function steadily, propelling the ongoing development of large language models and multimodal models. The AI industry has entered a phase of pursuing World Models. However, issues like current technical routes and data remain unresolved, likely requiring at least 3-5 more years of exploration.


On the other hand, breakthroughs in Agents are accelerating the deployment of AI in real-world scenarios. As Agents have reached a usable stage, the industry is advancing their application in areas like healthcare and meetings. To transition Agents from usable to useful, software-hardware co-design has become key. At the exhibition booths of the BAAI Conference, chip manufacturers occupied 'half the room,' with nearly all leading domestic AI chip companies present.



"We are standing at a new historical inflection point. Artificial intelligence is no longer just a tool transforming a specific industry but is becoming the underlying force reconstructing the world. AI Coding, autonomous agents, and model self-evolution are opening up possibilities for creating AI. World Models, embodied intelligence, and robotics are extending intelligence from the digital world to the physical world," said Wang Zhongyuan, President of the Beijing Academy of Artificial Intelligence (BAAI).


What exactly is happening within this wave of reconstruction by this underlying force?


On the first day of the BAAI Conference, the guests present offered this answer: AI is moving from 'being able to chat' to 'being able to work.' Scaling Law persists, World Models with unconverged technical directions become the focus of the next phase, while Agents have started transitioning from usable to useful, with many optimization challenges remaining.


AI Has Not Hit Its Technical Ceiling,


And Has Learned Self-Evolution


Over the past year, as high-quality internet text data was being exhausted, a pessimistic sentiment spread throughout the industry that 'Scaling Law is about to peak.'


In multiple forums at the BAAI Conference, the question 'Has the Scaling Law dividend diminished?' was frequently raised. Several guests denied this notion.


"I still firmly believe scaling is far from over," said Wang He, Founder and CTO of Galaxy Universal. "Looking back today, Scaling Law hasn't failed; it has just become more diversified."


Scaling continues to show its effect on a series of newly released large language models. Analyzing Anthropic's recently released Fable 5, Luo Fuli from Xiaomi suggested this model itself is a product of scientifically advancing scaling. It is the result of extending large models by combining three dimensions: parameter scale, synthetic data, and reinforcement learning.


"We speculate that Fable 5's parameter scale itself is likely several times that of the current largest open-source models. Additionally, it involved significant computational investment in Test-Time Scaling or reinforcement learning. Furthermore, synthetic data generated by humans and agents brought the data scale to a new order of magnitude," said Luo Fuli.


In the multimodal field, performance improvements brought by scaling are equally significant. Zhu Jun, Founder and Chief Scientist of Shengsheng Technology, stated that data quality, model size, and large-scale training all enhance model performance. With improved foundational model capabilities, models also learn physical laws and understand 3D scenes more efficiently.


While scaling continues to be effective, alongside the maturation of AI Coding and accelerated deployment of Agents, a trend of AI self-evolution is becoming evident, upgrading from writing code to autonomously completing product iteration updates.


"The foundation of the vast human digital world is largely constructed through code. With AI Coding making substantial progress and becoming mainstream, it means AI could gradually take over everything in the digital world," said Wang Zhongyuan.


Globally, using AI for product updates has become the norm.


"If the model determines an agent's capabilities, then the Harness determines the upper limit of those capabilities," said Li Jingqiu. "Its difficulty lies in further improving problem clarification, verification, and feedback on top of the model."


For example, relying solely on the model to understand a problem inevitably has limitations. The Harness needs to elaborate and enrich the user's simple one-sentence instruction so the model can better comprehend the requirement. This requires the Harness to leverage intent understanding. After receiving the task, it must design the subsequent workflow and then orchestrate the model to execute it. This process may require human intervention and correction, followed by checks before task completion.


World Models:


The Next Key Battleground for Large Models


Pushing outward along the boundaries of the digital world, World Models have become the next key battleground for large models.


"Currently, no single world model truly feels particularly impressive or solves all kinds of problems in the real physical world," said Wang Zhongyuan.


For World Models in their early developmental stage, the industry hasn't reached full consensus on the technologies involved. With technical routes not yet converged, a series of unresolved problems remain. Using data as an example, Wang Zhongyuan illustrated that whether video data, simulation data, or real-world physical data is needed, a clear methodological path hasn't been found yet.


Taking Galaxy Universal as an example, Wang He introduced their application of synthetic data at the event.


"Before the WAM (World Action Model) paradigm emerged, we conducted extensive experiments within the VLA paradigm using synthetic data, specifically for grasping tasks," said Wang He. "We used 1 billion frames of simulation data to prove: as long as you scale the data to this extent, you can achieve complete zero-shot learning. Give me any object in the real world, and it can handle the grasp."


Regarding the development progress of World Models, the BAAI predicts that 'at least several more years' are needed. The next three to five years will likely be a phase of continuous evolution and iteration for World Models.


Over the past few years, various world models with different technical routes have emerged in the industry, each progressing distinctively.


Taking multimodal world models as an example, Zhu Jun stated that video models and world models are closely related because world models need three capabilities: understanding and interpreting states, prediction, and action. Among currently accessible training data, video data is most relevant to world models.


With various technical routes diverging and industry consensus yet to form, the BAAI classifies world models into four categories:


First, language-centric world models, mapping other modalities and abilities into language space, including LLMs, VLMs, VLAs, etc.


Second, pixel-centric world models; video generation essentially predicts the next frame, but video generation models are not equivalent to world models, though they are related. The potentially very popular World Action Model (WAM) this year is evolving from a pixel-centric perspective.


Third, 3D structure-centric world models, including 3D reconstruction which focuses purely on the three-dimensional world.


Fourth, visual representation-centric world models.



Currently, BAAI is exploring a 'fifth' path – the fusion of language-centric and visual representation-centric approaches, namely latent space representation. This involves compressing information like text and images into a vector space to represent various states of the real physical world.


"Future unified latent space modeling will not be limited to visual space but encompass full-modal latent space. This is highly likely to be the true next possible path for world models," said Wang Zhongyuan.


At the conference, BAAI introduced the world model it is developing – WuJie · Physis-v0.1. Centered on physical space modeling to predict the next physical state, it is positioned as the world's first general-purpose world foundation model, emphasizing four key capabilities: 'physically correct, causally traceable actions, long-term temporal consistency, and general-purpose generalization.'



Currently, this model is still in the training phase. BAAI will continue to share progress in the second half of the year and will open-source the model upon training completion.


From 'Usable' to 'Useful':


Agents Face More Challenges


On the model side, progress in World Models drives the realization of physical AI; on the product side, Agents (Intelligent Agents) become the key products for AI to enter public life.


Since 2025, dubbed the 'Year of the Agent,' some impressive Agent products have emerged, showing signs of taking off. However, the unexpected surge in popularity of 'Lobsters' this year still came as a surprise.


Compared to last year when agents were mostly in an execution state, this year's agents have clearly become more proactive and capable, able to help users proactively execute more complex tasks.


At this year's BAAI Conference, BAAI also released four vertical-focused agents: BAAI Cardiac Agent, the world's first auxiliary diagnosis agent for cardiac magnetic resonance, aiding doctor decision-making by integrating multimodal capabilities and medical expertise; the autonomous research agent AREX for the scientific research field; SoulAgent, an agent helping users listen to meetings in real-time and capture key points; and a risk discovery agent targeting hazardous protein acquisition.


For example, regarding the meeting-listening agent, Guangzhui Intelligent tested its ability to summarize different meeting contents. SoulAgent did provide simple summaries of meeting content. While not as complete as minutes, the core viewpoints were accurate. This is particularly suitable for situations where parallel forum sessions overlap.



However, current agents still face numerous technical issues requiring further optimization. Yang An, President's Chair Professor at Nanyang Technological University, mentioned that to maintain and enhance agent capabilities, the most crucial aspects currently are related to context engineering, such as Memory, orchestration, etc.


At the agent sub-forum, Harness (literally meaning a horse's harness, referring to the entire engineering framework or environment built around an agent), which received little attention last year but gained significant popularity this year, became a high-frequency keyword mentioned on-site.


"If the model determines an agent's capabilities, then the Harness determines the upper limit of those capabilities," said Li Jingqiu. "Its difficulty lies in further improving problem clarification, verification, and feedback on top of the model."


For example, if relying solely on the model to understand a problem, limitations are inevitable. The Harness needs to elaborate and enrich the user's simple one-sentence instruction so the model can better comprehend the requirement. This requires the Harness to leverage intent understanding. After receiving the task, it must design the subsequent workflow and then orchestrate the model to execute it. This process may require human intervention and correction, followed by checks before task completion.


In short, like a real human assistant, every detailed step requires product refinement for the Harness to further improve the Agent's execution effectiveness.


Currently, Agents are still in the early stages of development. It is foreseeable that this industry has immense room for growth. Both improvements in model capabilities and solidification of engineering details will continue to enhance Agents' task-handling abilities.

This article is from WeChat Official Account: Guangzhui Intelligent , Author: Focus on Frontier Technology

Trending Cryptos

Related Questions

QWhat is the main theme of the 2026 Zhiyuan Conference according to the article?

AThe main theme is that AI is evolving from being a tool for chatting to becoming capable of performing practical tasks ('work'), with a focus on scaling laws, the pursuit of world models, and the advancement of Agents from being usable to good.

QWhat does the article suggest about the current status and future of Scaling Law in AI development?

AThe article states that the Scaling Law is far from reaching its limit and is still effectively driving progress. While its form has diversified, it continues to push advancements in large language models and multimodal models through increased parameters, synthetic data, and reinforcement learning.

QWhat are the four categories of world models outlined by the Zhiyuan Research Institute, and what is the 'fifth path' they are exploring?

AThe four categories are: 1) Language-centric models, 2) Pixel-centric models, 3) 3D-structure-centric models, and 4) Visual-representation-centric models. The 'fifth path' they are exploring is a fusion of language-centric and visual-representation-centric approaches, aiming for unified latent space modeling across all modalities.

QAccording to the article, what is 'Harness' in the context of AI Agents, and why is it important?

AHarness refers to the engineering framework or environment built around an AI Agent. It is crucial because it determines the upper limit of an Agent's capabilities. Its role is to clarify user intent, design task workflows, schedule model execution, and incorporate human intervention and verification to ensure tasks are completed correctly, going beyond what the model alone can understand and execute.

QWhat example does the article give to illustrate the current practical application and capability of AI Agents?

AThe article gives the example of SoulAgent, an AI Agent designed for listening to and summarizing meetings. It was tested at the conference and was able to provide simple, accurate summaries of the core points from different sessions, demonstrating its utility in situations where forum times overlap.

Related Reads

E Fund "Quits Drinking"

"Yi Fang Da 'Gives Up Alcohol'" - Summary The article analyzes the recent portfolio adjustments of star fund manager Zhang Kun, focusing on his flagship funds, the Yi Fang Da Blue Chip Select and Yi Fang Da Quality Select. The key trend highlighted is a significant and ongoing reduction in exposure to the consumer sector, particularly liquor (baijiu) stocks like Kweichow Moutai and Wuliangye, which were once core holdings. This shift is driven by several factors: a downturn in the consumer sector's business cycle post-pandemic, with slowing earnings growth for major players; a change in consumer behavior towards value over premium brands; and high channel inventory pressures, especially in the liquor industry. These fundamentals have led to a prolonged phase of valuation contraction for consumer stocks. Concurrently, market capital has rotated aggressively towards the high-growth technology sector, especially AI-related chains, which have dominated fund inflows. This style shift has further diminished the attractiveness of traditional consumer "core assets" for many funds. As a result, Zhang Kun's funds have drastically cut their consumer stock allocations. For example, the Blue Chip Select's baijiu holding proportion has fallen from over 40% to around 20%, and its overall top 10 holdings concentration dropped sharply. This mirrors a broader trend in the active equity fund industry, where allocations to consumer staples have plummeted to multi-year lows while technology holdings have surged. The article concludes that while the consumer sector's valuations are now at historical lows, offering some margin of safety, the fundamental outlook remains challenging due to weak macro consumption recovery and intense internal competition. Therefore, the sector is likely to see only sporadic rebounds rather than a sustained recovery, continuing to weigh on the performance of funds still heavily invested in it. Zhang Kun's timely reduction of Korean semiconductor holdings in his Asia Select fund amid a market crash is also noted as a contrast to his domestic strategy.

marsbit10m ago

E Fund "Quits Drinking"

marsbit10m ago

Before the Second Attempt at Hong Kong IPO, Topstar 'Hands Out' 548 Million in Related-Party Orders

Tosda, an industrial robotics firm, is making its second attempt at a Hong Kong listing shortly after its first application lapsed. The company has undergone a significant strategic shift, pivoting away from its lower-margin intelligent energy and environmental management system (IEEMS) business to focus on core robotics and machinery. This "revenue-reduction, profit-growth" strategy saw revenue nearly halve from 2022 to 2025, but resulted in a sharp turnaround to profitability, with a 1147% year-on-year increase in Q1 2026 net profit. This restructuring involved spinning off the IEEMS business, leading to a substantial jump in related-party transactions. Following the departure of a former executive, two newly associated companies are expected to handle up to 548 million RMB worth of IEEMS orders in 2026, with Tosda acting as a platform charging a 3% management fee. While profitability has improved, several financial concerns remain. Accounts receivable collection cycles have lengthened, inventory—particularly for robots—has surged significantly despite falling revenue, and the company faces challenges in overseas market expansion. Notably, overseas business currently has a lower gross margin than domestic operations. The company was also recently reprimanded by regulators for accounting inaccuracies related to revenue recognition and bad debt provisions. The success of its strategic pivot and its second listing attempt will depend on managing these financial risks, the transparency of its new business model, and the long-term performance of its core robotics segment.

marsbit20m ago

Before the Second Attempt at Hong Kong IPO, Topstar 'Hands Out' 548 Million in Related-Party Orders

marsbit20m ago

Is the AI Bull Market in Jeopardy? Philadelphia Semiconductor Index Breaks Below Key Support at 11,200 Points

The Philadelphia Semiconductor Index (SOX) faces a critical technical and fundamental test. It recently breached a key short-term support level around 11,200 points—a level that held in mid-July—and closed below it. Technical indicators have deteriorated, with the 21-day moving average crossing below the 50-day, signaling bearish momentum. Should the 11,200 support fail, the next major support is the 200-day moving average. The sell-off, which saw memory chip stocks like Micron and Western Digital plunge over 10%, stems from growing concerns about the cyclicality and quality of AI demand. Market participants question the sustainability of the AI capital expenditure boom, highlighted by divergences between the capital expenditures of mega-cap tech firms and semiconductor index performance. However, some positive structural shifts are emerging. Institutional positioning has normalized, unwinding the previously crowded "long semiconductors, short Mag 7" trade, leaving cleaner positioning. Seasonal tailwinds are also approaching, as the NDX historically performs well after this period, potentially supporting a rebound if retail inflows resume. The core question is whether the SOX can hold the 11,200 support. A failure could trigger further declines toward the 200-day moving average, while a successful defense, combined with lighter positioning and seasonal support, could offer a recovery window for the sector.

marsbit20m ago

Is the AI Bull Market in Jeopardy? Philadelphia Semiconductor Index Breaks Below Key Support at 11,200 Points

marsbit20m ago

Qualcomm and Arm Fall Together: The Bill for Memory Price Hikes Finally Arrives at Mobile Chip Companies

After posting Q2 FY2026 results, Qualcomm and Arm both saw their shares decline, reflecting the impact of memory price increases on the smartphone chip sector. Qualcomm's revenue of $9.95B slightly beat expectations, but EPS of $2.21 fell short. More concerning was its guidance for next quarter, with EPS projections below analyst estimates. The company directly attributed a >$1.50 per share annual EPS headwind to rising memory costs and supply constraints in Android phones, prompting planned price hikes. While automotive revenue grew 61% and is approaching one-third of phone revenue, the mobile segment declined 20%. Qualcomm also confirmed a significant reduction in its modem share for the upcoming iPhone and outlined a plan for data center revenue to replace all Apple-related income by FY2027. Arm's results surpassed expectations with revenue of $1.29B and EPS of $0.45, and its guidance was also strong. However, its stock fell. Key concerns included royalty revenue failing to set a new record and a downward revision to full-year royalty growth guidance from ~20% to the high-teens, citing weak smartphone demand and high memory prices. Despite robust growth in its data center business, Arm's premium valuation (over 100x forward P/E) means even beating expectations isn't enough to push the stock higher, as any sign of uncertainty is magnified. Ongoing global antitrust investigations add another risk factor. The situation highlights a broader shift. Memory price surges, which boosted Samsung's profits, are now pressuring chip designers. Meanwhile, the semiconductor sector saw significant corrections in July, with the hardest-hit stocks often being those with the biggest AI-driven gains year-to-date. Qualcomm, lacking such a premium, was an exception.

marsbit35m ago

Qualcomm and Arm Fall Together: The Bill for Memory Price Hikes Finally Arrives at Mobile Chip Companies

marsbit35m ago

Trading

Spot

Hot Articles

Discussions

Welcome to the HTX Community. Here, you can stay informed about the latest platform developments and gain access to professional market insights. Users' opinions on the price of AI (AI) are presented below.

活动图片