Understanding the New Economic Model of Tokenization

marsbitОпубликовано 2026-05-19Обновлено 2026-05-19

Введение

Understanding the New Token Economics Model The commercialization of AI applications is evolving from selling software and subscriptions to selling token call capacity. Tokens, the fundamental unit of information processing for large language models (LLMs), have become the basis for API billing and consumption. With call volumes exploding, tokens themselves are now being traded—procured, routed, split, and resold—forming a new intermediary market. This layer connects upstream LLM providers with downstream developers and enterprises, acting as a global wholesale-to-retail liquidity network. The rise of this business is fueled by a massive surge in China's daily token call volume—growing over a thousandfold from 100 billion in early 2024 to over 140 trillion by March 2026—and significant improvements in domestic LLM capabilities, which are now competitive globally. The core value of token distribution platforms extends beyond simple arbitrage. Key functions include aggregating multiple models (like GPT, Claude, and domestic models such as Kimi and DeepSeek) under a unified API, lowering network and payment barriers, and providing enterprise services like model selection, prompt engineering, and system integration. Profit models are diversifying: (1) resale margins; (2) technical premiums from proprietary inference acceleration (e.g., reducing costs to 1/10 of the industry standard); and (3) enterprise value-added services. High-consumption scenarios like marketing, short-f...

Author: Zhao Ying

Source: Wall Street News

The commercialization of AI applications is extending from selling software and memberships to selling token-calling capabilities. Here, Tokens refer to the smallest units of information processed by large models, serving as the basis for API billing, settlement, and consumption. As the volume of calls increases, Tokens themselves are beginning to be procured, routed, split, and resold like a form of "inventory."

Chen Liangdong, an analyst at Huayuan Securities, summarized the core change in a recent media industry report: "Token operations are forming a new intermediary market, which involves exploring token distribution models to connect upstream large model providers with downstream developers, enterprises, and individuals. The essence is the liquidity infrastructure for a global network of token wholesale to retail."

The background for this business is not complex: On one hand, China's daily token call volume has surged rapidly, rising from 100 billion tokens per day at the beginning of 2024 to 100 trillion by the end of 2025, surpassing 140 trillion by March 2026. On the other hand, domestic large models have improved significantly, entering the global top tier in certain rankings and call volumes. With increasing demand and a growing number of models, the real barriers to transactions have become payment, network access, interfaces, compliance, distribution channels, and scenario implementation.

However, token distribution cannot be simply understood as "reselling API quotas." The thinnest layer of profit comes from resale margins, while the thicker portion comes from inference acceleration, unified interfaces, enterprise-level prompt engineering, Agent orchestration, model selection, and integration with business systems. Precisely because the entry barrier is not high, the risks in this market are equally direct: intensified competition, funding requirements for upfront payments, bad debts, and policy changes from upstream model providers can all squeeze the profits of the intermediary layer.

Tokens Now Have "Wholesalers" and "Retailers"

The basic chain of token distribution includes three types of roles.

Upstream are the model providers, including ByteDance's Seedance series, Alibaba's Qwen series, Zhipu's GLM series, Moonshot AI's Kimi series, DeepSeek series, etc. They are the original suppliers of tokens.

In the middle are agency platforms responsible for procuring resources from upstream model providers and distributing them to end-users. Their work is not just about reselling quotas; they also convert the interface protocols of different models into a unified API format, enabling downstream users to access multiple models through a single API Key.

Downstream are the actual consumers of tokens, including individual users, developers, enterprise clients, and possibly lower-tier distributors.

The value of this intermediary layer focuses on several areas: reducing network barriers through domestic direct connections; enabling a single codebase to adapt to multiple models; supporting both personal and corporate payments; potentially obtaining lower costs through bulk procurement; and aggregating models like GPT, Claude, DeepSeek, and Kimi on one platform to reduce the cost of repeated integration for developers.

Thus, token distribution appears to be asset-light, requiring neither the training of large models nor massive server clusters. The core assets become the API routing and scheduling system, upstream model resources, channel clients, and service capabilities.

The Surge in Call Volume is the Most Direct Fuel for This Business

For the token operation model to succeed, there must first be a sufficiently large consumption volume.

China's daily token call volume increased more than a thousandfold in two years, from 100 billion to over 140 trillion tokens. This expansion stems from the deployment of various vertical Agents and the embedding of generative AI into more business processes by enterprises.

IDC data presents an even more aggressive trajectory: the number of active intelligent agents in Chinese enterprises is expected to exceed 350 million by 2031, with a compound annual growth rate (CAGR) exceeding 135%. As the density and complexity of agent tasks increase, the annual growth rate in token consumption by agents is projected to exceed 30-fold.

This change is already visible in execution-oriented agents. The weekly token consumption of OpenClaw on the OpenRouter platform increased from 0.81T between February 2 and March 16, 2026, to 4.97T, with its share rising from 8.31% to 24.36%.

Once tokens become a mass-consumed commodity, their procurement, pricing, routing, and settlement naturally stratify. Model providers may not directly serve every client, and end customers may not be willing to integrate with each model individually, creating space for the intermediary layer.

The Cost-Effectiveness of Domestic Models Opens the Door for Token Export

The improvement in domestic large model capabilities is a key variable enabling token distribution to expand from domestic to cross-border markets.

Data from SuperCLUE shows that domestic models like ByteDance's Doubao and the DeepSeek series have achieved overall scores exceeding 70 points, narrowing the gap with leading overseas models like GPT-5.4 and Gemini. Models like Tongyi Qianwen, Kimi, and Zhipu GLM have also formed a relatively clear tiered structure.

According to OpenRouter data, for the week ending May 10, 2026, Tencent's Hy3 preview (free) topped the call volume list. Among the top 5, top 10, and top 20 models, there were 2, 6, and 9 domestic large models, respectively.

A more significant change occurred in Q1 2026. From February 9 to 15, the call volume of Chinese models on OpenRouter reached 4.12 trillion tokens, surpassing the 2.94 trillion tokens of US models for the first time in the same period. From February 16 to 22, the weekly call volume of Chinese models further increased to 5.16 trillion tokens. Among the top five models on the platform by call volume, four were from Chinese providers: MiniMax M2.5, Kimi K2.5, Zhipu GLM-5, and DeepSeek V3.2, collectively accounting for 85.7% of the total call volume of the top five.

The price advantage is also prominent. The input price for both MiniMax M2.5 and GLM-5 is $0.3 per million tokens, compared to $5 for Claude Opus 4.6. For output, MiniMax M2.5 is $1.1, GLM-5 is $2.55, and Claude Opus 4.6 is $25. The cost-effectiveness of domestic models becomes more pronounced in high-token-consumption scenarios like AI Agents and code development.

Global AI Resource Imbalance Makes Routing Platforms the "Transit Hubs"

Token distribution doesn't just solve price issues; it also addresses resource mismatches.

Leading overseas large models face barriers like regional access restrictions, compliance rules, and payment hurdles, preventing them from directly reaching certain user groups, including developers in mainland China. Similarly, high-quality domestic models expanding overseas encounter challenges in localization, channel development, and user acquisition.

This imbalance fuels the demand for cross-border flow, aggregated routing, and layered distribution.

OpenRouter is already a typical example. The volume of tokens processed on its platform increased from 5-7 trillion per week in 2025 to over 20 trillion per week by April 2026. Its annualized revenue in 2026 exceeded $50 million, a roughly fivefold increase from the over $10 million annualized revenue disclosed in October 2025.

Similar platforms exist domestically. Silicon Flow is a one-stop large model cloud service platform based on its self-developed inference engine for efficient inference acceleration, while also providing enterprise-grade large model services. As of December 2025, the platform had over 9 million registered users, more than 10,000 enterprise users, and over 150 models available.

Even politically connected capital in the US has entered this field. On May 5, 2026, WLFI, a cryptocurrency company closely linked to Trump and his family, partnered with WorldClaw to launch WorldRouter, integrating over 300 models including Claude, GPT, and Gemini. Settled in USD, its pricing is approximately 30% lower than official public rates.

Real Profits May Not Lie in "Resale Margins"

There are three ways to profit from token distribution.

The first is resale margins. Platforms purchase API quotas in bulk from upstream model providers and resell them at a markup to downstream clients. OpenRouter, which adds about a 5.5% premium to supplier costs, exemplifies this model.

The second is technological premium. Platforms use self-developed inference acceleration engines to reduce the cost per token. Even when selling at prices close to or lower than official rates, they can generate gross profit through computational efficiency advantages. Silicon Flow's SiliconLLM and OneDiff technologies improve language model inference speeds by 10 times and text-to-image efficiency by 3 times, reducing the cost of large model API calls to as low as 1/10th of the industry average.

The third is enterprise value-added services. The cost of deploying AI for enterprises isn't just in token unit prices; it also includes prompt engineering, multi-model selection, business system integration, workflow orchestration, operational scheduling, and employee AI skills development. As basic token prices decline, these hidden costs become more likely points for monetization.

Silicon Flow's enterprise-level MaaS platform follows this direction: providing enterprise users with capabilities across three layers—model training and fine-tuning, deployment and inference, and application development support—covering data processing, model fine-tuning, prompt engineering, and RAG, ultimately delivered in the form of standardized APIs to industries like energy, finance, and government.

Marketing, Short Videos, Gaming, and E-commerce Are Scenarios That Consume Tokens More Easily

To be profitable, token distribution must ultimately land in real-world scenarios.

Generative AI applications are entering industries like healthcare, transportation, and industrial manufacturing and are starting to participate in core processes like corporate decision support and strategic management. However, many enterprises have weak foundations for digital transformation, insufficient data asset accumulation, and limited computing power investment, making direct AI deployment challenging.

In contrast, marketing and advertising companies already possess clients and scenarios, especially in short videos, webtoons, gaming, and e-commerce. Their token consumption demand is more direct and sustained. For such companies, the opportunity isn't just about reselling model capabilities but embedding tokens into client workflows for content generation, ad placement, asset production, and video creation.

Investment leads also unfold along two main lines:

One category includes companies with strong model capabilities, such as Alibaba, Tencent Holdings, Kuaishou, Kunlun Tech, Zhipu, MiniMax, etc.

The other category includes companies with strong token consumption scenarios and quality client sources, especially those with overseas client resources and marketing scenarios, and a willingness to actively invest in AI marketing and AI videoization. Examples include EasyClick and BlueFocus.

Risks Are Also Concrete: Low Barriers, Upfront Funding Requirements, Upstream Dependence

The token distribution business model is asset-light, but its moat is not inherently deep.

Peer competition is the first risk. The technical barrier for distribution is relatively low. Once leading distributors with capital, client, and channel advantages enter, they can quickly replicate the model, compressing profit margins.

Upfront funding requirements and bad debts are the second risk. Distributors often offer monthly or quarterly settlements to downstream clients but need to fund the upfront purchase of API quotas from upstream providers. The larger the token consumption scale, the greater the funding pressure. If clients delay payment, bad debt risks amplify simultaneously.

Policy changes by upstream model providers are the third risk. Large model providers control API pricing and access rules and may adjust prices or tighten policies for third-party access. For the intermediary layer, this is the most difficult factor to control.

Связанные с этим вопросы

QWhat is the core change in the commercialization of AI applications as described in the article?

AThe core change is that AI application commercialization is extending from selling software and memberships to selling Token-calling capability. A new middle-layer market for Token distribution is forming, connecting upstream large model vendors with downstream developers, enterprises, and individuals, essentially creating a liquidity infrastructure for the wholesale-to-retail network of global Tokens.

QWhat are the three key roles involved in the Token distribution chain?

AThe three key roles are: 1) Upstream Model Providers (like ByteDance Seedance, Alibaba Qwen, etc.), who are the source of Tokens. 2) Middle-layer Agent Platforms, responsible for distributing resources to end-users and providing unified APIs. 3) Downstream Consumers, including individual users, developers, enterprise clients, and possibly sub-distributors.

QWhat are the three main profit models for Token distribution platforms mentioned in the article?

AThe three main profit models are: 1) Resale Margin: Buying API quotas in bulk from vendors and selling at a markup. 2) Technical Premium: Using proprietary inference acceleration engines to lower per-Token costs and profit from efficiency gains. 3) Enterprise Value-added Services: Offering services like Prompt engineering, multi-model selection, system integration, and workflow orchestration.

QAccording to the article, what is a key factor that has enabled the expansion of Token distribution from domestic to cross-border markets?

AThe key factor is the significant improvement in the capabilities and cost-effectiveness of domestic Chinese large models. Their performance scores have closed the gap with top overseas models, and their lower prices (e.g., $0.3 per million input Tokens for some Chinese models vs. $5 for Claude Opus) create a strong value proposition for high-Token consumption scenarios in the global market.

QWhat are the primary risks associated with the Token distribution business model?

AThe primary risks are: 1) Intensifying Competition due to low entry barriers. 2) Capital Commitment and Bad Debt risk, as distributors must prepay upstream vendors while offering credit terms to downstream clients. 3) Policy Changes by Upstream Model Vendors, who control API pricing and access rules, making this an uncontrollable variable for the middle layer.

Похожее

When Hyperliquid Steals Solana's 'Internet Capital Market' Script

The article "When Hyperliquid Steals Solana's 'Internet Capital Markets' Playbook" discusses Solana's struggles to maintain its "internet capital markets" narrative by 2026. Despite its initial success as a high-performance "Ethereum killer," SOL's price has underperformed, dropping significantly compared to other major cryptocurrencies. Solana's vision of a global, on-chain trading network for all assets is being challenged not primarily by Ethereum, but by Hyperliquid. Hyperliquid, evolving from a perpetual contracts platform into a dedicated financial infrastructure Layer 1, has become a major beneficiary of the shift of derivatives trading from centralized exchanges to on-chain. The article argues that for high-frequency financial trading, a specialized, performance-focused chain like Hyperliquid may be more suitable than a general-purpose ecosystem like Solana. Further compounding Solana's issues was a major $200+ million exploit on its key perpetual protocol, Drift, in April, which damaged market confidence. In response, Solana founder Anatoly Yakovenko heavily promoted the protocol Phoenix as a replacement, boosting its visibility but not its trading volume, which remains far behind leading platforms. Solana supporters have launched a public critique of Hyperliquid's decentralization, pointing to its limited validators and closed-source code. Critics, however, note Solana's own declining validator count and centralization metrics. This strategy has also caused internal friction, with developers of other Solana protocols expressing discontent over the foundation's perceived favoritism towards Phoenix. The conclusion is that Hyperliquid's rise represents a challenge to the "general-purpose blockchain" narrative, proving that the core of a capital market might be a specialized trading engine rather than a broad ecosystem. If Solana cannot regain dominance in derivatives, it risks remaining a "meme coin paradise" while its grand "internet capital markets" ambition slips away.

marsbit22 мин. назад

When Hyperliquid Steals Solana's 'Internet Capital Market' Script

marsbit22 мин. назад

When Hyperliquid Steals Solana's 'Internet Capital Markets' Playbook

The article discusses how Solana's grand vision of becoming an "Internet Capital Markets" platform is facing significant challenges in 2026, primarily from the unexpected rise of Hyperliquid. Solana's performance has weakened, with its token SOL experiencing the largest price decline among major cryptocurrencies. Its core narrative of building a global, chain-based marketplace for all assets is under pressure both internally and externally. Hyperliquid, originally a perpetual futures exchange, has evolved into a dedicated Layer 1 financial infrastructure network. Its focused, trading-centric approach is attracting capital and challenging the assumption that a "general-purpose" ecosystem like Solana is necessary for a capital market. Hyperliquid's success suggests that for high-frequency trading, superior performance, liquidity, and user experience may be more critical than a broad application ecosystem. Internally, Solana's strategy suffered a blow from a major hack on the Drift Protocol in April, resulting in over $200 million in losses. In response, Solana founder Anatoly Yakovenko has heavily promoted Phoenix as a new decentralized perpetual futures platform on Solana. While this boosted Phoenix's visibility, its trading volume remains far behind leading platforms. Solana's community has launched a rhetorical attack against Hyperliquid, questioning its decentralization due to its limited validator set and closed-source code. Critics, however, point out Solana's own decreasing validator count and increasing centralization of stake. This focus on "decentralization metrics" has also caused internal friction, with other Solana ecosystem developers expressing discontent over the foundation's perceived favoritism towards Phoenix. The article concludes that the rise of Hyperliquid represents a challenge to the "general-purpose blockchain" narrative, proving that an efficient trading engine might be more central to a capital market than a vast ecosystem. If Solana cannot regain dominance in the derivatives space, it risks remaining a "meme coin paradise" rather than achieving its ambition of hosting global assets.

链捕手29 мин. назад

When Hyperliquid Steals Solana's 'Internet Capital Markets' Playbook

链捕手29 мин. назад

Samsung Bets on Mobile HBM: AI Moves from Cloud to Palm, a New Frontier in Semiconductor Investment?

Samsung is betting on bringing high-bandwidth memory (HBM) technology from servers to mobile devices, aiming to enable powerful on-device AI features in smartphones and tablets. This move is driven by the booming AI market, where HBM demand from data centers has fueled Samsung's record profits, with HBM4 already in mass production. By integrating mobile HBM, Samsung seeks to transform user AI experiences—making tasks like image generation and real-time translation faster, seamless, and more private by processing data locally. Strategically, this allows Samsung to leverage its vertical integration in memory, advanced packaging, and Exynos processors to differentiate its Galaxy devices against competitors like Apple and Qualcomm. It also opens a new consumer growth avenue, reducing reliance on volatile server HBM demand alone. The initiative is expected to benefit the broader supply chain, boosting demand for advanced packaging materials, thermal solutions, and other components. While promising, risks include potential delays in mobile HBM mass production beyond 2027, high initial costs, and the cyclical nature of the memory market. Nonetheless, Samsung's push signals a broader industry shift toward hybrid cloud-edge AI computing, positioning it as a key player in defining the future of AI-powered devices and presenting a potential long-term investment theme in semiconductors.

marsbit40 мин. назад

Samsung Bets on Mobile HBM: AI Moves from Cloud to Palm, a New Frontier in Semiconductor Investment?

marsbit40 мин. назад

Trillion-Dollar Banking Giant Adjusts Portfolio: Buys XRP Heavily, Clears Out Solana

In a significant portfolio rebalancing move, Italian banking giant Intesa Sanpaolo, with $1.1 trillion in assets, has made a notable shift in its cryptocurrency holdings. According to disclosures from Q4 2025 to Q1 2026, the bank's total crypto exposure surged from $100 million to approximately $235 million. The most striking action was its first-time establishment of an XRP position, investing around $18 million through the Grayscale XRP Trust. This marks a milestone as one of the first major European banks to adopt XRP via a regulated investment vehicle. This move is part of a broader, systematic digital asset strategy. The bank also substantially increased its Bitcoin exposure via ETFs and initiated its first Ethereum investment through a staking trust. In a contrasting strategic pivot, Intesa Sanpaolo drastically reduced its Solana holdings by over 99%, nearly exiting its position in a Bitwise Solana staking ETF. This shift is interpreted as reflecting a institutional preference for assets perceived with lower regulatory and operational risk, especially following Solana's network stability issues and improved clarity for XRP after its legal settlements. The bank's actions highlight key drivers for institutional adoption: clearer regulations, the availability of compliant ETF products, and the search for portfolio diversification. This trend signifies crypto's evolving status from a niche experiment to a recognized component of mainstream asset allocation, with institutions favoring gradual, regulated entry points over direct token ownership.

marsbit1 ч. назад

Trillion-Dollar Banking Giant Adjusts Portfolio: Buys XRP Heavily, Clears Out Solana

marsbit1 ч. назад

Base Native Leveraged Prediction Market OmenX Officially Launches on Mainnet

Base-native leveraged prediction market platform OmenX has officially launched on mainnet. It currently supports up to 5x leverage, with plans to increase to 10x based on platform liquidity and market conditions. Unlike traditional prediction markets where users fully collateralize YES/NO positions and wait for settlement, OmenX aims to create a trading platform-like experience. Users can open leveraged positions on event outcomes, and actively trade, adjust, or hedge these positions before the event concludes for greater capital efficiency. Alongside the mainnet launch, OmenX introduced a "Hedge-to-Earn" campaign targeting existing users of other prediction markets (initially Polymarket). This initiative allows users to claim incentives or hedging benefits on OmenX based on their existing positions, aiming to introduce them to leveraged trading and active risk management. OmenX positions itself as a derivatives trading platform for prediction market assets. The team believes that as platforms like Polymarket mainstream prediction markets, event outcomes are becoming a new tradable asset class. The next phase of demand will focus on leverage, liquidity, and advanced trading tools. Post-launch, OmenX plans to expand supported market types, optimize liquidity, and develop APIs and additional trading tools. The team is also in discussions with investors and partners to secure resources for further development.

链捕手1 ч. назад

Base Native Leveraged Prediction Market OmenX Officially Launches on Mainnet

链捕手1 ч. назад

Торговля

Спот

Фьючерсы

Обсуждения

Добро пожаловать в Сообщество HTX. Здесь вы сможете быть в курсе последних новостей о развитии платформы и получить доступ к профессиональной аналитической информации о рынке. Мнения пользователей о цене на AI (AI) представлены ниже.