By mid-2026, key operational metrics in the AI industry, including revenue and API call volume, continue to grow at a high rate. However, a detailed examination of the financial and business structures of various companies reveals different business models underlying the same set of growth figures. This article outlines four structural paradoxes in the current AI business model: cost, hierarchy, responsibility, and the open/closed source model. These four paradoxes point to the same underlying trend: intelligence is rapidly becoming commoditized, while profits remain concentrated in only a few segments.
The Cost Paradox: The Cheaper the Token, the Heavier the Bill
When GPT-4 was released in March 2023, the price per million tokens was $30 for input and $60 for output. Companies building customer service systems with it could incur monthly API costs exceeding the entire salary bill of their support team. Three years later, the price for equivalent intelligence levels has dropped by over 95%.
Common sense would suggest that AI expenditures should shrink accordingly, but the opposite is true. In March 2026, global weekly token consumption reached 20.4 trillion; daily consumption in China exceeded 140 trillion tokens, representing a thousand-fold increase since early 2024. This phenomenon aligns with Jevons Paradox in economics: as steam engine efficiency increased and coal consumption per unit of work decreased, total coal consumption actually rose. Falling token prices have not led to a decrease in total bills.
The reason lies in the change in demand structure. When prices become low enough, a large number of tasks that were previously uneconomical begin to be handled by AI. A company's monthly expenditure might rise from $50,000 to $500,000, but the volume of tasks completed is also unimaginable compared to the past. More crucially, there is a shift in the usage paradigm: past conversational interactions were limited by human time, whereas AI Agents can run 24/7. One employee might have 10 Agents working behind them, potentially leading to dozens of times the usage of purely manual interaction. Research shows that token consumption for agent programming tasks can be about a thousand times that of simple Q&A. A one-order-of-magnitude drop in unit price can lead to a two-order-of-magnitude increase in total consumption.
The distribution of costs within the industry chain is also noteworthy. Taking a $20 monthly Claude subscription as an example, after deducting computing power, channel, and operational costs, the model company's net profit margin is relatively thin; sometimes the profit flowing to the chip layer is not less than that of the model company. Users pay for intelligence, but profits largely remain upstream.
Faced with bill pressure, the industry is shifting from saving on unit price to saving on tasks: routing simple tasks to smaller models, setting task budgets for Agents to prevent idle running; the accounting unit is also shifting from price per million tokens to the total cost of completing a task. Practice shows that with the same model and volume, companies with capabilities in task decomposition, result verification, and process integration achieve significantly higher output efficiency. Meanwhile, subsidy models are also fading: Claude Code exited its low-cost subscription, Copilot switched to usage-based billing. The Jevons effect hasn't disappeared: the industry ultimately finds that what needs optimization is not the price list, but the tasks themselves.
While bills increase, the efficiency gains for enterprises often remain invisible. Solow famously pointed out: computers are everywhere except in the productivity statistics. The electric motor serves as a ready reference: available since the 1880s, it took US factories about forty years to realize its efficiency gains, as most companies simply replaced steam engines with electric motors without changing the transmission structure, whereas the value of electricity precisely required a complete rearrangement of production lines.
The situation with today's AI is the same. Databricks founder Ghodsi shared a case: a data connector that historically took nine months to develop, saw only a month-and-a-half reduction in the first attempt using a large language model; it wasn't until someone deconstructed and rebuilt the process that delivery in one quarter was achieved. His conclusion: the bottleneck isn't the model, it's the organizational structure.
The cost paradox therefore has two layers: Jevons Paradox states that the cheaper the token, the heavier the total bill; Solow Paradox states that without changing processes, efficiency doesn't show up in the statistics. The former requires optimizing tasks, the latter requires re-engineering processes.
The Hierarchy Paradox: The Application is King vs. The Application is Dead
The consensus of the mobile internet era was that the application is king: whoever captures user time has the opportunity to monetize through advertising and recommendations, with high-frequency apps beating low-frequency ones; app opens equated to value.
This round of AI has seen a reversal. Jensen Huang has publicly stated that economic value will ultimately be realized at the application layer, and successful applications will pull the entire industry. According to this logic, capital should cluster at the application layer. But by the summer of 2026, the mainstream view in the primary market was that the application is dead: the rapid evolution of foundation models means that thin applications can be overridden at any time by natural extensions of model capabilities. Capital is flowing first to models and infrastructure. The application layer still receives funding, but far from being seen as a given as in the previous cycle.
The industry ledger partially supports this capital choice. Breaking down the generative AI ecosystem's annualized revenue of approximately $400 billion, the chip layer captures about 70% of revenue and nearly 80% of gross profit; the application layer revenue is about $60 billion, with gross margins mostly between 0-30%, while chip companies like NVIDIA can achieve gross margins around 70%. The pyramid structure where the application layer held the largest share in the cloud computing era has inverted in the AI industry.
There is also differentiation within the model layer. Companies focusing on enterprise APIs and relying on rolling snowballs of usage from existing customers see revenue skyrocket while their inference gross margins can recover from deep losses to around 70%. Meanwhile, companies burdened with a large free user base perform far worse operationally than their headline revenue suggests. Surplus is not evenly distributed across the entire model layer, but concentrates on a few high-retention, high-margin paths. The capital market's neglect of applications has both cognitive bias and financial justification.
AI products no longer compete for user dwell time, but for whether users are willing to entrust an important task to them. Universal ChatBots and universal Agents are likely the main battleground for leading model companies, leaving limited room for startups to compete head-on. The viable direction that remains is embedding intelligence into specific scenarios and redoing the product according to how problems actually occur. Users don't stay because a product integrates a certain model, but because the product can understand materials, remember context, and connect to follow-up services. Harvey has operated deeply in the legal field for years; its accumulated records of legal reviews and precedent data are not easily replicable with short-term investment.
This leads to a simple criterion: when the foundation model strengthens by one generation, does the product's value increase or decrease? If it decreases, it's likely just packaging model capabilities; if it increases, it typically holds private context, business processes, complete delivery capabilities, or channels. The point of value realization hasn't disappeared; the criteria for success have simply changed: who can make themselves a harder-to-replace link as the model gets stronger.
The Responsibility Paradox: Profit Follows Responsibility
First, consider two sets of data. Anthropic's Annual Recurring Revenue (ARR) grew from about $1 billion in December 2024 to over $47 billion by May 2026, a more than forty-fold increase in a year and a half. OpenAI's official figures are: $2 billion ARR in 2023, $6 billion in 2024, over $20 billion in 2025, and exceeding $25 billion by the end of February 2026.
But growth rates only prove the existence of demand, not the ability to retain profits. Also experiencing high growth, Anthropic's inference infrastructure gross margin improved from 38% a year ago to over 70%; while OpenAI's internal document updated in November 2025 postponed the timeline for positive free cash flow to around 2030. Within the same AI industry, the growth trajectories of these two types of companies have already diverged. Regarding this divergence, three main explanations exist in the market, each with some merit but also missing key points.
The first explanation is that time solves everything: inference costs are falling by an order of magnitude each year, and the computing bill suppressing margins today might not be a problem in two or three years. Cloud computing, storage, and bandwidth all followed a path of early burning cash for scale, with profits naturally appearing after cost curves flattened. However, this explanation struggles with a fact: the two years of fastest cost decline coincided with the most pronounced profit divergence. If cost decline was the cure, industry profits should have improved across the board, not half improving and half deteriorating. The reason is that cost declines are universal; all vendors benefit simultaneously, and also simultaneously pass on the benefits to users through price reductions. Relying on cost deflation for profit is equivalent to admitting a lack of pricing power.
The second explanation is that the application layer has moats: products like customer service and ticketing are deeply embedded in client systems, with accumulated data and migration costs forming defenses. This held true five years ago, but the core selling point of the new generation of AI tools is precisely that migration costs approach zero—they don't replace the client's systems, but directly take over existing processes. As for data accumulation, one major capability leap in a foundation model can significantly devalue industry knowledge bases built over years. The most illustrative case is Intercom: it launched Fin, charging per actual problem solved at $0.99 per ticket, with no charge if unresolved. This suggests, to some extent, that the old model of selling customer service tools per seat is becoming unsustainable.
The third explanation is that charging based on outcomes is infeasible: it's hard to attribute whether a ticket was resolved or a sales conversion is due to AI. Therefore, charging per token or per seat is the norm, while outcome-based pricing is just a marginal experiment. The attribution challenge is real, but it's more likely to be a source of pricing power rather than an obstacle: precisely because attribution is difficult to achieve, vendors who can reliably establish attribution in specific fields turn the attribution capability itself into a barrier. Moreover, attribution is naturally clear in some areas: whether a ticket is resolved, whether a collection call results in repayment, whether an insurance claim is settled—all can be contractually verified. Outcome-based pricing will coexist with other models for a long time, applied in areas with clear attribution.
After excluding these three explanations, the remaining structural variable is only one: how broad a scope and how strong a constraint of responsibility a vendor is willing to bear for the outcomes. The heavier the responsibility, the higher the gross margin, and the larger the budget source. Selling by Token competes for IT budgets, while selling by outcome targets labor budgets: if a department has a monthly labor cost of $100,000, and AI can achieve equivalent results for $10,000, the pricing space is completely different.
This also answers another question: why businesses that are easiest to scale are also the hardest to retain profits for. Services like customer support, product description generation, voice-to-ticket have large volumes, are easy to demonstrate and sign contracts for, but are precisely the ones with the lowest responsibility content and are most easily directly covered by foundation models. Conversely, fields like law and medicine are low-frequency, heavily regulated, with high error costs. They involve heavy delivery, long sales cycles, and less sexy narratives, but they sell processes, data, and responsibility, which cannot be erased by a single model upgrade. Harvey's choice of law is no coincidence.
Therefore, to assess an AI company's position, instead of focusing on growth rate and model capabilities, it's better to answer two questions: what portion of the outcomes is it willing to take responsibility for, and is the attribution for that portion solid enough. The first is willingness, the second is capability; lacking either prevents earning an outcome premium. This criterion doesn't guarantee accuracy, as attribution technology is still in its early stages, and corporate purchasing habits change slower than optimistic expectations. But directionally, profits will likely not flow to the fastest-growing companies, but to those willing to take responsibility and capable of bearing it.
The Open/Closed Source Paradox: Open Source Gains Traffic, Closed Source Wins Revenue
Looking solely at usage, open source already holds an advantage. The top five models by traffic on the OpenRouter platform are all open-weight models; in terms of developer adoption rate, open source is 79%, closed source 71%. Chinese models are the main force in this round of open source expansion, with their share on that platform rising from less than 2% a year ago to over 45% by April 2026, accounting for 61% of token usage among the top ten models.
Looking solely at revenue, closed source's advantage is equally clear. An a16z survey of Global 2000 enterprise CIOs shows: enterprise spending on open source models dropped from 19% last year to 11% this year, with closed source at 89%, averaging about $7 million in annual large model budget per enterprise. The same structure appears within the OpenRouter platform: Anthropic captures 46% of the revenue share with 12% of the token share.
With only single-digit performance gaps and prices 5-20 times cheaper, why do enterprises still choose closed source? Because what enterprises purchase is never just intelligence itself, but also reliability, technical support, compliance guarantees, and a responsible entity when problems arise. 'Good enough' and 'daring to use in core business' are two different things. Another detail from the survey is noteworthy: the Total Cost of Ownership (TCO) for open and closed source is converging, with the speed of closed source price reduction outpacing the speed of open source building trust. Hybrid deployment has thus become the norm: using open source for experimentation and edge scenarios, and closed source for critical paths.
The deeper change is that the open source camp itself doesn't rely on model sales for profit either. The motivations of different vendors for open-sourcing vary. Some aim for a platform ecosystem, trading open weights for developer entry points and industry standard influence, subsidized by their core business; others use open source models as a customer acquisition channel for their cloud business, driving developers onto their cloud and thereby boosting continuous cloud revenue growth. The model layer is rapidly commoditizing, with profit pools shifting to the two ends: upstream computing power, and downstream orchestration, data, and services.
The challenge for the closed source camp is that its premium is strongly tied to the capability gap; each step open source takes closer loosens the per-token pricing power. But scarcity hasn't disappeared, it's just layered: cutting-edge models remain scarce, while 'good enough' intelligence is becoming commoditized. Therefore, to assess a company's position, the question isn't whether it chooses open or closed source, but what it relies on for revenue, and whether its revenue increases or decreases if open source advances further.
This article is from the WeChat public account "Tencent Research Institute" (ID: cyberlawrc), author: Liu Qiong






