Token Devours One-Third of Payroll, Silicon Valley's AI Bill is Spinning Out of Control

marsbitPublished on 2026-07-06Last updated on 2026-07-06

Abstract

The article discusses the dual reality of AI token costs in Silicon Valley. While research firm SemiAnalysis reports spending 30% of its employee salary budget on internal LLM tokens—translating to massive productivity gains like converting complex Excel models in minutes—other giants are struggling with ballooning, uncontrolled AI bills. Uber exhausted its annual AI budget in months after rapid engineer adoption, and Microsoft is cutting third-party AI tools due to high costs. NVIDIA's CEO argues tokens are becoming "means of production" and plans substantial AI budgets per engineer. Despite current cost concerns, the analysis emphasizes that cost collapse is just beginning. Through software optimizations (like 14x throughput boosts) and next-gen hardware (e.g., GB300 NVL72 with 17-32x H100 performance), real token costs can fall far below list prices. Anthropic's gross margins reportedly soared as token prices dropped. Gartner predicts a >90% inference cost drop by 2030. The piece highlights a split: massive AI capex ($740B announced) contrasts with tech layoffs and minimal measured economic impact so far. The transition mirrors past infrastructure shifts—investment precedes widespread productivity. For early adopters like SemiAnalysis, tokens already deliver high leverage; for others, the choice is to adopt now or risk falling behind.

Only $0.99 per million Tokens.

This is the real cost on the bill of SemiAnalysis—Silicon Valley's most hardcore semiconductor research firm.

But what's even more shocking is this number: Internal large model Token expenditure already accounts for 30% of total employee salaries.

It sounds like a lot—but flip the calculation: the output bought with this money previously required several times the human resource cost to cover. Per capita consumption nears 5 billion Tokens monthly, over five times Meta's per capita level, with core contributors' monthly consumption exceeding 100 billion.

Tasks that used to take junior analysts hours to complete, like converting Excel models or creating financial report charts, are now done in minutes, costing just a few dollars.

SemiAnalysis's own assessment hits the nail on the head: This isn't a 10% efficiency boost; it's the unit economics of professional services being rewritten.

Research firms, hedge funds, law firms—for all industries reliant on human intellect, Token expenditure reaching 20-30% of payroll is only a matter of time.

NVIDIA CEO Jensen Huang is more anxious than anyone.

At this year's GTC conference, he put it bluntly: An engineer with a $500k salary spends less than $250k on Tokens by year-end?

"I would be absolutely furious."

He plans to give every NVIDIA engineer a Token budget equivalent to six months' salary, and have 75,000 employees work alongside 7.5 million AI agents.

Not using AI? Huang says it's no different than a chip designer insisting on paper and pencil.

Token is no longer just a tool; it's becoming the "means of production" of the new era.

But the Other Half of Silicon Valley is Furious Over the AI Bill

Interestingly, while SemiAnalysis is saving real money with Tokens, giants in the Valley are tearing their hair out over AI bills.

Uber is the classic case.

Late last year, the company promoted Claude Code to 5,000 engineers, even creating leaderboards—more usage meant higher rank, fueling internal competition.

It worked too well: Engineer adoption was 32% in February, skyrocketed to 84% in March, and by April, 95% of engineers used AI monthly, with 70% of submitted code AI-generated. The annual budget? Already spent.

The CTO said they had to "redo the budget from scratch." Later it got stricter—Bloomberg reported Uber set a $1,500 monthly Token cap per employee, requiring special approval to exceed.

But COO Andrew Macdonald admitted on a podcast: AI usage is indeed rising, but its connection to consumer feature innovation... isn't visible yet.

Microsoft's situation is even more bizarre. Last month, The Verge reported Microsoft is canceling most Claude Code licenses, shifting to its own GitHub Copilot CLI.

The reason is simple: Money was being spent faster than value was being produced.

NVIDIA's VP of Applied Deep Learning, Bryan Catanzaro, was more direct in April: "For my team, compute costs far exceed employee costs."

An MIT 2024 study: In jobs primarily involving visual content, AI automation is economically viable in only 23% of scenarios.

In the remaining 77%, hiring people is cheaper than using AI.

There are even engineers complaining about AI agents "destroying his database and network" during use—he called it the cost of "overuse."

Sky-high budgets, runaway usage, constant mishaps—Silicon Valley is in the most fractured phase of AI economics.

On one side, unprecedented productivity gains; on the other, bills inflating at an equally unprecedented pace.

The Cost Collapse Has Only Just Begun

But SemiAnalysis's core argument is: Don't focus on today's price; the cost collapse has just started.

First, the software side.

Running DeepSeek R1 on a B300, with pure software optimizations via wideEP, disaggregation, and MTP, single GPU throughput can jump from a baseline of 1,000 tokens/second to 14,000 tokens/second—a 14x boost, purely through code.

Now, the hardware side.

An optimally configured GB300 NVL72 has 17x the throughput of an H100, jumping to 32x when switching to FP4 precision.

Opus 4.7 is priced at $5 per million input, $25 per million output, which seems expensive.

But due to agent workloads having an input-to-output ratio as high as 300:1, plus over 90% cache hit rate, the actual blended cost is compressed to $0.99.

Less than one-fifth of the list price.

Combine software and hardware, and one conclusion is hard to avoid: The gross margin expansion of large models isn't a one-off pricing coincidence; it's a structural trend.

Anthropic's ARR surged from $9 billion to over $44 billion this year, with gross margins jumping from 38% to over 70%—Tokens are getting cheaper, but the sellers are getting richer.

A Gartner report from March corroborates this: By 2030, the inference cost for trillion-parameter models will be over 90% lower than in 2025.

SemiAnalysis's judgment is clear: If you want to predict Token prices in 2027, the answer is one word—down.

The Money is Spent. What's Next?

This is precisely the most fractured aspect of AI today: Global tech companies have announced $740 billion in AI capital expenditure this year, a 69% surge from last year; simultaneously, tech industry layoffs are already outpacing last year's total.

Money is burning, people are being laid off, but Goldman Sachs' chief economist told a blunt truth—The actual economic impact of AI has been essentially zero so far.

It's not that AI is ineffective, but the growing pains of every infrastructure revolution: First, spend to build the pipes, then wait for the water to flow.

It was true for the electrical grid, the internet, and AI is no exception.

The only difference is that the speed of pipe-laying and the speed of water arriving are on a scale previous generations never saw.

SemiAnalysis is already on the side where the water is flowing—30% of payroll is buying several-fold output leverage, and the cost curve is still plummeting.

As for other companies: Wade across the river now, or chase after the cities are already built on the other side.

References:

https://x.com/SemiAnalysis_/status/2070915305858007345

This article is from WeChat public account "New Zhiyuan", author: ASI Revelation, editor: Solomon

Trending Cryptos

Related Questions

QWhat percentage of employee salary does the internal large model token expenditure account for at SemiAnalysis?

AThe internal large model token expenditure accounts for 30% of total employee salary at SemiAnalysis.

QWhat is the key argument of SemiAnalysis regarding the current AI economic situation?

ASemiAnalysis argues that we should not focus on today's AI prices, as a cost collapse has just begun, driven by both software and hardware optimizations leading to structurally lower token costs.

QAccording to the article, what was a major issue Uber faced with its AI tool adoption among engineers?

AUber faced a major issue where the adoption of its AI tool, Claude Code, was so successful that usage surged from 32% of engineers in February to 95% in April, exhausting the annual budget within months.

QWhat does NVIDIA CEO Jensen Huang say about engineers who do not use AI?

AJensen Huang says that not using AI is equivalent to a chip designer insisting on using paper and pencil.

QWhat is the projected trend for the inference cost of trillion-parameter large models by 2030 according to Gartner?

AAccording to Gartner's report, by 2030, the inference cost for trillion-parameter large models is projected to decline by over 90% compared to 2025.

Related Reads

The 'Saving U.S. Treasuries' Baton Pass: Bessent Fumbled Last Week, This Week It's Wash's Turn

"Rescuing US Treasuries" Relay: After Bessent's Miss, All Eyes Are on Walsh Last week, US Treasury Secretary Bessent's announcement to at least double long-term Treasury buybacks failed to sustainably lower yields, which quickly rebounded. The market response saw a drop in the dollar alongside surges in gold and Bitcoin, interpreted as a "pressure release valve" for anxiety. The focus now shifts to Fed Chairman Walsh's upcoming Jackson Hole speech. Markets are highly sensitive to his message, seeking clarity on the Fed's policy response to stubborn inflation and worsening fiscal conditions. Analysts warn that a lack of new guidance could disappoint markets and worsen the sell-off in long-dated bonds. Analysts question the scale of Bessent's operations, noting they are too small relative to the overall debt market and do not constitute quantitative easing. A key issue is the Fed's massive holdings of long-term bonds, which distorts the market. With the Fed holding low-yielding short-term bonds that are losing money relative to its policy rate, discussion is growing around a potential Fed-led "Operation Twist." This would involve selling short-term bonds to buy long-term ones, aiming to lower long-end yields without expanding the balance sheet. The upcoming PCE inflation data will set the stage for Walsh's speech. However, the window for action is narrowing amid political pressures. A critical threshold is the 30-year yield at 5%; holding above it could increase stress on the dollar and leveraged sectors. Overall, the article suggests that without coordinated Fed action to anchor inflation expectations, Treasury interventions may ultimately fail, with investors increasingly looking to assets like gold as hedges.

marsbit17m ago

The 'Saving U.S. Treasuries' Baton Pass: Bessent Fumbled Last Week, This Week It's Wash's Turn

marsbit17m ago

Hyperliquid's Compliance Journey: From Permissionless to Permissioned via HIP-3

Hyperliquid’s Compliance Path: From Permissionless to Permissioned HIP-3 Hyperliquid currently blocks U.S. access because its permissionless, on-chain infrastructure conflicts with U.S. market structure laws, which restrict futures trading to registered exchanges, clearinghouses, and brokers. Through its Hyperliquid Policy Center (HPC), the project is advocating for regulatory modernization, proposing that regulated entities be allowed to build products on HyperCore (its exchange and clearing layer) while fulfilling their compliance obligations. The platform’s modular stack separates roles like a traditional exchange (DCM), clearinghouse (DCO), and broker (FCM), but reconstructs them on-chain with code. This enables permissionless access, self-custody, and 24/7 global trading, but clashes with U.S. rules requiring KYC, specific margin models, and custodial arrangements. To resolve this, HPC is engaging with U.S. regulators (CFTC, SEC) to seek clarity that deploying on-chain software does not itself trigger licensing, and to establish exemptions allowing non-custodial wallets to route users to regulated derivatives. Recent political signals suggest openness to this approach. On the technical side, Hyperliquid Labs has introduced permissioned HIP-3 deployers on testnet. These allow regulated entities to launch markets, perform KYC, and whitelist compliant users. While these create separate order books, whitelisted market makers can bridge liquidity between them, ensuring deep, shared liquidity across the same L1. Features like payload-based “PA” permissions enable DEX-level account controls (e.g., reduce-only orders), mirroring traditional broker authorities. The strategy is not to open the native, permissionless front-end to U.S. users, but to position Hyperliquid as neutral infrastructure that U.S. regulated firms can use while meeting their legal duties. This paves a compliant path for U.S. investor access while preserving the protocol’s core, permissionless nature.

marsbit42m ago

Hyperliquid's Compliance Journey: From Permissionless to Permissioned via HIP-3

marsbit42m ago

Two Funding Rounds in Three Months: The Chinese Version of Palantir is on Fire

Investment Community AI has learned that Beijing Zhongshu Ruizhi Technology Co., Ltd., a domestic industrial-grade causal intelligence and high-reliability decision-making AI company, has recently completed a strategic financing round worth hundreds of millions of RMB. This round saw participation from China Internet Investment Fund, Suzhou Chuangtou National Social Security Fund, Financial Street Capital, ICBC Capital, Kunlun Capital, among others, with existing shareholders also increasing their investment. This follows a Series B funding round in the hundreds of millions completed just three months prior. The rapid succession of two major funding rounds signifies strong market recognition of the company's underlying original technology and scaled commercial implementation. Often referred to as the "Chinese version of Palantir," Zhongshu Ruizhi is entering a new phase of accelerated technological iteration, widespread scenario replication, and scaled performance release, mirroring the explosive growth of China's AI market. Founded in April 2020 by Dr. Han Han, a Tsinghua University Ph.D. and former core drafter of national AI policies, the company is mission-driven to "move AI from the digital world to the physical world." It focuses on the high-reliability, strong-decision industrial AI track and enterprise-grade AI Agent full-stack infrastructure. The team tackles the challenge of applying AI to China's vast and complex industrial and energy systems by developing a new intelligent operating system from scratch. Its core technological breakthrough lies in three proprietary底层 technologies: meta-causal cognitive theory, causal models, and a dynamic ontology engine. These address critical pain points of generative large models in industrial settings—such as AI hallucinations, insufficient reasoning, lack of temporal logic, unverifiable decisions, and multi-source rule conflicts—thereby providing trustworthy, explainable, and executable智能决策 capabilities. Commercially, Zhongshu Ruizhi has achieved scaled deployment, serving over 50 central state-owned enterprises and industrial groups in sectors like power, petroleum, and aerospace, with implementations in more than 800 highly complex production scenarios. The company reported doubled revenue in 2025, demonstrating strong self-sufficiency and a viable business model—a rarity among new-generation AI firms. The latest funds will be allocated towards advancing foundational theoretical research, replicating successful application models to expand market presence (including overseas), and attracting top-tier talent. Lead investor China Internet Investment Fund highlighted that in the current shift from general AI capability contests to deep industrial empowerment, industrial-grade causal intelligence is crucial for building China's modern digital foundation and fostering new quality productive forces. They expressed support for the company's efforts to define decision-making paradigms and trustworthy standards for industrial intelligence, aiming to secure a rule-making voice in the global physical AI arena.

marsbit52m ago

Two Funding Rounds in Three Months: The Chinese Version of Palantir is on Fire

marsbit52m ago

The Biggest Political Economy Question in the AI Era: As Robots Become More Capable, How Do Humans Share the Value?

In the AI era, the most pressing political economy question is: as machines become increasingly capable, how can humanity share in the value they create? An article originally critiquing China's tech focus has sparked a deeper debate on this global challenge. Historically, industrial progress improved efficiency but still relied on human labor for wealth creation and distribution. AI is fundamentally different—it is now replacing cognitive and knowledge work. As AI and robots take over more tasks, economic growth may continue while direct human participation in value creation shrinks, creating a core tension between productivity gains and widespread income generation. The issue is not unique to China. While leading tech companies amass enormous wealth, labor's share of income is declining globally. The core problem is a broken link: technological innovation and corporate profits are not translating into sufficient consumer income and demand. Three potential paths forward are outlined: a traditional capitalist model where profits primarily go to capital owners; a state-capitalist approach with public investment in AI; and more innovative models like digital sovereign wealth funds, universal shareholding, or AI-era basic income schemes to directly distribute AI-generated value. The future competitive advantage may lie not just in technological supremacy, but in which society can build a new, inclusive distribution system for the intelligent economy. The ultimate challenge is ensuring that as AI creates value, humans have a means to obtain income and share in the resulting widespread social benefits.

marsbit1h ago

The Biggest Political Economy Question in the AI Era: As Robots Become More Capable, How Do Humans Share the Value?

marsbit1h ago

Generating Profits for Seven Consecutive Quarters, Emerging Markets Carry Trade Outperforms Everything

For the seventh consecutive quarter, dollar-funded emerging market carry trades have delivered positive returns, marking the longest winning streak since 2008. According to Bloomberg's index, this strategy has gained approximately 22% since late 2024, outperforming U.S. Treasuries, emerging market sovereign, and corporate dollar debt. The core of the trade involves borrowing low-interest currencies like the U.S. dollar, euro, or yen to invest in high-yielding emerging market assets, such as Turkish lira bonds offering over 40% returns. Returns were amplified by favorable currency moves, with the dollar weakening against most emerging market currencies and other traditional funding currencies. For instance, the trade gained 48% on the Colombian peso in the past year. A key test came in August 2024 with a historic joint U.S.-Japan currency intervention, which caused only a modest 1% dip in the carry trade risk premium as investors shifted funding from the yen to the euro and Swiss franc. Looking ahead, the primary risk is the timing of Federal Reserve policy changes. While persistent inflation allows the Fed to hold rates, a rapid rise in long-term U.S. yields could threaten the trade. Another concern is crowding, as massive inflows increase vulnerability to a sudden reversal. High interest rates in regions like Latin America and Eastern Europe, supported by external factors like Middle East tensions and energy prices, continue to sustain the opportunity. Major investors remain engaged, favoring currencies like the Mexican peso, South African rand, and Turkish lira.

marsbit1h ago

Generating Profits for Seven Consecutive Quarters, Emerging Markets Carry Trade Outperforms Everything

marsbit1h ago

Trading

Spot

Hot Articles

Discussions

Welcome to the HTX Community. Here, you can stay informed about the latest platform developments and gain access to professional market insights. Users' opinions on the price of S (S) are presented below.

活动图片