Stop Staring at GPUs: CPUs Are Becoming the 'New Bottleneck' in the AI Era

marsbitPublié le 2026-04-13Dernière mise à jour le 2026-04-13

Résumé

In the AI era, while GPUs have long been the focus for computational power, the narrative is shifting as CPUs are increasingly becoming the new bottleneck. By 2026, system performance is more dependent on execution and scheduling capabilities, with CPUs playing a critical role in enabling AI operations. A supply crisis is emerging, with server CPU prices rising about 30% in Q4 2025 due to high demand and production constraints, as GPU orders compete for limited semiconductor capacity. Companies like Google and Intel have deepened collaborations, and Elon Musk is investing in custom CPU solutions for his ventures, highlighting the strategic importance of CPU infrastructure. The shift is driven by the rise of agentic AI, where CPUs handle tasks such as multi-step reasoning, API calls, and data I/O, accounting for 50–90.6% of total latency in intelligent workloads. Expanding context windows in AI models further strain GPU memory, necessitating CPU offloading for key-value cache management. Major players are adopting varied strategies: Intel is strengthening its Xeon processor line and partnerships; AMD is benefiting from increased demand, with server CPU revenue surpassing 40%; and NVIDIA is designing CPUs like Grace to optimize GPU-CPU synergy through high-speed interconnects. The industry is witnessing a rebalancing of compute infrastructure, with CPUs gaining prominence as essential enablers of scalable AI agent systems. By 2030, the CPU market is projected to double to ...

In the years of AI's rapid advancement, the industry has been largely dominated by one logic: computing power determines the ceiling, and GPUs are the core of that computing power.

However, entering 2026, this logic is beginning to shift: model inference is no longer the sole bottleneck; system performance increasingly depends on execution and scheduling capabilities. GPUs remain important, but the key factor determining whether AI can 'run' is gradually shifting to the long-overlooked CPU.

On April 9th, US local time, Google and Intel reached a multi-year agreement to deploy Intel's 'Xeon processors' at scale in global AI data centers, precisely to break this bottleneck. Intel CEO Pat Gelsinger (Note: The original Chinese name 陈立武 is likely a misattribution; the current CEO is Pat Gelsinger) stated bluntly that AI runs on the entire system, and CPUs and IPUs are the key to performance, efficiency, and flexibility. In other words, the CPU, which has been treated as a 'supporting role' for the past two years, is now choking the 'neck' of AI scaling.

Intel CEO Pat Gelsinger stated on social media: Intel is deepening its collaboration with Google, expanding from traditional CPUs to AI infrastructure (such as IPUs), to jointly advance AI and cloud computing capabilities.

The CPU is no longer just a passive supporting component but is becoming one of the key variables in AI infrastructure.

01 A 'Silent' Supply Crisis

While everyone was watching GPU delivery cycles, the tension in the CPU market had already quietly peaked.

According to the latest reports from multiple IT distributors, in the fourth quarter of 2025, the average selling price of server CPUs increased by about 30%. Such an increase is very rare in the relatively mature CPU market.

AMD's Data Center Group head, Forrest Norrod, revealed that CPU demand growth over the past three quarters has been beyond imagination. Currently, AMD's delivery lead times have extended from the original eight weeks to over ten weeks, with some models even facing delays of up to six months.

This shortage is primarily caused by a 'secondary effect' triggering a resource crunch. Industry insiders indicate that due to the extreme tightness of TSMC's 3nm production lines, wafer capacity originally allocated for CPUs is constantly being squeezed out by more profitable GPU orders. This has led to an ironic situation: AI labs have enough GPUs but find they cannot buy enough top-tier CPUs on the market to 'drive' these graphics cards.

Among those caught in this wave of CPU buying frenzy is Elon Musk.

Intel CEO Pat Gelsinger confirmed on social platforms that Musk has commissioned Intel to design and manufacture custom chips for his 'Terafab' project in Texas. This massive project aims to provide a unified computing base for xAI, SpaceX, and Tesla.

Musk's trust in Intel is largely because Intel is trying to embed itself into every layer, from ground-based data centers to orbital computing in space.

For Intel, this is undoubtedly a shot in the arm. Some industry analysts predicted that AMD's revenue share in the server CPU market would surpass Intel's in 2026, but Intel's deep inertia and manufacturing capabilities within the x86 ecosystem remain chips that major customers like Musk cannot ignore.

This kind of deep cross-industry bundling is elevating the competition in the CPU market from a pure parameter contest to a game of ecosystem and supply chain stability.

02 Why Has the CPU Become the 'Bottleneck'?

The core reason the CPU has suddenly become a bottleneck is that the work it needs to handle has fundamentally changed in the age of agents.

In the traditional chatbot model, the CPU is primarily responsible for scheduling and data processing, while the GPU handles the core inference computation. Since compute-intensive tasks are concentrated on the GPU side, overall latency is usually dominated by the GPU, and the CPU rarely becomes a performance bottleneck.

But agent workloads are completely different. An agent needs to perform multi-step reasoning, call APIs, read and write databases, orchestrate complex business flows, and integrate intermediate results into a final output. Tasks like search, API calls, code execution, file I/O, and result orchestration mostly fall on the CPU and host system side. The GPU is responsible for token generation (i.e., 'thinking'), while the CPU is responsible for translating the 'thoughts' into actual actions.

A paper published by Georgia Tech scholars in November 2025, 'A CPU-Centric Perspective on Agentic AI,' quantified the latency distribution in agent workloads. The study found that the time consumed by tool processing on the CPU side accounts for 50% to 90.6% of the total latency. In some scenarios, the GPU is ready to process the next batch of tasks while the CPU is still waiting for tool calls to return.

Another key factor is the rapid expansion of context windows. In 2024, mainstream models mostly supported 128K to 200K tokens. Entering 2025, models like Gemini 2.5 Pro, GPT-4.1, and Llama 4 Maverick began supporting over 1 million tokens. The KV cache (Key-Value Cache, used to accelerate the Transformer model inference process) grows linearly with the number of tokens, reaching about 200GB at 1 million tokens, far exceeding the 80GB VRAM capacity of a single H100.

One solution to this problem is to offload part of the KV cache to CPU memory. This means the CPU must not only manage orchestration and tool calls but also help bear data that doesn't fit in VRAM. CPU memory capacity, memory bandwidth, and the interconnect speed between the CPU and GPU thus become critical to system performance.

Therefore, CPUs suited for the agent era require low latency, consistent memory access capabilities, and stronger system-level协同 capabilities, rather than单纯 core count expansion.

03 What Are the Vendors Doing? Some Grab Territory, Others Change Designs

Faced with this sudden surge in CPU demand, the major players have completely different strategies.

Intel is the traditional leader in server CPUs. Data from Mercury Research shows that in Q4 2025, Intel still held a 60% share of the server CPU market, AMD had 24.3%, and Nvidia had 6.2%. But Intel has been playing catch-up with new technologies in recent years; this CPU demand explosion is both an opportunity and a test for them.

Intel's current strategy is a two-pronged approach. On one hand, continue selling Xeon processors, deeply绑定 with hyperscale customers like Google; on the other hand, partner with SambaNova to launch a combined solution based on Xeon processors and their self-developed RDU accelerators,主打 the selling point of 'running agent inference without GPUs'. The roadmap for Xeon 6 Granite Rapids and the 18A process will be key tests of whether Intel can turn the tables.

AMD is one of the biggest beneficiaries of this CPU demand surge. In Q4 2025, AMD's Data Center revenue was $5.4 billion, a year-on-year increase of 39%. Fifth-gen EPYC Turin accounted for over half of server CPU revenue, and deployments of cloud instances running EPYC grew over 50% year-on-year. AMD's server CPU revenue share exceeded 40% for the first time.

AMD CEO Lisa Su directly attributed the growth to the development of 'agents'—agent workloads push tasks 'back' to traditional CPU tasks.

In February 2026, AMD also announced a potential deal with Meta worth over $100 billion to supply MI450 GPUs and Venice EPYC CPUs.

However, AMD still has room for improvement in system-level协同, lacking a mature high-speed CPU-GPU interconnect capability类似 NVLink C2C. As agent (Agent) systems place increasing demands on data interaction and协同 efficiency, the importance of this aspect is also gradually rising.

Nvidia's approach to CPU design is completely different from Intel's and AMD's.

The Nvidia Grace CPU has only 72 cores, while AMD EPYC and Intel Xeon typically have 128. Nvidia's AI Infrastructure VP, Dion Harris, explained: 'If you're a hyperscaler, you want to maximize the number of cores per CPU, which basically drives down the cost, the dollar-per-core cost. So it's a business model.'

In other words, in the AI computing体系, the CPU's role is no longer that of a general-purpose workhorse but rather a 'scheduling hub' serving the GPU. If the CPU can't keep up, the expensive GPUs are forced to wait, and overall efficiency drops.

Therefore, Nvidia prioritizes efficient协同 between the CPU and GPU in its design. For example, through the NVLink C2C interconnect, the bandwidth between the CPU and GPU is boosted to about 1.8TB/s, far higher than traditional PCIe, and the CPU can directly access GPU memory, greatly simplifying KV cache management.

Currently, Nvidia sells the Vera CPU as a standalone product. CoreWeave was the first customer. The deal with Meta is even more夸张; this is its first large-scale 'pure Grace deployment,' meaning CPUs deployed大规模 independently without GPUs paired.

Ben Bajarin, Principal Analyst at Creative Strategies, pointed out that in high-intensity system collaboration, the CPU's processing power must keep pace with the accelerator's iteration speed. If there is even a one percent delay in the data通道, the entire AI cluster's economic efficiency suffers significantly. This pursuit of极致 system efficiency is forcing all major players to re-evaluate CPU performance metrics.

Holger Mueller, VP and Principal Analyst at Constellation Research, stated that as AI workloads shift towards agent-driven architectures, the CPU's position is becoming more central. He noted: 'In the agent world, agents need to call APIs and various business applications, tasks most suitable for CPUs to complete.'

He added: 'Currently, there is no consensus on whether GPUs or CPUs are more suitable for handling inference tasks. GPUs have an advantage in model training, and custom ASICs like TPUs have their specialties. But one thing is clear: Google needs to adopt a hybrid processor architecture. Therefore, Google's choice to partner with Intel is reasonable.'

04 Conclusion: In the Agent Era, the Computing Power Balance is Swinging Back

In the latest industry observations, one data point deserves attention. In the massive $38 billion合作协议 between Amazon AWS and OpenAI, the official announcement also explicitly mentioned scaling 'tens of millions of CPUs'.

In recent years, the industry's focus has typically been on those 'hundreds of thousands of GPUs'. However, the fact that cutting-edge labs like OpenAI are proactively treating CPU scale as a key planning variable sends a clear signal: scaling agent workloads must be built upon a massive CPU infrastructure.

Bank of America predicts that by 2030, the global CPU market size could double from the current $27 billion to $60 billion. Almost all of this additional share will be driven by AI.

We are witnessing the expansion of a全新的 infrastructure: big tech is no longer just stacking GPUs but is simultaneously expanding an entire layer of 'CPU scheduling infrastructure' specifically to support the operation of AI agents.

The alliance between Intel and Google, as well as Musk's heavy investment in custom chips, all prove one fact: the winning point in the AI race is moving forward. When computing power is no longer scarce, whoever can solve the system-level 'bottleneck' first will have the last laugh in this trillion-dollar game.

*Special contributor Jin Lu also contributed to this article.

This article is from the WeChat public account 'Tencent Technology', author: Li Hailun, editor: Xu Qingyang

Cryptos en tendance

CitreaCTR

wrapped stUSDTWSTUSDT

Velodrome FinanceVELODROME

BrevisBREV

PancakeSwapCAKE

JUSTJST

Questions liées

QWhy is CPU becoming the new bottleneck in the AI era according to the article?

ACPU is becoming the bottleneck because AI workloads, especially in the agentic AI era, require extensive multi-step reasoning, API calls, database operations, and complex task orchestration. These tasks are primarily handled by the CPU, and studies show that CPU-side tool processing can account for 50% to 90.6% of total latency. Additionally, the expansion of context windows in models requires KV cache offloading to CPU memory, making CPU memory capacity, bandwidth, and interconnect speed critical.

QWhat significant partnership is mentioned in the article to address the CPU bottleneck?

AGoogle and Intel have entered into a multi-year agreement to deploy Intel's Xeon processors globally in AI data centers. This partnership aims to enhance system performance, efficiency, and flexibility by leveraging CPUs and IPUs (Infrastructure Processing Units) as key components in AI infrastructure.

QHow did the CPU market change in Q4 2025 as reported?

AIn Q4 2025, the average selling price of server CPUs increased by approximately 30%, which is rare in the mature CPU market. Delivery cycles extended, with AMD's wait times increasing from eight to over ten weeks, and some models facing delays of up to six months due to supply constraints and competition for wafer capacity from GPU production.

QWhat role does CPU play in agentic AI workloads compared to traditional AI models?

AIn traditional AI models, CPUs mainly handle scheduling and data processing while GPUs perform core inference tasks. In agentic AI, CPUs are responsible for executing multi-step reasoning, calling APIs, reading/writing databases, orchestrating complex workflows, and integrating results—tasks that constitute the majority of the latency. The GPU generates tokens ('thinking'), but the CPU turns those results into actionable outputs.

QHow are major companies like Intel, AMD, and NVIDIA adapting to the increased importance of CPUs in AI?

AIntel is deepening partnerships (e.g., with Google) and developing combinations like Xeon processors with accelerators. AMD is benefiting from increased demand, with its EPYC CPUs seeing significant growth and new large deals (e.g., with Meta). NVIDIA is designing CPUs like Grace with a focus on high-efficiency coordination with GPUs through technologies like NVLink C2C, prioritizing system-level synergy over core count.

Lectures associées

Liste des altcoins les plus populaires selon les recherches des dernières heures publiée !

La plateforme CoinGecko a publié une liste des cryptomonnaies les plus recherchées par les utilisateurs au cours des dernières heures. Le jeton Pudgy Penguins ($PENGU) est en tête, suivi de Catecoin (CATE) et de Bless ($BLESS). Sur les 24 dernières heures, CATE a enregistré une hausse de prix impressionnante de 126,2%, tandis que $BLESS a augmenté de 86,1% et $PENGU de 3,9%. What IF (IF) a également progressé de 41,9%. Le classement complet des actifs les plus consultés sur CoinGecko, avec leur capitalisation boursière, est le suivant : 1. Pudgy Penguins ($PENGU) – 389,13 millions de dollars 2. Catecoin (CATE) – 19,62 millions de dollars 3. Bless ($BLESS) – 32,72 millions de dollars 4. Aerodrome Finance (AERO) – 385,03 millions de dollars 5. Hyperliquid (HYPE) – 11,43 milliards de dollars 6. Ethereum (ETH) – 224,17 milliards de dollars 7. Chainlink (LINK) – 6,17 milliards de dollars 8. Aave (AAVE) – 1,42 milliard de dollars 9. What IF (IF) – 31,24 millions de dollars 10. Polkadot (DOT) – 1,34 milliard de dollars 11. Bitcoin (BTC) – 1,27 trillion de dollars 12. Virtual Protocol (VIRTUAL) – 366,19 millions de dollars 13. Algorand (ALGO) – 758,15 millions de dollars 14. Cash Cat (CASHCAT) – 41,81 millions de dollars 15. Solana (SOL) – 42,38 milliards de dollars. *Ceci ne constitue pas un conseil en investissement.

cryptonews.ruIl y a 1 h

Liste des altcoins les plus populaires selon les recherches des dernières heures publiée !

cryptonews.ruIl y a 1 h

Pour 100 000 $ par mois : Truth Social vend l'accès aux publications de Trump à des sociétés d'investissement

Le groupe Trump Media and Technology Group (TMTG) a lancé le 1er août 2026 « Truth API », un service d’accès payant en temps réel aux publications des comptes les plus influents de Truth Social, notamment celui de l’ancien président Donald Trump. Destiné aux investisseurs institutionnels et aux firmes de trading haute fréquence, l’abonnement pourrait coûter jusqu’à 100 000 dollars par mois. TMTG justifie cette initiative par la création d’une source de revenus stable et la monétisation de ses actifs. Cette commercialisation d’un accès prioritaire aux posts présidentiels a suscité des critiques de la part de législateurs américains, dont des sénateurs démocrates et républicains, qui demandent des enquêtes sur d’éventuelles violations des règles de marché et dénoncent un accès privilégié vendu à prix d’or. L’analyse pointe un risque systémique similaire à celui observé en 2013, lorsqu’un tweet piraté avait provoqué une chute brutale des marchés. Le service Truth API, sans mécanisme avéré de vérification en temps réel, pourrait transformer le compte de Trump en une cible pour des manipulations, soulevant la question de la responsabilité en cas de diffusion de fausses informations influençant les marchés financiers.

cryptonews.ruIl y a 2 h

Pour 100 000 $ par mois : Truth Social vend l'accès aux publications de Trump à des sociétés d'investissement

cryptonews.ruIl y a 2 h

La stratégie maintient le dividende privilégié du STRC à 12 % alors que le prix reste encore en dessous du pair

Les actions préférées STRC de Strategy, dont le prix de clôture était de 89,46 $ fin juillet (bien en dessous de leur valeur nominale de 100 $), maintiendront leur dividende à 12 % pour le mois d'août. Le président exécutif Michael Saylor a confirmé cette information, notant que le dividende est désormais versé deux fois par mois. Malgré une perte nette importante au deuxième trimestre (8,22 milliards de $), principalement due à une perte non réalisée sur ses réserves de Bitcoin, Strategy a constitué une réserve de trésorerie de 3,75 milliards de $ pour garantir le paiement des dividendes préférés. La direction réitère son objectif de faire remonter le cours de STRC vers 99-100 $ à terme et continue de racheter ces titres tant qu'ils se négocient en dessous du pair. Parallèlement, Saylor a évoqué une annonce potentielle concernant les avoirs en Bitcoin de l'entreprise, laissant entendre une possible évolution de sa stratégie de trésorerie.

cointelegraphIl y a 3 h

La stratégie maintient le dividende privilégié du STRC à 12 % alors que le prix reste encore en dessous du pair

cointelegraphIl y a 3 h

Les retraits de Bitcoin se poursuivent : 8 ans de stockage en portefeuille froid Coldcard se sont terminés par un solde nul

Le portefeuille matériel Coldcard a été compromis, entraînant une nouvelle vague de retraits depuis les appareils vulnérables. Selon Galaxy Research, environ 1 367,05 BTC (88,6 millions de dollars) ont été dérobés à partir de 4 585 adresses. Le problème ne réside pas dans le firmware, qui a été corrigé, mais dans les phrases seed générées entre mars 2021 et les mises à jour correctives. Ces phrases, créées en raison d'une erreur de programmation ayant conduit à l'utilisation d'un générateur de nombres aléatoires logiciel (Yasmarang) au lieu du générateur matériel STM32, sont prévisibles et vulnérables à une attaque par force brute hors ligne. Les propriétaires concernés doivent impérativement générer une nouvelle phrase seed sur un firmware corrigé et transférer leurs actifs, sous peine de rester exposés. L'histoire d'un investisseur de 39 ans illustre l'impact dévastateur : après avoir accumulé 2 BTC (130 000 dollars) sur huit ans via un travail physique, en les conservant comme protection contre l'hyperinflation dans son pays, il a tout perdu en quelques minutes. Son cas montre que même les stratégies de conservation à long terme les plus prudentes ("cold storage") ne sont pas infaillibles. D'un point de vue historique, cet incident rappelle les faiblesses passées des générateurs de nombres aléatoires dans la cryptographie. Il remet en question l'idée reçue selon laquelle le stockage hors ligne garantit automatiquement une sécurité absolue. La communauté espère que le fabricant pourra aider à récupérer les fonds volés.

cryptonews.ruIl y a 3 h

Les retraits de Bitcoin se poursuivent : 8 ans de stockage en portefeuille froid Coldcard se sont terminés par un solde nul

cryptonews.ruIl y a 3 h

En Corée du Sud, les volumes d'échanges de 15 altcoins explosent !

Les principales plateformes d'échange de cryptomonnaies sud-coréennes, Upbit et Bithumb, rapportent une forte augmentation du volume des transactions pour plusieurs altcoins. Sur les dernières 24 heures, le volume total des altcoins les plus populaires a atteint environ 347,7 millions de dollars. MetaDAO (META) arrive en tête, avec un volume de 65,84 millions de dollars uniquement sur Upbit, représentant 12,39% du volume spot total de la bourse. Euler ($EUL) suit avec 47,65 millions de dollars, et le $XRP, toujours populaire auprès des investisseurs sud-coréens, a atteint 38,11 millions de dollars. La liste complète des 15 altcoins montre une activité intense, notamment pour ThunderCore (TT, 35,64M$), Babylon (BABY, 25,15M$) et Geodnet (GEOD, 20,28M$). Cet engouement marqué pour des actifs numériques au-delà du Bitcoin illustre la dynamique spéculative sur le marché sud-coréen. *Ceci n'est pas un conseil en investissement.

cryptonews.ruIl y a 5 h

En Corée du Sud, les volumes d'échanges de 15 altcoins explosent !

cryptonews.ruIl y a 5 h

Trading

Spot

Articles tendance

Comment acheter ERA

Bienvenue sur HTX.com ! Nous vous permettons d'acheter Caldera (ERA) de manière simple et pratique. Suivez notre guide étape par étape pour commencer votre parcours crypto.Étape 1 : Création de votre compte HTXUtilisez votre adresse e-mail ou votre numéro de téléphone pour ouvrir un compte sur HTX gratuitement. L'inscription se fait en toute simplicité et débloque toutes les fonctionnalités.Créer mon compteÉtape 2 : Choix du mode de paiement (rubrique Acheter des cryptosCarte de crédit/débit : utilisez votre carte Visa ou Mastercard pour acheter instantanément Caldera (ERA).Solde ：utilisez les fonds du solde de votre compte HTX pour trader en toute simplicité.Prestataire tiers ：pour accroître la commodité d'utilisation, nous avons ajouté des modes de paiement populaires tels que Google Pay et Apple Pay.P2P ：tradez directement avec d'autres utilisateurs sur HTX.OTC (de gré à gré) : nous offrons des services personnalisés et des taux de change compétitifs aux traders.Étape 3 : stockage de vos Caldera (ERA)Après avoir acheté vos Caldera (ERA), stockez-les sur votre compte HTX. Vous pouvez également les envoyer ailleurs via un transfert sur la blockchain ou les utiliser pour trader d'autres cryptos.Étape 4 : tradez des Caldera (ERA)Tradez facilement Caldera (ERA) sur le marché Spot de HTX. Il vous suffit d'accéder à votre compte, de sélectionner la paire de trading, d'exécuter vos trades et de les suivre en temps réel. Nous offrons une expérience conviviale aux débutants comme aux traders chevronnés.

712 vues totalesPublié le 2025.07.17Mis à jour le 2026.06.02

Discussions

Bienvenue dans la Communauté HTX. Ici, vous pouvez vous tenir informé(e) des derniers développements de la plateforme et accéder à des analyses de marché professionnelles. Les opinions des utilisateurs sur le prix de ERA (ERA) sont présentées ci-dessous.

Catégories populaires

Indepth Research1,444 actualités