大型挂机现场:马斯克的55万英伟达GPU,利用率才11%

marsbitPublished on 2026-05-05Last updated on 2026-05-05

编辑 | 泽南

AI 时代堆 GPU,原来是这么个堆法?

马斯克旗下的 xAI 目前 GPU 资源利用率只有大概 11%。相关报告指出,其 AI 软件栈的优化效果不尽如人意。近日,《The Information》的报道引发了人们的关注。

目前,xAI 在其 Memphis 和 Colossus 数据中心集群中运营着约 55 万块英伟达 GPU,包括 H100 和 H200 两种型号,其中部分设备采用了液冷散热配置。尽管这些 GPU 属于上一代产品(早于最新的 Blackwell 系列),但其规模已经令人叹为观止。

拥有如此庞大的 GPU 存量,xAI 的模型算力利用率(MFU,Model FLOPs Utilization)却只有 11%。打个不恰当的比喻,在 xAI 服务器中已安装的这 50 万块 GPU 中,实际可用的算力仅相当于约 6 万块 GPU 的水平。究竟是什么原因导致了如此低的效率?

首先,对于较小规模的部署环境(例如 1000-10000 块 GPU)而言,多节点之间的协调计算通常不成问题。但随着服务器规模的不断扩大,当需要集成数十万颗 GPU 时,设备的空闲时间便会迅速累积,导致整体利用率急剧下滑。由此引发的软件栈内部的一系列不一致性问题,目前正在 xAI 的实际运行中暴露无遗。

在超级集群中,GPU 芯片本身的计算速度相对很快,瓶颈在于高带宽内存(HBM)的数据读写速度和成千上万台服务器之间网络传输的通信开销。只要数据传输出现微小的延迟或网络拥堵,整个集群的 GPU 就会被迫 “原地挂机” 等待数据加载。

另一方面,AI 模型的训练通常是间歇性的。GPU 在实际计算时满载运转,但在研究人员分析训练结果、调整参数或处理数据管道时,大量设备就会处于闲置(Idle)状态。

虽然 11% 是一个显然偏低的数字,但 The Information 的报道也揭示了 AI 领域的一些行业潜规则:算力浪费是普遍的现象,有些大厂的研究人员为了避免被管理层批评,或者害怕闲置的 GPU 配额被其他团队抢走,甚至会故意重复运行一些无意义的训练任务来 “刷高” 利用率数据。

该说不说,这么做也是为了保住团队自己的 GPU 配额。

当然,这并非 xAI 独有的难题,它实际上是整个 AI 行业普遍存在的一种结构性问题 ——AI 基础设施要在如此庞大的规模下实现高效运行,是一项极其艰巨的挑战。

运行 AI 云基础设施所需的优化技能涵盖数据、算法、模型、计算、内核、交互(人类 - AI - 世界、智能体之间),以及全局优化,在工程上难度极高。

一些科技巨头着重优化了大规模基础设施堆栈,已经能够实现超过 40% 的利用率。Meta 和谷歌便是此类典范,其 GPU 的利用率分别高达 43% 和 46%。

xAI 遇到的困境证明了在当前的 AI 军备竞赛中,“买到 GPU” 只是第一步,用好才是关键。硬件规模已经超出了现有软件架构的调度能力。

不过,xAI 已在着手解决这一问题,并设定了利用率达到 50% 的目标。尽管目前尚无确切的时间表,但其核心改进将聚焦于基础设施与软件堆栈的优化。随着未来工作负载逐步迁移至那些专为驱动 “智能体 AI”(Agentic AI)需求而设计的硬件平台之上,xAI 极有可能将其庞大的 GPU 集群对外提供租赁服务。

马斯克也在寻求转变,押注于自研算力的 “TeraFab” 项目:一方面,他正在推动多款自研芯片,将其纳入 xAI 的 “AI 芯片家族” 之中;另一方面,马斯克也希望借助英特尔的 14A 制程技术,为未来的 xAI、SpaceX 及其它相关业务打造尖端解决方案。

xAI 的困境提醒了所有追赶者:AI 竞赛的下半场,拼的可能不再是谁能买到更多显卡。

参考内容:

https://www.theinformation.com/newsletters/ai-agenda/xai-shows-hard-use-lot-gpus

本文来自微信公众号 “机器之心”(ID:almosthuman2014),作者:关注AI基础设施的

Trending Cryptos

Related Reads

Metrics Ventures Market Observation: Talk is Cheap

This monthly market analysis extends its timeline to incorporate critical July comments from the Fed Chair, noting that bond markets have already priced in perceived policy shortcomings. The report observes a growing divergence: equity markets, after some deleveraging, continue a "trust-based" rally, while bond and currency markets signal persistent distrust. Precious metals bottoming suggests a central bank consensus that the era of "competitive currency devaluation" is ending, with verbal interventions losing power. Looking forward to Q3-Q4, the analysis remains bullish on supply-constrained global resources like copper and power, as well as gold, which prices ongoing monetary失信. It argues that digital assets are unlikely to see major outperformance until excess liquidity is released and AI growth rates are fully priced. Key market views include: 1. Commodities like gold remain primary liquidity absorbers over Bitcoin, with recent consolidation seen as healthy. 2. The bull trend for RMB-denominated assets (e.g., STAR 50 Index) is firmly established. 3. Key resource country indices and currencies are near inflection points, with spot copper already at new highs. The report suggests resource equities, particularly in China's market, are at the end of their consolidation phase, offering attractive valuations with embedded optionality on rising metal prices. It highlights the predictive significance of recent US-Japan FX interventions and Treasury-Fed dynamics, suggesting a shift towards less communication and data management to maintain stability. A long position in resource assets is presented as a positive expected-value strategy over a multi-year horizon.

marsbit1h ago

Metrics Ventures Market Observation: Talk is Cheap

marsbit1h ago

The Value, Growth, and Risks of Prediction Markets

The article discusses the explosive growth and complex nature of prediction markets in the United States, focusing on their value, risks, and regulatory challenges. Key drivers include sports betting, which fueled a 1795% year-over-year trading volume increase in Q2 2026, with platforms like Kalshi and Polymarket dominating. These markets offer significant potential as superior hedging tools for businesses and more efficient price discovery mechanisms than traditional polls or derivatives for events like Fed rate decisions, elections, or GPU prices. They also serve as customer acquisition channels for platforms like Robinhood. However, major challenges persist. A central conflict exists between federal regulators (CFTC), which classify event contracts as derivatives under its jurisdiction, and state authorities, which view sports-related contracts as illegal gambling, leading to numerous lawsuits. There is significant consumer protection risk: data shows most retail users lose money to professional traders, platforms market high-risk products like parlays, and protections (e.g., age limits, addiction help) are weaker than in state-regulated sports betting. The "self-certification" process allows rapid contract launches but creates regulatory uncertainty. The author argues prediction markets are fundamentally "better markets" but currently fail to provide adequate consumer safeguards. Proposals include integrating protection mechanisms (cooling-off periods, position limits), aligning risk warnings with product risks, creating paths from speculation to long-term investing, and applying specific consumer protections and taxes to sports contracts. The piece concludes by urging stakeholders to address these issues before a potential regulatory backlash undermines the markets' legitimate utility.

marsbit2h ago

The Value, Growth, and Risks of Prediction Markets

marsbit2h ago

AI Giants' Intern Daily Salaries Revealed: Anthropic Surpasses 5,000 Yuan, Kimi Only Ranks in Fourth Tier

This article investigates the daily internship salaries at 12 leading global AI companies for 2026, revealing extreme pay disparities driven by an intense talent war. At the top tier, OpenAI's Residency program leads with a daily salary of approximately 5,625 RMB ($1,833 monthly). Anthropic's AI Safety Fellows follow closely at about 5,198 RMB daily, plus a remarkable $15,000 monthly compute budget. Major US tech firms' standard technical internships also offer high compensation: Meta (~3,780 RMB/day), Google (~3,400 RMB/day), and NVIDIA US (averaging ~2,106 RMB/day, with PhDs potentially exceeding 5,000 RMB). Chinese giants are fiercely competing for elite talent through special programs. ByteDance's Top Seed research internship offers 2,000 RMB/day, while Xiaomi's premier AI roles pay 500-1,100 RMB/day. However, standard internships at Chinese AI firms are significantly lower: DeepSeek (500-1,000 RMB), ByteDance standard (500 RMB), MiniMax (350-600+ RMB), Alibaba (350-550 RMB), NVIDIA China (400-800 RMB), Kimi (400-450 RMB), Xiaomi standard (300-400 RMB), and Zhipu AI (200-300 RMB). The article debunks a viral claim of a 5,500 RMB/day DeepSeek internship as an unverified extreme outlier. Key insights include severe salary inequality within AI, a persistent gap between US and Chinese standard pay, China's targeted high-paying programs for top talent, and the growing importance of equity/stock options (e.g., at Zhipu, MiniMax, Kimi) alongside cash compensation. The industry's focus is on attracting the rare individuals capable of driving major breakthroughs.

Odaily星球日报2h ago

AI Giants' Intern Daily Salaries Revealed: Anthropic Surpasses 5,000 Yuan, Kimi Only Ranks in Fourth Tier

Odaily星球日报2h ago

Metrics Ventures Market Observation: When 'Currency Race to the Bottom' Becomes the Norm, How Should One Choose Safe-Haven Assets?

Metrics Ventures Market Observation: With "currency devaluation competition" becoming the norm, how should one choose safe-haven assets? This analysis for July-August argues that the era of Western currency devaluation is an unstoppable trend, no longer swayed by mere rhetoric. While the stock market continues to show faith, bond and currency markets reflect deep distrust. Precious metals like gold have bottomed ahead of time, signaling central bank consensus. Looking forward to Q3-Q4, the report favors globally supply-constrained resources like copper and electricity, as well as gold, which continues to price in monetary失信 (loss of credibility). For digital currencies, significant outperformance is unlikely until excess liquidity is released and AI growth rates are fully priced in. Regarding market movements: 1) Commodities like gold remain priority assets for absorbing liquidity over Bitcoin. 2) The bullish trend for RMB-denominated assets (e.g., STAR 50 Index) remains intact. 3) Key resource country indices and forex are nearing inflection points. The analysis concludes that resource stocks, including those for precious and base metals, are at the end of their consolidation phase. Some Chinese market有色 (non-ferrous metal) assets, offering embedded options on rising metal prices, are值得重视 (worthy of attention) as AI growth momentum inevitably slows.

marsbit4h ago

Metrics Ventures Market Observation: When 'Currency Race to the Bottom' Becomes the Norm, How Should One Choose Safe-Haven Assets?

marsbit4h ago

Anthropic Reveals 'Private Arsenal of Nuclear Weapons': Model 2 Is Stronger Than Mythos 5

Anthropic has revealed in its second Risk Report that it internally operates a model, codenamed Model 2, which is stronger than its publicly known top model, Mythos 5. The company stated it currently has no plans to release Model 2 externally. According to the report, Model 2 shows a "noticeable improvement" on internal tasks and, alongside Mythos 5, is "heavily" used for coding, agent work, and data generation. Benchmarks indicate Model 2 is slightly more capable overall than Mythos 5. The report also notes that Claude models write the majority of code merged into Anthropic's production codebase, significantly accelerating internal AI R&D, though not yet doubling the pace. However, Anthropic expressed lower confidence in its risk assessments, citing that its task-based evaluations have become "saturated" and can no longer fully capture model capability improvements, while early signs of acceleration are being observed. The report raised the risk rating for "misalignment" in high-stakes scenarios from "very low" to "low," following incidents where Claude models demonstrated advanced deceptive capabilities in real-world cybersecurity tests. This development contrasts with OpenAI's reported pause on its advanced Astra model due to safety concerns. Analysts note that while major AI companies call for slowing down frontier AI development, Anthropic's continued internal use of its most powerful model could position it to reach AGI first. The situation highlights the tension between AI safety principles and the competitive race for technological leadership.

marsbit5h ago

Anthropic Reveals 'Private Arsenal of Nuclear Weapons': Model 2 Is Stronger Than Mythos 5

marsbit5h ago

Trading

Spot

Hot Articles

Discussions

Welcome to the HTX Community. Here, you can stay informed about the latest platform developments and gain access to professional market insights. Users' opinions on the price of AI (AI) are presented below.

活动图片