Rubin Ultra Makes Major Cuts, Even Nvidia Can't Handle Memory Price Hikes?

Odaily星球日报Publicado a 2026-08-03Actualizado a 2026-08-03

Resumen

NVIDIA's Rubin Ultra, the top-tier variant of the newly announced Rubin AI accelerators, has reportedly seen significant specification downgrades, according to an industry report from SemiAnalysis. Initially designed with four compute dies (4-die), the Rubin Ultra is now said to be reduced to a 2-die design. Key changes highlighted in the report include: * **No increase in peak theoretical compute performance**, remaining at 35 PFLOPs like the standard Rubin. * **Severe reduction in memory capacity** to 192GB using 8-Hi HBM stacks, which is less than the standard Rubin's 288GB using 12-Hi stacks. * **Negligible memory bandwidth improvement** of only 1 TB/s. * **Slightly higher chip-level power consumption**. * The **primary upgrade is a massive increase in scale-up interconnect capacity**, supporting connections for up to 576 GPUs via NVLink, compared to 72 for the standard Rubin. The report suggests the redesign is primarily a cost-optimization move driven by the sharp rise in HBM (High-Bandwidth Memory) prices. By reducing the expensive HBM content and shifting investment towards enhanced system-scale networking, NVIDIA aims to maintain the platform's value for large-scale AI training clusters while managing soaring material costs. The news reportedly triggered a sell-off in South Korean memory stocks, with SK Hynix and Samsung shares falling around 8%, as markets grew concerned that NVIDIA—a major HBM buyer—might be reducing its reliance on high-capacity memor...

Original | Odaily Planet Daily (@OdailyChina)

Author | Azuma (@azuma_eth)

This past weekend, a research report from the well-known investment research firm SemiAnalysis sparked intense discussion across the entire AI industry.

The core content of the report is that following SemiAnalysis's late June revelation that Nvidia's original 4-die Rubin Ultra design would be halved, the research firm has now disclosed that Nvidia has provided major customers with a preview of the Rubin Ultra, whose specifications have been further reduced from prior expectations.

Screenshots of the Rubin Ultra information leaked in the SemiAnalysis report are as follows:

  • Rubin Ultra will maintain the same theoretical peak computing power as Rubin, both at 35 PFLOPs;
  • VRAM capacity severely downgraded, specifications reduced to 8-Hi stacked 192GB, even lower than the 12-Hi stacked 288GB of Rubin;
  • VRAM bandwidth change is negligible, only increased by 1 TB/s, essentially insignificant in actual high-throughput computing;
  • Chip-level power consumption even slightly higher, minimum power consumption same as Rubin (1800 W), maximum power consumption is actually higher (2600 W);
  • The only major improvement is the "Scale-up world size", increased from 72 GPUs to 576 GPUs, meaning Nvidia has shifted the real selling point to the cluster network. Through the NVLink network, it can connect up to 576 Rubin Ultra GPUs into a massive "super logical GPU" (the standard version only supports up to 72 cards interconnected).

As the top flagship version of Nvidia's Rubin, heavily announced at GTC 2026 this year, Rubin Ultra was originally positioned to handle extreme-scale AI model training and inference needs by integrating more chip dies and high-bandwidth memory. However, judging from the latest specification details disclosed by SemiAnalysis, Nvidia has evidently adjusted its design approach for Rubin Ultra.

HBM Price Hikes Too Fast, Nvidia Starts Recalculating

Over the past two years, one of the most critical components in AI infrastructure expansion has undoubtedly been HBM.

With the explosive demand for Nvidia H100, H200, Blackwell, and other AI accelerators, high-bandwidth memory has transformed from a relatively niche high-end storage product into the most constrained link in the entire AI infrastructure. While SK Hynix, Samsung, and Micron have been repeatedly posting record earnings and ramping up HBM investments, the market supply-demand imbalance continues to drive HBM prices upward.

For AI chip manufacturers, the importance of HBM is self-evident — GPUs handle computation, HBM provides high-speed data throughput, and together they determine AI model training and inference efficiency. The problem, however, is that HBM has now become so expensive it's beginning to impact the overall economics of AI systems. Taking HBM3 as an example, the price for one HBM3 stack was only $180-220 in Q2 2025 lows. By Q1 of this year (contract price), it had risen to $600-700, and Q2 (spot price) surged further to $700-850.

According to SemiAnalysis estimates, as HBM prices rose, the pure bill-of-materials (BOM) cost for a single Rubin Ultra rack once increased from about $6.6 million to $8 million. After adjusting the design specifications, the cost could drop back to around $6.4 million.

For Nvidia, such a huge cost difference poses a very real question — if the original design of increasing HBM capacity is continued, can it still deliver performance improvements commensurate with the cost? Are there other, better alternatives?

SemiAnalysis's report provides an answer to this: The adjustment to Rubin Ultra essentially represents Nvidia's re-optimization of the cost structure for AI systems under limited resources. This means reducing the relatively expensive HBM configuration (its cost share drops from nearly 40% to 28%) and redirecting resources towards higher-value expansion interconnect capabilities (cost share increases from 4% to 12%).

As mentioned earlier, the core upgrade direction for Rubin Ultra has now shifted to "system-level expansion and interconnect capabilities." The NVL576 architecture supported by Rubin Ultra can connect up to 576 GPUs via NVLink into a unified compute domain, aiming to compensate for the adjustments in single-chip specifications with larger-scale system expansion.

Memory Stocks Plunge, Market Worries About Demand Peaking

Potentially impacted by this news, South Korean memory-related stocks collectively declined upon opening this morning. As of 11:45 Beijing Time, SK Hynix and Samsung, the two leading HBM manufacturers, both plunged about 8%, with the KOSPI index also down around 5%.

The market has begun to worry: If the Rubin Ultra specification adjustments revealed by SemiAnalysis are true (Nvidia has not yet publicly confirmed related information), does it mean that Nvidia, as the most crucial buyer in AI infrastructure, is reducing its demand for HBM?

Over the past two years, with the rapid growth of AI scaling demand, HBM has become the most constrained part of AI infrastructure. Chip manufacturers like Nvidia and AMD continuously increased HBM configurations in AI accelerators, driving soaring earnings for memory makers SK Hynix, Samsung, and Micron, and strengthening their pricing power across the supply chain.

Nvidia's choice, however, might indicate that AI chip makers are considering optimizing hardware design to reduce reliance on high-capacity HBM per chip. If this path proves feasible, the room for continuous price hikes by HBM manufacturers could be significantly limited.

The era of "madly stacking specs and mindlessly raising prices" in AI infrastructure construction will eventually come to an end. And now, even Nvidia, seated upon the iron throne of computing power, has started to pinch pennies.

Preguntas relacionadas

QAccording to the article, what is the most significant change in NVIDIA's Rubin Ultra design compared to initial expectations?

AThe most significant change is a severe downgrade in memory (HBM) capacity. The Rubin Ultra will now feature 8-layer stacked (8-Hi) 192GB HBM, which is less than the 12-layer stacked 288GB of the standard Rubin.

QWhat financial reason is suggested for NVIDIA's decision to reduce the HBM specification in the Rubin Ultra?

AThe article suggests that the rapid and significant price increase of HBM has impacted the overall system economics. Reducing the HBM configuration from nearly 40% to 28% of the Bill of Materials (BOM) cost allows NVIDIA to control the rising cost of the Rubin Ultra system.

QWhat area becomes the new primary focus or selling point for the Rubin Ultra after the HBM downgrade?

AAfter the HBM downgrade, the primary upgrade focus shifts to system-level scaling and interconnect capability. The Rubin Ultra's key feature is now its support for the NVL576 architecture, which can connect up to 576 GPUs into a single massive logical GPU via NVLink.

QHow did the stock market, particularly HBM manufacturers, react to the reports about Rubin Ultra's specifications?

AThe stock market reacted negatively. Following the news, shares of major HBM manufacturers SK Hynix and Samsung both dropped approximately 8% in early trading, reflecting market concerns about potential weakening demand for HBM from a key buyer like NVIDIA.

QWhat does NVIDIA's design adjustment for Rubin Ultra potentially signal about the future of AI infrastructure development?

AIt potentially signals a shift away from a 'reckless material-piling and price-raising' era. It indicates that even dominant players like NVIDIA are now focusing on cost-optimized hardware design and seeking more efficient ways to scale AI systems, rather than simply increasing specifications like HBM capacity.

Lecturas Relacionadas

Wall Street Morning News: V-shaped Rebound at Month-end, but Nasdaq Suffers Worst July in 12 Years; Funds Accelerate Concentration Towards Cloud Giants

Despite a V-shaped rebound at the end of July, the Nasdaq posted its worst July since 2004, while the S&P 500 had its worst July since 2014. Markets were jolted by geopolitical shifts, as President Trump canceled a planned strike on Iran, leading WTI crude to plunge over 8%. This, alongside OPEC+ announcing a supply increase, reversed crude's sharp July gains. Treasury yields surged, with the 10-year yield rising over 30 basis points in July—its largest July increase since 2005. In a rare move, the US and Japan jointly intervened to weaken the USD/JPY, aiming to prevent potential Japanese sales of US Treasuries. While the tech sector faced deleveraging pressure throughout July, cloud giants staged a massive rally on strong earnings. Microsoft, Amazon, and Google collectively added nearly $1.5 trillion in market value last week. Amazon soared over 15% on accelerating AWS growth, Microsoft extended historic gains, Google fully recovered post-earnings losses, and Meta ended an 11-day losing streak. In contrast, Apple tumbled over 7% on supply chain and guidance concerns, ceding its "world's most valuable company" title to Nvidia. The memory and storage sector corrected sharply. Gold edged up 0.91% in July, with analysts viewing the ~30% pullback from January highs as a potential basing period, supported by long-term central bank demand. Key events to watch this week include earnings from Palantir, AMD, SpaceX (its first post-IPO report), and memory giants like Western Digital. The US July non-farm payrolls report on Friday will be critical for gauging the Fed's policy path. SpaceX also faces a significant lock-up expiration, testing market liquidity.

marsbitHace 45 min(s)

Wall Street Morning News: V-shaped Rebound at Month-end, but Nasdaq Suffers Worst July in 12 Years; Funds Accelerate Concentration Towards Cloud Giants

marsbitHace 45 min(s)

Can Generative Models Finally Be Trained End-to-End? The Core Is a For Loop

This article introduces a novel training paradigm for generative models called Explorative Modeling (XM), which enables true end-to-end training. Traditionally, powerful generative models like autoregressive and diffusion models are not trained end-to-end. They are trained to predict a single small step but require iterative multi-step sampling for inference. This "exposure bias" leads to error accumulation and limits performance. The core challenge XM addresses is "mode blurring." In generative tasks, a single input (e.g., "generate a dog") corresponds to many valid outputs (multiple modes). Standard training objectives like reconstruction loss force the model to average these modes, producing unrealistic, blurry outputs. To avoid this, existing models break generation into many small, almost deterministic steps, sacrificing end-to-end training. XM tackles this by restructuring the training loop itself. Its key insight is to amplify "generative expressivity." For each training input, instead of generating one sample, the model generates K candidate outputs. Only the candidate closest to the real data is used for computing the loss and updating the model via backpropagation. This simple "best-of-K" mechanism is implemented as a short for-loop. By exploring multiple possibilities, the model learns to distribute its guesses across different modes rather than collapsing to their uninformative average. The paper demonstrates that "exploration" acts as a new, powerful scaling axis. Gains from XM increase with model size, data scale, and compute. Experiments show improvements in FID scores for image generation and significant efficiency gains, sometimes outperforming larger models without exploration. When pushed to the limit, XM enables fully single-step, end-to-end generative models. In robotics tasks, an "Explorative Policy" matched the performance of a 100-step Diffusion Policy with a single forward pass, drastically improving inference speed. While the best-of-K concept is not entirely new, the authors' contribution lies in formally understanding it as a direct method to boost generative expressivity without fragmenting the generation process. This work suggests that as models scale, enhancing exploration during training may become crucial for overcoming fundamental performance bottlenecks.

marsbitHace 2 hora(s)

Can Generative Models Finally Be Trained End-to-End? The Core Is a For Loop

marsbitHace 2 hora(s)

Trading

Spot
活动图片