Rubin Ultra Makes Major Cuts, Even Nvidia Can't Handle Memory Price Hikes?

Odaily星球日报Pubblicato 2026-08-03Pubblicato ultima volta 2026-08-03

Introduzione

NVIDIA's Rubin Ultra, the top-tier variant of the newly announced Rubin AI accelerators, has reportedly seen significant specification downgrades, according to an industry report from SemiAnalysis. Initially designed with four compute dies (4-die), the Rubin Ultra is now said to be reduced to a 2-die design. Key changes highlighted in the report include: * **No increase in peak theoretical compute performance**, remaining at 35 PFLOPs like the standard Rubin. * **Severe reduction in memory capacity** to 192GB using 8-Hi HBM stacks, which is less than the standard Rubin's 288GB using 12-Hi stacks. * **Negligible memory bandwidth improvement** of only 1 TB/s. * **Slightly higher chip-level power consumption**. * The **primary upgrade is a massive increase in scale-up interconnect capacity**, supporting connections for up to 576 GPUs via NVLink, compared to 72 for the standard Rubin. The report suggests the redesign is primarily a cost-optimization move driven by the sharp rise in HBM (High-Bandwidth Memory) prices. By reducing the expensive HBM content and shifting investment towards enhanced system-scale networking, NVIDIA aims to maintain the platform's value for large-scale AI training clusters while managing soaring material costs. The news reportedly triggered a sell-off in South Korean memory stocks, with SK Hynix and Samsung shares falling around 8%, as markets grew concerned that NVIDIA—a major HBM buyer—might be reducing its reliance on high-capacity memor...

Original | Odaily Planet Daily (@OdailyChina)

Author | Azuma (@azuma_eth)

This past weekend, a research report from the well-known investment research firm SemiAnalysis sparked intense discussion across the entire AI industry.

The core content of the report is that following SemiAnalysis's late June revelation that Nvidia's original 4-die Rubin Ultra design would be halved, the research firm has now disclosed that Nvidia has provided major customers with a preview of the Rubin Ultra, whose specifications have been further reduced from prior expectations.

Screenshots of the Rubin Ultra information leaked in the SemiAnalysis report are as follows:

  • Rubin Ultra will maintain the same theoretical peak computing power as Rubin, both at 35 PFLOPs;
  • VRAM capacity severely downgraded, specifications reduced to 8-Hi stacked 192GB, even lower than the 12-Hi stacked 288GB of Rubin;
  • VRAM bandwidth change is negligible, only increased by 1 TB/s, essentially insignificant in actual high-throughput computing;
  • Chip-level power consumption even slightly higher, minimum power consumption same as Rubin (1800 W), maximum power consumption is actually higher (2600 W);
  • The only major improvement is the "Scale-up world size", increased from 72 GPUs to 576 GPUs, meaning Nvidia has shifted the real selling point to the cluster network. Through the NVLink network, it can connect up to 576 Rubin Ultra GPUs into a massive "super logical GPU" (the standard version only supports up to 72 cards interconnected).

As the top flagship version of Nvidia's Rubin, heavily announced at GTC 2026 this year, Rubin Ultra was originally positioned to handle extreme-scale AI model training and inference needs by integrating more chip dies and high-bandwidth memory. However, judging from the latest specification details disclosed by SemiAnalysis, Nvidia has evidently adjusted its design approach for Rubin Ultra.

HBM Price Hikes Too Fast, Nvidia Starts Recalculating

Over the past two years, one of the most critical components in AI infrastructure expansion has undoubtedly been HBM.

With the explosive demand for Nvidia H100, H200, Blackwell, and other AI accelerators, high-bandwidth memory has transformed from a relatively niche high-end storage product into the most constrained link in the entire AI infrastructure. While SK Hynix, Samsung, and Micron have been repeatedly posting record earnings and ramping up HBM investments, the market supply-demand imbalance continues to drive HBM prices upward.

For AI chip manufacturers, the importance of HBM is self-evident — GPUs handle computation, HBM provides high-speed data throughput, and together they determine AI model training and inference efficiency. The problem, however, is that HBM has now become so expensive it's beginning to impact the overall economics of AI systems. Taking HBM3 as an example, the price for one HBM3 stack was only $180-220 in Q2 2025 lows. By Q1 of this year (contract price), it had risen to $600-700, and Q2 (spot price) surged further to $700-850.

According to SemiAnalysis estimates, as HBM prices rose, the pure bill-of-materials (BOM) cost for a single Rubin Ultra rack once increased from about $6.6 million to $8 million. After adjusting the design specifications, the cost could drop back to around $6.4 million.

For Nvidia, such a huge cost difference poses a very real question — if the original design of increasing HBM capacity is continued, can it still deliver performance improvements commensurate with the cost? Are there other, better alternatives?

SemiAnalysis's report provides an answer to this: The adjustment to Rubin Ultra essentially represents Nvidia's re-optimization of the cost structure for AI systems under limited resources. This means reducing the relatively expensive HBM configuration (its cost share drops from nearly 40% to 28%) and redirecting resources towards higher-value expansion interconnect capabilities (cost share increases from 4% to 12%).

As mentioned earlier, the core upgrade direction for Rubin Ultra has now shifted to "system-level expansion and interconnect capabilities." The NVL576 architecture supported by Rubin Ultra can connect up to 576 GPUs via NVLink into a unified compute domain, aiming to compensate for the adjustments in single-chip specifications with larger-scale system expansion.

Memory Stocks Plunge, Market Worries About Demand Peaking

Potentially impacted by this news, South Korean memory-related stocks collectively declined upon opening this morning. As of 11:45 Beijing Time, SK Hynix and Samsung, the two leading HBM manufacturers, both plunged about 8%, with the KOSPI index also down around 5%.

The market has begun to worry: If the Rubin Ultra specification adjustments revealed by SemiAnalysis are true (Nvidia has not yet publicly confirmed related information), does it mean that Nvidia, as the most crucial buyer in AI infrastructure, is reducing its demand for HBM?

Over the past two years, with the rapid growth of AI scaling demand, HBM has become the most constrained part of AI infrastructure. Chip manufacturers like Nvidia and AMD continuously increased HBM configurations in AI accelerators, driving soaring earnings for memory makers SK Hynix, Samsung, and Micron, and strengthening their pricing power across the supply chain.

Nvidia's choice, however, might indicate that AI chip makers are considering optimizing hardware design to reduce reliance on high-capacity HBM per chip. If this path proves feasible, the room for continuous price hikes by HBM manufacturers could be significantly limited.

The era of "madly stacking specs and mindlessly raising prices" in AI infrastructure construction will eventually come to an end. And now, even Nvidia, seated upon the iron throne of computing power, has started to pinch pennies.

Domande pertinenti

QAccording to the article, what is the most significant change in NVIDIA's Rubin Ultra design compared to initial expectations?

AThe most significant change is a severe downgrade in memory (HBM) capacity. The Rubin Ultra will now feature 8-layer stacked (8-Hi) 192GB HBM, which is less than the 12-layer stacked 288GB of the standard Rubin.

QWhat financial reason is suggested for NVIDIA's decision to reduce the HBM specification in the Rubin Ultra?

AThe article suggests that the rapid and significant price increase of HBM has impacted the overall system economics. Reducing the HBM configuration from nearly 40% to 28% of the Bill of Materials (BOM) cost allows NVIDIA to control the rising cost of the Rubin Ultra system.

QWhat area becomes the new primary focus or selling point for the Rubin Ultra after the HBM downgrade?

AAfter the HBM downgrade, the primary upgrade focus shifts to system-level scaling and interconnect capability. The Rubin Ultra's key feature is now its support for the NVL576 architecture, which can connect up to 576 GPUs into a single massive logical GPU via NVLink.

QHow did the stock market, particularly HBM manufacturers, react to the reports about Rubin Ultra's specifications?

AThe stock market reacted negatively. Following the news, shares of major HBM manufacturers SK Hynix and Samsung both dropped approximately 8% in early trading, reflecting market concerns about potential weakening demand for HBM from a key buyer like NVIDIA.

QWhat does NVIDIA's design adjustment for Rubin Ultra potentially signal about the future of AI infrastructure development?

AIt potentially signals a shift away from a 'reckless material-piling and price-raising' era. It indicates that even dominant players like NVIDIA are now focusing on cost-optimized hardware design and seeking more efficient ways to scale AI systems, rather than simply increasing specifications like HBM capacity.

Letture associate

China's Tech Industry Is Crossing the 'Visibility Threshold' in Bulk

China's tech industries are successively crossing a "visibility threshold" into the global market spotlight. This wave, exemplified by AI, robotics, and innovative pharmaceuticals (the "new-new three"), differs from past waves like garments/appliances or EVs/batteries/solar panels. Earlier industries met existing demand, but these new sectors are helping to *define* emerging markets and applications. Crossing this threshold signifies an industry has achieved scale, competitive advantages, rapid growth, and the ability to drive its supply chain. Recent data shows China's AI sector growing over 30%, producing 80% of certain intelligent robots, and a surge in new drug approvals. This acceleration stems from China's industrial evolution from "completeness" to "density": different sectors now share technologies, talent, and infrastructure, enabling capabilities (e.g., from EVs) to rapidly migrate to adjacent fields (e.g., robotics). This dense ecosystem, combined with improved system integration and faster market access, acts as an infrastructure for generating new industries. The nature of competition is also evolving. The visible product is just an interface; long-term advantage lies in the underlying patents, standards, and open ecosystems. While building this "moat," China's approach emphasizes openness—through open-source models and collaborative R&D—to foster wider adoption and iterative improvement. Crossing the visibility line marks the start of market validation. The consecutive emergence of these industry waves signals China is developing a sustained capacity to generate new, globally relevant innovations, shifting from being defined by global narratives to participating in shaping them.

marsbit29 min fa

China's Tech Industry Is Crossing the 'Visibility Threshold' in Bulk

marsbit29 min fa

Trading

Spot
活动图片