Original | Odaily Planet Daily (@OdailyChina)
Author | Azuma (@azuma_eth)

This past weekend, a research report from the well-known investment research firm SemiAnalysis sparked intense discussion across the entire AI industry.
The core content of the report is that following SemiAnalysis's late June revelation that Nvidia's original 4-die Rubin Ultra design would be halved, the research firm has now disclosed that Nvidia has provided major customers with a preview of the Rubin Ultra, whose specifications have been further reduced from prior expectations.

Screenshots of the Rubin Ultra information leaked in the SemiAnalysis report are as follows:
- Rubin Ultra will maintain the same theoretical peak computing power as Rubin, both at 35 PFLOPs;
- VRAM capacity severely downgraded, specifications reduced to 8-Hi stacked 192GB, even lower than the 12-Hi stacked 288GB of Rubin;
- VRAM bandwidth change is negligible, only increased by 1 TB/s, essentially insignificant in actual high-throughput computing;
- Chip-level power consumption even slightly higher, minimum power consumption same as Rubin (1800 W), maximum power consumption is actually higher (2600 W);
- The only major improvement is the "Scale-up world size", increased from 72 GPUs to 576 GPUs, meaning Nvidia has shifted the real selling point to the cluster network. Through the NVLink network, it can connect up to 576 Rubin Ultra GPUs into a massive "super logical GPU" (the standard version only supports up to 72 cards interconnected).
As the top flagship version of Nvidia's Rubin, heavily announced at GTC 2026 this year, Rubin Ultra was originally positioned to handle extreme-scale AI model training and inference needs by integrating more chip dies and high-bandwidth memory. However, judging from the latest specification details disclosed by SemiAnalysis, Nvidia has evidently adjusted its design approach for Rubin Ultra.
HBM Price Hikes Too Fast, Nvidia Starts Recalculating
Over the past two years, one of the most critical components in AI infrastructure expansion has undoubtedly been HBM.
With the explosive demand for Nvidia H100, H200, Blackwell, and other AI accelerators, high-bandwidth memory has transformed from a relatively niche high-end storage product into the most constrained link in the entire AI infrastructure. While SK Hynix, Samsung, and Micron have been repeatedly posting record earnings and ramping up HBM investments, the market supply-demand imbalance continues to drive HBM prices upward.
For AI chip manufacturers, the importance of HBM is self-evident — GPUs handle computation, HBM provides high-speed data throughput, and together they determine AI model training and inference efficiency. The problem, however, is that HBM has now become so expensive it's beginning to impact the overall economics of AI systems. Taking HBM3 as an example, the price for one HBM3 stack was only $180-220 in Q2 2025 lows. By Q1 of this year (contract price), it had risen to $600-700, and Q2 (spot price) surged further to $700-850.
According to SemiAnalysis estimates, as HBM prices rose, the pure bill-of-materials (BOM) cost for a single Rubin Ultra rack once increased from about $6.6 million to $8 million. After adjusting the design specifications, the cost could drop back to around $6.4 million.
For Nvidia, such a huge cost difference poses a very real question — if the original design of increasing HBM capacity is continued, can it still deliver performance improvements commensurate with the cost? Are there other, better alternatives?
SemiAnalysis's report provides an answer to this: The adjustment to Rubin Ultra essentially represents Nvidia's re-optimization of the cost structure for AI systems under limited resources. This means reducing the relatively expensive HBM configuration (its cost share drops from nearly 40% to 28%) and redirecting resources towards higher-value expansion interconnect capabilities (cost share increases from 4% to 12%).
As mentioned earlier, the core upgrade direction for Rubin Ultra has now shifted to "system-level expansion and interconnect capabilities." The NVL576 architecture supported by Rubin Ultra can connect up to 576 GPUs via NVLink into a unified compute domain, aiming to compensate for the adjustments in single-chip specifications with larger-scale system expansion.
Memory Stocks Plunge, Market Worries About Demand Peaking
Potentially impacted by this news, South Korean memory-related stocks collectively declined upon opening this morning. As of 11:45 Beijing Time, SK Hynix and Samsung, the two leading HBM manufacturers, both plunged about 8%, with the KOSPI index also down around 5%.
The market has begun to worry: If the Rubin Ultra specification adjustments revealed by SemiAnalysis are true (Nvidia has not yet publicly confirmed related information), does it mean that Nvidia, as the most crucial buyer in AI infrastructure, is reducing its demand for HBM?
Over the past two years, with the rapid growth of AI scaling demand, HBM has become the most constrained part of AI infrastructure. Chip manufacturers like Nvidia and AMD continuously increased HBM configurations in AI accelerators, driving soaring earnings for memory makers SK Hynix, Samsung, and Micron, and strengthening their pricing power across the supply chain.
Nvidia's choice, however, might indicate that AI chip makers are considering optimizing hardware design to reduce reliance on high-capacity HBM per chip. If this path proves feasible, the room for continuous price hikes by HBM manufacturers could be significantly limited.
The era of "madly stacking specs and mindlessly raising prices" in AI infrastructure construction will eventually come to an end. And now, even Nvidia, seated upon the iron throne of computing power, has started to pinch pennies.





