Rubin Ultra Makes Major Cuts, Even Nvidia Can't Handle Memory Price Hikes?

Odaily星球日报2026-08-03 tarihinde yayınlandı2026-08-03 tarihinde güncellendi

Özet

NVIDIA's Rubin Ultra, the top-tier variant of the newly announced Rubin AI accelerators, has reportedly seen significant specification downgrades, according to an industry report from SemiAnalysis. Initially designed with four compute dies (4-die), the Rubin Ultra is now said to be reduced to a 2-die design. Key changes highlighted in the report include: * **No increase in peak theoretical compute performance**, remaining at 35 PFLOPs like the standard Rubin. * **Severe reduction in memory capacity** to 192GB using 8-Hi HBM stacks, which is less than the standard Rubin's 288GB using 12-Hi stacks. * **Negligible memory bandwidth improvement** of only 1 TB/s. * **Slightly higher chip-level power consumption**. * The **primary upgrade is a massive increase in scale-up interconnect capacity**, supporting connections for up to 576 GPUs via NVLink, compared to 72 for the standard Rubin. The report suggests the redesign is primarily a cost-optimization move driven by the sharp rise in HBM (High-Bandwidth Memory) prices. By reducing the expensive HBM content and shifting investment towards enhanced system-scale networking, NVIDIA aims to maintain the platform's value for large-scale AI training clusters while managing soaring material costs. The news reportedly triggered a sell-off in South Korean memory stocks, with SK Hynix and Samsung shares falling around 8%, as markets grew concerned that NVIDIA—a major HBM buyer—might be reducing its reliance on high-capacity memor...

Original | Odaily Planet Daily (@OdailyChina)

Author | Azuma (@azuma_eth)

This past weekend, a research report from the well-known investment research firm SemiAnalysis sparked intense discussion across the entire AI industry.

The core content of the report is that following SemiAnalysis's late June revelation that Nvidia's original 4-die Rubin Ultra design would be halved, the research firm has now disclosed that Nvidia has provided major customers with a preview of the Rubin Ultra, whose specifications have been further reduced from prior expectations.

Screenshots of the Rubin Ultra information leaked in the SemiAnalysis report are as follows:

  • Rubin Ultra will maintain the same theoretical peak computing power as Rubin, both at 35 PFLOPs;
  • VRAM capacity severely downgraded, specifications reduced to 8-Hi stacked 192GB, even lower than the 12-Hi stacked 288GB of Rubin;
  • VRAM bandwidth change is negligible, only increased by 1 TB/s, essentially insignificant in actual high-throughput computing;
  • Chip-level power consumption even slightly higher, minimum power consumption same as Rubin (1800 W), maximum power consumption is actually higher (2600 W);
  • The only major improvement is the "Scale-up world size", increased from 72 GPUs to 576 GPUs, meaning Nvidia has shifted the real selling point to the cluster network. Through the NVLink network, it can connect up to 576 Rubin Ultra GPUs into a massive "super logical GPU" (the standard version only supports up to 72 cards interconnected).

As the top flagship version of Nvidia's Rubin, heavily announced at GTC 2026 this year, Rubin Ultra was originally positioned to handle extreme-scale AI model training and inference needs by integrating more chip dies and high-bandwidth memory. However, judging from the latest specification details disclosed by SemiAnalysis, Nvidia has evidently adjusted its design approach for Rubin Ultra.

HBM Price Hikes Too Fast, Nvidia Starts Recalculating

Over the past two years, one of the most critical components in AI infrastructure expansion has undoubtedly been HBM.

With the explosive demand for Nvidia H100, H200, Blackwell, and other AI accelerators, high-bandwidth memory has transformed from a relatively niche high-end storage product into the most constrained link in the entire AI infrastructure. While SK Hynix, Samsung, and Micron have been repeatedly posting record earnings and ramping up HBM investments, the market supply-demand imbalance continues to drive HBM prices upward.

For AI chip manufacturers, the importance of HBM is self-evident — GPUs handle computation, HBM provides high-speed data throughput, and together they determine AI model training and inference efficiency. The problem, however, is that HBM has now become so expensive it's beginning to impact the overall economics of AI systems. Taking HBM3 as an example, the price for one HBM3 stack was only $180-220 in Q2 2025 lows. By Q1 of this year (contract price), it had risen to $600-700, and Q2 (spot price) surged further to $700-850.

According to SemiAnalysis estimates, as HBM prices rose, the pure bill-of-materials (BOM) cost for a single Rubin Ultra rack once increased from about $6.6 million to $8 million. After adjusting the design specifications, the cost could drop back to around $6.4 million.

For Nvidia, such a huge cost difference poses a very real question — if the original design of increasing HBM capacity is continued, can it still deliver performance improvements commensurate with the cost? Are there other, better alternatives?

SemiAnalysis's report provides an answer to this: The adjustment to Rubin Ultra essentially represents Nvidia's re-optimization of the cost structure for AI systems under limited resources. This means reducing the relatively expensive HBM configuration (its cost share drops from nearly 40% to 28%) and redirecting resources towards higher-value expansion interconnect capabilities (cost share increases from 4% to 12%).

As mentioned earlier, the core upgrade direction for Rubin Ultra has now shifted to "system-level expansion and interconnect capabilities." The NVL576 architecture supported by Rubin Ultra can connect up to 576 GPUs via NVLink into a unified compute domain, aiming to compensate for the adjustments in single-chip specifications with larger-scale system expansion.

Memory Stocks Plunge, Market Worries About Demand Peaking

Potentially impacted by this news, South Korean memory-related stocks collectively declined upon opening this morning. As of 11:45 Beijing Time, SK Hynix and Samsung, the two leading HBM manufacturers, both plunged about 8%, with the KOSPI index also down around 5%.

The market has begun to worry: If the Rubin Ultra specification adjustments revealed by SemiAnalysis are true (Nvidia has not yet publicly confirmed related information), does it mean that Nvidia, as the most crucial buyer in AI infrastructure, is reducing its demand for HBM?

Over the past two years, with the rapid growth of AI scaling demand, HBM has become the most constrained part of AI infrastructure. Chip manufacturers like Nvidia and AMD continuously increased HBM configurations in AI accelerators, driving soaring earnings for memory makers SK Hynix, Samsung, and Micron, and strengthening their pricing power across the supply chain.

Nvidia's choice, however, might indicate that AI chip makers are considering optimizing hardware design to reduce reliance on high-capacity HBM per chip. If this path proves feasible, the room for continuous price hikes by HBM manufacturers could be significantly limited.

The era of "madly stacking specs and mindlessly raising prices" in AI infrastructure construction will eventually come to an end. And now, even Nvidia, seated upon the iron throne of computing power, has started to pinch pennies.

İlgili Sorular

QAccording to the article, what is the most significant change in NVIDIA's Rubin Ultra design compared to initial expectations?

AThe most significant change is a severe downgrade in memory (HBM) capacity. The Rubin Ultra will now feature 8-layer stacked (8-Hi) 192GB HBM, which is less than the 12-layer stacked 288GB of the standard Rubin.

QWhat financial reason is suggested for NVIDIA's decision to reduce the HBM specification in the Rubin Ultra?

AThe article suggests that the rapid and significant price increase of HBM has impacted the overall system economics. Reducing the HBM configuration from nearly 40% to 28% of the Bill of Materials (BOM) cost allows NVIDIA to control the rising cost of the Rubin Ultra system.

QWhat area becomes the new primary focus or selling point for the Rubin Ultra after the HBM downgrade?

AAfter the HBM downgrade, the primary upgrade focus shifts to system-level scaling and interconnect capability. The Rubin Ultra's key feature is now its support for the NVL576 architecture, which can connect up to 576 GPUs into a single massive logical GPU via NVLink.

QHow did the stock market, particularly HBM manufacturers, react to the reports about Rubin Ultra's specifications?

AThe stock market reacted negatively. Following the news, shares of major HBM manufacturers SK Hynix and Samsung both dropped approximately 8% in early trading, reflecting market concerns about potential weakening demand for HBM from a key buyer like NVIDIA.

QWhat does NVIDIA's design adjustment for Rubin Ultra potentially signal about the future of AI infrastructure development?

AIt potentially signals a shift away from a 'reckless material-piling and price-raising' era. It indicates that even dominant players like NVIDIA are now focusing on cost-optimized hardware design and seeking more efficient ways to scale AI systems, rather than simply increasing specifications like HBM capacity.

İlgili Okumalar

Bill CLARITY Proposes Adding Requirements for Cryptoplatform Developers

A new amendment to the CLARITY bill proposes creating a category called "non-decentralized financial trading protocols." This includes individuals or groups with direct or indirect control over a protocol's functions, operations, or consensus rules. These entities would be required to register with the Commodity Futures Trading Commission (CFTC). The bill also mandates the CFTC and Treasury Department to develop specific rules and standards for these platforms. The stated goal is to distinguish truly decentralized projects from those that are effectively controlled, for example, through a majority of governance tokens or a project treasury. The amendment aims to resolve previous debates by clearly assigning compliance responsibilities, including for customer funds, to developers and governing communities even if they don't hold the assets directly. Additionally, the updated bill proposes legal protections for stablecoin issuers and exchanges that freeze assets suspected of being involved in illicit activity, shielding them from civil lawsuits demanding unfreezing. All crypto platforms would also be subject to bank-like regulations. The senator sponsoring the amendments hopes they will garner broader support for the CLARITY bill in an upcoming procedural Senate vote. The bill retains an "ethics provision" prohibiting government officials and their spouses from issuing, promoting, or engaging in crypto business, though a separate proposal could allow certain individuals, like former President Trump, a tax deferral on crypto asset sales.

cryptonews.ru29 dk önce

Bill CLARITY Proposes Adding Requirements for Cryptoplatform Developers

cryptonews.ru29 dk önce

Novo Holdings points to lack of investment in quantum applications

Novo Holdings reports that from 2014-2025, about 70% of the $13.9 billion in private investment in the quantum sector went to hardware and components. Investment in business applications remains insufficient relative to their potential value. By early September 2026, disclosed private funding for quantum companies had already surpassed the total for all of 2025. The report argues that while initial hardware investment was necessary, the next phase requires funding for specialized algorithms and services alongside computer development. These solutions should help customers with specific tasks like selecting molecules for research. Developers who create hard-to-replicate services—combining algorithms, proprietary data, and client integration—will gain an advantage. The life sciences, including pharmaceuticals, are identified as a highly promising field for early quantum applications. However, even early fault-tolerant computers are not expected to handle entire drug development workflows. Instead, they will tackle critical, small-scale tasks where classical approximations fail, such as calculating chemical bond behavior. More routine calculations will still be performed by classical computers. For a quantum service to be commercially valuable, its results must tangibly influence a client's next decision, like which experiment to run. Useful applications in materials science, energy, and chemistry may emerge earlier than in pharma. The report outlines two scenarios: early fault-tolerant systems in the late 2020s/early 2030s and more scalable systems from the mid-2030s. Key risks include progress in classical computing and AI potentially reducing quantum advantages, and the possibility of in-house development by pharmaceutical firms or cloud providers.

cryptonews.ru30 dk önce

Novo Holdings points to lack of investment in quantum applications

cryptonews.ru30 dk önce

U.S. Treasury's $5 Billion+ Buyback of 10-Year Bonds Failed to Stop Their Sell-Off

On September 10th, the U.S. Treasury, led by Scott Bessent, conducted a significantly larger-than-usual buyback of long-term government bonds, purchasing $5.2 billion worth. However, this intervention failed to halt a continued sell-off, with the yield on the benchmark 10-year Treasury note rising above 4.85% for the first time since 2023 and hitting nearly 4.98% the next day. Investors viewed the scale of the operation as insufficient for the massive $32 trillion U.S. debt market. They cited persistent pressures including rising oil prices, inflation risks, a large budget deficit, and expectations of further Federal Reserve rate hikes as reasons the buyback couldn't reverse the trend. The Treasury had announced it was prepared to buy up to $6 billion in long-term debt, triple its standard $2 billion operation. Yet, it only accepted $5.2 billion of the over $10 billion in investor offers. Market participants, including strategists from Bank of America and Amundi, criticized the move as too timid and potentially counterproductive, signaling government "nervousness" without addressing core issues. Key factors driving yields higher include the U.S. national debt exceeding $40 trillion, soaring debt servicing costs, strong new borrowing needs, and inflationary pressures from high energy prices. Additional concerns were raised by proposals for large new government spending. Analysts concluded the market is effectively "fighting the U.S. Treasury," with the intervention unable to overcome these fundamental headwinds.

cryptonews.ru32 dk önce

U.S. Treasury's $5 Billion+ Buyback of 10-Year Bonds Failed to Stop Their Sell-Off

cryptonews.ru32 dk önce

Fly Brain Successfully Made to Engage in Crypto Trading

Researchers have conducted an experiment using a digital connectome (brain map) of a fruit fly, the Drosophila MaleCNS v1.0, containing 166,700 neurons and 25.6 million synaptic connections. In this project, real-time Bitcoin ($BTC) to US Dollar Coin ($USDC) price data from Coinbase was converted into an RGB visual signal. This signal was fed into the fly brain model, with "buy" actions encoded as "sweet" and "sell" as "bitter." A fixed neural reader then translated the brain's activity into trading commands—buy, sell, or hold—which were routed through an intermediary service to place spot orders on Coinbase Advanced. The experiment ran in a safe, simulated environment with virtual $100, not real money. The programmer noted that real-money trading is technically possible by creating a limited portfolio (max $100 USDC) with restricted API keys. However, the fly's model would be capped at trades of $10 each, with a maximum of 24 trades per day, no leverage, no shorting, and trading halted after a $20 portfolio loss. The system incorporates a learning mechanism: profitable outcomes stimulate 15 dopamine-like neurons (PAM11), while losses activate two aversive neurons (PPL101), encouraging the model to avoid unprofitable trades. Despite this, the developer acknowledges that the fly brain has not yet demonstrated an ability to trade profitably; synaptic changes do not equate to a successful trading strategy. This remains a research experiment. In related news, a Solana-based bot recently turned $0.227 USDC into roughly $696,000 in seconds by capitalizing on a sharp price drop of the memecoin ANB.

cryptonews.ru32 dk önce

Fly Brain Successfully Made to Engage in Crypto Trading

cryptonews.ru32 dk önce

İşlemler

Spot
活动图片