It's mind-blowing.
Just now, NVIDIA for the first time ever announced the initial on-chip test data of its next-generation flagship cabinet, the Vera Rubin NVL72.
And, for this test, they directly pulled in DeepSeek -V4-Pro, running the most realistic 'Agent Coding' task!
The results are astonishing: compared to the current reigning champion GB300 NVL72, Vera Rubin's throughput per megawatt has skyrocketed up to 30x, and token cost has plummeted up to 35x!

Someone exclaimed, 'I wouldn't even dare write a PPT to fool investors like this!'
Another netizen commented brilliantly, 'I used to think the H200 at thirty thousand dollars a card was expensive. Looking back now, the H200 doesn't even qualify as the budget version anymore.'

Let's take a closer look.
From H200 to GB300, the throughput for real Agent workloads increased up to 30x, while cost roughly doubled.
From GB300 to Vera Rubin, throughput soared another 30x, with the whole rack price roughly doubling again.
According to Jensen's Law of Moore, it turns out NVIDIA's biggest technological breakthrough in the past two years isn't the GPU itself, but making the price 4 times more expensive while delivering a whopping 900x speed leap!

This time, NVIDIA announced a counterintuitive truth to everyone: The LLM era is over, the Agent era has arrived, all previous AI benchmark standards are obsolete!

Meanwhile, the Vera CPU, built specifically for agents, has been installed overnight by Musk's SpaceXAI, and is even preparing to be launched directly into space.
Also today, the NVIDIA Groq 3 LPX has entered full-scale production.
Gemma 4 31B running on it is simply stunning, hitting 3400 Tokens per second. Trillion-parameter model throughput has surged 35x.


These new records directly rewrite the commercial landscape of the Agent ecosystem, starting tonight!
Why Must NVIDIA Launch Vera Rubin?
OpenRouter's latest data provides the answer: In the real world, a single Agent AI task consumes 15x the tokens of a regular chat conversation!
For example, if an Agent is tasked with researching a company for an investment decision, during this process, the agent and sub-agents will continuously reason, and the accumulated tokens become the input for the next steps. Long context handling thus becomes the most critical limiting factor for agentic AI.

The more interactions an agent has, the greater the throughput — the number of agents increased 10x, tool calls increased 2x, the contradiction between computing power and algorithms has spawned new demands.

Don't Mistake Chat for Agents, Old Benchmarks Are All Useless!
At this point, traditional AI benchmarks (like fixed-length 8K/1K sequence tests) become completely invalid.
NVIDIA officially clarified: Performance measurement must evolve! We can no longer test single inference requests; we must capture the complete Agent workflow.
To this end, they used the AgentX benchmark from SemiAnalysis.
This test is no longer rigid Q&A, but replays real 'code-writing sessions' featuring contextual growth, tool calls, and sub-agent generation.

This is why the H200 appears inadequate on the new Agent battlefield, and why Vera Rubin is destined to reign.
Vera Rubin NVL72 × DeepSeek-V4-Pro:
The Compute Monster Has Arrived
To demonstrate the power of Vera Rubin NVL72, NVIDIA directly used the open-source king — DeepSeek -V4-Pro (1.6T) for on-chip testing.
Once the data was out, the results were staggering —
Under the AgentX workload, Vera Rubin NVL72's throughput per megawatt was a full 30x higher than GB300 NVL72 at its peak!

Note, this comparison isn't against the old H200, but the hot, mainstream GB300.
In the same DeepSeek -V4-Pro test, the GB300's throughput per megawatt was already 15x higher than the H200's.
Yet Vera Rubin, standing on GB300's shoulders, raised it another 30x, lifting the entire Pareto curve!
What does this mean?
For 'AI factories' constrained by power supply, this directly expands the boundaries of physical laws.
Under the same power budget, Vera Rubin can handle 30x more agentic work.
For megawatt-scale or even gigawatt-scale data centers, this is like conjuring 30x more compute assets out of thin air.
Furthermore, NVIDIA DSX MaxLPS technology enables power management at the GPU, rack, and workload levels, allowing up to 40% more GPUs to be configured within the same megawatt budget, further increasing AI factory throughput per megawatt!

Token Cost Plummets 35x! The Agent Industry's 'Ledger' Is Completely Rewritten
Throughput per megawatt directly impacts the cost per generated token. The performance surge brings the most immediate consequence: a nuclear explosion in business models.
NVIDIA announced that Vera Rubin NVL72's cost to produce one million tokens is up to 35x lower than GB300 NVL72!

When inference costs plummet 35x, 24/7 digital employees will become reality, and super-apps in the ToC sector will experience an explosion.
Jensen's move has directly sliced open the gateway to the Agent application era.

Extreme Co-design: How Did Jensen Do It?
You might ask: How can it achieve a 30x speedup? Is it due to single-chip process technology?
Obviously not.
The astounding performance leap of Vera Rubin NVL72 is powered by 'Extreme Co-design'.

For example, disaggregated serving, distributed KV cache, KV-aware routing, MegaMoE, and more.

Additionally, NVFP4 quantization directly compresses model weights to 4-bit precision, dramatically reducing memory footprint without sacrificing output quality, making throughput take off.
Meanwhile, the sixth-generation NVLink and large MoE models provide an interconnect network 10x faster and with 3x lower latency than standard Ethernet, allowing large models like DeepSeek , based on MoE architecture, to fluidly call different 'expert' sub-networks across 72 GPUs.
This is no longer about selling graphics cards; Jensen is selling an entire 'AI power plant'.

NVIDIA Groq 3 LPX in Full Production:
3400 Tokens/Sec, Coding Agents Are Transformed!
At the Hot Chips 2026 conference, NVIDIA also dropped a bombshell: Groq 3 LPX has begun full-scale production!

The Groq 3 LPX is the exclusive extension for the Vera Rubin NVL72 data center platform.
It is 'tailor-made' for this system, specializing in low-latency inference acceleration.
This black tech was acquired by NVIDIA last December from the startup Groq for a cool $20 billion.

AI agents encounter the issue of decoding latency. To end this pain point, the Groq 3 LPX cleverly separates massive context processing from token generation entirely.
NVIDIA's solution: let the Rubin GPU handle large-scale context, and delegate the ultra-fast token generation task to the Groq 3 LPX.
One handles 'reading,' the other handles 'writing,' each doing what they do best.

In a full rack-scale deployment, up to 256 LP30 accelerators can collaborate with GPUs via ultra-high-bandwidth interconnects, constructing an enterprise-scale inference engine.

Consequently, the Groq 3 LPX achieves record-breaking output speeds.
In Artificial Analysis's benchmarks, the Groq 3 LPX running Gemma 4 31B with a 100K token ultra-long context window achieved a staggering output speed of 3,400 Tokens/sec!
This makes its response speed 4x faster than competitors' for latency-sensitive workloads.

Multi-step agent tasks that used to take hours can now be completed in minutes!

For generating 5000 tokens, the time taken dropped from 50 seconds to 1.5 seconds, a 34x difference.

And in coding tasks, the Groq 3 LPX reached a maximum output speed of 6981 token/s.

Jensen Huang stated directly that by advancing ultra-high-speed token generation via LPX, they have achieved 'yet another huge leap' in AI throughput and responsiveness.

Nebius has already deployed the chip in its Nebius Token Factory, delivering an ultimate experience of instant response for every step of the agentic loop.

Following closely behind is Groq itself.

First Agent-Dedicated CPU Released,
Musk Installs It Overnight 'Into Space'!
The most ambitious move in this announcement is the Vera CPU.
For Agents, NVIDIA specifically built a CPU.
The Vera CPU is designed specifically to feed agents.
It is equipped with 88 in-house designed Olympus cores and high-bandwidth LPDDR5X memory, with bandwidth hitting a staggering 1.2TB/s.

Why does the LLM era need a brand new CPU?
On the Hot Chips stage, NVIDIA's Vera CPU executive stated directly, 'Agentic AI is the most complex computational task in history.'

A single task can involve an agent running hundreds of steps behind the scenes: calling tools, executing Python code, searching context, processing massive data...
All these coordination and dispatch tasks have been borne by the CPU. Hence, NVIDIA had to build a separate CPU for Agents.
Recognizing this point, Musk's SpaceXAI overnight announced the formal, scaled deployment of NVIDIA's Vera CPU!
Back in May, NVIDIA quietly sent the first Vera samples to 'test the waters' at A社, OpenAI, and SpaceXAI.
Just three months later, SpaceXAI couldn't wait any longer and moved into full deployment.
Even more insane, SpaceXAI is building a gigawatt-scale compute factory based on the Vera Rubin platform to power Grok.
Moreover, this computing power is set to be launched directly into space.
Musk stated, 'The two powerhouses have joined forces to design an optimized version of the Vera Rubin NVL72, to be deployed at scale in 2028.'

SpaceXAI's first-generation 'Starmind' AI satellite plans to adopt the Vera Rubin NVL72 rack-scale system.

From a Single GPU to an Entire Agent Factory
On the same day, two major chips, the Vera CPU and Groq 3 LPX, entered full-scale production.
This is highly significant for NVIDIA.

In the LLM era, NVIDIA captured training and inference with GPUs.
In the Agent era, its territory has already expanded to CPUs, inference accelerators, NVLink, and more.
Jensen is no longer selling just a GPU.
He's selling an AI factory that can continuously produce tokens.
This time, Jensen is re-laying the foundation for the entire 'Agent Era'.
References:
https://x.com/MinLiBuilds/status/2091915873661686204
https://developer.nvidia.com/blog/nvidia-vera-rubin-and-blackwell-set-a-new-standard-for-agentic-ai-performance-per-watt/
https://developer.nvidia.com/blog/inside-nvidia-groq-3-lpx-the-low-latency-inference-accelerator-for-the-nvidia-vera-rubin-platform/
https://developer.nvidia.com/blog/maximizing-ai-factory-performance-per-watt-with-nvidia-dsx-maxlps/
This article is from WeChat public account "New Zhiyuan", author: ASI Revelation





