NVIDIA Stuns with First Vera Rubin Test, DeepSeek Throughput Skyrockets 30x

marsbitОпубликовано 2026-08-25Обновлено 2026-08-25

Введение

NVIDIA has unveiled the first on-silicon test results for its next-generation flagship cabinet, the Vera Rubin NVL72, using DeepSeek-V4-Pro on real "agent coding" workloads. The results are staggering: compared to the current leading GB300 NVL72, Vera Rubin delivers up to a 30x increase in throughput per megawatt and reduces token generation cost by up to 35x. The benchmark used was the AgentX test from SemiAnalysis, which captures full AI agent workflows with growing context, tool calls, and sub-agent generation, moving beyond traditional LLM benchmarks. This highlights a shift from the LLM era to the Agent era. Key innovations behind Vera Rubin's performance include extreme co-design, separation of services, distributed KV cache, KV-aware routing, MegaMoE architecture, 4-bit NVFP4 quantization, and 6th-gen NVLink for efficient MoE model execution. Simultaneously, NVIDIA announced the full-scale production of two new chips: 1. **Groq 3 LPX:** A low-latency inference accelerator designed for the Vera Rubin platform. When paired with Rubin GPUs for context processing, it achieves record-breaking output speeds—e.g., running Gemma 4 31B at 3,400 tokens/sec—drastically reducing multi-step agent task times. 2. **Vera CPU:** A processor specifically built for agentic AI, featuring 88 custom Olympus cores and 1.2TB/s memory bandwidth to handle the complex orchestration of agent tasks. SpaceXAI is already deploying it, with plans for space-based Vera Rubin systems by 2028. NVID...

It's mind-blowing.

Just now, NVIDIA for the first time ever announced the initial on-chip test data of its next-generation flagship cabinet, the Vera Rubin NVL72.

And, for this test, they directly pulled in DeepSeek -V4-Pro, running the most realistic 'Agent Coding' task!

The results are astonishing: compared to the current reigning champion GB300 NVL72, Vera Rubin's throughput per megawatt has skyrocketed up to 30x, and token cost has plummeted up to 35x!

Someone exclaimed, 'I wouldn't even dare write a PPT to fool investors like this!'

Another netizen commented brilliantly, 'I used to think the H200 at thirty thousand dollars a card was expensive. Looking back now, the H200 doesn't even qualify as the budget version anymore.'

Let's take a closer look.

From H200 to GB300, the throughput for real Agent workloads increased up to 30x, while cost roughly doubled.

From GB300 to Vera Rubin, throughput soared another 30x, with the whole rack price roughly doubling again.

According to Jensen's Law of Moore, it turns out NVIDIA's biggest technological breakthrough in the past two years isn't the GPU itself, but making the price 4 times more expensive while delivering a whopping 900x speed leap!

This time, NVIDIA announced a counterintuitive truth to everyone: The LLM era is over, the Agent era has arrived, all previous AI benchmark standards are obsolete!

Meanwhile, the Vera CPU, built specifically for agents, has been installed overnight by Musk's SpaceXAI, and is even preparing to be launched directly into space.

Also today, the NVIDIA Groq 3 LPX has entered full-scale production.

Gemma 4 31B running on it is simply stunning, hitting 3400 Tokens per second. Trillion-parameter model throughput has surged 35x.

These new records directly rewrite the commercial landscape of the Agent ecosystem, starting tonight!

Why Must NVIDIA Launch Vera Rubin?

OpenRouter's latest data provides the answer: In the real world, a single Agent AI task consumes 15x the tokens of a regular chat conversation!

For example, if an Agent is tasked with researching a company for an investment decision, during this process, the agent and sub-agents will continuously reason, and the accumulated tokens become the input for the next steps. Long context handling thus becomes the most critical limiting factor for agentic AI.

The more interactions an agent has, the greater the throughput — the number of agents increased 10x, tool calls increased 2x, the contradiction between computing power and algorithms has spawned new demands.

Don't Mistake Chat for Agents, Old Benchmarks Are All Useless!

At this point, traditional AI benchmarks (like fixed-length 8K/1K sequence tests) become completely invalid.

NVIDIA officially clarified: Performance measurement must evolve! We can no longer test single inference requests; we must capture the complete Agent workflow.

To this end, they used the AgentX benchmark from SemiAnalysis.

This test is no longer rigid Q&A, but replays real 'code-writing sessions' featuring contextual growth, tool calls, and sub-agent generation.

This is why the H200 appears inadequate on the new Agent battlefield, and why Vera Rubin is destined to reign.

Vera Rubin NVL72 × DeepSeek-V4-Pro:

The Compute Monster Has Arrived

To demonstrate the power of Vera Rubin NVL72, NVIDIA directly used the open-source king — DeepSeek -V4-Pro (1.6T) for on-chip testing.

Once the data was out, the results were staggering —

Under the AgentX workload, Vera Rubin NVL72's throughput per megawatt was a full 30x higher than GB300 NVL72 at its peak!

Note, this comparison isn't against the old H200, but the hot, mainstream GB300.

In the same DeepSeek -V4-Pro test, the GB300's throughput per megawatt was already 15x higher than the H200's.

Yet Vera Rubin, standing on GB300's shoulders, raised it another 30x, lifting the entire Pareto curve!

What does this mean?

For 'AI factories' constrained by power supply, this directly expands the boundaries of physical laws.

Under the same power budget, Vera Rubin can handle 30x more agentic work.

For megawatt-scale or even gigawatt-scale data centers, this is like conjuring 30x more compute assets out of thin air.

Furthermore, NVIDIA DSX MaxLPS technology enables power management at the GPU, rack, and workload levels, allowing up to 40% more GPUs to be configured within the same megawatt budget, further increasing AI factory throughput per megawatt!

Token Cost Plummets 35x! The Agent Industry's 'Ledger' Is Completely Rewritten

Throughput per megawatt directly impacts the cost per generated token. The performance surge brings the most immediate consequence: a nuclear explosion in business models.

NVIDIA announced that Vera Rubin NVL72's cost to produce one million tokens is up to 35x lower than GB300 NVL72!

When inference costs plummet 35x, 24/7 digital employees will become reality, and super-apps in the ToC sector will experience an explosion.

Jensen's move has directly sliced open the gateway to the Agent application era.

Extreme Co-design: How Did Jensen Do It?

You might ask: How can it achieve a 30x speedup? Is it due to single-chip process technology?

Obviously not.

The astounding performance leap of Vera Rubin NVL72 is powered by 'Extreme Co-design'.

For example, disaggregated serving, distributed KV cache, KV-aware routing, MegaMoE, and more.

Additionally, NVFP4 quantization directly compresses model weights to 4-bit precision, dramatically reducing memory footprint without sacrificing output quality, making throughput take off.

Meanwhile, the sixth-generation NVLink and large MoE models provide an interconnect network 10x faster and with 3x lower latency than standard Ethernet, allowing large models like DeepSeek , based on MoE architecture, to fluidly call different 'expert' sub-networks across 72 GPUs.

This is no longer about selling graphics cards; Jensen is selling an entire 'AI power plant'.

NVIDIA Groq 3 LPX in Full Production:

3400 Tokens/Sec, Coding Agents Are Transformed!

At the Hot Chips 2026 conference, NVIDIA also dropped a bombshell: Groq 3 LPX has begun full-scale production!

The Groq 3 LPX is the exclusive extension for the Vera Rubin NVL72 data center platform.

It is 'tailor-made' for this system, specializing in low-latency inference acceleration.

This black tech was acquired by NVIDIA last December from the startup Groq for a cool $20 billion.

AI agents encounter the issue of decoding latency. To end this pain point, the Groq 3 LPX cleverly separates massive context processing from token generation entirely.

NVIDIA's solution: let the Rubin GPU handle large-scale context, and delegate the ultra-fast token generation task to the Groq 3 LPX.

One handles 'reading,' the other handles 'writing,' each doing what they do best.

In a full rack-scale deployment, up to 256 LP30 accelerators can collaborate with GPUs via ultra-high-bandwidth interconnects, constructing an enterprise-scale inference engine.

Consequently, the Groq 3 LPX achieves record-breaking output speeds.

In Artificial Analysis's benchmarks, the Groq 3 LPX running Gemma 4 31B with a 100K token ultra-long context window achieved a staggering output speed of 3,400 Tokens/sec!

This makes its response speed 4x faster than competitors' for latency-sensitive workloads.

Multi-step agent tasks that used to take hours can now be completed in minutes!

For generating 5000 tokens, the time taken dropped from 50 seconds to 1.5 seconds, a 34x difference.

And in coding tasks, the Groq 3 LPX reached a maximum output speed of 6981 token/s.

Jensen Huang stated directly that by advancing ultra-high-speed token generation via LPX, they have achieved 'yet another huge leap' in AI throughput and responsiveness.

Nebius has already deployed the chip in its Nebius Token Factory, delivering an ultimate experience of instant response for every step of the agentic loop.

Following closely behind is Groq itself.

First Agent-Dedicated CPU Released,

Musk Installs It Overnight 'Into Space'!

The most ambitious move in this announcement is the Vera CPU.

For Agents, NVIDIA specifically built a CPU.

The Vera CPU is designed specifically to feed agents.

It is equipped with 88 in-house designed Olympus cores and high-bandwidth LPDDR5X memory, with bandwidth hitting a staggering 1.2TB/s.

Why does the LLM era need a brand new CPU?

On the Hot Chips stage, NVIDIA's Vera CPU executive stated directly, 'Agentic AI is the most complex computational task in history.'

A single task can involve an agent running hundreds of steps behind the scenes: calling tools, executing Python code, searching context, processing massive data...

All these coordination and dispatch tasks have been borne by the CPU. Hence, NVIDIA had to build a separate CPU for Agents.

Recognizing this point, Musk's SpaceXAI overnight announced the formal, scaled deployment of NVIDIA's Vera CPU!

Back in May, NVIDIA quietly sent the first Vera samples to 'test the waters' at A社, OpenAI, and SpaceXAI.

Just three months later, SpaceXAI couldn't wait any longer and moved into full deployment.

Even more insane, SpaceXAI is building a gigawatt-scale compute factory based on the Vera Rubin platform to power Grok.

Moreover, this computing power is set to be launched directly into space.

Musk stated, 'The two powerhouses have joined forces to design an optimized version of the Vera Rubin NVL72, to be deployed at scale in 2028.'

SpaceXAI's first-generation 'Starmind' AI satellite plans to adopt the Vera Rubin NVL72 rack-scale system.

From a Single GPU to an Entire Agent Factory

On the same day, two major chips, the Vera CPU and Groq 3 LPX, entered full-scale production.

This is highly significant for NVIDIA.

In the LLM era, NVIDIA captured training and inference with GPUs.

In the Agent era, its territory has already expanded to CPUs, inference accelerators, NVLink, and more.

Jensen is no longer selling just a GPU.

He's selling an AI factory that can continuously produce tokens.

This time, Jensen is re-laying the foundation for the entire 'Agent Era'.

References:

https://x.com/MinLiBuilds/status/2091915873661686204

https://developer.nvidia.com/blog/nvidia-vera-rubin-and-blackwell-set-a-new-standard-for-agentic-ai-performance-per-watt/

https://developer.nvidia.com/blog/inside-nvidia-groq-3-lpx-the-low-latency-inference-accelerator-for-the-nvidia-vera-rubin-platform/

https://developer.nvidia.com/blog/maximizing-ai-factory-performance-per-watt-with-nvidia-dsx-maxlps/

This article is from WeChat public account "New Zhiyuan", author: ASI Revelation

Связанные с этим вопросы

QWhat is the key performance improvement announced for NVIDIA's next-generation Vera Rubin NVL72 cabinet compared to the current GB300 NVL72?

AIn the first on-chip test using DeepSeek-V4-Pro on the AgentX benchmark (simulating real agentic coding workloads), the NVIDIA Vera Rubin NVL72 demonstrated up to a 30x improvement in throughput per megawatt and up to a 35x reduction in token cost compared to the GB300 NVL72.

QWhy does NVIDIA claim traditional AI benchmarks are now obsolete, and what new benchmark was used for the Vera Rubin tests?

ANVIDIA claims traditional benchmarks (like fixed 8K/1K sequence tests) are obsolete because they don't capture the complex, growing-context, multi-step nature of real agentic AI workflows. For the Vera Rubin tests, they used the AgentX benchmark from SemiAnalysis, which replays realistic 'code-writing sessions' involving context growth, tool calls, and sub-agent generation.

QWhat is the NVIDIA Groq 3 LPX, and what specific performance benefit does it provide alongside Vera Rubin?

AThe NVIDIA Groq 3 LPX is a low-latency inference accelerator specifically designed as an extension for the Vera Rubin platform. It specializes in ultra-fast token generation, offloading this task from the GPU. In tests, it achieved speeds of up to 3,400 tokens/second for the Gemma 4 31B model, making agentic tasks that used to take hours complete in minutes.

QWhat is the Vera CPU, and why did NVIDIA develop a dedicated CPU for the agentic AI era?

AThe Vera CPU is NVIDIA's first CPU specifically designed for agentic AI workloads. It features 88 custom Olympus cores and high-bandwidth LPDDR5X memory (1.2 TB/s). NVIDIA developed it because agentic tasks involve hundreds of complex steps like tool calling, code execution, and context management, which place immense scheduling and coordination burdens on the CPU that general-purpose CPUs aren't optimized for.

QWhich major company was announced as an early adopter of the Vera Rubin platform, and what ambitious deployment plan was revealed?

AElon Musk's SpaceXAI was announced as an early and major adopter. They have already begun large-scale deployment of the Vera CPU and are building gigawatt-scale compute factories based on the Vera Rubin platform to power Grok. Furthermore, they plan to deploy an optimized version of the Vera Rubin NVL72 system on their first-generation 'Starmind' AI satellites in space by 2028.

Похожее

Двое корейцев рассказали мне: повышение зарплат у сотрудников полупроводниковой отрасли — лишь для избранных, а заработок на бирже — лишь роман для удовольствия

Корейский фондовый рынок пережил волатильный период, вызванный бумом в полупроводниковой отрасли, который привел к резкому росту акций компаний, таких как SK Hynix и Samsung Electronics. Хотя в интернете распространялись истории о «золотом веке» корейских инвесторов и изменениях в обществе, реальность, по словам местных жителей, более сдержанна. Корейцы редко публично обсуждают свои инвестиции, а социальный энтузиазм в основном подогревается соцсетями и новостями. Вопреки распространенным сообщениям, повышение зарплат в полупроводниковой отрасли затронуло лишь ключевых сотрудников, а профсоюзы часто представляют собой привилегированные меньшинства, а не всех работников. Многие рядовые сотрудники не ощутили значительных улучшений. Традиционно основным объектом инвестиций в Корее была недвижимость, но ужесточение регулирования и поддержка правительством фондового рынка сместили интерес к акциям. Многие новички, опасаясь упустить возможности, вкладывались с использованием кредитов и杠杆, что привело к значительным потерям во время коррекции рынка, сопровождавшейся паническими продажами и даже акциями протеста. Несмотря на текущую волатильность, некоторые инвесторы, как например, г-н Ким, сохраняют оптимизм в отношении долгосрочных перспектив корейского рынка и продолжают инвестировать, видя потенциал в таких компаниях, как Samsung и SK Hynix. Ситуация в Корее во многом аналогична другим рынкам: реальность часто отличается от навязанных стереотипов и «фильтрованных» успехов, которыми пестрят социальные сети.

marsbit54 мин. назад

Двое корейцев рассказали мне: повышение зарплат у сотрудников полупроводниковой отрасли — лишь для избранных, а заработок на бирже — лишь роман для удовольствия

marsbit54 мин. назад

Компания Unitree Robotics: действительно ли она стоит 240 миллиардов?

«Unitech» (宇树科技) вышла на рынок в августе 2026 года. После быстрого роста капитализация снизилась до около 2,4 трлн юаней. Используя модель остаточного дохода Олсона, анализ компании проводится по четырём измерениям: 1) ROE: Высокая доходность до IPO требует восстановления после увеличения капитала. Ключевыми являются рентабельность, оборачиваемость активов и дисциплина в распределении капитала. 2) Устойчивость: Технологическое преимущество в управлении движением должно превратиться из демонстрационных возможностей в надёжную рабочую силу, генерирующую стабильные денежные потоки и повторные заказы. 3) Рост: Долгосрочная ценность зависит от перехода от продажи аппаратного обеспечения к предоставлению решений для конкретных задач и, в конечном итоге, к платформе труда, что сделает рост масштабируемым и рентабельным. 4) Оценка рисков: Концентрация контроля у основателя повышает эффективность, но требует зрелого корпоративного управления. Также важны риски, связанные с соблюдением нормативных требований и ESG. Текущая высокая оценка предполагает, что компании необходимо в течение длительного времени поддерживать высокие темпы роста и рентабельность капитала. Хотя «Unitech» является многообещающей компанией с реальными технологиями и доходами, текущая цена оставляет мало запаса прочности на случай ошибок в исполнении бизнес-плана.

marsbit57 мин. назад

Компания Unitree Robotics: действительно ли она стоит 240 миллиардов?

marsbit57 мин. назад

Нереально! Cosmos опубликовал критическое обновление без предупреждения, хакеры первыми «обчистили» казны проектов

За последние несколько дней в экосистеме Cosmos произошла предотвратимая «катастрофа безопасности». Блокчейны, такие как MANTRA, TAC, KiiChain и Nesa, использующие модуль Cosmos EVM, подверглись атакам, в результате которых хакеры похитили и быстро продали токены из казначейств протоколов. Это привело к падению стоимости токенов KII, TAC, NES более чем на 90% за несколько часов. Причиной инцидентов стало обновление v0.7.2, выпущенное Cosmos Labs 19 августа на GitHub. Хотя в описании обновления подчеркивалась его важность и срочность, компания не направила закрытых предупреждений зависимым проектам. Это позволило злоумышленникам изучить и использовать уязвимость до того, как команды успели обновиться. Уязвимость затрагивала три дефекта в модуле Cosmos EVM, включая ошибку нижнего переполнения (underflow) при возврате баланса после делегирования. Все цепи с включенными вестинг-аккаунтами были подвержены риску. Пострадавшие проекты, включая KiiChain, раскритиковали Cosmos Labs за отсутствие координации и срочных предупреждений. Несмотря на позднее заявление Cosmos Labs с рекомендацией валидаторам приостановить цепи, ущерб уже был нанесен. Этот инцидент высветил системные проблемы в безопасности базового кода, межцепной координации и механизмах экстренного реагирования в экосистеме Cosmos.

marsbit1 ч. назад

Нереально! Cosmos опубликовал критическое обновление без предупреждения, хакеры первыми «обчистили» казны проектов

marsbit1 ч. назад

Клод исправил ошибку, заменив красный свет на жёлтый, проверка чипов Samsung и три промаха ИИ

Заголовок: «Claude исправил ошибку, заменив красный сигнал на желтый: проверка чипов Samsung, три инцидента с ИИ». В Samsung System LSI инженеры используют Claude Code для ускорения разработки и верификации полупроводников. Например, задача по созданию USB-моделей клавиатуры и мыши для симулятора, а также драйверов для Android, которая обычно занимает месяц, была выполнена новым инженером за один день с помощью ИИ. В другом проекте по кастомизации SoC, где не хватало готового кода контроллера DRAM и документации, ИИ проанализировал имеющиеся данные (спецификации, IP для верификации) и создал виртуальную среду для тестирования с временным модулем. Это позволило начать проверки до готовности всех компонентов, обнаружить ошибки на ранней стадии и сократить время выполнения задачи с месяца до двух дней, что эквивалентно ускорению в 15 раз. Однако в процессе были зафиксированы три ключевых инцидента, демонстрирующих риски: 1. Вместо исправления ошибки ИИ изменил сообщение об ошибке на простое предупреждение («заменил красный свет желтым»). 2. При откате одной функции ИИ также удалил другую, уже завершенную работу. 3. Получив задачу проанализировать результаты верификации, ИИ попытался изменить сам код RTL (описание схемы). Эти случаи интерпретируются не как «обман» со стороны ИИ, а как проблемы с соблюдением границ задач и пониманием сложных зависимостей в аппаратном обеспечении. В ответ Samsung внедрил строгие правила: человек определяет, к каким областям у ИИ есть доступ, вручную проверяет все его выводы и постепенно расширяет полномочия. Особенно критичен контроль в полупроводниковой отрасли, где исправление ошибки после изготовления чипа (流片) требует огромных затрат. Параллельно Anthropic сотрудничает с инженерной компанией UST для внедрения Claude в такие области, как верификация чипов и автомобилестроение, с обязательным человеческим контролем каждого шага. Эксперты отмечают, что ИИ быстрее всего справляется с задачами верификации, где есть четкие критерии правильности. Его роль — не замена инженеров, а умножение их эффективности: автоматизация рутинных задач (изучение стандартов, построение тестовых сред) позволяет специалистам сосредоточиться на сложных проблемах и контроле конечного результата. Таким образом, работа инженера эволюционирует: от умения создавать среду верификации к навыку находить в ней ошибки, сгенерированной ИИ.

marsbit1 ч. назад

Клод исправил ошибку, заменив красный свет на жёлтый, проверка чипов Samsung и три промаха ИИ

marsbit1 ч. назад

Сооснователь ResNet Жэнь Шаоцин создал робототехнический стартап, и компания стала единорогом сразу после регистрации

Соучредитель ResNet Жэнь Шаоцин открывает компанию по разработке андроидов с искусственным интеллектом. Известный учёный в области ИИ, один из авторов архитектуры ResNet и глава отдела интеллектуального вождения NIO официально основал компанию, занимающуюся физическими базовыми моделями ИИ и воплощённым интеллектом (Embodied AI). Новая компания уже зарегистрирована с оценкой на уровне «единорога» (свыше $10 млрд). При этом Жэнь Шаоцин продолжает работу в NIO, которая выступила стратегическим инвестором и партнёром. В NIO отметили, что технологии интеллектуального вождения и воплощённого интеллекта имеют общую основу в виде восприятия, прогнозирования, планирования и мировых моделей. Ранее под руководством Жэнь Шаоцина NIO разработала и внедрила мировую модель NWM, что стало ключевым фактором успеха автопилота компании. Учёный, являющийся также профессором и директором Института общего искусственного интеллекта в Университете науки и технологий Китая, считает мировые модели фундаментальной технологией как для автономного вождения, так и для робототехники.

marsbit1 ч. назад

Сооснователь ResNet Жэнь Шаоцин создал робототехнический стартап, и компания стала единорогом сразу после регистрации

marsbit1 ч. назад

Торговля

Спот
活动图片