NVIDIA Stuns with First Vera Rubin Test, DeepSeek Throughput Skyrockets 30x

marsbitPublicado a 2026-08-25Actualizado a 2026-08-25

Resumen

NVIDIA has unveiled the first on-silicon test results for its next-generation flagship cabinet, the Vera Rubin NVL72, using DeepSeek-V4-Pro on real "agent coding" workloads. The results are staggering: compared to the current leading GB300 NVL72, Vera Rubin delivers up to a 30x increase in throughput per megawatt and reduces token generation cost by up to 35x. The benchmark used was the AgentX test from SemiAnalysis, which captures full AI agent workflows with growing context, tool calls, and sub-agent generation, moving beyond traditional LLM benchmarks. This highlights a shift from the LLM era to the Agent era. Key innovations behind Vera Rubin's performance include extreme co-design, separation of services, distributed KV cache, KV-aware routing, MegaMoE architecture, 4-bit NVFP4 quantization, and 6th-gen NVLink for efficient MoE model execution. Simultaneously, NVIDIA announced the full-scale production of two new chips: 1. **Groq 3 LPX:** A low-latency inference accelerator designed for the Vera Rubin platform. When paired with Rubin GPUs for context processing, it achieves record-breaking output speeds—e.g., running Gemma 4 31B at 3,400 tokens/sec—drastically reducing multi-step agent task times. 2. **Vera CPU:** A processor specifically built for agentic AI, featuring 88 custom Olympus cores and 1.2TB/s memory bandwidth to handle the complex orchestration of agent tasks. SpaceXAI is already deploying it, with plans for space-based Vera Rubin systems by 2028. NVID...

It's mind-blowing.

Just now, NVIDIA for the first time ever announced the initial on-chip test data of its next-generation flagship cabinet, the Vera Rubin NVL72.

And, for this test, they directly pulled in DeepSeek -V4-Pro, running the most realistic 'Agent Coding' task!

The results are astonishing: compared to the current reigning champion GB300 NVL72, Vera Rubin's throughput per megawatt has skyrocketed up to 30x, and token cost has plummeted up to 35x!

Someone exclaimed, 'I wouldn't even dare write a PPT to fool investors like this!'

Another netizen commented brilliantly, 'I used to think the H200 at thirty thousand dollars a card was expensive. Looking back now, the H200 doesn't even qualify as the budget version anymore.'

Let's take a closer look.

From H200 to GB300, the throughput for real Agent workloads increased up to 30x, while cost roughly doubled.

From GB300 to Vera Rubin, throughput soared another 30x, with the whole rack price roughly doubling again.

According to Jensen's Law of Moore, it turns out NVIDIA's biggest technological breakthrough in the past two years isn't the GPU itself, but making the price 4 times more expensive while delivering a whopping 900x speed leap!

This time, NVIDIA announced a counterintuitive truth to everyone: The LLM era is over, the Agent era has arrived, all previous AI benchmark standards are obsolete!

Meanwhile, the Vera CPU, built specifically for agents, has been installed overnight by Musk's SpaceXAI, and is even preparing to be launched directly into space.

Also today, the NVIDIA Groq 3 LPX has entered full-scale production.

Gemma 4 31B running on it is simply stunning, hitting 3400 Tokens per second. Trillion-parameter model throughput has surged 35x.

These new records directly rewrite the commercial landscape of the Agent ecosystem, starting tonight!

Why Must NVIDIA Launch Vera Rubin?

OpenRouter's latest data provides the answer: In the real world, a single Agent AI task consumes 15x the tokens of a regular chat conversation!

For example, if an Agent is tasked with researching a company for an investment decision, during this process, the agent and sub-agents will continuously reason, and the accumulated tokens become the input for the next steps. Long context handling thus becomes the most critical limiting factor for agentic AI.

The more interactions an agent has, the greater the throughput — the number of agents increased 10x, tool calls increased 2x, the contradiction between computing power and algorithms has spawned new demands.

Don't Mistake Chat for Agents, Old Benchmarks Are All Useless!

At this point, traditional AI benchmarks (like fixed-length 8K/1K sequence tests) become completely invalid.

NVIDIA officially clarified: Performance measurement must evolve! We can no longer test single inference requests; we must capture the complete Agent workflow.

To this end, they used the AgentX benchmark from SemiAnalysis.

This test is no longer rigid Q&A, but replays real 'code-writing sessions' featuring contextual growth, tool calls, and sub-agent generation.

This is why the H200 appears inadequate on the new Agent battlefield, and why Vera Rubin is destined to reign.

Vera Rubin NVL72 × DeepSeek-V4-Pro:

The Compute Monster Has Arrived

To demonstrate the power of Vera Rubin NVL72, NVIDIA directly used the open-source king — DeepSeek -V4-Pro (1.6T) for on-chip testing.

Once the data was out, the results were staggering —

Under the AgentX workload, Vera Rubin NVL72's throughput per megawatt was a full 30x higher than GB300 NVL72 at its peak!

Note, this comparison isn't against the old H200, but the hot, mainstream GB300.

In the same DeepSeek -V4-Pro test, the GB300's throughput per megawatt was already 15x higher than the H200's.

Yet Vera Rubin, standing on GB300's shoulders, raised it another 30x, lifting the entire Pareto curve!

What does this mean?

For 'AI factories' constrained by power supply, this directly expands the boundaries of physical laws.

Under the same power budget, Vera Rubin can handle 30x more agentic work.

For megawatt-scale or even gigawatt-scale data centers, this is like conjuring 30x more compute assets out of thin air.

Furthermore, NVIDIA DSX MaxLPS technology enables power management at the GPU, rack, and workload levels, allowing up to 40% more GPUs to be configured within the same megawatt budget, further increasing AI factory throughput per megawatt!

Token Cost Plummets 35x! The Agent Industry's 'Ledger' Is Completely Rewritten

Throughput per megawatt directly impacts the cost per generated token. The performance surge brings the most immediate consequence: a nuclear explosion in business models.

NVIDIA announced that Vera Rubin NVL72's cost to produce one million tokens is up to 35x lower than GB300 NVL72!

When inference costs plummet 35x, 24/7 digital employees will become reality, and super-apps in the ToC sector will experience an explosion.

Jensen's move has directly sliced open the gateway to the Agent application era.

Extreme Co-design: How Did Jensen Do It?

You might ask: How can it achieve a 30x speedup? Is it due to single-chip process technology?

Obviously not.

The astounding performance leap of Vera Rubin NVL72 is powered by 'Extreme Co-design'.

For example, disaggregated serving, distributed KV cache, KV-aware routing, MegaMoE, and more.

Additionally, NVFP4 quantization directly compresses model weights to 4-bit precision, dramatically reducing memory footprint without sacrificing output quality, making throughput take off.

Meanwhile, the sixth-generation NVLink and large MoE models provide an interconnect network 10x faster and with 3x lower latency than standard Ethernet, allowing large models like DeepSeek , based on MoE architecture, to fluidly call different 'expert' sub-networks across 72 GPUs.

This is no longer about selling graphics cards; Jensen is selling an entire 'AI power plant'.

NVIDIA Groq 3 LPX in Full Production:

3400 Tokens/Sec, Coding Agents Are Transformed!

At the Hot Chips 2026 conference, NVIDIA also dropped a bombshell: Groq 3 LPX has begun full-scale production!

The Groq 3 LPX is the exclusive extension for the Vera Rubin NVL72 data center platform.

It is 'tailor-made' for this system, specializing in low-latency inference acceleration.

This black tech was acquired by NVIDIA last December from the startup Groq for a cool $20 billion.

AI agents encounter the issue of decoding latency. To end this pain point, the Groq 3 LPX cleverly separates massive context processing from token generation entirely.

NVIDIA's solution: let the Rubin GPU handle large-scale context, and delegate the ultra-fast token generation task to the Groq 3 LPX.

One handles 'reading,' the other handles 'writing,' each doing what they do best.

In a full rack-scale deployment, up to 256 LP30 accelerators can collaborate with GPUs via ultra-high-bandwidth interconnects, constructing an enterprise-scale inference engine.

Consequently, the Groq 3 LPX achieves record-breaking output speeds.

In Artificial Analysis's benchmarks, the Groq 3 LPX running Gemma 4 31B with a 100K token ultra-long context window achieved a staggering output speed of 3,400 Tokens/sec!

This makes its response speed 4x faster than competitors' for latency-sensitive workloads.

Multi-step agent tasks that used to take hours can now be completed in minutes!

For generating 5000 tokens, the time taken dropped from 50 seconds to 1.5 seconds, a 34x difference.

And in coding tasks, the Groq 3 LPX reached a maximum output speed of 6981 token/s.

Jensen Huang stated directly that by advancing ultra-high-speed token generation via LPX, they have achieved 'yet another huge leap' in AI throughput and responsiveness.

Nebius has already deployed the chip in its Nebius Token Factory, delivering an ultimate experience of instant response for every step of the agentic loop.

Following closely behind is Groq itself.

First Agent-Dedicated CPU Released,

Musk Installs It Overnight 'Into Space'!

The most ambitious move in this announcement is the Vera CPU.

For Agents, NVIDIA specifically built a CPU.

The Vera CPU is designed specifically to feed agents.

It is equipped with 88 in-house designed Olympus cores and high-bandwidth LPDDR5X memory, with bandwidth hitting a staggering 1.2TB/s.

Why does the LLM era need a brand new CPU?

On the Hot Chips stage, NVIDIA's Vera CPU executive stated directly, 'Agentic AI is the most complex computational task in history.'

A single task can involve an agent running hundreds of steps behind the scenes: calling tools, executing Python code, searching context, processing massive data...

All these coordination and dispatch tasks have been borne by the CPU. Hence, NVIDIA had to build a separate CPU for Agents.

Recognizing this point, Musk's SpaceXAI overnight announced the formal, scaled deployment of NVIDIA's Vera CPU!

Back in May, NVIDIA quietly sent the first Vera samples to 'test the waters' at A社, OpenAI, and SpaceXAI.

Just three months later, SpaceXAI couldn't wait any longer and moved into full deployment.

Even more insane, SpaceXAI is building a gigawatt-scale compute factory based on the Vera Rubin platform to power Grok.

Moreover, this computing power is set to be launched directly into space.

Musk stated, 'The two powerhouses have joined forces to design an optimized version of the Vera Rubin NVL72, to be deployed at scale in 2028.'

SpaceXAI's first-generation 'Starmind' AI satellite plans to adopt the Vera Rubin NVL72 rack-scale system.

From a Single GPU to an Entire Agent Factory

On the same day, two major chips, the Vera CPU and Groq 3 LPX, entered full-scale production.

This is highly significant for NVIDIA.

In the LLM era, NVIDIA captured training and inference with GPUs.

In the Agent era, its territory has already expanded to CPUs, inference accelerators, NVLink, and more.

Jensen is no longer selling just a GPU.

He's selling an AI factory that can continuously produce tokens.

This time, Jensen is re-laying the foundation for the entire 'Agent Era'.

References:

https://x.com/MinLiBuilds/status/2091915873661686204

https://developer.nvidia.com/blog/nvidia-vera-rubin-and-blackwell-set-a-new-standard-for-agentic-ai-performance-per-watt/

https://developer.nvidia.com/blog/inside-nvidia-groq-3-lpx-the-low-latency-inference-accelerator-for-the-nvidia-vera-rubin-platform/

https://developer.nvidia.com/blog/maximizing-ai-factory-performance-per-watt-with-nvidia-dsx-maxlps/

This article is from WeChat public account "New Zhiyuan", author: ASI Revelation

Preguntas relacionadas

QWhat is the key performance improvement announced for NVIDIA's next-generation Vera Rubin NVL72 cabinet compared to the current GB300 NVL72?

AIn the first on-chip test using DeepSeek-V4-Pro on the AgentX benchmark (simulating real agentic coding workloads), the NVIDIA Vera Rubin NVL72 demonstrated up to a 30x improvement in throughput per megawatt and up to a 35x reduction in token cost compared to the GB300 NVL72.

QWhy does NVIDIA claim traditional AI benchmarks are now obsolete, and what new benchmark was used for the Vera Rubin tests?

ANVIDIA claims traditional benchmarks (like fixed 8K/1K sequence tests) are obsolete because they don't capture the complex, growing-context, multi-step nature of real agentic AI workflows. For the Vera Rubin tests, they used the AgentX benchmark from SemiAnalysis, which replays realistic 'code-writing sessions' involving context growth, tool calls, and sub-agent generation.

QWhat is the NVIDIA Groq 3 LPX, and what specific performance benefit does it provide alongside Vera Rubin?

AThe NVIDIA Groq 3 LPX is a low-latency inference accelerator specifically designed as an extension for the Vera Rubin platform. It specializes in ultra-fast token generation, offloading this task from the GPU. In tests, it achieved speeds of up to 3,400 tokens/second for the Gemma 4 31B model, making agentic tasks that used to take hours complete in minutes.

QWhat is the Vera CPU, and why did NVIDIA develop a dedicated CPU for the agentic AI era?

AThe Vera CPU is NVIDIA's first CPU specifically designed for agentic AI workloads. It features 88 custom Olympus cores and high-bandwidth LPDDR5X memory (1.2 TB/s). NVIDIA developed it because agentic tasks involve hundreds of complex steps like tool calling, code execution, and context management, which place immense scheduling and coordination burdens on the CPU that general-purpose CPUs aren't optimized for.

QWhich major company was announced as an early adopter of the Vera Rubin platform, and what ambitious deployment plan was revealed?

AElon Musk's SpaceXAI was announced as an early and major adopter. They have already begun large-scale deployment of the Vera CPU and are building gigawatt-scale compute factories based on the Vera Rubin platform to power Grok. Furthermore, they plan to deploy an optimized version of the Vera Rubin NVL72 system on their first-generation 'Starmind' AI satellites in space by 2028.

Lecturas Relacionadas

Dos surcoreanos me cuentan: los aumentos salariales en los semiconductores son minoritarios, y las ganancias en bolsa también son solo 'historias felices'

Dos coreanos (un directivo de Samsung y el propietario de una clínica de estética) comparten su perspectiva sobre el reciente boom y la volatilidad del mercado bursátil surcoreano, impulsado por el sector de los semiconductores. Contrariamente a los relatos exagerados y "demasiado dramáticos" que circulan en las redes sociales sobre una "edad de oro" de euforia colectiva, la realidad social es más contenida. Los coreanos no suelen mostrar abiertamente sus ganancias o pérdidas. El aumento del interés por la bolsa es real (más cuentas, más inversión), pero no un frenesí público. En cuanto al sector semiconductor, un ejecutivo de Samsung aclara que los rumores sobre grandes aumentos salariales generalizados son inexactos. Los sindicatos coreanos ("sindicatos nobles") a menudo representan a una minoría privilegiada, y las mejoras competitivas entre empresas como Samsung y SK suelen limitarse a empleados clave o investigadores, no a la plantilla general. El crecimiento del sector se traduce más en nueva contratación que en redistribución de salarios. Históricamente, la inversión inmobiliaria era la principal vía de acumulación de riqueza en Corea. Solo recientemente, debido a la regulación del mercado inmobiliario y las políticas de estímulo bursátil, la bolsa ha ganado atractivo, llevando a muchos novatos a invertir, a veces con préstamos o apalancamiento. Esta oleada de nuevos inversores, combinada con la toma de beneficios por parte de capitales extranjeros e institucionales, ha contribuido a la alta volatilidad y a las recientes interrupciones ("circuit breakers"). Muchos inversores minoristas, sobreendeudados, han sufrido grandes pérdidas, reduciendo gastos y vendiendo activos. Prevalece una sensación de desilusión ("al final, con acciones no se gana dinero"). A pesar de la corrección, uno de los entrevistados mantiene el optimismo a largo plazo sobre el valor de empresas líderes como Samsung o SK Hynix y sigue invirtiendo. La conclusión es que, más allá del desconocimiento cultural, la dinámica especulativa en Corea no difiere mucho de la de otros mercados: la percepción de que "los demás viven mejor" suele ser un espejismo.

marsbitHace 52 min(s)

Dos surcoreanos me cuentan: los aumentos salariales en los semiconductores son minoritarios, y las ganancias en bolsa también son solo 'historias felices'

marsbitHace 52 min(s)

¿Vale Unitree Technology realmente 240.000 millones de yuanes?

A mediados de agosto de 2026, Unitree Tech cotizó en la STAR Market con una capitalización bursátil que alcanzó brevemente los 444.900 millones de yuanes antes de estabilizarse alrededor de los 243.900 millones. Esto plantea la pregunta: ¿justifica una valoración tan elevada una empresa con unos ingresos de 17.000 millones de yuanes en 2025? El análisis utiliza el modelo de ingresos residuales de Ohlson, evaluando cuatro dimensiones clave: ROE, sostenibilidad, crecimiento y riesgo. **ROE:** Unitree demostró alta rentabilidad antes de la OPV, pero la inyección masiva de capital tras la salida a bolsa diluirá temporalmente el ROE. El reto es reconstruirlo mediante márgenes, rotación de activos y disciplina en la asignación de capital. **Sostenibilidad:** Su ventaja reside en el control motor y capacidades de ingeniería integral. Sin embargo, la verdadera barrera comercial requiere transformar los "gestos" demostrativos en una "capacidad laboral" estable, fiable y con costes controlados para los clientes, generando flujos de caja recurrentes. **Crecimiento:** El camino a largo plazo implica evolucionar de vender hardware a ofrecer soluciones por escenarios y, finalmente, convertirse en una plataforma de "trabajo" con un ecosistema de datos y modelos que se refuerce a sí mismo (más despliegues → más datos → modelos más inteligentes). **Evaluación de Riesgos:** La estructura de gobierno con derechos de voto especiales del fundador acelera la toma de decisiones, pero requiere sólidos contrapesos institucionales. Los riesgos de ESG, como la seguridad física, la privacidad de datos y la normativa exterior (ej. nuevas restricciones de la FCC en EE.UU.), son factores de riesgo crecientes que afectan a la tasa de descuento. **Conclusión sobre la valoración:** Con una capitalización de ~243.900 millones de yuanes (~144 veces los ingresos de 2025), el mercado está anticipando un crecimiento excepcional y sostenido durante años. Unitree es una empresa destacada con tecnología, producto e ingresos reales, pero el precio actual deja un margen de seguridad muy estrecho, exigiendo una ejecución casi impecable en todos los frentes para justificar las elevadas expectativas implícitas. La inversión se basa en su capacidad para convertir la ventaja técnica en un ROE alto y sostenible.

marsbitHace 53 min(s)

¿Vale Unitree Technology realmente 240.000 millones de yuanes?

marsbitHace 53 min(s)

¡Increíble! Cosmos publica un parche de alta gravedad sin avisar antes, y los hackers vacían las arcas de los proyectos

En los últimos días, el ecosistema Cosmos ha experimentado una serie de ataques devastadores en cadenas como MANTRA, TAC, KiiChain y Nesa, todas ellas construidas con el módulo Cosmos EVM. Los atacantes drenaron las reservas de tesorería de los protocolos, causando caídas de más del 90% en los precios de tokens nativos como KII, TAC y NES en cuestión de horas. La causa raíz fue una vulnerabilidad crítica corregida en la actualización v0.7.2 del módulo Cosmos EVM, publicada por Cosmos Labs en GitHub el 19 de agosto. Aunque la publicación destacaba la importancia y urgencia de la corrección de seguridad, Cosmos Labs no proporcionó advertencias privadas ni notificaciones obligatorias a los equipos de los proyectos dependientes antes de hacerla pública. Esta falta de coordinación permitió a los atacantes estudiar el parche público y explotar la vulnerabilidad antes de que muchas cadenas pudieran actualizar. El error explotaba una combinación de tres defectos en el módulo, incluido un desbordamiento por debajo (*underflow*), afectando a todas las cadenas con cuentas de liquidación gradual (*vesting accounts*) habilitadas. KiiChain y otros proyectos criticaron duramente a Cosmos Labs por su manejo irresponsable, señalando que la crisis "se podría haber evitado" con una comunicación y una respuesta de emergencia adecuadas. A pesar de que los ataques continuaron durante días, la declaración pública y las recomendaciones de Cosmos Labs para pausar las cadenas llegaron demasiado tarde, después de que se produjeran pérdidas masivas. El incidente ha generado una amplia crítica en la comunidad, acusando a Cosmos Labs de fallar en sus deberes de coordinación y de carecer de un proceso robusto para desplegar parches de seguridad críticos para todo su ecosistema. Este evento ha puesto de manifiesto graves deficiencias en la auditoría de código, los mecanismos de coordinación y los protocolos de respuesta a emergencias de Cosmos, erosionando aún más la confianza en una red que ya enfrentaba desafíos de adopción y éxodo de proyectos.

marsbitHace 58 min(s)

¡Increíble! Cosmos publica un parche de alta gravedad sin avisar antes, y los hackers vacían las arcas de los proyectos

marsbitHace 58 min(s)

Los capitalistas de riesgo comienzan a predecir el futuro con IA

DigClaw, empresa especializada en inteligencia predictiva, ha logrado posicionar su marco de predicción Rhizome v1 en los puestos #1, #3 y #7 en la plataforma de evaluación FutureX, utilizando tres modelos base distintos (como Kimi-K3 y DeepSeek-V4-Pro). Este resultado valida su hipótesis central: la capacidad predictiva puede residir fuera del modelo base, en un sistema que combina de forma estructurada búsqueda, razonamiento causal e inferencia probabilística. El sistema Rhizome aborda las limitaciones de los LLMs para la predicción (falta de causalidad, razonamiento contrafáctico y calibración) mediante tres decisiones de diseño clave: 1) Desacopla la búsqueda de información (optimizada para relevancia) del razonamiento causal. 2) Registra cada predicción en una trayectoria completa para una calibración probabilística continua, creando un activo de datos único. 3) Implementa un marco de actualización bayesiana consciente de las cadenas causales para evitar el recuento doble de evidencias y el exceso de confianza. Esta infraestructura predictiva, probada externamente en FutureX, es el núcleo de Newborn Ventures, un fondo de capital de riesgo impulsado por IA que aplica estas capacidades para identificar tendencias (Beta) y oportunidades de inversión. DigClaw ofrece este sistema para la toma de decisiones estratégicas en empresas, instituciones financieras y fondos gubernamentales.

marsbitHace 1 hora(s)

Los capitalistas de riesgo comienzan a predecir el futuro con IA

marsbitHace 1 hora(s)

Trading

Spot
活动图片