Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents

New SemiAnalysis AgentX benchmarks show NVIDIA Vera Rubin NVL72 delivers 30x higher throughput per megawatt and 45x lower token cost than GB300 NVL72 on agentic

AI infrastructure teams running agentic workloads at scale can now serve up to 30x more work for every megawatt consumed, and cut token costs by as much as 45x, on NVIDIA’s Vera Rubin NVL72 platform. New measurements from SemiAnalysis on real-world agentic coding trajectories put Vera Rubin NVL72 well ahead of the previous-generation GB300 NVL72 on the metrics that determine AI factory economics: throughput per watt and cost per million tokens. The benchmark, called SemiAnalysis AgentX, reflects recorded agentic coding sessions with authentic context growth, tool calls, and sub-agent spawning, capturing the way agentic AI actually consumes compute rather than how a single chat request does.

Why agentic AI workloads stress infrastructure differently

OpenRouter data cited in the analysis shows agentic AI workloads consume about 15x more tokens than a simple chat request. The reason is structural. When an AI agent researches a company for an investment decision, it queries financial databases, searches news and filings, spawns sub-agents to run peer comparisons and model valuations, then synthesizes everything into a recommendation. Each step adds tokens that become the input to the next, so long-context handling becomes central to performance. The same pattern repeats across software development, customer service, and deep research.

Unlike chat or document summarization, where input and output sequences typically run from 1K to 8K tokens, agentic sessions accumulate context across steps and can reach hundreds of thousands of input tokens, with wide variability between requests. As agentic AI moves into production across industries, the infrastructure running it has to meet that token demand efficiently.

What the SemiAnalysis AgentX measurements show

On the SemiAnalysis AgentX workload, the NVIDIA Blackwell platform delivered leading performance across multiple agentic models, including Kimi K3, GLM5.3, Qwen3.5, and DeepSeek V4 Pro. On DeepSeek V4 Pro, GB300 NVL72 already delivered up to 15x better throughput per megawatt than the NVIDIA Hopper architecture.

Vera Rubin extends that advantage across the entire performance-versus-power curve, reaching as much as 30x higher throughput per megawatt than GB300 NVL72 on the DeepSeek V4 Pro model. The same measurement shows up to 45x lower cost per million tokens than GB300 NVL72, which lets operators run agents continuously at scale across the full breadth of customer workloads.

For power-constrained AI factories, throughput per megawatt determines revenue and cost per million tokens determines profit margin. NVIDIA DSX MaxLPS power-management technology adds further headroom by provisioning up to 40% more GPUs within the same megawatt budget.

How the platform reaches those numbers

The gains come from co-designed optimization across every layer of the Vera Rubin NVL72 stack. Several techniques matter most for agentic inference:

  • Disaggregated serving separates prefill from decode so each scales independently.
  • Rate matching synchronizes prefill and decode GPU speeds to maximize efficiency.
  • Large-scale expert parallelism distributes mixture-of-experts sub-networks across the scale-up GPU domain.
  • Distributed KV-caching extends memory across the scale-up GPU domain, while KV-cache offloading tiers less-active context to host and storage.
  • KV-aware routing directs requests to GPUs that already hold the relevant cached context, cutting redundant computation across long sessions.
  • Fused CUDA kernels such as MegaMoE combine many computation and inter-GPU communication operations into a single execution pass.
  • Fifth-generation Tensor Cores and a third-generation Transformer Engine on Rubin GPUs accelerate both prefill and decode.
  • NVFP4 quantization compresses weights to 4-bit precision, reducing memory footprint and increasing throughput without sacrificing output quality.

The NVL72 scale-up domain, a defining architecture across Vera Rubin and Grace Blackwell, enables the high-bandwidth and low-latency inter-GPU communication that these techniques depend on. Sixth-generation NVIDIA NVLink interconnect and NVLink Switches deliver 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet alternatives. Optimized CUDA kernels, inference runtimes like NVIDIA TensorRT LLM, and serving frameworks like NVIDIA Dynamo complete the software stack, all co-designed with the hardware.

The wider Vera Rubin system

The early AgentX results do not yet reflect Vera CPU performance for tool calling. The full Vera Rubin platform is a seven-chip architecture that also includes the Vera CPU, Groq 3 LPU, NVLink 6 Switch, BlueField-4 DPU, Spectrum-6 SPX, and ConnectX-9 SuperNIC, all built for AI factories deploying agents at scale. Vera Rubin is in full production and is scaling across the ecosystem.

For teams evaluating infrastructure for agentic AI, the practical takeaway is straightforward: throughput per megawatt and cost per token, measured on workloads that mirror real agentic behavior, now separate one generation of hardware from the next by an order of magnitude or more.

FAQ

What does the SemiAnalysis AgentX benchmark measure?

SemiAnalysis AgentX measures performance on recorded real-world agentic coding sessions, preserving actual context growth, tool calls, and sub-agent spawning. It captures the full agent workflow rather than a single inference request.

How much more efficient is NVIDIA Vera Rubin NVL72 than GB300 NVL72 on agentic workloads?

On the SemiAnalysis AgentX workload using the DeepSeek V4 Pro model, Vera Rubin NVL72 delivers up to 30x higher throughput per megawatt and up to 45x lower cost per million tokens than GB300 NVL72.

Why are agentic AI workloads harder on infrastructure than chat?

Agentic sessions accumulate context across steps, reaching hundreds of thousands of input tokens, and OpenRouter data shows they consume about 15x more tokens than a simple chat request. Long-context handling and sustained token throughput become the limiting factors.

Related coverage


This article summarizes reporting from blogs.nvidia.com. See our editorial disclaimer for how our articles are produced.

🤖
Is your business visible to AI assistants?

Run a free scan to see your AI Visibility Score, SEO rating, and local citation accuracy.

Check Your Score →