NVIDIA’s new Vera Rubin NVL72 is here: AI agents don’t just chat — they research, code, and reason through dozens of steps before delivering an answer. That process burns through tokens fast, and NVIDIA’s newest data centre platform is built specifically to handle it. The Vera Rubin NVL72 delivers up to 30x higher throughput per megawatt than its predecessor, according to fresh performance data NVIDIA just published.
Why Agentic AI Needs Different Infrastructure
Here’s the core problem NVIDIA is solving: according to OpenRouter data, agentic AI workloads consume 15x more tokens than a simple chat request. Think about an AI agent researching a company for an investment call — it queries databases, searches filings, spins up sub-agents to run comparisons, then synthesizes everything into a recommendation.
Every step’s output becomes the next step’s input, so context keeps growing throughout the task. That makes handling massive, expanding context efficiently central to how well agentic AI actually performs at scale. For more on how AI infrastructure is evolving, check our AI and technology coverage.

The Headline Numbers
| Metric | Vera Rubin NVL72 vs. GB300 NVL72 |
|---|---|
| Throughput per megawatt | Up to 30x higher |
| Cost per million tokens | Up to 35x lower |
| GPUs per megawatt budget | Up to 40% more (via DSX MaxLPS power management) |
NVIDIA measured these figures using SemiAnalysis AgentX, a benchmark built from recorded real-world agentic coding sessions — meaning actual context growth, tool calls, and sub-agent spawning, not a simplified synthetic test. The results are currently pending independent SemiAnalysis review.

Why This Actually Matters for AI Companies
For data centres running on fixed power budgets, throughput per megawatt directly determines how much revenue an AI factory can generate, while cost per million tokens determines the profit margin on top of that revenue. A 30x efficiency jump isn’t just a benchmark bragging right — it’s the difference between running agentic AI profitably at scale or not.
How NVIDIA Got There: Extreme Codesign

Rather than one single breakthrough, Vera Rubin’s gains come from stacking multiple optimization techniques together across the entire platform:
- Disaggregated serving — separates context processing from response generation so each scales independently
- Distributed KV-caching — spreads memory across GPUs and offloads inactive context to storage, avoiding recomputation
- Large-scale expert parallelism — distributes sub-networks in mixture-of-experts models across the GPU cluster
- NVFP4 quantization — compresses model weights to 4-bit precision without sacrificing output quality
Underpinning all of this is sixth-generation NVLink interconnect technology, delivering 10x higher packet rates and 3x lower latency than standard Ethernet — critical for keeping GPUs talking to each other fast enough to make these techniques work at scale. Full technical details are available via NVIDIA’s Vera Rubin platform page.
What’s Next
Vera Rubin NVL72 is now in full production and scaling across NVIDIA’s partner ecosystem, with the company noting continuous software optimizations should push performance even higher over time. Keep following our AI infrastructure coverage section as more real-world deployments and independent benchmarks emerge.





