NVIDIA Vera Rubin NVL72 Delivers Up to 30x More Work Per Watt for AI Agent Workloads

NVIDIA has detailed the Vera Rubin NVL72 system's efficiency profile specifically for agentic AI inference, claiming up to 30x improvement in work-per-watt compared to prior generations. The NVL72 rack-scale system combines Vera CPUs with Rubin GPUs and is designed end-to-end for the latency and throughput demands of multi-step agent pipelines. For developers deploying agent frameworks at scale, this efficiency gain directly translates to lower per-token costs and the ability to run larger context windows within practical power budgets. NVIDIA positions this as the new baseline for production agentic deployments, meaning cloud providers and enterprises building on NVIDIA infrastructure will have a credible path to cost-competitive agent serving. Developers evaluating inference infrastructure for agent workloads should factor these efficiency numbers into provider and hardware selection.
Read original source ↗Part of the 2026-08-25 briefing→