NVIDIA Vera CPU: Inside the 88-Core Processor Purpose-Built for the Agentic AI Era

Written by

in

At Hot Chips 2026, NVIDIA delivered the most detailed technical breakdown yet of its Vera CPU, a purpose-built Arm server processor featuring 88 custom Olympus cores designed specifically for agentic AI workloads. Presented on August 25, 2026, the session revealed new benchmarks, memory architecture decisions, and platform-level integration details that underscore NVIDIA’s ambitions to challenge Intel and AMD in the data center CPU market. The Vera CPU is NVIDIA’s first processor built entirely around its own custom core design, a significant departure from the Grace CPU, which used a stock Arm Neoverse N2 core. Its release as part of the broader Vera Rubin platform marks a strategic bet that the agentic AI era demands fundamentally different silicon from the ground up.

What Was Announced

At Hot Chips 2026, NVIDIA engineers presented comprehensive technical details about the Vera CPU, the compute heart of the company’s next-generation Vera Rubin AI platform. The chip features 88 custom Olympus cores across a monolithic compute die, supported by eight 128-bit LPDDR5X memory controllers capable of delivering up to 1.2 TB/s of memory bandwidth through the new SOCAMM2 form factor.

Unlike the Grace CPU, which used a standard Arm Neoverse N2 core, Vera marks the first time NVIDIA has built a fully custom Arm-based server CPU core from scratch. The Olympus core architecture was designed to prioritize single-threaded execution speed and low memory latency over raw core count, traits that matter most in orchestration-heavy agentic workloads.

NVIDIA’s Hot Chips presentation also revealed that Vera ships in a split-die configuration: a single monolithic compute die houses all 88 cores, while memory and I/O functions are handled by separate chiplets. These components connect via NVIDIA’s NVLink-C2C interconnect, which also links two Vera CPUs in a dual-socket configuration or connects the CPU to Rubin GPUs in the tightly integrated Vera Rubin AI factory system.

The Vera Rubin platform as a whole spans seven distinct chips and five rack configurations, encompassing the Vera CPU, the Rubin GPU, the BlueField-4 networking card, the Spectrum-6 Ethernet switch, and Groq LPUs for inference acceleration. NVIDIA described it as a full-stack AI factory platform designed from end to end for large-scale agentic AI deployment.

Technical Details

Spatial multithreading is one of Vera’s most distinctive design features. NVIDIA’s implementation splits core execution resources across two parallel pipelines, but allows data and cache to move freely between threads as workloads shift. This architecture is well-suited to agentic AI tasks, where a CPU must simultaneously manage code execution, memory I/O, tool call scheduling, and multi-step orchestration loops without stalling on any single pipeline.

In benchmarks presented at Hot Chips, NVIDIA reported close to 1.8x performance improvement on agentic workloads compared to traditional rack-scale CPUs, with data-processing workloads showing a more modest 1.5x improvement. NVIDIA also claimed roughly 2x efficiency gains across the board, a metric reflecting compute delivered per watt rather than raw throughput.

The memory subsystem uses LPDDR5X connected via eight 128-bit memory controllers, delivering low-latency, high-bandwidth access suited to the scatter-gather memory patterns typical in agentic pipelines. NVIDIA deliberately avoided high-bandwidth memory (HBM) for the CPU, a tradeoff that prioritizes energy efficiency and lower fabrication cost. This places Vera in a distinct niche from GPU-class accelerators, even within the Vera Rubin platform itself.

Industry Impact and Reactions

The Vera CPU puts NVIDIA in direct competition with AMD’s EPYC Zen 6 and Intel’s Xeon 7 series for data center CPU deployments. NVIDIA’s positioning, however, is differentiated from both: rather than competing on core count or general-purpose throughput, the company is framing Vera as a specialized AI orchestration processor for the agentic era.

The move mirrors a broader industry pattern of purpose-built silicon for AI. Just as GPUs displaced CPUs for AI training workloads over the past decade, NVIDIA is betting that CPU architectures must similarly evolve to handle the next wave of inference and agentic tasks at scale. The company stated that major cloud providers and enterprise infrastructure vendors are planning to adopt Vera as part of Vera Rubin platform deployments.

From a competitive standpoint, Intel and AMD have both introduced AI-optimized cores in their server processor lines, but neither offers the tight CPU-to-GPU integration that NVLink-C2C enables in the Vera Rubin system. That coupling is particularly important for agentic AI applications where the CPU and GPU must coordinate at low latency to execute multi-step AI pipelines with minimal overhead.

What Comes Next

NVIDIA has indicated that Vera Rubin system deployments will begin ramping through the second half of 2026, following the platform’s production readiness announcement earlier this year. Enterprises and cloud providers are expected to receive early allocations through the remainder of 2026 as NVIDIA scales manufacturing in partnership with TSMC.

Additional technical details and partner announcements related to Vera and the Vera Rubin platform are expected to emerge through the remainder of the Hot Chips 2026 conference, which continues through August 26.

Conclusion

NVIDIA’s Hot Chips 2026 presentation on the Vera CPU marks an inflection point in the AI hardware landscape. By building a CPU from the ground up for agentic AI, NVIDIA is not only expanding its addressable data center market but signaling a broader design philosophy: the infrastructure of the next AI wave will need to be rearchitected at every level, from the GPU up through the CPU and interconnects. The Vera Rubin platform represents NVIDIA’s most vertically integrated AI system to date, and the technical details unveiled today confirm it is built for a world where autonomous AI agents are the primary computational workload.

Stay updated on the latest AI news at Evolve Digital.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *