NVIDIA released internal performance data for its Vera CPU, an Arm-based processor designed specifically for agentic AI workloads. This release matters because early adopters like OpenAI, Anthropic, SpaceX, and Perplexity are already integrating the chip into their infrastructure. Buyers and developers should watch this closely as the industry shifts toward specialized hardware for complex AI agent runtimes.

Arm-based chip targets agentic AI with massive memory bandwidth
The Vera processor sits on a monolithic die and features 88 custom Olympus cores. Each core supports two threads, giving the chip a total of 176 threads. NVIDIA designed the Olympus architecture to handle workloads with large instruction footprints and frequent control-flow changes. This design choice targets agent runtimes, interpreters, compilers, and graph analytics frameworks that often leave execution resources underutilized on traditional architectures.
NVIDIA Vera CPU Specifications
- Core Count: 88 custom Olympus cores
- Thread Count: 176 threads
- L3 Cache: 164MB
- Memory Type: SOCAMM LPDDR5X 9600 MT/s
- Memory Capacity: 1.5TB
Technical details highlight a significant memory overhaul for the Vera CPU. The chip includes 164MB of L3 cache and pairs with 1.5TB of SOCAMM LPDDR5X memory running at 9600 MT/s. This configuration delivers 1.2TB/s of memory bandwidth. An NVLink-C2C integration provides 1.8TB/s of bandwidth between the CPU and GPU. NVIDIA states that this setup delivers up to 3X the bandwidth with 40% lower latency compared to x86 solutions.
Performance claims show Vera outperforming AMD EPYC Turin in specific metrics. The CPU offers up to a 1.9X improvement in IPC over Turin when testing various AI workflows. Branch prediction performance increases by up to 2.3X compared to the same AMD competitor. Vera also achieves 3.5X higher taken-branches per cycle than x86 rivals. In practical Agentic AI tasks, the chip runs 1.8X faster in Python and 1.5X faster in data processing.
NVIDIA explained the engineering behind these gains in a direct quote. The company stated that Olympus addresses execution challenges with advanced branch prediction and a 10-wide decode engine. A neural branch predictor sits at the center of this approach to improve accuracy on difficult branch patterns. We touched on Nvidia and Microsoft Bet AI Can in our earlier Nvidia coverage to track similar vendor strategies.
Vera represents a concrete step in NVIDIA's expansion into CPU hardware for AI. The processor combines high core counts with massive memory bandwidth to support demanding workloads. Early adopters are already deploying the chip for agentic AI applications. The performance data confirms significant improvements in branch prediction and memory efficiency over existing x86 competitors.
Source: TweakTown




Discussion
0 comments