NVIDIA Vera CPU: Olympus Cores Built for Maximum Single-Thread Performance in Agentic AI
Performance Claim:
- NVIDIA claims the Vera CPU delivers up to 1.8x higher performance on agentic workloads. This is based on internal measurements from July 2026 using the SPEC CPU® 2026 benchmark, with performance measured as single-thread throughput under full socket load.
Olympus Core Architecture:
- Front End: Designed for branch-heavy code, it features advanced branch prediction, including a neural branch predictor, a 10-wide decode engine, and the ability to handle up to two taken branches per cycle.
- Mid Core: Focuses on exposing parallelism in dependent work. It uses a wide rename and allocation engine, a large reorder buffer, and deep out-of-order execution. It also includes dependency-breaking features like memory renaming, value prediction, and critical-path acceleration.
- Execution Engine: A wide backend dynamically schedules operations across a broad set of integer, branch, vector, floating-point, cryptographic, load, and store units, prioritizing high IPC over high clock frequency.
- Cache Subsystem: Optimized for irregular memory access patterns common in agent workflows. It includes a deep cache hierarchy, memory disambiguation, multiple hardware prefetch engines, and a specialized graph prefetcher for analytics workloads.
System Architecture and Memory:
| Component | Specification | Benefit |
|---|---|---|
| Multithreading | NVIDIA Spatial Multithreading (SMT) on 88 Olympus cores (176 threads total). | Partitions core resources between two threads to reduce noisy-neighbor effects and provide more predictable performance compared to traditional SMT. |
| Coherency Fabric | NVIDIA Scalable Coherency Fabric (SCF) on a monolithic die. | Provides up to 3.4 TB/s of bisectional bandwidth and a 164 MB unified L3 cache, avoiding die-hop latency common in chiplet designs. |
| Memory | SOCAMM2 LPDDR5X modules. | Delivers up to 1.2 TB/s aggregate bandwidth (14 GB/s per core) at lower power than DDR5. Modules are field-replaceable for enterprise serviceability (RAS). |
Connectivity and Scalability:
- Dual-Socket Scaling: Uses second-generation NVLink-C2C for coherent socket-to-socket connectivity. Each socket presents as a single NUMA domain, creating a simple two-NUMA-node topology in a dual-socket system to improve predictability.
- I/O: A dual-socket configuration provides 176 lanes of PCIe 6.4 and supports CXL 3.1 for memory expansion and pooling.
- Security: Integrates Confidential Computing based on Arm CCA/RME with per-VM encryption keys. It supports secure assignment of TDISP-capable devices and uses authenticated C2C encryption to protect data across links as part of the Vera Rubin NVL72 platform.
NVIDIA's Vera CPU is a direct assault on the x86 server CPU market, built on the conviction that agentic AI is not just another workload, but the *defining* workload for the next generation of data centers. The design philosophy is a clear rebuke of the 'more cores, more chiplets' trend, arguing instead for a monolithic, software-first design that prioritizes per-thread performance, predictable latency, and a simple NUMA topology. This is NVIDIA betting that the future of computing is less about handling millions of uniform web requests and more about orchestrating thousands of complex, stateful AI agents. The key question is whether this agentic profile is distinct and dominant enough to justify a specialized CPU, or if Vera will be a niche product in a world still dominated by general-purpose computing tasks.
Don't read this site daily. Get it in your inbox.
The daily brief and Sunday deep dive — distilled, scored, and opinionated. For builders only.