Nvidia Vera CPU Specs: Olympus Cores, SPEC Benchmarks
Nvidia detailed its Vera CPU — 88 custom Olympus cores, 1.2 TB/s memory, and SPEC CPU 2026 scores that edge AMD's Epyc dual-socket flagship.
Nvidia is best known for GPUs, but the chip it detailed this week is a CPU — and the pitch is that the processor, not the accelerator, is becoming the bottleneck for agentic AI. On July 21, 2026, Nvidia published a technical blog post and an accompanying white paper laying out the architecture of Vera, its first fully custom server CPU core, along with a set of SPEC CPU 2026 benchmark results the company says put Vera ahead of AMD’s flagship data-center part.
Vera is the CPU half of the Vera Rubin platform, the successor to the Grace Blackwell generation. Where the previous Grace CPU leaned on standard Arm Neoverse cores, Vera is built around Olympus, a core Nvidia designed in-house — its first ground-up CPU core in years. The company’s argument is blunt: as AI workloads shift from single model calls to long-running agents that orchestrate tools, browse, and manage other models, the serial, latency-sensitive work lands on the CPU, and off-the-shelf cores were not tuned for it.
What’s inside Vera
According to Nvidia’s disclosures, each Vera CPU packs 88 Olympus cores and 176 threads, using simultaneous multithreading to run two threads per core. That is a deliberately modest core count next to the density race AMD and others have been running — AMD’s top Epyc parts push well past a hundred cores per socket. Nvidia’s stated priority is the opposite of density: maximum single-thread performance and memory bandwidth, the two things that gate how fast one agent can think through a chain of steps.
The Olympus core is built wide. Nvidia describes a front end fed by a neural branch predictor, a 64 KB, four-way L1 instruction cache, and a 48-instruction decode queue. The core can fetch up to 16 instructions per cycle and feeds a 10-wide decoder capable of processing ten fused instructions per cycle — an aggressive front end aimed at keeping the execution units busy on the branchy, unpredictable code that agent orchestration tends to produce.
Behind the cores sits a 164 MB unified L3 cache and an on-die coherency fabric that Nvidia rates at up to 3.4 TB/s of internal bandwidth. Main memory is LPDDR5X delivered through SOCAMM2 modules — a socketed, serviceable form factor rather than soldered-down memory — for up to 1.2 TB/s of aggregate memory bandwidth. The whole thing is a monolithic die, not a chiplet package, which simplifies latency and coherency at the cost of some manufacturing flexibility.
Vera is built on the Arm instruction set and, like the rest of the Vera Rubin generation, is fabricated on TSMC’s 3 nm process. For readers keeping score on the broader architecture debate, the Arm-versus-x86 contest is exactly the terrain Vera is fighting on: an Arm-based custom core taking direct aim at x86 incumbents in the one market — the data center — where x86 has been most entrenched. The choice of a leading-edge process node is table stakes for that fight.
The benchmark claim
The headline number Nvidia released is a SPEC CPU 2026 comparison against AMD’s Epyc 9755, with both chips running in a dual-socket configuration. In the SPECrate integer suite — a throughput test that runs many copies of integer workloads in parallel — Nvidia reported an overall base score of 925 for Vera versus 898 for the dual-socket Epyc 9755 system, a roughly 3% margin.
Two caveats belong on that figure, and Nvidia’s own framing includes the first. These are unofficial results the company published itself, not independently audited submissions to SPEC’s public database, and vendor-run benchmarks are chosen to flatter the vendor. The second caveat is the more interesting one for Nvidia’s story: it says Vera reached that score with a smaller thread count than the Epyc system it beat. If that holds up under independent testing, it means Vera is extracting more throughput per thread — consistent with the single-thread-first design philosophy — rather than simply brute-forcing the result with more parallelism.
SPECrate is a throughput benchmark, so an edge there is notable precisely because it is not where a low-core-count chip is supposed to win. The single-threaded SPECspeed numbers, where Olympus’s wide core should shine most, will be the ones to watch when third parties get hardware in hand.
Why the CPU suddenly matters
For most of the AI boom, the CPU in an accelerated server was plumbing: it fed data to the GPUs, handled the operating system, and otherwise stayed out of the way. The division of labor between CPUs, GPUs, and TPUs treated the GPU as the star and the CPU as support.
Agentic workloads scramble that hierarchy. An AI agent that plans a task, calls tools, waits on results, parses them, decides what to do next, and dispatches sub-agents spends a lot of its wall-clock time on serial, latency-bound logic — exactly the work a GPU is bad at and a fast CPU core is good at. Nvidia’s framing is that in a rack full of Rubin GPUs, the Vera CPUs orchestrating them can become the pacing item: if the CPU stalls, expensive accelerators sit idle. A large L3 cache and high memory bandwidth are aimed squarely at keeping that orchestration loop fed.
That is why Nvidia is talking up single-thread performance rather than core count. The company is not trying to win a general-purpose server-CPU spec sheet; it is trying to win the specific job of driving its own accelerators as fast as possible. Vera connects to the Rubin GPUs over Nvidia’s high-bandwidth chip-to-chip interconnect, forming the Vera Rubin superchip at the heart of the platform. The CPU and GPU are designed to be sold, and optimized, together.
The competitive picture
Vera lands in a data-center CPU market that has spent a decade as an AMD-versus-Intel duel, with Arm-based challengers — Amazon’s Graviton, Ampere, and Nvidia’s own Grace — nibbling at the edges. By building a fully custom core and benchmarking it head-to-head against Epyc, Nvidia is signaling it intends to be a first-class CPU vendor, not just a GPU company that ships a companion chip.
The timing is pointed. Nvidia disclosed the Vera details as rival AMD was preparing its own AI-focused messaging, and the move reads as a preemptive strike — get the architecture and the benchmark narrative out first. It also arrives while the broader chip trade is whipsawing: semiconductor stocks rallied hard this week after a bruising early-July selloff, with investors trying to price how much of the AI data-center buildout is durable demand versus froth. A credible Nvidia CPU widens the company’s share of every server it ships into — more silicon content per rack, sold by one vendor.
Vera Rubin has been in production at TSMC and is on track for partner availability in the second half of 2026, with the platform having entered full production earlier this year. The white paper is the technical case Nvidia wants system builders and hyperscalers weighing over the next two quarters, and it is the same Vera Rubin hardware now anchoring Microsoft’s expanded European AI infrastructure deal with Mistral.
What it means
Nvidia is telling the market that owning the CPU is now strategic, not optional — and the Vera disclosure is the argument, laid out in benchmarks, for why.
Who wins. Nvidia, if the story holds. Every Vera CPU that ships alongside a Rubin GPU is silicon content Nvidia captures instead of AMD or Intel, and a tighter CPU-GPU pairing is a genuine performance lever for the agentic workloads Nvidia is betting the next cycle on. Customers building large-scale AI agents-style orchestration get a CPU actually tuned for the job rather than a repurposed general-purpose part.
Who should be uneasy. AMD and Intel. The data-center CPU has been their fortress even as Nvidia dominated accelerators; a custom Nvidia core that benchmarks competitively against Epyc — on a throughput test, no less — is a direct incursion. Arm-based rivals like Ampere and the hyperscalers’ in-house cores now have a much larger, better-funded competitor validating the same thesis they’ve been selling.
What to watch next. First, independent SPEC results — vendor numbers are a starting gun, not a finish line, and the single-thread SPECspeed figures matter more than the throughput headline for Vera’s actual pitch. Second, who buys: whether hyperscalers standardize on Vera Rubin as a unit or keep mixing Nvidia GPUs with third-party CPUs. Third, real agentic benchmarks — the case for Vera rests on a claim about how AI agents spend their time, and the proof will be end-to-end agent throughput on production workloads, not synthetic integer suites. If Nvidia is right that the CPU is the new bottleneck, Vera is how it plans to own that bottleneck too.
Keep reading
Chisato · · 4 min read Nvidia–SK Group $500B Deal: HBM Supply, AI Factory
Nvidia and SK Group unveiled a $500B+ AI partnership locking in SK Hynix HBM4 supply and a 2GW AI factory in Korea. Here are the details and what to watch.
Kurumi · · 6 min read Nvidia's $500B AI Compute Financing: What to Know
Nvidia lined up $500B from BlackRock, Blackstone, Apollo, KKR, Brookfield and Goldman to finance AI compute — and to make chips an asset class.
Kurumi · · 5 min read OLIX Raises $312M for Photonic AI Inference Chips
UK startup OLIX raised $312M at a $3.3B valuation for its optical AI inference chips, backed by Arm and Reed Hastings. What the photonic bet means.