Garp Independent AI & technology journalism
Saturday, September 26, 2026 Sign In · Join Subscribe
Latest Ando wants to take on Slack with a team messaging app that lets humans and agents work together

AI news, research, models, robotics, chips, startups, and infrastructure coverage.

Updated daily

Home  /  Chips & Semiconductors  /  NVIDIA Vera CPU at Hot Chips 2026

Chips

NVIDIA Vera CPU at Hot Chips 2026

NVIDIA Vera CPU at Hot Chips 2026

For the third and final presentation in this morning’s first CPU session at Hot Chips 2026, NVIDIA is taking to the lectern to present on Vera, its next-generation server CPU. Based on their new Olympus CPU core architecture, Vera is an in-house Arm server CPU with 88 cores that serves as a critical part of NVIDIA’s upcoming Vera Rubin AI systems.

It is also their most ambitious effort yet to expand their overall presence in the server CPU market, as NVIDIA is looking to capture a larger piece of the pie with their specialized-but-powerful CPU. Today’s presentation follows the company’s release of their Vera/Olympus whitepaper last month. Please note, we are covering this live, so please excuse typos. At a high level, Vera is designed to be a very modern and monstrous CPU in virtually every regard. The large chip focuses heavily on IPC over core count, leaving it with 88 CPU cores altogether, and flanked by eight 128-bit LPDDR5X memory controllers. The NVLink-C2C interface on the chip is designed to pair well with NVIDIA’s own Rubin GPUs, but it can also plug in to anything else that implements NVLink-C2C, or to another Vera chip for a 2P pure CPU setup. The successor to NVIDIA’s Grace CPUs, which have backed their server systems for both the Grace Hopper (GH) and Grace Blackwell (GB) generations, Vera is designed to be a significant advancement from Grace in every possible way. NVIDIA is starting things off by framing this talk about how AI is the most complex computational workload yet. And it is not one task or query, but it is a full workflow that touches multiple tools and types of computation. You cannot just optimize for one use case and win. Vera Rubin has been built with extreme co-design across hardware and even software. This spans the NVL72, LPX3, Bluefield racks, and others. NVIDIA has shown off this curve (and similar) a lot over the past year: their performance frontier, the trade-off between token throughput and user interactivity. High batching can maximize token generation, but it makes the response time quite poor. Vera Rubin is meant to push this frontier out in every direction, moving the sweet spot higher in both throughput an interactivity.