Garp Independent AI & technology journalism
Saturday, September 26, 2026 Sign In · Join Subscribe
Latest Ando wants to take on Slack with a team messaging app that lets humans and agents work together

AI news, research, models, robotics, chips, startups, and infrastructure coverage.

Updated daily

Home  /  Chips & Semiconductors  /  NVIDIA’s Groq 3 LPU Accelerators for Heterogeneous AI Compute at Hot Chips 2026

Chips

NVIDIA’s Groq 3 LPU Accelerators for Heterogeneous AI Compute at Hot Chips 2026

NVIDIA’s Groq 3 LPU Accelerators for Heterogeneous AI Compute at Hot Chips…

The second AI presentation of the afternoon comes from NVIDIA, who besides doing talks on the Vera CPU and Rubin GPU, are also presenting a talk on the use of language processing units (LPUs) in their Vera Rubin racks. This talk is arguably a bit of an unusual one, because NVIDIA does not currently produce their own LPUs (though they are under development).

The LPUs being used in Vera Rubin generation racks – and specifically the dedicated LPX racks – come from Groq, whom NVIDIA is buying the chips from as part of a broader acquihire of the company. None the less, with the bulk of Groq’s talent now at NVIDIA, it falls to NVIDIA to promote the chips. Groq’s LPUs are designed to fill a weak spot in NVIDIA’s Vera Rubin stack. GPUs are great for pre-fill, and depending on how the design of the chip is optimized, so-so at decode. Whereas LPUs – essentially purpose-built chips with large amounts of on-die SRAM to keep model data as close to the compute hardware as possible. Fittingly, the biggest rationale for the use of LPUs in a Vera Rubin server cluster is to boost performance at low latencies, taking advantage of that SRAM to get results back ASAP so that AI models can move on to the next token. This article is being written live from the presentation, so please excuse any typos. NVIDIA’s Groq 3 LPU Accelerators for Heterogeneous AI Compute at Hot Chips 2026 Yesterday NVIDIA released the first third-party benchmarks of an LPX rack, so the timing of that and this talk is not coincidental. Nor is the fact that NVIDIA has been promoting the use of LPUs in both their Vera and Rubin talks. Today’s talk is reiterating parts of that for the Hot Chips crowd, as well as laying out the case for using LPUs with Vera Rubin and how they fit in to the broader ecosystem. Recapping a common refrain through all of NVIDIA’s talks at Hot Chips, the company considers agenetic AI to be the most complex computing workload in history. And to that end, it has required new approaches to hardware development – as well as a whole lot more hardware. LPUs are the building block of the LPX rack, which is one of several racks that make up the larger Vera Rubin hardware ecosystem. And once again showing off NVIDIA’s performance curves/frontiers for the Vera Rubin ecosystem, plotting tokens/sec/watt versus tokens/sec/user. The Rubin GPU hardware has its own performance curve, but overall performance can drop pretty hard if you try to push higher user interactivity (i.e.