Garp Independent AI & technology journalism
Saturday, September 26, 2026 Sign In · Join Subscribe
Latest Ando wants to take on Slack with a team messaging app that lets humans and agents work together

AI news, research, models, robotics, chips, startups, and infrastructure coverage.

Updated daily

Home  /  Chips & Semiconductors  /  Google’s TPUv8s for Training and Inference at Hot Chips 2026

Chips

Google’s TPUv8s for Training and Inference at Hot Chips 2026

Google’s TPUv8s for Training and Inference at Hot Chips 2026

Google’s rate of development for their tensor processor units (TPUs) has been nothing short of remarkable. The company is already preparing their eighth generation of TPUs, which were announced a bit earlier this year.

The TPUv8 family includes both the TPU 8t for training, as well as the TPU 8i for inference, making Google’s hardware offerings unique among the hyperscalers. Whereas the other operators have been focused solely on developing inference hardware while relying on commodity hardware for training, Google has continued to develop their own hardware for both tasks right up to the current day. This experience, Google has argued in the past, is part of what has given them an edge in this market. And the TPUv8 family is meant to extend that. Besides providing their own internal chip design for inference, the TPU 8i is also notable for being a major consumer of Google’s in-house developed CPU, the Axion. TPU 8i chips are paired with Axion CPUs in a 2-to-1 ratio inside Google’s nodes. Previously, the company was using x86 CPUs for this task, so the switch to Axion has both moved Google to their own hardware and has moved them over to the Arm architecture in the process. Briefly, here is a quick look at the history of Google’s TPU development, which is now on their eighth generation. The first TPU was a PCIe card optimized for inference. The next TPU was a design optimized for training with shared memory. The company has (roughly) alternated between inference-optimized chips and training-optimized chips. This has fed into each other, as roughly one-third of the forward pass on training is also a type of inference. Over the years, Google has been able to increase performance by a million-fold. At this point Google is also designing these chips to build for the world. Rather than just building for their internal workloads, there is a bit more consideration for cloud customer workloads (though Google’s internal workloads are still very important to them).