Garp Independent AI & technology journalism
Saturday, September 26, 2026 Sign In · Join Subscribe
Latest Ando wants to take on Slack with a team messaging app that lets humans and agents work together

AI news, research, models, robotics, chips, startups, and infrastructure coverage.

Updated daily

Home  /  Chips & Semiconductors  /  AMD MI400 GPU at Hot Chips 2026

Chips

AMD MI400 GPU at Hot Chips 2026

AMD MI400 GPU at Hot Chips 2026

AMD is on stage at Hot Chips 2026 detailing the AMD Instinct MI400 series GPU architecture, the compute that will drive the Helios rack-scale systems the company has been showing around the industry. AMD’s talk covers the MI455X silicon, the cache, memory, and compute upgrades over the MI355X, and the ROCm software stack AMD is pushing to stay competitive with CUDA.

If you have not seen it yet, Ryan did an awesome in-depth MI400 deep-dive that goes into tons of detail. This one is running live, so please excuse any typos while AMD presents. AMD opens with the argument that AI work is broadening from single-model training into a mix of frontier training, enterprise fine-tuning, and always-on inference. This scale curve runs from the roughly 65 million parameter Transformer in 2017, through GPT-4 class models around a trillion parameters in 2023, and on to the 10 trillion-plus agentic and reasoning models AMD expects this year, with the takeaway that infrastructure has to move data as fast as models scale. AMD frames the MI400 family around Helios, the rack-scale AI infrastructure it expects to ship. Headline numbers are a 2.9 exaflop rack with 31 TB of HBM4 memory and 1.7 PB/s of HBM4 bandwidth across 72 GPUs, along with 260 TB/s of scale-up and 43 TB/s of scale-out bandwidth per rack. We have covered AMD’s double-wide Helios racks before, and this fills in the silicon behind them. This basic building block is a compute tray that combines compute, host CPU, memory, and networking. Each tray holds four AMD Instinct MI455X EAMs fed by a single-socket AMD EPYC 9006 SP7 server CPU over Infinity Fabric, with UALoE links carrying scale-up traffic at 1.8 TB/s per direction per GPU and up to three AMD Pensando Vulcano 800 AI NICs per EAM handling scale-out. Here you can see the Helios node from Advancing AI 2026: At the center of the tray is the AMD Instinct MI455X, an enhanced modular chiplet design built from eight accelerator complex dies on N2 flanked by fabric and cache dies plus I/O dies on N3P. It packages 256 total active work group processors with 192 MB of global L2 and 12x HBM4 stacks running 432GB at 23.3 TB/s, and it connects through PCIe Gen 6 as well as 72 UALoE lanes pushing 3.6 TB/s. We have a full AMD Instinct MI455X and CDNA 5 deep dive for more on the GPU. This packaging split is notable because each die type moves to the node that best fits its job. Compute dies on N2 sit under 3D hybrid-bonded XCDs for higher density per watt, while the N3P fabric, cache, and I/O dies, plus CoWoS-L packaging, tie the whole package around the twelve HBM4 stacks. AMD boils the MI400 changes into three buckets: bigger memory and cache, faster compute, and less data movement.