Garp Independent AI & technology journalism
Saturday, September 26, 2026 Sign In · Join Subscribe
Latest Ando wants to take on Slack with a team messaging app that lets humans and agents work together

AI news, research, models, robotics, chips, startups, and infrastructure coverage.

Updated daily

Home  /  Chips & Semiconductors  /  Micron Evolving Memory Architectures for AI at Hot Chips 2026

Chips

Micron Evolving Memory Architectures for AI at Hot Chips 2026

Micron Evolving Memory Architectures for AI at Hot Chips 2026

ServeTheHome — kicking off Hot Chips 2026 Day 0, which has become a big day in itself, Micron is talking about evolving memory architectures for AI. From what it can tell, Micron will be walking through how high-bandwidth memory and advanced packaging have moved to the center of AI system design.

Ok let us get started. These are being done live, so please excuse typos. Micron Evolving Memory Architectures for AI at Hot Chips 2026 Micron’s talk opened with the OpenAI compute-efficient frontier to frame why memory has become central to language model scaling. Micron points to the power-law relationship between model size, dataset size, and training compute, arguing that all three must scale together for LLM performance to improve. Memory is the resource that ties those three factors together. Now we are moving onto the memory wall. I wonder how many of these we are going to see today or even this week at Hot Chips! Micron is saying that AI accelerator TFLOPS roughly grow by 3x every two years while the bandwidth of 2.5D attached memory such as HBM climbs at under 2x per two years. That widening gap is why memory technology has become a key factor limiting AI system performance. Micron frames HBM as the compute-memory bridge, tracking rising cube capacity, bandwidth, and energy efficiency across generations from HBM2e through HBM5. You have to enjoy the technical conference with unlabeled Y axis. Here, Micron is showing a typical GPU at over 12,000 square millimeters when counting the package with base die and eight HBM4 instances, meaning memory silicon can exceed 8x the surface of the GPU die itself when using 12-high HBM4. This is actually a neat view to show HBM raises the memory bandwidth ceiling, shifting that boundary and allowing memory-intensive AI workloads to run faster before they hit a memory limit. Micron’s talk is now walking through what HBM is and how each generation has advanced. A product table traces HBM1 through HBM4.