Garp Independent AI & technology journalism
Saturday, September 26, 2026 Sign In · Join Subscribe
Latest Ando wants to take on Slack with a team messaging app that lets humans and agents work together

AI news, research, models, robotics, chips, startups, and infrastructure coverage.

Updated daily

Home  /  Chips & Semiconductors  /  OpenAI Jalapeno Custom AI ASIC at Hot Chips 2026

Chips

OpenAI Jalapeno Custom AI ASIC at Hot Chips 2026

OpenAI Jalapeno Custom AI ASIC at Hot Chips 2026

OpenAI took the Hot Chips 2026 stage on Day 2 to detail Jalapeño, an in-house inference ASIC and system built with Broadcom and designed to be the best compute platform for OpenAI’s own inference workloads. Richard Ho, Ravi Narayanaswami, and Chris Leary walked through the chip’s roughly nine-month path from initial RTL to tapeout, its performance positioning against NVIDIA GB200 and GB300, and an architecture built around HBM4 and a spatial programming model.

We are doing this one live from the session, so please excuse typos. OpenAI Jalapeño is framed as an inference platform rather than a raw accelerator. OpenAI is talking about this in terms of the silicon, together with its host and accelerator rack pair, targeting state-of-the-art performance per watt at low latency for multi-chip workloads, aided by AI-accelerated hardware and software co-design. This project moved quickly once OpenAI concluded that inference and agentic workloads needed a purpose-built design. This timeline shows an architecture concept in late 2024, an RTL freeze in 2025, a late 2025 tapeout, Codex running in early 2026, with ChatGPT on the chip not long after. OpenAI frames the design around two metrics, time to last token for user experience and tokens per joule for inference efficiency. Across those, it compares systems along the full Pareto frontier of request latency versus energy per token rather than chasing raw chip counts, throughput per chip, or time to first token. For comparisons, OpenAI uses InferenceX, a public, power-normalized benchmark across a basket of open-source models that spans the full prefill-to-decode spectrum. Runs normalize to the package TDP, with Jalapeño at 700 watts against the GB200 at 1.2 kilowatts, and the GB300 and MI355X at 1.4 kilowatts, and OpenAI measured against the July 2026 Pareto frontier across other accelerators. Jalapeño uses single-token prediction, whereas the NVIDIA baselines it is compared against use multi-token prediction. This figure shows how MTP uses seven tiny draft-model turns plus one batched large-trunk pass to produce up to eight output tokens, reducing the number of expensive large-model passes by up to 8x. OpenAI ran GPT-OSS to stress latency limits, DeepSeek R1 is a good draft-model case, and the 1-trillion-parameter Kimi K2.5 shows scaling across many devices. OpenAI notes none of the three were co-designed for Jalapeño and that it got all of them running between when A0 silicon returned to the lab and now. Here is the throughput-per-kilowatt frontier for the GPT-OSS 120B model. Jalapeño sits on the Pareto frontier at 700 watts, compared to the GB200 at 1,200 watts, meaning it holds a lead in useful throughput per kilowatt at equal latency.