Garp Independent AI & technology journalism
Saturday, September 26, 2026 Sign In · Join Subscribe
Latest Ando wants to take on Slack with a team messaging app that lets humans and agents work together

AI news, research, models, robotics, chips, startups, and infrastructure coverage.

Updated daily

Home  /  Chips & Semiconductors  /  OpenAI’s first custom chip “Jalapeño” reportedly beats Nvidia’s Blackwell and Rubin in inference benchmarks

Chips

OpenAI’s first custom chip “Jalapeño” reportedly beats Nvidia’s Blackwell and Rubin in inference benchmarks

OpenAI’s first custom chip “Jalapeño” reportedly beats Nvidia’s Blackwell and Rubin in…

OpenAI showed off the first benchmarks for its in-house inference chip at the Hot Chips conference. “Jalapeño” reportedly outperforms both Nvidia’s Blackwell and Rubin in throughput per watt and token latency.

The chip handles inference only, meaning it runs AI models but doesn’t train them. Jalapeño isn’t tuned to OpenAI’s own models either. It’s a general-purpose LLM inference accelerator. OpenAI claims Jalapeño delivers 1.5x to 1.9x more AI work per watt at peak throughput across all three tested models, with 1.7x to 3.6x lower end-to-end latency than the best commercially available systems. For interactive workloads, the company says performance is 2.1x to 4.1x higher.Ad The results come from tests using SemiAnalysis’s public InferenceX benchmark. OpenAI provided the numbers. SemiAnalysis verified some runs on-site in the lab. The models tested were GPT-OSS 120B, Deepseek R1 670B, and Kimi K2.5 1T. On GPT-OSS, Jalapeño hit about 1,400 tokens per second per user. On Deepseek R1, it topped 700 tokens per second on a single concurrent request.Ad Jalapeño posted these numbers without using techniques like multi-token prediction or speculative decoding, while some of the comparison systems did rely on those optimizations, so there’s still room to improve. In its headline performance-per-watt comparison, “Jalapeño smokes every other chip,” SemiAnalysis writes. SemiAnalysis CEO Dylan Patel added, “Usually first generation chips aren’t competitive, but OpenAI is beating Nvidia Blackwell and even Rubin.”Ad SemiAnalysis points out that the fairer comparison isn’t Blackwell but Nvidia’s newer Vera Rubin platform, since both use HBM4 memory. Even here, Jalapeño squeezes out more output tokens per megawatt than Vera Rubin, even though Nvidia’s accelerator uses the multi-token prediction optimization that Jalapeño hasn’t adopted yet. On total cost of ownership per token, the two come out roughly even. There are caveats, though.