Garp Independent AI & technology journalism
Sunday, September 27, 2026 Sign In · Join Subscribe
Latest Don’t be fooled by this summer of AI hype 

AI news, research, models, robotics, chips, startups, and infrastructure coverage.

Updated daily

Home  /  AI News  /  GPT-5.6 Sol goes 14x faster as OpenAI launches Ultrafast mode powered by Cerebras

AI News

GPT-5.6 Sol goes 14x faster as OpenAI launches Ultrafast mode powered by Cerebras

GPT-5.6 Sol goes 14x faster as OpenAI launches Ultrafast mode powered by…

OpenAI is launching a preview of its new “Ultrafast” mode, which delivers up to 750 output tokens per second from its flagship model, GPT-5.6 Sol. OpenAI says Ultrafast is designed to combine the speed of smaller models with the full capabilities of a large reasoning model, enabling what the company calls “more useful work per second.” During incident response, for example, engineers could have logs, code changes, and reports analyzed while an outage is still happening, helping them pinpoint the cause and prepare a fix in real time.

OpenAI says it’s already using the model internally for this purpose.Ad OpenAI pitches several other scenarios. In finance, the model could evaluate market signals and flag suspicious transactions while conditions are still shifting. In customer support, complex inquiries could be resolved in real time, even when finding the answer requires multiple steps or systems. In e-commerce, it could answer product questions, check inventory levels, and personalize recommendations before a hesitant buyer abandons their cart.Ad Video: Generating a 3D warehouse simulator, shown on the left in Ultrafast mode Research is another area OpenAI highlights. Experiments that previously ran overnight as batch jobs could turn into interactive work sessions, letting teams test an idea, review results, adjust their approach, and kick off another run without breaking their workflow.Ad OpenAI already monetizes inference speed in tiers. Through the API, it offers a “Fast Mode” that promises up to 2.5x speed with lower latency for GPT-5.6 Sol at roughly double the price. Ultrafast adds a third, faster, and likely pricier tier. The logic mirrors cloud providers like AWS, which have long charged more for the same service at higher performance levels. OpenAI is applying that to AI inference. If speed becomes a bottleneck across industries, this tiered model gives OpenAI a direct cut of the revenue gains that faster inference creates.Ad Follow The Decoder for AI news, background stories and expert analyses.