GPT-6 Astra is the first model making OpenAI willing to declare the “AGI era”
OpenAI has shipped GPT-6 Astra, its most capable model to date. President Greg Brockman says it might already qualify as “AGI” or is at least within reach, meaning an AI system that outperforms humans at most economically valuable work by OpenAI’s own definition.
GPT-6 Astra is rolling out first to select organizations through OpenAI’s Daybreak program. Over the coming days, it will become available to ChatGPT Plus, Pro, Business, and Enterprise customers, as well as through the API and cloud platforms like AWS Bedrock and Microsoft Azure. Astra was pretrained on more than 100,000 GPUs at the Stargate facility in Texas. OpenAI researcher Aidan Clark called it the company’s largest training run ever. The jump from Sol to Astra represents a bigger capability gain than the jump to Sol from earlier models, Clark said, in part because previous AI models played a role in monitoring training.Ad In benchmarks OpenAI published, Astra scores well above its predecessor GPT-5.6 Sol and Anthropic’s Fable models. Astra hits top marks across a range of disciplines: logical reasoning (98.6 percent on ARC-AGI-3, though under its own test conditions), math (97.6 percent on FrontierMath Tier 4 v2), software engineering (74.1 percent on DeepSWE v1.1), expert knowledge (96 percent on GPQA Diamond), engineering (95.9 percent on BenchCAD), and cybersecurity (100 percent on ExploitBench).Ad BenchCAD cost: ~43% below Sol, ~86% below Fable 5.1. Terminal-Bench 4.0 cost: ~9% below Sol, ~63% below Fable 5.1.Ad Lower-cost settings: Terminal-Bench Science 61.1% at ~27% lower cost; GPQA Diamond 94.9% at ~37% lower cost. Prime gaps improved from 240 to 186, and a large-gap bound term improved for the first time in over 80 years. Fable 5 and 5.1 are not included in LifeSciBench, GeneBench Pro, and MedChemBench because they reject most questions.Ad SRE-Bench within four attempts: 99.2% versus 68.7% for Sol. Astra found two previously unknown zero-days during evaluation.Ad Alignment (lower is better except for Impossible ExploitGym) Impossible-task scope test: Sol exceeded its authorized target 48% of the time, Astra 0%. Astra is 3x less likely to misstate its own capabilities. In scientific work, the model reportedly improved a mathematical result on prime gaps and set new records in biology, chemistry, medicine, and physics evaluations. OpenAI also positions Astra as a model that can reliably operate a computer the way a human would, and on OSWorld 2.0, which measures that ability, Astra scored 72.6 percent at about 40 minutes per task compared to Sol’s 65.7 percent at roughly 75 minutes.