Garp Independent AI & technology journalism
Sunday, September 27, 2026 Sign In · Join Subscribe
Latest Don’t be fooled by this summer of AI hype 

AI news, research, models, robotics, chips, startups, and infrastructure coverage.

Updated daily

Home  /  AI News  /  GPT-6 Astra gives mathematicians a breather, and OpenAI says that’s by design

AI News

GPT-6 Astra gives mathematicians a breather, and OpenAI says that’s by design

GPT-6 Astra gives mathematicians a breather, and OpenAI says that’s by design

OpenAI’s GPT-6 Astra tops the ErdosBench for open math problems, even though chief scientist Jakub Pachocki says the company didn’t prioritize math. That says something about where “AGI” actually stands.

Mathematicians can take a brief (likely very brief) breather after Astra’s release. The field has been grappling with existential questions lately, and while the new model does push math research forward, it’s not moving at the pace some had hoped for and others had feared. On ulam.ai’s ErdosBench, Astra took first place. The benchmark covers 226 open math problems inspired by the famous Erdős problems. Astra scored 3.23 and solved 106 problems, 43 of them completely. It disproved 27 more. It also disproved 27 others. Compared to Sol, which solved 78 problems at maximum reasoning, Astra showed stronger scientific writing and was less prone to overblown claims. In some cases, it actually understated its own results. Overall, benchmark developer Przemek Chojecki called it “a solid 5%-10% gain on various math-research skills tested,” but the benchmark is far from saturated. OpenAI could have made Astra much stronger in math research but decided against it, even though the company had put math wins front and center in its first announcement. In his essay “An Alien Mind,” OpenAI chief scientist Jakub Pachocki writes, “[…] we believe we could make the models better at specifically mathematics research with additional focus, but we do not prioritize this direction because of the urgency we feel about RSI and automated alignment research, as I will discuss later.” That means the most capable math model right now isn’t the result of targeted optimization. It’s a byproduct of other priorities. OpenAI is pouring its resources into recursive self-improvement and securing future AI systems, since “we believe it is the only way to remain at the frontier of AI research moving forward,” Pachocki writes. (Note: Since the essay was published, OpenAI has reportedly trained better internal math models, a complicated story, and the topics of RSI and AI safety have exploded.) Pachocki’s statement is interesting for another reason, too: it shows that hard optimization trade-offs are being made at the jagged frontier of AI development.