Claude Sonnet 5 continues Anthropic’s pattern of hiding price increases behind unchanged token rates
Anthropic — claude Sonnet 5 ranks fifth in the Artificial Analysis Intelligence Index with 53 points and even beats the pricier Opus 4.8 on some agent-based tasks. But the model chews through about 40 percent more tokens per task than its predecessor, nearly doubling real costs despite identical list prices.
Artificial Analysis evaluated Claude Sonnet 5 before its release and added it to its Intelligence Index. Sonnet 5 scored 53 points at peak performance, tying with GPT-5.5 (high) for fifth place. Four models rank higher: GPT-5.5 (xhigh) at 55, Opus 4.7 at 54, Opus 4.8 at 56, and Claude Fable 5, once again generally available as of today, at 60 points. That’s a six-point jump over Sonnet 4.6 (47 points), but Sonnet 5 chews through far more tokens to get there. On paper, Sonnet 5 keeps the same token prices as its predecessor: $3 per million input tokens and $15 per million output tokens, while Opus 4.8 sits at $5 and $25. Yet according to Artificial Analysis, an average task in the Intelligence Index costs $2.29 with Sonnet 5, versus about $1.97 with Opus 4.8. At the maximum performance setting (“max”), Sonnet 5 burns through about 40 percent more output tokens per task than Sonnet 4.6. In agent-based knowledge work benchmarks like AA-Briefcase and GDPval-AA, it runs about three times as many agent loops as its predecessor. Sonnet 4.6 cost about $1.20 per task. That’s nearly doubled, even though Sonnet 5 beats Opus 4.8 on some of these tasks. Anthropic is running a promotional rate of $2 or $10 per million tokens through September 1, but Artificial Analysis based its results on regular prices. Sonnet 5 still falls short of larger models on reasoning- and knowledge-heavy benchmarks. On CritPt, a frontier physics reasoning test from Argonne National Labs and the University of Illinois, it scored 17 percent. That’s 14 points above its predecessor but below GLM-5.2, Claude Opus, Fable, and GPT-5.5 in their higher configurations. Elsewhere, Sonnet 5 shows solid gains over Sonnet 4.6: a 9-point jump on Terminal-Bench v2.1, 10 points on Humanity’s Last Exam, and 7 points on SciCode.