Anthropic’s Claude Sonnet 5.5 nearly matches Opus 5.5 on benchmarks while costing up to 30 percent less per task
Anthropic has released Claude Sonnet 5.5, the second model in its Claude 5.5 family. It generates output more than 30 percent faster, costs up to 30 percent less per task, and nearly matches Opus 5.5 on several benchmarks.
Opus 5.5 is built for complex tasks that demand careful judgment. Sonnet 5.5 targets well-defined everyday work like fixing bugs, writing docs, building presentations, and creating spreadsheets. Anthropic also announced Claude Haiku 5.5 for the coming weeks. That model will focus on high-throughput, low-cost use cases. With Fable, Opus, and Sonnet, Anthropic already has counterparts to OpenAI’s GPT-6 Astra, Sol, and Luna, which kicked off the latest pricing battle. Roughly speaking, Opus sits a bit above Sol, Sonnet above Luna, and Fable above Astra, though Anthropic charges more across the board. Performance differences between matched tiers are small enough that cost may end up being the deciding factor. Haiku could help Anthropic close that gap.Ad The performance gap between Sonnet 5.5 and its predecessor is most significant in coding, according to Anthropic. On Terminal-Bench 4.0, a test for agentic coding, Sonnet 5.5 hits 70.6 percent compared to Sonnet 5’s 10.3 percent. On CursorBench 4.0, which recreates real coding sessions from the Cursor editor, Sonnet 5.5 scores 55.5 percent versus 34.1 percent, landing just two points below Opus 5.5 (57.8 percent).Ad On FrontierCode 1.1 at the “High” setting, Sonnet 5.5 scores ten points above Sonnet 5 at roughly one-fifteenth the cost per task, Anthropic says. Early testers praised how quickly the model grasps a codebase. Sonnet 5.5 also batches tool calls more often than its predecessor, which cuts the number of steps needed. Reasoning intensity has one odd wrinkle. At the highest effort level, “Max,” Sonnet 5.5 actually scores worse on FrontierCode than at “Xhigh.” Anthropic says that at maximum effort, the model more frequently triggers a code-review function that splits work across multiple sub-agents. In some cases, that led to timeouts or changes outside the task scope, both of which FrontierCode penalizes.Ad On GDPval-AA, an OpenAI-developed knowledge-work benchmark covering tasks from 44 professions and nine industries, Sonnet 5.5 scores 1,844 points.