Anthropic’s Claude Fable 5 dominates new industry benchmarks at a steep premium
Anthropic’s Claude Fable 5 tops all six new industry-specific performance indices from Artificial Analysis, covering finance, law, and medicine. But that lead comes at a steep cost.
How well do current AI models perform on industry-specific tasks in fields like finance, law, or medicine? Artificial Analysis has introduced six new Capability Indices that compare AI models across Finance & Accounting, Legal, Healthcare & Medical, Strategy & Ops, Engineering, and Economics. The new indices expand on the platform’s existing Agentic and Coding indices. According to Artificial Analysis, the methodology is based on occupational classifications from the US O*NET system. Domain-specific skills are derived from the job tasks defined there, covering things like financial modeling, legal research and contract analysis, or clinical decision support. For each domain, the benchmark suite is assembled fresh and weighted by how often a given skill shows up in that industry. Artificial Analysis says all benchmarks are run independently.Ad Anthropic’s Claude Fable 5 (with Opus 4.8 fallback) takes first place across all eight indices. Claude Opus 4.8 (max) comes in second in six of eight categories, according to Artificial Analysis, while OpenAI’s GPT-5.5 (xhigh) grabs second in the remaining two. GPT-5.6, set to launch tomorrow, could close the gap between Anthropic’s current models and GPT-5.5.AdDEC_D_Incontent-1 Below the top tier, rankings shift considerably by domain. Google’s Gemini 3.5 Flash, Gemini 3.1 Pro Preview, OpenAI’s GPT-5.5 (xhigh), Anthropic’s Claude Sonnet 5 (max), and the Chinese GLM-5.2 (max) trade places depending on the task. Among open-weights models, GLM-5.2 (max) leads in five of the six industry indices. With 53 points, GLM-5.2 (max) reaches fifth place overall in the Engineering benchmark, just two points behind Claude Sonnet 5 (max) and GPT-5.5 (xhigh), which both score 55. In the Strategy & Ops Index, Deepseek V4 Pro (max) takes the open-weights lead with 38 points.Ad The Artificial Analysis results track with current data from Arena, an independent platform where millions of real users rank models through blind comparisons. As of the July 7 leaderboard snapshot, Claude Fable 5 holds first place in the Text Arena, Code Arena, and Agent Arena. Anthropic is the only lab leading all three main categories. In the Agent Arena, Fable 5 scores 16.58 percent above the model average, well ahead of OpenAI’s GPT-5.5 xHigh at 8.66 percent and Z.ai’s GLM 5.2 at 6.62 percent.