Claude Fable 5: The first Mythos model is powerful, expensive, and heavily filtered
Anthropic has released Claude Fable 5, the first model in its new Mythos class. It leads nearly every benchmark, including SWE-bench Verified at 95 percent, but costs twice as much as Opus 4.8 at 10 or 50 dollars per million tokens.
With Claude Fable 5, Anthropic has shipped a model that tops nearly every benchmark. Fable 5 is the first publicly available version of the “Mythos class.” According to Anthropic, Fable shares its base model with Claude Mythos 5 but adds strict guardrails that block potentially harmful requests related to cybersecurity, biology, chemistry, and model distillation. Mythos 5 is also available but limited to a small group of users. What “Mythos” actually means on a technical level is mostly guesswork. Every CEO Dan Shipper, whose team had early access, reports that Anthropic staff told him there’s nothing special about the architecture. Within the Haiku, Sonnet, and Opus family, Mythos simply refers to the largest and most capable model. Developer Simon Willison suspects the same, that it’s the biggest Anthropic model publicly available to date. Fable just feels “big,” Willison writes, “not just in terms of speed and cost, but also in how much it knows.” Artificial Analysis backs this up: on its AA-Omniscience knowledge and hallucination benchmark, Fable scores 40 points, seven more than the previous leader, Gemini 3.1 Pro. Among open-weight models, that kind of gap typically tracks with model size. Fable 5 sits atop nearly every leaderboard. On the Artificial Analysis Intelligence Index, it hits 64.9 points, roughly five ahead of GPT-5.5 as the closest competitor. On GDPval-AA, an agentic benchmark for real-world work tasks, it posts an Elo score of 1,932. On Humanity’s Last Exam, Fable reaches 53 percent, more than seven points above Opus 4.8. A single run of that test cost about $2,200, including fallback costs. The evaluation service Vals ranks Fable 5 first on its overall index and across all coding benchmarks, including SWE-bench Verified at 95 percent and Vibe Code Bench at 90.35 percent.