|

Claude Fable 5.1 Explained: Anthropic’s New Frontier Model

On September 1, 2026, Anthropic released Claude Fable 5.1 alongside a second model, Claude Mythos 5.1, positioning both as the top of the Claude lineup for the hardest reasoning and the longest, most autonomous agentic work the company builds for. Two days later, OpenAI answered with GPT-6 Astra at an identical headline price. That timing was not a coincidence, and it set up the closest, most genuinely contested frontier-model comparison either company has produced in some time.

This is a plain-language walkthrough of what Fable 5.1 actually is, what it costs, where independent testing says it leads, where it does not, and what any of this means if you are deciding what to actually use it for.

For where the rest of the Claude lineup and the wider assistant landscape stood going into this release, see our Claude AI Review 2026.


What Claude Fable 5.1 Is

Claude Fable 5.1 is Anthropic’s generally available frontier model for demanding reasoning and long-horizon agentic work, the kind of task that unfolds over many tool calls and an extended session rather than resolving in a single exchange. It is a point release built on top of the earlier Fable 5, and Anthropic’s own documentation is direct about where it sits in the lineup: it is positioned above Claude Opus 5, which itself sits above Claude Sonnet 5, as the most expensive and most capable model Anthropic currently offers to the general public.

Anthropic’s own guidance to developers is worth repeating because it runs against the instinct to always reach for the newest, priciest model. The company recommends starting most production workloads on Opus 5, and moving specific tasks up to Fable 5.1 only when Opus 5 at its highest effort setting falls short. Fable 5.1 trades a real price premium for extra headroom on the hardest reasoning, agentic coding, and multistep research tasks, not as a wholesale replacement for everyday work.

Claude Mythos 5.1 launched the same day and shares the same underlying model, specifications, and pricing as Fable 5.1. What differs is access. Mythos is not a generally available public API tier; it is positioned for vetted organizations through Anthropic’s Project Glasswing and other trusted-access programs, aimed at cyberdefense and life-science research contexts where Anthropic wants tighter control over who can use the model’s full capabilities. For the overwhelming majority of developers and businesses, Fable 5.1 is the relevant model of the two.


The Specs That Matter

Fable 5.1 runs a 1 million token context window with a maximum output of 128,000 tokens, unchanged from its predecessor. Thinking is adaptive and always on, meaning there is no way to fully disable the model’s internal reasoning process. Developers control depth through an effort parameter spanning five levels, from low up to max, rather than switching reasoning off entirely. The default effort on the Claude API is set to high.

The knowledge cutoff is June 2026, which Anthropic describes as the freshest reliable cutoff of any Claude model at time of release. Text generated by both Fable 5.1 and Mythos 5.1 carries Anthropic’s statistical text watermark automatically across every platform where the model is available, a provenance feature that runs in the background without requiring any setup from the developer.

Fable 5.1 is available day one across the Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry, and it appears directly inside Claude Code, with launch-day reporting also confirming availability inside Cursor. One technical detail worth flagging for anyone doing cost comparisons against older logs: the tokenizer introduced with a prior Claude model generation, and carried into Fable 5.1, produces roughly 30 percent more tokens for the same piece of text than older tokenizers did. If you are comparing historical cost data against Fable 5.1 usage, that shift alone can make the newer model look more expensive than it actually is unless you account for it.


Pricing

Fable 5.1 lists at $10 per million input tokens and $50 per million output tokens, the same headline rate as GPT-6 Astra and unchanged from the earlier Fable 5. The number that actually moved, and the one buried deepest in the pricing page, is cache reads, which dropped from $1.00 to $0.25 per million tokens, a 75 percent cut.

That reduction sounds like a footnote and is not one. Anthropic estimates the cache-read cut reduces costs by roughly 25 percent for typical workloads billed by token, and by as much as 45 percent for highly agentic workloads that repeatedly reuse large chunks of context across many tool calls. For any standing agent or long-running coding session that leans heavily on cached context, that line item matters more to the actual monthly bill than the unchanged headline price does. It is also worth noting as a practical caveat that zero data retention is not generally available for Fable 5.1, and the model remains a standard 30-day retention window by default, a detail worth checking before routing sensitive data through it.


Where Fable 5.1 Leads

Running both Fable 5.1 and GPT-6 Astra through the same neutral evaluation harness, independent testing firm Artificial Analysis found Fable 5.1 taking the lead on both of its flagship composite scores in the days immediately following launch: the Intelligence Index and the Coding Agent Index. Those specific composite numbers moved considerably within the following week as the evaluator revised its own methodology, and the most current reading has the two models much closer together, in some versions a near tie. What has held up more consistently across every revision is Fable 5.1’s advantage on Humanity’s Last Exam with tools, a benchmark built around the hardest, most sustained expert-level reasoning tasks currently available, where it measured meaningfully ahead of Astra.

Fable 5.1 also shows a clear, specific jump over its own predecessor on long-horizon agentic and scientific work. Anthropic’s own reported numbers put it at 55.8 percent on Terminal-Bench 4.0, up from 42.0 percent for Fable 5, and 52.6 percent on Terminal-Bench-Science 0.1, more than double the earlier model’s score on the same test.

Safety and adversarial resistance is the other area where Fable 5.1 has a genuine, sustained edge. Independent red-teaming and prompt-injection resistance testing has continued to favor current Claude models over GPT-6 Astra, even accounting for real safety improvements OpenAI made in its own release. Anthropic has built a large part of its public identity around this kind of safety research for longer than most competitors, and that sustained investment continues to show up in third-party adversarial testing.

The tradeoff behind all of this is token usage. Fable 5.1 tends to use more output tokens to reach a given answer than Astra does for a comparable task, which means that even at an identical headline price per token, the actual dollar cost of completing a specific task can end up higher on Fable 5.1 than on Astra, despite Fable 5.1 sometimes scoring higher on the underlying capability measure. Sustained reasoning quality and lower cost per task are not the same axis, and this generation of models makes that distinction unusually visible.


How This Fits Against GPT-6 Astra

There is no single winner here, and treating this as a simple ranking misses the more useful answer. Fable 5.1 and GPT-6 Astra ship at identical headline pricing and land within 48 hours of each other, and independent testing this month has consistently found each model leading in different, genuinely distinct categories rather than one model dominating across the board.

Fable 5.1’s strengths cluster around sustained, expert-level reasoning, long-horizon agentic coding, and safety under adversarial testing. Astra’s strengths cluster around token efficiency, structured math and hard science benchmarks, and autonomous computer use, the kind of task where an agent operates real software interfaces with minimal human guidance. If your workload is heavy on complex, multi-step reasoning where getting the answer right the first time matters more than shaving tokens off the bill, Fable 5.1’s advantages point in your direction. If your workload leans on high-volume automation, structured math and science problems, or cost per completed task as the binding constraint, Astra’s advantages point the other way.

For a full side-by-side breakdown of both models across coding, computer use, cybersecurity, and safety, see our detailed comparison. And for the full picture of what Astra itself actually is, including OpenAI’s own AGI framing and how it holds up under scrutiny, see our ChatGPT Review 2026 for where OpenAI’s flagship product stands.


What It Means for Everyday Users and Businesses

For a typical Claude subscriber, Fable 5.1 is not the model you should assume you need by default. Anthropic’s own guidance is to keep standard production work on Opus 5, which costs half as much, and reach for Fable 5.1 selectively on the specific tasks where its extra reasoning depth or long-horizon agentic stamina measurably improves the outcome. Reflexively routing every request to the most expensive model in the lineup is likely to raise your bill without a proportional improvement in most day-to-day work.

For businesses building on the API, three practical points matter more than the headline benchmark race. First, the 75 percent cut to cache-read pricing is the most consequential line item in this release for any standing agent or long-running session architecture, and it deserves more attention in cost modeling than the unchanged input and output rates. Second, Fable 5.1’s higher token usage per task means the identical sticker price against Astra does not guarantee an identical bill, and workload-specific testing is the only reliable way to know which model is actually cheaper for your specific use case. Third, zero data retention is not currently available for Fable 5.1, so organizations with strict data handling requirements should confirm the current retention terms before routing sensitive material through it.

For anyone deciding whether to switch platforms entirely, the more useful question is not which model wins some abstract intelligence contest, but which model’s specific strengths, reasoning depth and safety for Fable 5.1, token efficiency and computer use for Astra, line up with the actual work you need done.


Frequently Asked Questions

Is Claude Fable 5.1 the same thing as Claude Mythos 5.1, and which one should I actually use?

They share the same underlying model, specifications, and pricing, but they are not interchangeable in practice because access differs. Fable 5.1 is generally available through the standard Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry, and it is what almost every developer and business should be using. Mythos 5.1 is restricted to vetted organizations through Anthropic’s trusted-access programs, primarily aimed at cyberdefense and life-science research contexts. You should not assume access to one implies access to the other, and ordinary Fable availability does not mean Mythos access has also been granted.

Why does Fable 5.1 sometimes cost more per task than GPT-6 Astra even though both list at the exact same price per token?

Because price per token and cost per completed task are different measurements, and this generation of models makes that gap unusually visible. Fable 5.1 tends to use more output tokens to work through a given task than Astra does, so even though both models charge an identical $10 per million input tokens and $50 per million output tokens, the total token consumption for a comparable task can differ enough that the actual dollar cost diverges from what the identical sticker price would suggest. The practical fix is to test both models against your specific representative workload rather than assuming pricing parity means cost parity.

Should I switch all my work from Claude Opus 5 to Fable 5.1 now that it’s out?

Anthropic’s own guidance says no, at least not wholesale. The company explicitly recommends starting most production workloads on Opus 5, which costs half as much per token, and moving only specific tasks up to Fable 5.1 when Opus 5 at high effort measurably falls short, typically the hardest reasoning problems, the most demanding agentic coding sessions, and multistep research tasks that benefit from the extra headroom. Routing everything to Fable 5.1 by default is likely to increase costs without a proportional improvement on the large majority of everyday requests, where Opus 5 already performs well.


Note: Benchmark figures in this explainer were current as of September 2026 and have already shifted within days of launch as independent evaluators revised their testing methodology. Verify current scores directly at the source, Artificial Analysis and Anthropic’s own published documentation, before making any decision based on a specific number cited here.

Related Articles