Anthropic has officially introduced Claude Fable 5.1 and Claude Mythos 5.1, the latest evolution of its most advanced models for coding, knowledge work, scientific research and long-running agentic tasks.
The distinction between the two models is unusual: Fable 5.1 and Mythos 5.1 share the same underlying model but use different safeguard configurations. Fable 5.1 is generally available, while Mythos 5.1 is reserved for vetted organizations and professionals through trusted-access programs for cybersecurity and the life sciences.
Anthropic describes Fable 5.1 as its most capable generally available model for ambitious and long-running work. The official Anthropic announcement highlights agentic coding, long-horizon problem solving, scientific research and improved economics for workloads that heavily reuse context.
Built to work for hours
The most important change in Fable 5.1 is not simply better answers to individual prompts. The model is designed to maintain control across much longer sequences of work, using tools, verifying its own output and recovering when individual steps fail.
Anthropic describes jobs lasting hours and spanning several applications, including backlog work, browser operation, managed agents and software changes that touch an entire codebase.
For coding, Fable 5.1 can write tests to verify its own modifications, use vision to compare results against an original design and continue through autonomous verification loops without requiring constant human supervision.
Effort configuration is also important. Anthropic says that at Low or Medium effort, Fable 5.1 can achieve results similar to or better than Fable 5 at substantially lower cost. Claude Code defaults to High effort, while Claude Cowork and Claude.ai default to Medium.
Claude Fable 5.1 benchmarks
Anthropic published comparisons between Fable 5.1, Fable 5, Opus 5 and GPT-5.6 Sol across agentic scientific research, coding, computer use, multidisciplinary reasoning and business workflows.
| Benchmark | Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| Terminal-Bench-Science 0.1 Agentic scientific research | 52.6% | 24.7% | 29.0% | 22.4% |
| Terminal-Bench 4.0 Agentic coding | 55.8% 60.9% Mythos 5.1 | 42.0% | 52.3% | 37.3% |
| GDPval-AA v2 Knowledge work | 1853 | 1723 | 1824 | 1711 |
| OSWorld 2.0 Computer use — partial | 77.9% | 72.9% | 75.4% | — |
| OSWorld 2.0 Computer use — strict | 41.7% | 36.1% | 39.6% | — |
| Humanity's Last Exam Multidisciplinary reasoning — no tools | 60.9% | 57.8% | 56.6% | — |
| Humanity's Last Exam Multidisciplinary reasoning — with tools | 65.0% | 63.8% | 63.6% | — |
| AutomationBench Business workflows | 31.4% | 17.1% | 26.9% | 19.6% |
| CursorBench 3.2.0 Agentic coding | 73.4% | 70.5% | 70.0% | 67.2% |
Where the generational leap appears
The most striking result comes from Terminal-Bench-Science 0.1. In Anthropic's evaluation setup, Fable 5.1 reaches 52.6%, compared with 24.7% for Fable 5.
Anthropic cautions that the benchmark has a standard error of approximately ±3.5–4.5 points per model. The public leaderboard, using three trials per task and the Claude Code harness, reports 30.0% for Opus 5 and 21.4% for Fable 5, while Anthropic's setup reproduces them at 29.0% and 24.7% respectively, both within the expected noise.
The improvement is also significant on AutomationBench, where Fable 5.1 scores 31.4% compared with 17.1% for Fable 5. On Terminal-Bench 4.0, Fable 5.1 reaches 55.8%, while Mythos 5.1 reaches 60.9%.
The Fable-Mythos gap is particularly noteworthy because both use the same underlying model. Anthropic explains that the difference reflects tasks where Fable's cybersecurity safeguards intervened.
On CursorBench 3.2.0, Fable 5.1 reaches 73.4%, compared with 70.5% for Fable 5, 70.0% for Opus 5 and 67.2% for GPT-5.6 Sol.
An important note about the benchmarks
Fable 5.1 was evaluated with its production safeguards enabled. Anthropic notes that when safeguards intervened on specific tasks, Fable 5.1 and Fable 5 received a zero on OSWorld 2.0, while Fable 5 received a zero on AutomationBench.
In other safeguard interventions, cybersecurity tasks were completed by Claude Opus 4.8 and biology tasks by Claude Opus 5. This means safeguard behavior can reduce Fable's measured performance on benchmarks containing sensitive tasks.
The OSWorld 2.0 results also use the benchmark authors' August 2026 task release. Because those task files differ from earlier versions, Anthropic warns that the new figures are not directly comparable with previously published OSWorld results.
Real-world testing
Anthropic also published results and experiences reported by early-access partners.
Millennium describes an extremely rare crash, occurring roughly once every million runs, that its engineers had been unable to explain for four to five years. According to the company, Fable 5.1 traced the issue through an external library and a core dump and identified a bug in the vendor library.
MongoDB says Fable 5.1 built a complex prototype in around three days after researching service code and documentation, then spending hours implementing it through unattended verification loops.
Ramp reported an unattended 38-hour session on a machine-learning problem, during which the model identified a label artifact, corrected it, launched six parallel experiments overnight and returned with results and next steps.
Browserbase says Fable 5.1 completed 82% of tasks on its hardest internal browser-agent benchmark, compared with 74% for Opus 5 and 57% for Fable 5.
These figures are partner-reported internal evaluations rather than standardized independent benchmarks and should be interpreted accordingly.
Scientific research becomes a major focus
Anthropic dedicates a significant portion of the announcement to the scientific capabilities of Fable 5.1 and Mythos 5.1.
Fable 5.1 trained a neural network to create a new high-resolution elevation map covering approximately one third of Venus, using radar data collected by NASA's Magellan mission more than 30 years ago and an existing map covering part of the planet.
According to Anthropic, the new map reveals details at a scale of approximately 2–3 kilometers instead of 10–20 kilometers and improves height accuracy by up to 25%. The map has been released under a Creative Commons license ahead of NASA's VERITAS and ESA's EnVision missions.
Mythos 5.1 was also tested in protein design. On three targets from Adaptyv Bio competitions, Anthropic reports binding affinities up to ten times higher than the best competition submissions. Across 12 targets, the model achieved a hit rate approaching 50%.
In computational biology, Mythos 5.1 wrote custom GPU kernels and cached intermediate results to accelerate seven open-source deep-learning models by up to 2.5x while maintaining identical outputs. Anthropic estimates GPU cost reductions of between 30% and 60% for some genome-scale analyses.
Fable 5.1 vs Mythos 5.1
The difference between the two products is therefore not the base intelligence. Fable 5.1 and Mythos 5.1 are the same underlying model, but Mythos uses more permissive safeguards for vetted professionals.
Anthropic operates two trusted-access programs: the Cyber Verification Program for defensive security work and the Life Sciences Verification Program, developed with the US government for professional life-sciences research.
Claude Mythos remains restricted to vetted organizations. Claude Security, Anthropic's product for scanning codebases for vulnerabilities and suggesting patches for human review, is now powered by Mythos 5.1.
More precise cybersecurity safeguards
Fable 5.1 expands what developers can do with the generally available model. Anthropic now allows Fable 5.1 to be used for identifying software vulnerabilities, enabling a broader range of defensive security work.
Anthropic says Claude Code users should see roughly 60% fewer cybersecurity safeguard interventions per session compared with the previous Fable 5 safeguards.
Some dual-use activities remain restricted or redirected, including penetration testing, exploit generation and binary-based vulnerability scanning.
Biology safeguards have also become more precise. Anthropic says they now trigger 85% less often on benign elementary biology and medical queries than the safeguards originally launched with Fable 5.
Enterprise Frontier Safeguards
Anthropic is also introducing Enterprise Frontier Safeguards (EFS), an architecture intended to combine automated misuse prevention with enterprise privacy and zero-data-retention requirements.
With EFS, customer data remains in cloud infrastructure controlled by the customer rather than Anthropic's systems. Human review is also performed by the customer by default.
Anthropic says EFS was developed with more than 100 customers across financial services, healthcare, manufacturing, telecom, law, retail and the public sector, alongside AWS, Google Cloud and Microsoft Azure.
Stronger anti-distillation protections
Fable 5.1 also introduces stronger protections against model distillation, where the capabilities of an advanced model are systematically extracted to train another system.
For new API accounts created from launch day onward, users can no longer manually edit Claude's earlier context in a multi-turn conversation while preserving the transcript of Claude's previous reasoning.
Anthropic says this closes a publicly documented distillation technique. Existing accounts are initially unaffected, although the company says similar restrictions will apply to future model releases.
EU AI Act text watermarking
The release also includes an important transparency change. Anthropic signed the EU AI Act's Code of Practice on Transparency of AI-Generated Content in July 2026.
Models released after August 2, 2026 therefore include an invisible text watermark. Rather than adding visible text or metadata, it provides a numerical mechanism for estimating the likelihood that Claude contributed to writing a piece of text.
Anthropic says the watermark contains no information about the user, their organization or their conversations and has no practical impact on output quality.
A detection API is also being rolled out in private preview to eligible organizations, including regulators, law enforcement, media organizations, fact-checkers, independent researchers, educational institutions and EU civil-society groups.
Pricing: cache reads are the major change
Claude Fable 5.1 keeps Fable 5's base pricing at $10 per million input tokens and $50 per million output tokens.
The major change concerns cache reads, where the model reuses context it has already processed. Cache-read pricing falls by 75% to $0.25 per million tokens.
This matters especially for coding agents and long-running workflows. An agent working for hours on the same repository may repeatedly reuse source files, instructions and earlier context. When cache reads account for most of the workload, reducing their price can materially change the economics of the entire task.
Anthropic estimates that typical Fable 5.1 workloads cost around 25% less than Fable 5. For complex coding and highly agentic workloads, savings can reach approximately 45%.
Why Fable 5.1 matters for developers
Fable 5.1 makes the shift in coding-model evaluation increasingly clear. Generating an isolated function or code snippet is no longer enough to describe what a frontier model should be able to do.
A real coding agent must understand a repository, trace dependencies between services, modify multiple files, run tests, interpret failures, verify results and keep working until the task is actually complete.
That is why the improvements on Terminal-Bench, AutomationBench and CursorBench may be more meaningful than small gains on traditional reasoning tests: they measure capabilities closer to those required for genuinely autonomous software agents.
Availability
Claude Fable 5.1 is available across Anthropic's platforms and major cloud providers, including Amazon Web Services, Google Cloud and Microsoft Azure.
Developers can access the model through the Claude API using claude-fable-5-1. Developer resources are available through the Claude Platform.
Anthropic also lists Fable 5.1 as available to Pro, Max, Team and Enterprise users.
Conclusion
Claude Fable 5.1 does not look like an update built simply to add a few points to conventional benchmarks. Its focus is much more clearly on models that can work longer, use tools, verify their own results and complete complex workflows with less supervision.
The jump from 24.7% to 52.6% on Terminal-Bench-Science and from 17.1% to 31.4% on AutomationBench illustrates where a significant part of the generational improvement lies. At the same time, cheaper cache reads make the large-context, high-iteration workloads required by such agents substantially more economical.
The frontier-model race may therefore increasingly move away from a relatively simple question — which model produces the best answer? — toward a much more practical one: which model can receive an objective, work on it autonomously for hours and return a verified, genuinely usable result?
Sources and further reading: Anthropic — Claude Fable 5.1 and Mythos 5.1, Anthropic — Claude Fable, Anthropic — Claude Mythos, Claude Platform.