The short answer: choose Claude Opus 5.5 for most new coding, agentic and knowledge-work deployments. It leads Fable 5.1 and Opus 5 across Anthropic's latest headline benchmark table while charging $4/$20 per million input/output tokens. Choose Fable 5.1 when your own evaluation shows that its premium research ability, open-ended problem solving or long-horizon orchestration justifies $10/$50 pricing. Keep Opus 5 when an existing production workflow has not yet been validated on 5.5.
This guide compares the three names readers are asking about. Pricing is API list pricing at launch, not the price of a Claude subscription. Benchmark results come from Anthropic's published evaluations and should be read with the test conditions belowânot as a guarantee for your own workload.
At a glance
| Model | Launch position | Input / output | Cache read | Best fit now |
|---|---|---|---|---|
| Fable 5.1 | Premium frontier model, September 2026 | $10 / $50 per 1M tokens | $0.25 per 1M | Frontier research and open-ended work |
| Opus 5 | Value-focused Opus, July 2026 | $5 / $25 per 1M tokens | $0.50 per 1M | Existing validated Opus 5 deployments |
| Opus 5.5 | Leading general model, September 2026 | $4 / $20 per 1M tokens | $0.20 per 1M | Default for new work |
These are standard API prices at launch. Batch processing, fast mode and subscription access use different economics. Anthropic estimates Fable 5.1 typical workloads cost about 25% less than Fable 5 because of lower cache-read pricing.
Benchmark comparison
Higher is better in all four rows. These figures appear together in Anthropic's Opus 5.5 launch table and use production safeguards. They are the cleanest direct comparison Anthropic has published for these three models.
| Benchmark | Measures | Fable 5.1 | Opus 5 | Opus 5.5 |
|---|---|---|---|---|
| Terminal-Bench 4.0 | Agentic terminal coding | 55.8% | 52.3% | 66.4% |
| Terminal-Bench-Science 0.1 | Agentic scientific research | 52.6% | 29.0% | 58.7% |
| AutomationBench | Multi-step business workflows | 31.4% | 26.9% | 40.0% |
| OSWorld 2.0 | Computer use, partial score | 80.7% | 74.0% | 81.8% |
Do not overread small gaps. Anthropic gives Terminal-Bench 4.0 a standard error of roughly ±1.6â2 points for the earlier Claude models and ±2.6 points for Opus 5.5. The OSWorld values are partial scores. Safeguard interventions also counted as failures in some evaluations or triggered fallback models in others. Your own prompts, tools and effort level can change the ranking.
Fable 5.1 versus Opus 5
Fable 5.1 is Anthropic's premium model for coding, knowledge work, scientific research and long-running problem solving. It replaced Fable 5 in September 2026 with higher performance and cache reads reduced to $0.25 per million tokens. Base input and output prices remain $10 and $50 per million tokens.
Opus 5 is the lower-price model at $5/$25. Fable 5.1 beats it across the four benchmarks above, with the largest difference on scientific research. The choice is not simply âbest score winsâ: Fable 5.1 costs twice as much per uncached token, so teams should reserve it for tasks where its extra capability reduces failures, review time or repeated runs enough to repay that premium.
What changed from Opus 5 to Opus 5.5?
Opus 5.5 improves both sides of the value equation. Its list price falls from $5/$25 to $4/$20 per million input/output tokens. Cache reads fall from $0.50 to $0.20, which matters disproportionately for coding agents that repeatedly reuse a large repository context. Anthropic says a typical workload costs 40% less overall because 5.5 also uses fewer tokens per task.
The performance gain is substantial in the published numbers: 52.3% to 66.4% on Terminal-Bench 4.0, 29.0% to 58.7% on Terminal-Bench-Science, 26.9% to 40.0% on AutomationBench, and 74.0% to 81.8% on the partial OSWorld 2.0 score. Anthropic also says Opus 5.5 output is more than 30% faster and that its writing puts the key information first, uses less jargon and follows style rules more reliably.
Coding and software agents
Opus 5.5 is the strongest choice of these three for repository-wide migrations, audits and long-running Claude Code sessions. Anthropic reports an early tester auditing and fixing a 200,000-line codebase in under three hours, compared with more than 20 hours for Opus 5 and 2.5 times as many tokens. That result is a customer example, not a universal speed estimate, but it illustrates where the newer model is aimed: fewer retries and more complete edits across sprawling work.
For a production agent already stable on Opus 5, test 5.5 before switching. Re-run your own success, regression, latency and cost evaluations; inspect whether outputs or tool-use patterns changed; then migrate deliberately. A model upgrade can expose assumptions in prompts, parsers and approval gates even when the model itself is better.
Research and knowledge work
Fable 5.1 is built for frontier research, long-context analysis and difficult professional work. Opus 5 offered a cheaper route to much of that capability. Opus 5.5 now leads Anthropic's reported knowledge-work comparisons while approaching Fable 5.1 on many tasks at a lower operating cost.
For reports, financial analysis and document-heavy work, choose based on verified output rather than model prestige. Create a representative test set with hard-to-find sources, conflicting evidence, tables and calculations. Score citation accuracy, numerical correctness, coverage, editing time and total cost. Anthropic's own announcement warns that benchmark margins at this capability level can exaggerate real-world differences.
Computer use, business workflows and safety
Opus 5.5 posts the highest computer-use and automation scores of the three. It also adds stronger prompt-injection defenses and, according to Anthropic, is less likely than recent models to take hard-to-reverse actions or move outside assigned boundaries. Those improvements reduce risk; they do not make unrestricted automation safe.
Keep tools least-privileged, isolate untrusted web content, log actions and require approval before external messages, deployments, purchases, deletions or other irreversible steps. When an agent reads web pages or documents and can also change systems, prompt injection becomes a workflow security problemânot merely a model-quality problem.
Which model should you choose?
Choose Fable 5.1 ifâŠ
Your hardest research, scientific or open-ended tasks improve enough to justify premium pricing, or you need a high-capability orchestrator for long-horizon work.
Choose Opus 5 ifâŠ
Your production prompts, tools and governance have already been validated on claude-opus-5, and stability matters more than immediate savings.
Choose Opus 5.5 ifâŠ
You are starting new work or can run a migration evaluation. It offers the strongest published performance here with the lowest list and cache-read prices.
A fair way to run your own comparison
- Collect 20â50 representative tasks, including difficult and ordinary cases.
- Use the same tools, context, stopping rules and scoring rubric for each model.
- Record the exact model ID, effort setting, token counts, latency and fallback behavior.
- Grade correctness before style. For code, run tests; for research, open sources; for numbers, recalculate.
- Include human review time in total cost. Cheap tokens are not cheap if they create more correction work.
Verdict
Opus 5.5 wins this comparison for most users. It costs half as much as Fable 5.1 per uncached token, costs less than Opus 5, and leads both in Anthropic's published results for agentic coding, scientific research, business workflows and computer use. Fable 5.1 remains the model to test on frontier research and difficult open-ended work where the premium can be justified. Opus 5 is the stability choice for existing validated systems, not the best starting point for a new one.
Primary sources
- Anthropic: Introducing Claude Opus 5.5
- Anthropic: Introducing Claude Opus 5
- Anthropic: Introducing Claude Fable 5.1 and Claude Mythos 5.1
Prices and availability can change. Check Anthropic's current pricing and model documentation before committing a production budget.