Anthropic · Model guide
Claude Fable 5.1
Anthropic's most capable generally available model, for demanding reasoning and long-horizon agentic work: long-running agentic coding, multistep research, and document, spreadsheet, and slide work. Thinking is always on and cannot be turned off, so effort is the only depth control, and forced tool choice is rejected on this model. Claude Mythos 5.1 is the same model with different safeguards, available to Project Glasswing participants only.
What this model is good at
What the vendor positions it for, and which of our models to reach for instead.
VendorFor demanding reasoning and long-horizon agentic work.source
Agentic coding · Anthropic positions it for long-horizon agentic work and reports stronger long-running agentic coding than Claude Fable 5.(vendor claim)source
Multistep research · Anthropic calls its research capabilities an early glimpse of how AI models will contribute to scientific progress, and reports 52.6% on Terminal-Bench-Science 0.1 against 24.7% for Claude Fable 5.(vendor claim)source
Professional knowledge work · Anthropic reports gains on document, spreadsheet, and slide work, and the highest GDPval-AA v2 rating in its own comparison table.(vendor claim)source
- anthropic/claude-opus-5cheaper —Anthropic recommends starting with Claude Opus 5 for most workloads, and reaching for Claude Fable 5.1 only for demanding reasoning or when Opus 5 at higher effort still falls short.(vendor claim)source
- anthropic/claude-fable-5predecessor —Claude Fable 5 remains available if you depend on forced tool use, which this model rejects.source
- anthropic/claude-sonnet-5cheaper —Claude Sonnet 5 is a fifth of the input price for agent work that does not need the top model.
Pricing and billing
Billed per token. Cached input is charged at the cache-read rate.
| Price / 1M tokens | You pay |
|---|---|
| Input | $ 10.00 |
| Output | $ 50.00 |
| Cache read | $ 0.25 |
| Cache write | $ 12.50 |
+ 2.0K × $ 50.00
Capabilities and limits
What CrossModel guarantees across every route this model can take right now.
- Context window
- 1.0M tokens
- Max output
- 128.0K tokens
- Input / output modalities
- Text + Image → Text
- Streaming
- Supported
- Tool calling
- Supported
- Structured output (JSON)
- Supported
- Image input
- Supported
- Thinking mode
- Supported· always on
- Reasoning effort
- low · medium · high · xhigh · max
- Available endpoints
- /v1/chat/completions · /v1/responses · /v1/messages
Published benchmarks
Scores the vendor reported, with the evaluation setup each one came from.
- CursorBench 3.2.0score·Anthropic · 2026-09-01setup not fully disclosed
Anthropic writes the version as 3.2.0. Kept separate from the CursorBench v3.2 rows in this catalog, which come from a different reporter and may not be the same patch revision. Anthropic evaluated Claude Fable 5.1 with production safeguards enabled and states this likely lowers its score: tasks where the safeguards intervened were scored zero on OSWorld 2.0, and elsewhere cybersecurity tasks were completed by Claude Opus 4.8 and biology tasks by Claude Opus 5.
73.4% - Terminal-Bench 4.0pass_rate·Anthropic · 2026-09-01setup not fully disclosed
Claude Mythos 5.1, the same underlying model with different safeguards, scores 60.9%; Anthropic attributes the gap to its earlier, less precise cyber safeguards and expects it to narrow. Anthropic evaluated Claude Fable 5.1 with production safeguards enabled and states this likely lowers its score: tasks where the safeguards intervened were scored zero on OSWorld 2.0, and elsewhere cybersecurity tasks were completed by Claude Opus 4.8 and biology tasks by Claude Opus 5.
55.8%
- AutomationBenchscore·Anthropic · 2026-09-01setup not fully disclosed
Business workflows. Anthropic evaluated Claude Fable 5.1 with production safeguards enabled and states this likely lowers its score: tasks where the safeguards intervened were scored zero on OSWorld 2.0, and elsewhere cybersecurity tasks were completed by Claude Opus 4.8 and biology tasks by Claude Opus 5.
31.4%
- OSWorld 2.0 (partial credit)success_rate·Anthropic · 2026-09-01setup not fully disclosed
Partial-credit scoring. Kept separate from the plain OSWorld 2.0 rows in this catalog, whose numbers come from the Claude Opus 5 launch table and do not reconcile with either column here. Anthropic evaluated Claude Fable 5.1 with production safeguards enabled and states this likely lowers its score: tasks where the safeguards intervened were scored zero on OSWorld 2.0, and elsewhere cybersecurity tasks were completed by Claude Opus 4.8 and biology tasks by Claude Opus 5.
77.9% - OSWorld 2.0 (strict)success_rate·Anthropic · 2026-09-01setup not fully disclosed
Strict scoring. Kept separate from the plain OSWorld 2.0 rows in this catalog, whose numbers come from the Claude Opus 5 launch table and do not reconcile with either column here. Anthropic evaluated Claude Fable 5.1 with production safeguards enabled and states this likely lowers its score: tasks where the safeguards intervened were scored zero on OSWorld 2.0, and elsewhere cybersecurity tasks were completed by Claude Opus 4.8 and biology tasks by Claude Opus 5.
41.7%
- GDPval-AA v2elo·Anthropic · 2026-09-01setup not fully disclosed
Elo is relative to the field, so it moves as models are added: this table rates Claude Opus 5 at 1824 while the Claude Opus 5 launch table rated it 1861. Compare within one table, not across. Anthropic evaluated Claude Fable 5.1 with production safeguards enabled and states this likely lowers its score: tasks where the safeguards intervened were scored zero on OSWorld 2.0, and elsewhere cybersecurity tasks were completed by Claude Opus 4.8 and biology tasks by Claude Opus 5.
1853
- Humanity's Last Examaccuracy·Anthropic · 2026-09-01setup not fully disclosed
No tools. Anthropic evaluated Claude Fable 5.1 with production safeguards enabled and states this likely lowers its score: tasks where the safeguards intervened were scored zero on OSWorld 2.0, and elsewhere cybersecurity tasks were completed by Claude Opus 4.8 and biology tasks by Claude Opus 5.
60.9% - Humanity's Last Exam (with tools)accuracy·Anthropic · 2026-09-01setup not fully disclosed
Anthropic does not disclose which tools the run had. Anthropic evaluated Claude Fable 5.1 with production safeguards enabled and states this likely lowers its score: tasks where the safeguards intervened were scored zero on OSWorld 2.0, and elsewhere cybersecurity tasks were completed by Claude Opus 4.8 and biology tasks by Claude Opus 5.
65%
- Terminal-Bench-Science 0.1accuracy·Anthropic · 2026-09-01setup not fully disclosed
Anthropic reports a standard error of 3.5-4.5 points per model. Its own harness reproduces the public leaderboard (3 trials/task, Claude Code harness) at 29.0% for Claude Opus 5 and 24.7% for Claude Fable 5, against 30.0% and 21.4% published. Anthropic evaluated Claude Fable 5.1 with production safeguards enabled and states this likely lowers its score: tasks where the safeguards intervened were scored zero on OSWorld 2.0, and elsewhere cybersecurity tasks were completed by Claude Opus 4.8 and biology tasks by Claude Opus 5.
52.6%
Vendor-reported numbers, not CrossModel measurements. Scores are only comparable when the benchmark version, metric and evaluation setup match, so nothing here is averaged or ranked.
Use it in your tools
Point the base URL at CrossModel and paste this model ID — every tool below has a setup guide.
Frequently asked questions
What is Claude Fable 5.1?
How much does Claude Fable 5.1 cost?
Does Claude Fable 5.1 support tool calling and structured output?
Which endpoint do I call?
Can I try it without writing code?
Chat first, integrate later
Chat first, then wire it in once you like the answers.