Anthropic’s Fable 5.1 Takes Benchmark Lead, but Its Top Setting Costs More Per Task
Artificial Analysis found Anthropic’s newest general model reached its highest measured intelligence score, while the token use required at maximum effort raised the cost of its own evaluation tasks.
Listen to this story
The audio brief
Story brief
3 key pointsFable 5.1’s benchmark lead comes with a configurable cost curve, not a blanket efficiency win. Artificial Analysis measured 58–66 across five effort levels, with maximum effort costing an estimated $3.76 per task versus $3.14 for Fable 5. Lowering effort to xhigh produced 65 at $2.72, while the evaluator found overlap or ties on several subtests. Enterprises can now test the general model through AWS, but should...
- 01
Maximum effort used 143.7 million output tokens per Intelligence Index task, versus 13.1 million at the lowest setting.
- 02
Anthropic cut cache-read pricing to $0.25 per million tokens; Artificial Analysis estimated $1.40 saved, outweighed by heavier output use.
- 03
Fable 5.1 and Opus 5 were not decisively separated: GDPval-AA v2 intervals overlapped, and AA-Briefcase was effectively tied.
Anthropic’s Claude Fable 5.1 has taken the top measured position on Artificial Analysis’ Intelligence Index, scoring 66 at its maximum effort setting. The independent evaluator’s accompanying cost analysis, however, puts that setting at $3.76 per benchmark task—20% above the prior Fable 5—making the release a test of whether its performance edge justifies the additional compute.
Artificial Analysis said the 66 score was the highest it had measured. The result placed Fable 5.1 ahead of Claude Opus 5 at 63, Fable 5 at 62, GPT-5.6 Sol at 61 and Grok 4.6 at 61 in its cited comparison. The model’s score also sits six points above Moonshot AI’s Kimi K3 and Z.ai’s GLM-5.3 in the reported comparison with leading Chinese models.
The score depends on how hard the model is asked to work
Fable 5.1 has five effort settings. Across them, Artificial Analysis measured scores from 58 to 66, while output-token use ranged from 13.1 million to 143.7 million tokens per Intelligence Index task. The top score is therefore not a default measure of the model’s cost or behavior in every application; it is the result at the most token-intensive setting in this evaluation.
Lower cache pricing does not erase the output bill
Anthropic kept Fable 5’s listed prices for input, output and cache-write tokens, but cut cache-read pricing to $0.25 per million tokens. Artificial Analysis attributes the higher maximum-effort task cost chiefly to Fable 5.1 using about 1.7 times as many output tokens as Fable 5; the cache reduction saved an estimated $1.40 per task in its evaluation.
A less intensive option narrows the trade-off. At xhigh effort, Fable 5.1 scored 65 at an estimated $2.72 per Intelligence Index task—$1.04 below its maximum-effort run, though still above the $2.34 estimated for Opus 5 at maximum effort. Those figures compare a particular benchmark workload, not universal prices for customer deployments.
What the benchmark results do and do not show
- Fable 5.1’s 66 is the strongest Intelligence Index result Artificial Analysis has recorded.
- Its lead on agentic knowledge-work tests is not cleanly decisive against Opus 5: Artificial Analysis said their GDPval-AA v2 confidence intervals overlap and called their AA-Briefcase results effectively tied.
- On one accuracy-focused evaluation, Fable 5.1 attempted more answers than Fable 5 and also responded more often when it was wrong, leaving their overall AA-Omniscience Index scores level.
A general release with a controlled counterpart
The benchmark update concerns Fable 5.1, Anthropic’s general model for coding and knowledge work. Anthropic says the separately restricted Mythos 5.1 uses the same underlying model, but access remains limited to vetted organizations because it retains fuller cybersecurity and biology capabilities. Fable applies safeguards in those areas, and some safety-flagged requests are routed to other Claude models; Artificial Analysis said this fallback served about 4% of output tokens in its Intelligence Index evaluation.
For buyers, the immediate distinction is not simply between a leading score and lower cost. It is between effort settings that trade more output tokens for incremental benchmark performance, with safeguards and fallback behavior also part of the product’s measured result. Fable 5.1 is generally available on AWS, giving enterprises another route to test that trade-off on their own workloads.
Editorial analysis
Our Read
Our read: Fable 5.1’s result sharpens a buying question that raw leaderboard positions cannot settle: how much extra work is worth paying for when the model’s highest score uses substantially more output tokens. The next useful evidence will be deployment data that compares completion quality, retries and total cost on real workloads at different effort settings. That matters especially as model-serving platforms have gained business customers looking for lower-cost and adaptable alternatives, while frontier-model spending has remained concentrated among major providers.
Sources
- artificialanalysis.aiClaude Fable 5.1 tops the Artificial Analysis Intelligence Index
- scmp.comFrontier AI at a cost: what Anthropic’s Fable 5.1 means for US-China model race
- anthropic.comwww.anthropic.com
- aws.amazon.comClaude Fable 5.1, Anthropic