Daily issuepublished

Fable 5.1 leads, but its top setting costs more

Its highest score came at maximum effort, while a lower setting narrowed the cost trade-off and AWS adds a test route.

By 7 min read
Fable 5.1 leads, but its top setting costs more
Fable 5.1 leads, but its top setting costs more

The audio edition

Listen to this newsletter

About 1:42
0:001:42
Read transcript
The clearest signal today is that AI performance is still arriving with a bill attached. Anthropic’s Fable 5.1 took the top measured spot on Artificial Analysis’s Intelligence Index, scoring 66 at its maximum effort setting—the highest score that evaluator has recorded. But that run was estimated at three dollars and 76 cents per benchmark task, versus three dollars and 14 cents for Fable 5, a 20 percent increase. The crucial caveat is that this is not one fixed model behavior. Across five effort settings, Fable 5.1 scored from 58 to 66, while output-token use ranged from 13.1 million to 143.7 million tokens per task. At the less intensive xhigh setting, it scored 65 for an estimated two dollars and 72 cents. Anthropic kept its input, output, and cache-write prices, while cutting cache-read pricing to 25 cents per million tokens. Artificial Analysis says that saving reduced the task cost by about a dollar and 40 cents, but Fable 5.1 still used roughly 1.7 times as many output tokens. And the lead is not decisive everywhere. Fable 5.1 and Opus 5 had overlapping confidence intervals on GDPval-AA v2, and were effectively tied on AA-Briefcase. About 4 percent of evaluation output tokens came from fallback to other Claude models, so safeguards and routing are part of the observed result too. Fable 5.1 is generally available on AWS, giving enterprises a practical way to test whether the extra reasoning pays off on their own workloads—or simply creates more work and cost downstream. That downstream cost is already visible in freelance work. Freelancer.com says listings for correcting AI output rose 87 percent to 10,760 worldwide between August 2025 and June 2026. The jobs include rebuilding images, repairing robotic voiceovers and malformed 3D models, and revising prose that feels repetitive or emotionally flat. Upwork reported a 70 percent rise in AI-remediation gigs, while Fiverr said searches for AI-cleanup services increased more than twentyfold since 2023. Those are listings and searches, not completed projects or earnings. And the repair is not always quick: illustrator Todd Van Linda said cleanup can take hours or days. The broader pattern is a trade—faster first drafts, followed by human verification and reconstruction. The same verification problem is emerging in hiring. After five AI interviews produced no follow-up, a job seeker identified as Christopher used ChatGPT Voice to impersonate him in a ten-minute call with Riley, an AI recruiter used by Everforth Apex Systems. A fictional résumé-matched applicant also completed a 23-minute interview without a visible next step. That is one reported experiment, and Everforth did not respond to WIRED. It does not establish the company’s criteria or how widespread synthetic applications are. But Greenhouse reported that 63 percent of surveyed U.S. job seekers had faced an AI interview, while 51 percent of candidates who completed one never heard back. Automation is now screening automation, without necessarily improving accountability. And in robotaxis, the useful safeguard is explanation close to the decision itself. MIT and Motional’s CW-Net places readable concepts—such as “close to cyclist”—inside the planner’s final trajectory choice, producing an explanation in real time rather than after the fact. In a private-track test, it revealed that a vehicle had selected a path that would have hit a cyclist; emergency braking prevented the impact. The system was trained on 130 million labeled scenes, and reported driving capability changed by less than 1 percent. Across these stories, the thing to watch is not just whether AI gets better. It is whether effort settings, cleanup labor, feedback loops, and interpretable safeguards make that improvement usable in the real world.
Its highest score came at maximum effort, while a lower setting narrowed the cost trade-off and AWS adds a test route.
Daily issue / Second-Order Effects Thursday, September 3, 2026
Our tools Superpower ChatGPT/WFH.team/Snipman

Today's briefing

What matters today

Inside today's briefing
01
02
03
04
Anthropic’s Fable 5.1 Takes Benchmark Lead, but Its Top Setting Costs More Per Task

Lead story / benchmark

Fable 5.1 leads a benchmark at a higher cost

Read full story  ↗
A tool for your workflow WFH.team A focused feed of carefully selected remote roles and practical work-from-home resources. Remote work, without the noisy job-board scroll
 
Browse remote roles  ↗
Freelancer.com Sees AI-Cleanup Listings Rise 87% as Creatives Shift to Repair Work

research

Freelancers are being hired to repair AI output

Continue reading  ↗
Job Seeker Sends ChatGPT to AI Recruiter After Five Interviews Without Follow-Up

culture

A job seeker sent ChatGPT to an AI recruiter

Continue reading  ↗
MIT and Motional Build AI That Exposes Robotaxi Planning Errors

research

MIT and Motional built a tool to expose robotaxi errors

Continue reading  ↗
Second-Order Effects themed section header

A benchmark lead raises a practical question: what does a higher top-setting cost change in production?

Abliteration Removes Refusal Mechanisms for Offensive Cyber WorkRead story ↗
Fable Cache Becomes Cheaper After 24 Reuses When Gemini MissesRead story ↗
 

Daily tool drop

5 AI tools worth knowing today

Selected for fit, not rank
Monid One key connects AI agents to more than 1,800 APIs across data and content services. Best for / Builders wiring agents to external tools Open ↗
Dial Gives agents phone numbers for calls, SMS, iMessage, and inbound verification codes. Best for / Teams building phone-capable agents Open ↗
HydraDB OSS Open-source graph database for AI memory, ontologies, and agent context on object storage. Best for / Engineers building context-rich AI apps Open ↗
Doop A multiplayer canvas where Claude, Codex, and MCP agents design and review work live. Best for / Product teams co-designing with agents Open ↗
Orato Scores short speaking drills for pacing, fluency, vocabulary, and coherence. Best for / Professionals sharpening spoken delivery Open ↗
 
Anthropic’s 80-Environment Reward-Hacking Test Produced Cyber and Safety Evasions Anthropic’s reward-hacking test produced cyber evasions ↗research
Meta Blocks AI Glasses From Recording When Their Warning Light Is Covered Meta’s AI glasses stop recording if the warning light is covered ↗security risk
New York City Public Schools Will Block Student AI Use Through Eighth Grade New York City schools plan to bar student AI use through eighth grade ↗government action
 

The Internet Had a Point

 
From our network. Tools built for the way you work. Useful products from the team behind Superpower Daily.

Reader check-in

Help shape tomorrow's briefing

One click tells us what to keep, improve, or tighten.

Prefer one email a week? Get the essential AI moves in the Sunday Weekly Digest.

Superpower Daily tracks the companies, models, products, tools, policy decisions, and cultural shifts moving AI. Follow Superpower Daily Email preferences