Blog

Claude Fable 5.1: Cheaper Cache, Pricier Per Task

September 2026 · 7 min read · AI Strategy

Line drawing of a thin inlet pipe and a much wider outlet pipe on one tank
← Back to all posts

Claude Fable 5.1 landed in September 2026 with a price cut attached and, for once, the price cut is not the whole story. Anthropic dropped the cache read rate by three quarters. Independent measurement from Artificial Analysis says the effective cost of finishing a task went up, not down, because the model spends more tokens thinking.

Both things are true at the same time. Which one shows up on your bill depends on the shape of your workload, and most Australian teams running Claude Code or Cowork have never actually looked at that shape.

What does Claude Fable 5.1 actually cost to run?

Anthropic cut the Claude Fable 5.1 cache read rate from $1.00 to $0.25 per million tokens, a drop of 75 per cent, and says typical workloads land about 25 per cent cheaper with heavy agent workloads up to 45 per cent cheaper. Artificial Analysis measured something different at the task level: at maximum effort the model costs $3.69 per task, roughly 20 per cent more than Fable 5, because it produces about 1.7 times the output tokens. Cached input got cheaper. Generated output got larger. A bill moves with whichever of those two dominates the work.

Where the cache discount actually lands

Cache reads cover the part of a prompt that does not change between calls. A long system prompt, a repository map, a policy document, a schedule of rates. If your agent re-reads the same context on every turn, the discount is real and it is large. If your work is mostly short prompts producing long answers, the discount barely touches you.

Sorting your own usage into those two buckets takes about an afternoon. The workloads that benefit most tend to share a few traits.

  • Long-lived agent sessions that carry the same project context across dozens of turns.

  • Codebase work where the same files and conventions are re-read on every call.

  • Document review runs where one large reference set is checked against many short inputs.

  • Scheduled jobs that fire the same prompt scaffold every morning against fresh data.

  • Anything where the reusable context is bigger than the answer you get back.

The inverse is the pattern that gets more expensive. Short brief in, long reasoning and long answer out, repeated all day. That describes a lot of research, drafting and analysis work, and it is exactly where the extra output tokens bite.

Reported changes and what each one does to a monthly bill

Reported Claude Fable 5.1 cost changes and what each one does to a monthly bill
Line itemWhat was reportedWhat it means for the bill
Cache read rateCut from $1.00 to $0.25 per million tokensRepeated context gets much cheaper to send
Typical workloadAbout 25 per cent less, per AnthropicA real saving when prompts repeat
Heavy agent workloadUp to 45 per cent less, per AnthropicThe best case, and it needs long sessions
Output tokens per taskAbout 1.7 times Fable 5, per Artificial AnalysisMore thinking, more billed output
Cost at maximum effort$3.69 per task, about 20 per cent above Fable 5A net increase for generation-heavy work
Cost at xhigh effort$2.72 per task, scoring 65 on the Intelligence IndexUsually the better value setting

The effort setting is the lever most teams ignore. Artificial Analysis put Fable 5.1 at 66 on its Intelligence Index at maximum effort and 65 at xhigh, while cost per task fell from $3.69 to $2.72. One index point for roughly a quarter off the run rate is not a close call for routine work. Save maximum effort for jobs where being wrong is expensive, which is the same reasoning behind our guide to choosing between Fable, Opus, Sonnet and Haiku.

Quota burn is the operational problem, not the invoice

Subscription users reported the sharper version of this. One builder on a 20x Max plan said a single session reached half the weekly quota before the task was finished. Another burned half a week's allowance on a handful of tasks. The pattern repeats what happened at the Fable 5 launch: a stronger model that consumes its allowance faster.

For a business, that is a scheduling problem more than a pricing one. If your team hits a quota wall on a Wednesday, the cost of Claude that week is not the subscription. It is the two days of work that stopped. We see this often enough that we treat quota headroom as a capacity metric rather than a billing detail, and it belongs in the same view as everything else covered in the AI cost dashboard every owner should have.

  • Track tokens per completed task, not tokens per day, so a model change shows up in one number.

  • Set effort levels per job type rather than globally, and write the rule down.

  • Watch the gap between session start and quota exhaustion, weekly, per user.

  • Keep one fallback model configured so a quota wall degrades the work instead of stopping it.

What we tell Australian clients to do before switching

A model upgrade is not a procurement event, so it rarely gets a review, and that is how AI spend drifts. Our advice to Sydney and Melbourne clients is to spend a fortnight measuring before moving production work across. A structured spend review of this kind runs between $6,000 and $12,000 for a mid-market team, and it usually pays for itself on the first quarter's invoice because it finds work sitting on an effort setting nobody chose deliberately.

Two other details matter for Australian buyers. The context window is one million tokens, which changes what fits in a single session and therefore how much of it is cacheable. Enterprise Frontier Safeguards keep enterprise data inside a customer-controlled cloud, which is the question our clients in regulated sectors ask first. To model the trade before committing, our ROI calculator is built for this comparison, and our services page sets out how we run the measurement.

What not to read into this

None of this says Fable 5.1 is expensive. It says the headline discount and the effective cost move in opposite directions, and a 75 per cent cut on one line item is not a 75 per cent cut on a bill. The Artificial Analysis figures are per-task measurements on their own test set, not a quote for your workload.

It also does not say you should stay on the older model. A model that finishes a task in one pass at a higher token cost can still beat one that needs three attempts, and per-task measurement captures that where per-token pricing does not. That distinction is the whole argument in cost per outcome as a better AI metric than cost per seat. Measure your own work, then decide.

FAQ

Frequently asked questions

How much did Anthropic cut the Claude Fable 5.1 cache read price?

The cache read rate fell from $1.00 to $0.25 per million tokens, a drop of 75 per cent. Anthropic says typical workloads cost around 25 per cent less and heavy agent workloads up to 45 per cent less.

Why does Claude Fable 5.1 cost more per task despite the cheaper cache?

Artificial Analysis measured about 1.7 times the output tokens compared with Fable 5, because the model spends longer reasoning. At maximum effort that works out to $3.69 per task, roughly 20 per cent above Fable 5.

Is maximum effort worth the extra cost on Claude Fable 5.1?

Artificial Analysis scored the model 66 at maximum effort and 65 at xhigh, while cost per task dropped from $3.69 to $2.72. For routine work the lower setting is usually the better trade.

How large is the Claude Fable 5.1 context window?

One million tokens. That changes how much reference material fits in a single session, which in turn changes how much of a prompt is cacheable and therefore how much the cheaper cache rate saves.

Why are subscription users burning through quota so quickly?

Because the model produces more output tokens per task, allowances deplete faster than on the previous version. Users on high-tier plans reported reaching half their weekly quota within a small number of sessions.

Ready to move from AI pilot to production?

We help mid-market Australian businesses deploy AI automations that actually reach production and deliver measurable ROI.