A$1,800 a month. That was the gap one Australian client expected to save by moving a document-summarisation workload from Claude to DeepSeek V4 Flash, worked out from two pricing pages. Once we ran their real traffic through both, the gap was under A$400. Nothing about either vendor's price was wrong. The comparison was.
Two price moves in the same month
Claude Fable 5.1 became generally available in September 2026 with cache reads cut to US$0.25 per million tokens, down from US$1.00, a change we unpacked in our Fable 5.1 pricing breakdown. Over the same period DeepSeek V4 Flash has been charging roughly US$0.14 per million input tokens on a cache miss and US$0.28 per million output tokens.
Put those side by side and DeepSeek looks cheaper on every line. For a lot of real workloads it still is. The trouble is that the list-price comparison quietly assumes every token is a fresh, uncached input, and very few production workloads look like that.
Is Claude Fable 5.1 cheaper than DeepSeek V4 Flash for Australian workloads?
Per token, no: DeepSeek V4 Flash lists lower input and output prices than Claude Fable 5.1. Per month, it depends on how much of your prompt is repeated context. Workloads that resend the same system prompt, policy set or retrieved documents on every call pay Claude's US$0.25 cache-read rate on most of their input, which shrinks the gap sharply. Output-heavy workloads with little reuse keep most of DeepSeek's price advantage.
That is why two businesses in the same Sydney office block can look at the same pricing pages and reach opposite, correct conclusions. The deciding variable is not the model. It is the shape of the traffic.
Where the gap actually sits
Split any bill into three streams and the comparison becomes much easier to reason about. The table uses only the per-million rates listed above, and treats DeepSeek conservatively at its cache-miss rate.
| Traffic stream | Claude Fable 5.1 | DeepSeek V4 Flash | How the gap behaves |
|---|---|---|---|
| Repeated context (cache read) | US$0.25 | About US$0.14 on a miss | Narrow: about 11 cents per million |
| Fresh, uncached input | Standard input rate | About US$0.14 | Wider, varies with volume |
| Generated output | Standard output rate | About US$0.28 | Widest; output stays DeepSeek's strongest line |
| Retries and human review | Depends on first-pass quality | Depends on first-pass quality | Often larger than the token bill |
Take an illustrative workload that resends one billion tokens of repeated context a month. Under the old Claude cache rate that stream cost US$1,000. At US$0.25 it costs US$250, against about US$140 on DeepSeek at the miss rate. A US$860 gap on that stream has become roughly US$110. The remaining difference lives almost entirely in fresh input and generated output.
The client whose A$1,800 gap shrank below A$400
The client above runs a summarisation pipeline where every request carries the same long instruction set and a shared reference pack, then adds a short new document. Most of their input tokens were repeated context. On list prices alone the apparent monthly difference was about A$1,800. Modelled against their actual cached-to-uncached ratio after the Fable 5.1 update, the effective gap came in under A$400 at their volume.
Under A$400 a month is not nothing. It is about A$4,800 a year. But it sits in a different decision bracket from A$21,600, and it has to be weighed against the cost of moving: prompt rework, a rebuilt evaluation set and a new set of failure modes to learn. We covered why switching costs routinely outweigh a headline price cut in an earlier piece, and the same arithmetic applies here.
Four numbers to model before you compare
You can do a first pass of this in a spreadsheet in an afternoon, using a week of real logs rather than a guess.
The share of input tokens that are repeated context and would be served from cache, measured per workload rather than averaged across the business
The ratio of output to input tokens, because output is where DeepSeek's price advantage survives
First-pass acceptance rate on your own documents: how often each model's answer is used without a retry or edit
Staff minutes spent reviewing or correcting output, priced at a loaded hourly rate
The last two are where cheaper models most often give back their saving. A model that needs a second pass on one request in five has quietly added a fifth to its token bill before anyone has looked at the review time. If you want a structured way to price that, our ROI calculator walks through the staff-time side.
When DeepSeek V4 Flash still wins on cost
This is not an argument that open-weight pricing is a myth. There are workloads where Flash is the cheaper, sensible choice even after a proper model:
High-volume classification, tagging and routing, where prompts are short and each call is mostly fresh input
Output-heavy bulk generation with low stakes and cheap review, such as internal first drafts nobody sends externally
Batch jobs with little shared context, where a cache discount has nothing to act on
The fast, large-context design behind Flash, which we explained in our V4 Flash explainer, suits that kind of throughput work. Cost is also only one axis. For anything carrying personal information, the Privacy Act and where the data is processed can settle the question before price does, which is the subject of our Claude vs DeepSeek data questions piece.
A caution on reading this comparison
Every figure here is a per-token list rate as at September 2026, and both vendors have changed prices more than once this year. Exchange rates move too: an AUD budget built on US-dollar token prices carries currency risk that a pricing page never shows. Re-run the model whenever either vendor announces a change, rather than treating one comparison as settled.
If you are weighing a platform switch off the back of a pricing announcement, book a session with us and bring a week of usage logs. We will model both options against your actual workload before anyone rewrites a prompt.



