US$0.14 per million input tokens. That is roughly what DeepSeek V4 Flash charges as at September 2026, and it is the kind of number now turning up in five-year business cases across Sydney and Melbourne. The number is real. The problem is the assumption quietly stapled to it: that it will still be true in 2030.
A widely circulated 2026 industry analysis on inference economics argues that the current open-weight price war is being propped up by something temporary. Not clever engineering, not generous investors, but a supply glut in high-bandwidth memory, the chip component that sets how much a GPU can serve at once.
Why does high-bandwidth memory decide what AI inference costs?
High-bandwidth memory, or HBM, is the fast memory stacked beside a GPU's processor. It holds the model's weights and the working state of every request being served, so it caps how large a model a card can run and how many users it can handle at once. When HBM is plentiful, providers can pack more work onto each card and cut prices. When it tightens, every served token costs more.
That is why the analysis matters more than a typical chip-market story. The scenario modelling suggests that as memory supply tightens toward 2027 to 2030, the sub-US$0.20-per-million-token pricing that made open-weight models look so attractive this year is unlikely to hold. The price war rests on a cost input nobody in it controls.
The flat-line assumption in most business cases
Most vendor comparisons we see from Australian businesses treat today's open-weight API pricing as a flat line for the next three to five years. That assumption is the single biggest risk in a build-versus-buy decision right now, and it hides in plain sight because nobody writes it down. It is simply what happens when a spreadsheet multiplies this month's price by 60 months.
A more defensible model splits the horizon in two:
Next 12 to 18 months: open-weight API pricing is genuinely cheap and worth using for high-volume, low-stakes tasks like classification and bulk summarisation.
Two to four years out: treat memory-driven price rises as a real scenario rather than a tail risk, and build a 20 to 40 per cent cost buffer into any total-cost-of-ownership comparison against a managed platform like Claude.
The short-term half is not a hedge. If you have a large, dull, low-risk workload, cheap tokens today are a genuine saving and there is no reason to leave it on the table. The mistake is locking in hardware, contracts or architecture on the assumption that the discount is permanent.
Self-hosting carries the same exposure, just less visibly
Some teams read this and conclude the answer is to own the GPUs. It does not escape the problem. A typical Australian mid-market business self-hosting a mid-size open model on local cloud GPUs is already paying $2,500 to $6,000 a month in infrastructure alone, before the engineering time to keep it patched and monitored. We have costed that DevOps labour line for an Australian team separately, and the premium for hosting in Sydney sits on top again.
Cloud GPU rental is priced off the same hardware market. If memory costs rise as the scenario models suggest, that infrastructure number moves in one direction only. Here is what the scenario stress-test looks like against that monthly range.
| Scenario | Low end | High end | Extra per year at high end |
|---|---|---|---|
| Today's baseline | $2,500 | $6,000 | $0 |
| Memory-driven rise of 20% | $3,000 | $7,200 | $14,400 |
| Memory-driven rise of 40% | $3,500 | $8,400 | $28,800 |
| Memory-driven rise of 60% | $4,000 | $9,600 | $43,200 |
These are straight percentage rises applied to the infrastructure line only. They leave engineering time untouched, which is conservative, and they are scenarios rather than forecasts. The point is not that 60 per cent will happen. It is that a business case which only works at 0 per cent is fragile.
How to build the memory risk into a decision
When we run this for clients weighing self-hosting against a managed API, the stress-test covers three things:
The current API or hosting cost baseline, measured from real invoices rather than list prices.
A memory-price-rise scenario at 20, 40 and 60 per cent increases.
A break-even comparison against Claude's published pricing at each scenario.
The third step usually changes the conversation. At baseline, self-hosting may clear the token-volume break-even comfortably. At a 40 per cent rise, the same workload can land on the wrong side of it. If the decision flips inside a plausible scenario, that is a signal to keep commitments short and portable rather than signing a multi-year hardware lease.
A fair caveat about managed pricing
Managed platforms buy from the same memory market, so nobody should pretend Claude pricing is immune to hardware costs over a decade. The difference is who carries the volatility month to month. With a managed service, that risk sits on the vendor's books and shows up, if at all, as a published price change you can plan around. With self-hosting, it arrives as a cloud bill that grew between quarters with no notice.
None of this is an argument against open-weight models. It is an argument against pricing a five-year infrastructure decision off this week's spot price. If you want your current AI cost model stress-tested before you commit, start with our AI readiness assessment or book a time to walk through the numbers.



