Blog

The HBM Shortage and Open-Source AI Cost: Plan Ahead

September 2026 · 6 min read · ROI & Business Case

A memory chip at the centre with a flat price line on the left rising steeply on the right
← Back to all posts

US$0.14 per million input tokens. That is roughly what DeepSeek V4 Flash charges as at September 2026, and it is the kind of number now turning up in five-year business cases across Sydney and Melbourne. The number is real. The problem is the assumption quietly stapled to it: that it will still be true in 2030.

A widely circulated 2026 industry analysis on inference economics argues that the current open-weight price war is being propped up by something temporary. Not clever engineering, not generous investors, but a supply glut in high-bandwidth memory, the chip component that sets how much a GPU can serve at once.

Why does high-bandwidth memory decide what AI inference costs?

High-bandwidth memory, or HBM, is the fast memory stacked beside a GPU's processor. It holds the model's weights and the working state of every request being served, so it caps how large a model a card can run and how many users it can handle at once. When HBM is plentiful, providers can pack more work onto each card and cut prices. When it tightens, every served token costs more.

That is why the analysis matters more than a typical chip-market story. The scenario modelling suggests that as memory supply tightens toward 2027 to 2030, the sub-US$0.20-per-million-token pricing that made open-weight models look so attractive this year is unlikely to hold. The price war rests on a cost input nobody in it controls.

The flat-line assumption in most business cases

Most vendor comparisons we see from Australian businesses treat today's open-weight API pricing as a flat line for the next three to five years. That assumption is the single biggest risk in a build-versus-buy decision right now, and it hides in plain sight because nobody writes it down. It is simply what happens when a spreadsheet multiplies this month's price by 60 months.

A more defensible model splits the horizon in two:

  • Next 12 to 18 months: open-weight API pricing is genuinely cheap and worth using for high-volume, low-stakes tasks like classification and bulk summarisation.

  • Two to four years out: treat memory-driven price rises as a real scenario rather than a tail risk, and build a 20 to 40 per cent cost buffer into any total-cost-of-ownership comparison against a managed platform like Claude.

The short-term half is not a hedge. If you have a large, dull, low-risk workload, cheap tokens today are a genuine saving and there is no reason to leave it on the table. The mistake is locking in hardware, contracts or architecture on the assumption that the discount is permanent.

Self-hosting carries the same exposure, just less visibly

Some teams read this and conclude the answer is to own the GPUs. It does not escape the problem. A typical Australian mid-market business self-hosting a mid-size open model on local cloud GPUs is already paying $2,500 to $6,000 a month in infrastructure alone, before the engineering time to keep it patched and monitored. We have costed that DevOps labour line for an Australian team separately, and the premium for hosting in Sydney sits on top again.

Cloud GPU rental is priced off the same hardware market. If memory costs rise as the scenario models suggest, that infrastructure number moves in one direction only. Here is what the scenario stress-test looks like against that monthly range.

Monthly self-hosting infrastructure cost under memory-price-rise scenarios (AUD)
ScenarioLow endHigh endExtra per year at high end
Today's baseline$2,500$6,000$0
Memory-driven rise of 20%$3,000$7,200$14,400
Memory-driven rise of 40%$3,500$8,400$28,800
Memory-driven rise of 60%$4,000$9,600$43,200

These are straight percentage rises applied to the infrastructure line only. They leave engineering time untouched, which is conservative, and they are scenarios rather than forecasts. The point is not that 60 per cent will happen. It is that a business case which only works at 0 per cent is fragile.

How to build the memory risk into a decision

When we run this for clients weighing self-hosting against a managed API, the stress-test covers three things:

  • The current API or hosting cost baseline, measured from real invoices rather than list prices.

  • A memory-price-rise scenario at 20, 40 and 60 per cent increases.

  • A break-even comparison against Claude's published pricing at each scenario.

The third step usually changes the conversation. At baseline, self-hosting may clear the token-volume break-even comfortably. At a 40 per cent rise, the same workload can land on the wrong side of it. If the decision flips inside a plausible scenario, that is a signal to keep commitments short and portable rather than signing a multi-year hardware lease.

A fair caveat about managed pricing

Managed platforms buy from the same memory market, so nobody should pretend Claude pricing is immune to hardware costs over a decade. The difference is who carries the volatility month to month. With a managed service, that risk sits on the vendor's books and shows up, if at all, as a published price change you can plan around. With self-hosting, it arrives as a cloud bill that grew between quarters with no notice.

None of this is an argument against open-weight models. It is an argument against pricing a five-year infrastructure decision off this week's spot price. If you want your current AI cost model stress-tested before you commit, start with our AI readiness assessment or book a time to walk through the numbers.

FAQ

Frequently asked questions

What is HBM and why do AI models need it?

HBM is high-bandwidth memory stacked next to a GPU's processor. It holds model weights and live request state, so it limits both the size of model a card can run and how many requests it serves at once.

Will open-weight model API prices go up?

Nobody can promise either way, but 2026 scenario analysis suggests today's very low prices depend on a memory supply glut that may tighten toward 2027 to 2030, so a rise is a planning scenario worth modelling.

Does self-hosting an open-weight model protect you from price rises?

Not really. Cloud GPU rental is priced off the same hardware market, so a memory-driven increase flows through to your monthly infrastructure bill, often with less notice than a published API price change.

How much buffer should a five-year AI cost model include?

For open-weight comparisons, a 20 to 40 per cent buffer on medium-term costs is a reasonable starting point, with a separate check at 60 per cent to see whether the decision still holds.

Ready to move from AI pilot to production?

We help mid-market Australian businesses deploy AI automations that actually reach production and deliver measurable ROI.