Cloud GPU spot prices have been drifting down through 2026 as more capacity comes online and providers compete harder for AI workloads. If you have a spreadsheet somewhere that says self-hosting an open-weight model does not pay yet, it is reasonable to wonder whether that answer has quietly expired.
For most Australian small and mid-market teams it has not. The price cut is real, but it lands on the smaller half of the bill. This post walks through what moved, what did not, and how to re-run the sum for your own workload.
Do falling GPU prices change the self-hosting break-even?
Falling GPU spot prices lower the self-hosting break-even only slightly, because compute was never the largest cost for a small team. Spot rates for mid-tier GPUs in Australian cloud regions fell roughly 20 to 30% over the past year, while engineering, patching and failover costs stayed flat. Self-hosting still tends to pay only above about $8,000 to $10,000 a month in equivalent managed API spend, held consistently.
Below that level the fixed overhead of running the thing outweighs whatever the cheaper compute gives back. The threshold is a monthly-usage figure, not a hardware figure, which is the same point we made in the token-volume break-even maths.
What moved, and why
The drop applies to the GPU class suited to a 20B to 70B parameter open-weight model, which is the range most business document and drafting workloads sit in. It was driven mostly by hyperscalers adding capacity, not by the hardware itself becoming cheaper to make. That matters for planning: a price set by competition for capacity can go back up when demand catches the supply.
It is also a discount on a local price. Australian regions carry their own uplift over US regions, and we showed how to cost the Sydney premium yourself. A percentage cut on spot rates does not remove that gap.
The four costs that stayed put
Engineering time to build, monitor and maintain the deployment: typically 10 to 20 hours a month, at $100 to $180 an hour for a competent Sydney-based contractor.
Security patching and vulnerability management, which does not get cheaper because GPUs did.
The opportunity cost of engineers working on infrastructure instead of the product.
Redundancy and failover, if the workload has to stay up during business hours.
Run the first line alone and you get $1,000 to $3,600 a month before a single token is processed. That labour figure is the one most calculators leave out, and we costed it in detail in the DevOps labour line item.
| Cost line | Moved with GPU prices? | Typical size |
|---|---|---|
| GPU compute (spot) | Yes, down roughly 20 to 30% | Varies with workload |
| Engineering maintenance | No | 10 to 20 hours a month |
| Security patching | No | Ongoing |
| Redundancy and failover | No | Depends on uptime needs |
| Opportunity cost | No | Product work not done |
A worked example: Perth logistics, document processing
A Perth logistics company running a self-hosted document processing model switched providers to capture spot pricing. Its GPU bill fell from roughly $3,200 to $2,300 a month. Engineering maintenance stayed at around $1,800 a month.
| Line | Before | After | Change |
|---|---|---|---|
| GPU compute | $3,200 | $2,300 | Down $900 (28%) |
| Engineering maintenance | $1,800 | $1,800 | No change |
| Total | $5,000 | $4,100 | Down $900 (18%) |
A 28% cut on the GPU line became an 18% cut on the total. That is $10,800 a year, worth having, and nothing like the shift a headline about falling GPU prices implies. The larger the labour share of your bill, the more the saving shrinks. A team whose maintenance cost equals its compute cost sees a 30% GPU cut turn into 15% overall.
Spot pricing brings its own risk
Spot capacity is cheap because the provider can take it back. For a batch job that can pause and resume overnight, that is a fair trade. For a workload staff rely on between nine and five, an interruption is an outage, and the usual fix is reserved instances. Reserved pricing erodes much of the spot saving, so the number in the headline is often not the number you can actually buy at the availability you need.
Volatility cuts the other way too. If you budget on today's spot rate and it rises next quarter, the break-even you calculated moves against you with no change in your own usage.
How to re-run the sum
Model total cost of ownership with labour included, not just the GPU invoice.
Price the availability you need. If it must stay up in business hours, use reserved rates in the model, not spot.
Compare against what the same monthly volume would cost on a managed model such as Claude.
Re-run it whenever GPU pricing shifts meaningfully, because the answer does change over time.
If the managed-equivalent figure sits under about $8,000 a month, cheaper GPUs are a nice-to-have and not a reason to move off a managed platform. If you are well above $10,000 and steady, self-hosting deserves a proper look, and the lower compute price makes a case that already worked a little better.
What not to conclude
None of this says self-hosting is a bad idea, or that GPU prices do not matter. Falling prices are a real tailwind for open-weight economics, and for teams already past the threshold they improve a case that already stood up. The narrower point is that as at October 2026 the price drop has not pulled many Australian SMBs across the line. Keep watching, and do not switch yet.
If you want a second pair of eyes on the numbers, our ROI calculator gives a first pass, or book a short call and we will work through your workload with you.



