Claude will save you money compared with building your own model. That's the pitch most consultants lead with, and for the vast majority of Australian small and medium businesses it's true. But it isn't true because self-hosting is a bad idea in principle. It's true because of a number: how many tokens you actually push through a model each month. Below a certain volume, a managed subscription wins outright. Above it, the maths flips. This isn't a matter of taste or vendor loyalty. It's arithmetic, and once you've done it for your own business, you never need to guess again.
Every few months a Sydney business owner asks us whether they should stop paying per token and run their own model instead. It's a fair question, and one worth taking seriously rather than waving away. The answer just happens to depend entirely on volume, not on preference.
This isn't unique to Sydney either. The same question comes up from Melbourne law practices, Brisbane logistics operators and Perth accounting firms, and the answer doesn't change with the postcode. It's a volume question, not a location question.
Where the crossover actually sits
Current industry analysis puts the crossover points in a fairly consistent range, and it's worth having these numbers in front of you before any vendor conversation.
Under 20 million tokens a month, a managed API wins on total cost once operations are counted
Between 20 million and 100 million tokens a month it's genuinely contested, and the answer depends on how variable your load is
Above 100 million tokens a month, self-hosting starts to win on unit economics, and above 500 million it usually wins clearly
Those thresholds hold roughly steady across providers and use cases, which is part of why they're useful. They aren't marketing numbers from one vendor trying to keep you on its platform. They reflect the basic economics of running inference at scale: fixed costs that don't move with usage, and marginal costs that do. Once you know which side of the range your business sits on, the rest of the decision mostly makes itself.
Here's the part that surprises most owners: almost no Australian SMB is anywhere near the bottom of that range, let alone the top. A 40-person professional services firm running document summarisation, email drafting and a customer-facing assistant typically processes somewhere between 3 million and 12 million tokens a month. At that volume, a Claude subscription plus API usage runs around $900 to $3,500 a month, all in. That's the entire managed column for most businesses reading this.
What the self-hosted column actually contains
The instinct is to compare GPU hardware cost against a software subscription and call it a day. That's the wrong comparison, and it's the one that gets businesses into trouble. The GPU line is the small part. Downloaded model weights account for roughly 2 to 5 per cent of total deployment cost. The rest is people and process, and it adds up fast.
A minimum viable production team of 1.5 to 2 full-time engineers, which in Sydney runs $280,000 to $420,000 a year
Model refresh cycles, because the model you deploy in August is already behind by November
Monitoring, evaluation harnesses and incident response for a system your customers now depend on
Idle capacity, since accelerators bill whether or not anyone is actually asking questions
Add those line items together and the realistic floor for a production self-hosted deployment in Australia sits above $150,000 a year. For the professional services firm above, spending $900 to $3,500 a month on Claude, that's roughly forty times what they're currently paying. Forty times, for a system that still needs to be maintained, patched and refreshed by the same 1.5 to 2 engineers you just budgeted for.
There's also an opportunity cost that rarely makes it onto the spreadsheet. Every dollar and every engineering hour spent standing up and babysitting a self-hosted model is a dollar and an hour not spent on the product or service the business actually sells. For most SMBs, that trade-off alone settles the question well before the invoice does.
When the number genuinely flips
None of this means self-hosting is always the wrong call. There are legitimate cases where it is exactly right, and it's worth naming them honestly rather than pretending the answer is always stay on the API.
Businesses selling an AI-powered product where inference is a direct cost of goods sold, not an internal tool
High-volume document processing, the kind you see in insurance or conveyancing, where token volume genuinely clears the threshold
Workloads where a regulator or a client contract requires onshore processing that a managed option can't satisfy
These aren't hypothetical categories. We see them, just rarely, and almost never in the form business owners expect when they first raise the question.
If you think you might be in one of those three groups, the exercise that settles it is boring, which is exactly why it works.
How to run the sums yourself
You don't need a consultant to tell you which side of the line you're on. You need a month of real usage data and the discipline to price both columns honestly.
Pull one full month of real token usage from your API dashboard, not an estimate
Price the managed column as it actually bills, subscription plus usage, in AUD
Price the self-hosted column with salaries included, not just hardware, using Sydney engineering rates as your floor
Compare the two totals against the thresholds above, not against a vendor's marketing claim
Revisit the comparison annually, since your usage will grow and the thresholds occasionally move
Run that five-step version and you'll get an answer in an afternoon, not a quarter. You don't need procurement, you don't need a working group, and you don't need an outside opinion until the numbers actually put you inside the contested band.
We've run that exercise for firms convinced they needed their own cluster and found they were sitting at roughly 4 per cent of the break-even volume. Not close. Not getting there. Four per cent. The conversation that follows isn't really about GPUs at all. It's about what the business was actually trying to solve, which usually turns out to be something a well-configured Claude workflow handles for a few hundred dollars a month rather than a few hundred thousand a year.
What this means for your business
The self-hosting question isn't really a technology decision. It's a spreadsheet with your real numbers in it. Most Australian SMBs will look at that spreadsheet and stay on a managed model for years, and that's not a compromise. It's the correct answer for the volume they're running. A smaller number will look at the same spreadsheet, confirm they've genuinely crossed the line, and find the maths supports the investment either way.
There's a compliance dimension worth separating from the cost dimension, because the two get conflated constantly. Running your own infrastructure doesn't automatically satisfy the Privacy Act, and a managed Claude deployment doesn't automatically fail it either. Data residency, contractual terms and audit trails are usually the actual requirement, not the location of the GPU. If a regulator or an enterprise client genuinely mandates onshore inference, that's a real constraint and belongs in the self-hosted column regardless of volume. But it's a different question from the break-even maths above, and the two shouldn't be allowed to blur together when a board is deciding where to spend $150,000 a year.
Bring us a month of your own token usage and we'll run the comparison with you, honestly, including the line items most vendors leave out. Get in touch through our contact page and we'll work through the numbers together.



