Most self-hosting cost comparisons stop at hardware and electricity. A 2026 total-cost-of-ownership analysis puts a single GPU server for a 7B model at $6,000 to $8,000 in hardware, amortising to a few dollars a day, with electricity adding a few hundred dollars a year on top. Those numbers are real and they are not misleading. They are simply the smaller half of the bill.
The larger half is a person, and people do not appear in vendor calculators.
The line item that does not show up in the spreadsheet
A self-hosted model deployment needs 10 to 20 hours a month of ongoing engineering time. Not setup time, which everybody budgets for, but the recurring work: patching the inference server, monitoring throughput and memory, handling the failure modes that never show up in a demo, and re-tuning after a dependency update breaks something quietly enough that nobody notices for a week.
At a market rate of $130 to $220 an hour for a senior DevOps or machine learning engineer in Sydney or Melbourne, that is $1,300 to $4,400 a month in labour alone, before you have served a single token beyond what the hardware already covers. Over a year it comes to $15,600 to $52,800, sitting on top of the hardware and electricity figures that are the only numbers most comparisons show.
Put differently: on a $7,000 GPU, the labour to keep it useful costs more in the first year than the hardware did, and it costs that again every year after.
When the maths still works
None of which means self-hosting is a bad idea. It breaks even against managed API pricing at real volume, and above the threshold it wins decisively. Industry analysis puts the crossover at roughly 500,000 tokens a day for a 7B model on a single GPU, and around 2 million tokens a day for a 70B model on larger hardware. Above those levels self-hosting can save 60 to 80 percent on pure token costs even after the labour is properly counted.
Signs the arithmetic is in your favour:
You are already running consistent, predictable volume well above the break-even threshold, rather than spiky or seasonal demand that leaves the hardware idle half the month.
You already employ, or can genuinely justify hiring, the engineering capacity the maintenance work needs.
Data residency is a real contractual requirement, for a client bound by the Privacy Act or a specific sovereignty clause, rather than a preference someone expressed in a meeting.
The workload is batch rather than interactive, so a slow hour is an inconvenience rather than an outage.
Signs it does not work
Your volume sits under the break-even threshold, which describes most Australian small and mid-sized businesses we work with, often by an order of magnitude.
The 10 to 20 hours a month would come out of a role that is already stretched. The maintenance is not free just because nobody invoices for it separately; it displaces something else.
You are evaluating self-hosting mainly because of a headline cost-per-token figure, without having costed the labour against it.
Nobody on the team has run production inference before, which turns the first six months into a learning exercise billed at senior engineer rates.
The opportunity cost nobody prices
The subtlest version of this mistake is the business that decides the maintenance is free because it will be absorbed by an existing engineer. That is not free, it is unbudgeted. The work still happens, it still takes 10 to 20 hours, and those hours come out of whatever that engineer was going to build instead. If that engineer is one of three people shipping your actual product, the real cost of self-hosting is measured in delayed roadmap, not in dollars, and it is usually larger.
The reverse case is also worth naming honestly. A business with a platform team that already maintains GPU infrastructure for other reasons genuinely does absorb this at close to zero marginal cost, and for them the calculators are roughly right. The question is which of those two situations you are actually in, and most teams answer it optimistically.
What we tell clients
Self-hosting is a real option for a specific volume profile, not a default and not a mark of sophistication. Before recommending it we cost the full picture, hardware plus electricity plus labour at real rates, against actual token volume from the last 90 days rather than a vendor's best-case comparison or a projection someone made in a planning session.
Roughly two thirds of the time that exercise ends with the client staying on a managed API, and being glad they checked rather than disappointed. The other third get a clear, defensible business case for hardware, which is worth having when it goes to the board.
If you are weighing self-hosting an open model against staying on a managed API, we will run the full-cost comparison against your real numbers rather than a template, book a session and you will get the answer either way.



