Claude gets compared to open-weight models constantly this year, and the comparison is almost always framed on price. That framing falls apart the moment you look at what the open model actually needs in order to run. The frontier open-weight releases topping the leaderboards in mid 2026 share a property that rarely makes the highlight reel: they are enormous. Moonshot recommends at least 64 accelerators to serve its Kimi K3 model. GLM-5.2 runs 744 billion parameters, with 40 billion active per token. Neither of these is something a business downloads onto a server in the back office and switches on before lunch. They are cluster-scale systems wearing an open licence.
That distinction matters for any Sydney, Melbourne or Brisbane business that heard the word open and assumed it meant free, or at least theirs to control outright. Open licence and runnable are two separate questions, and the gap between them is where the real cost of self-hosting a frontier model actually lives. Get the two confused and you can spend months chasing a deployment that was never realistic for a business your size.
What 64 accelerators costs in Australia
Reserved capacity of that size in an Australian cloud region prices in the range of A$80,000 to A$140,000 a month, depending on accelerator generation and commitment term. That figure covers compute only, before a single customer query has been answered and before anyone has written a line of serving code. Buying the hardware outright is worse for most buyers, not better. Enterprise inference cards still land between A$15,000 and A$40,000 each, which puts a 64-card cluster past A$1.2 million before you have racked a single unit, run the power to the room, or paid for the cooling that keeps it alive. Set that beside a managed model subscription priced on usage, and the self-hosted path needs a genuinely large, sustained inference workload before the economics start to make sense at all, let alone favour it.
The bill doesn't stop at the hardware
That spend buys capacity, not a working system. Getting a 64-accelerator cluster from racked hardware to something a business can actually rely on for production traffic still requires:
One or two engineers who genuinely understand expert-parallel serving and can debug it under load at 2am, not just stand it up once and hope
An evaluation harness built around your own data, because vendor benchmarks will not tell you how the model behaves on your customers' questions and edge cases
A standing plan for the next model, because a better one will land in roughly three months and reopen the entire question of whether the cluster was worth it
None of that shows up in the headline accelerator count, and none of it is a one-off cost. It is recurring, it needs specialist headcount, and it competes for the same engineering time that should be building the product your business actually sells. Most Australian SMBs and mid-market firms do not have this capacity sitting idle, waiting to be assigned to a serving cluster.
The gap between open and runnable
There is a widening split in the open-weight field, and Australian buyers should know which side of it a model sits on before getting excited about its licence:
Small and mid-size open models, roughly 7 billion to 70 billion parameters, that a competent engineering team can genuinely run and maintain on one or two cards
Frontier open models in the hundreds of billions to trillions of parameters, which are open in licence and effectively closed in practice to anyone without a cluster
The second group is still useful, and it is not a wasted release. It keeps pressure on closed-model pricing and gives researchers something to inspect line by line, which is valuable for the field as a whole. What it is not is a procurement option for a Perth engineering firm or a Melbourne health service trying to decide what to run next quarter. Reading the licence tells you what you are legally allowed to do with a model. It tells you nothing about whether your business can afford, staff and operate it, and that second question is the one that actually decides the project.
The comparison that actually matters
If the appeal of an open model is cost control, the model you can realistically afford to serve is a mid-size one, not the name sitting at the top of the leaderboard. The honest comparison is that mid-size open model against Claude, not the leaderboard leader against Claude. Run that comparison properly, on quality per dollar rather than on parameter count, and it usually favours the managed option for Australian businesses under a few hundred staff. The leaderboard-topping open model was never really in the running once the true accelerator count, the specialist engineering hours, and the three-month refresh cycle are all sitting on the same page as the invoice.
If the driver is control, not cost
Sometimes the pull toward self-hosting has nothing to do with the invoice at all. It is a clause requiring Australian data processing, a client's audit right, or an internal policy on how customer data is handled. Each of those has a cheaper answer than a cluster. Claude running in an Australian region through Amazon Bedrock satisfies most of these obligations at a fraction of the operating cost of a 64-accelerator build, and it does so through contract terms and a documented architecture rather than one your own engineers have to design, defend and keep patched indefinitely. Before assuming a self-hosted cluster is the only way to tick that box, go back to the actual clause and confirm what it requires, rather than what a leaderboard post implied it required.
A quick gut-check before you commit
Before signing off on a self-hosted deployment of a frontier open model, it is worth working through a short list:
Do you have a sustained inference workload large enough to justify A$80,000 to A$140,000 a month in reserved capacity, or would usage-based pricing suit the real traffic pattern better
Do you already employ engineers who can run expert-parallel serving under load, or would this be the first project they have attempted at this scale
Is the obligation driving this decision actually about cost, or is it a data-residency or audit clause that a properly configured managed deployment can satisfy
Have you priced the mid-size open model against Claude specifically, rather than benchmarking the leaderboard leader against a managed model it was never a fair substitute for
Are you prepared to revisit this build in roughly three months when a better model lands and the whole comparison resets
The firms that regret self-hosting a frontier open model are rarely the ones that ran the numbers first. They are the ones that read a leaderboard, saw the word open, and assumed it meant free. Running the comparison properly, accelerator count, engineering hours, refresh cycle and all, usually settles the question before any hardware is ordered.
None of this means open weights are a bad idea. A mid-size open model, properly scoped to what your engineering team can actually operate, is a perfectly sound choice for plenty of Australian businesses. The mistake is treating a leaderboard headline as a procurement decision, when the real decision sits three or four steps downstream, in accelerator counts, monthly cloud invoices, and whether anyone on staff can be woken up to fix a serving cluster at 2am.
We price both paths properly for Australian businesses before anyone commits to a cluster or a subscription. Get in touch and we will work through the real numbers for your workload.



