Every few months a new open-weight model lands and someone declares the closed-model era over. In 2026, for the first time, the benchmark data actually backs that up. And yet enterprise adoption of open-weight models just fell. If you're an Australian business owner weighing a build vs buy call on your AI stack, that contradiction is worth sitting with before you pick a side.
The quality gap has effectively closed
At the end of 2023, the best closed model was scoring roughly 17.5 percentage points above the best open alternative on standard knowledge benchmarks. By early 2026 that gap had shrunk to near zero. Independent benchmark trackers through mid-2026 now show leading open-weight models matching proprietary systems on coding tasks and sitting within 3 to 8 percentage points on the hardest composite reasoning and safety evaluations, where Claude, GPT and Gemini's frontier models still hold a real but narrowing edge.
On cost, the picture is more lopsided again. Open models are now closing 70 to 90% of the capability gap at roughly 5 to 10 times lower per-token inference cost. Some, like MiniMax M2.7, are landing within striking distance of Claude Opus on real coding workloads at around 50 times less per million output tokens. On paper, that's the kind of number that should be redrawing every vendor shortlist in the country.
Adoption went the other way
According to Menlo Ventures, open-source models fell from 19% of enterprise usage in 2024 to just 11% in 2025, despite the quality gains above. That's not a rounding error, and it isn't about capability. The teams pulling back aren't doing it because the models got worse. They're doing it because the layers underneath model quality, the parts that don't show up on a leaderboard, got harder to manage.
Production gap: only 53% of open-model teams reach production, against 63% for teams running closed models. The rest stall in pilot.
Total cost of ownership: cheaper tokens don't offset the hosting, fine-tuning, and MLOps headcount required to run open weights safely at scale.
Runtime governance: closed vendors ship guardrails, audit logs, and support contracts bundled in. Open deployments have to build all of that themselves.
Concentration risk: many enterprises bet their whole open strategy on one model family, so that vendor's release cadence became the entire segment's growth rate.
There's also a geographic split worth knowing. Greater China and East Asia lead the world in open-weight adoption at around 89%, while South America and Western Europe are the only regions where closed models still outpace open ones. Australia doesn't yet have a clean equivalent figure, but our clients' buying patterns track closer to the Western Europe end: preference for a managed, governed platform over a self-hosted open stack.
Why this matters more for Australian buyers
Most of the commentary on this gap is written for US enterprises with in-house MLOps teams and a security function that can sign off on self-hosting. That's not most of the Australian SMB and mid-market businesses we work with in Sydney and across the country. If your business is regulated under APRA, handles data covered by the Privacy Act, or simply doesn't have a platform team, the total cost of ownership problem above isn't a nuisance, it's the whole decision.
Self-hosting an open-weight model well enough to trust it with customer or financial data is a genuine build. Realistic budgets for a properly governed self-hosted deployment, including fine-tuning, monitoring, and a support runbook, start around $35,000 and climb fast from there. Compare that against a managed Claude deployment through Automata AI, where the governance, audit trail, and vendor support are already built in, and the calculus shifts quickly for anyone without an existing platform team.
The build vs buy read for 2026
The honest 2026 answer isn't open or closed. It's a mix, and where you land on that mix should follow your actual constraints, not the headline benchmark number.
If you have a platform team and a cost-sensitive, high-volume workload, an open-weight model fine-tuned for that specific task can be the right call.
If you're running customer-facing or regulated workflows without in-house MLOps, a managed model like Claude removes the governance burden the Menlo data shows most teams underestimate.
If you're already committed to one open model family, treat that as a concentration risk to manage, not a strategy to defend.
Re-run your build vs buy numbers now. A gap that looked like $80,000 in favour of open-weight a year ago may look very different once governance and production-readiness are priced in.
The quality gap closing is real, and it's good news for the industry generally. But it doesn't automatically mean your business should switch. Model choice is one line item in a much bigger decision about who carries the operational risk. If you want a second opinion on where your stack actually sits, book a brainstorm and we'll walk through it against your specific constraints. Book a brainstorm and we'll map it out together.



