We get this question roughly once a quarter, usually right after a new open-weight model tops a benchmark leaderboard: why hasn't Automata AI switched? It's a fair question from a client watching headlines about DeepSeek or Qwen matching frontier model scores at a fraction of the training cost. Here's the honest answer, from a Sydney consultancy that has to make production systems work reliably for Australian businesses, not just win an argument on a benchmark table.
Benchmarks aren't the job
A model that scores well on a coding benchmark or a reasoning leaderboard is answering a different question than the one our clients actually ask us. Our clients ask whether an AI system will correctly parse a messy invoice from a Perth supplier, hold context across a 40-page contract without losing track of a clause, and do it reliably enough that a project manager can trust the output without checking every line. Benchmark scores tell you about performance on a fixed test set. They tell you very little about how a model behaves on the specific, messy, real documents an Australian business actually produces.
We test every serious model release against our own internal suite of real client-style tasks, not published benchmarks, and Claude has consistently been the model that handles long, messy, ambiguous real-world documents most reliably. That's the metric that matters when we're the ones on the hook if a system misreads a contract clause.
The infrastructure question nobody puts on the leaderboard
Support and accountability: when something breaks in production, we need a vendor relationship, not a GitHub issue thread.
Data handling terms: Australian clients in regulated industries need clear, reviewable agreements, not a self-hosted model where the compliance burden sits entirely with us.
Consistency across releases: an open-weight model swap can silently change output behaviour; a stable subscription relationship with clear versioning is easier to build a business on.
Running an open-weight model in production means Automata AI becomes the infrastructure team: GPU hosting, fine-tuning, security patching, and the liability if the model behaves unexpectedly on a client's data. That's a real cost, roughly $45,000 to $60,000 a year in hosting and engineering time for a small deployment, and it's a cost our clients would ultimately pay for through us, for a capability Claude already delivers as a subscription.
Where we're honest about the trade-off
This isn't a claim that Claude is best at literally everything, or that open-weight models have no place anywhere. A research team with genuine reason to fine-tune on proprietary data and the engineering capacity to run that safely might get real value from a private deployment. For the client-facing production systems we build, contract review tools, AP automation, agent workflows that touch real business data, reliability and accountability matter more than shaving a percentage point off a benchmark score, and that's the trade-off we make deliberately, not by default.
We also don't treat this as a permanent, closed decision. If a model genuinely outperforms Claude on the tasks our clients actually need, on real documents, with the support and data handling terms an Australian business can stand behind, that's worth revisiting. It hasn't happened yet across the releases we've tested.
A concrete example from a recent build
A Melbourne accounting client asked us to compare Claude against an open-weight alternative for a tax-document extraction pipeline earlier this year. The open model matched Claude on a clean, well-formatted sample PDF. It fell over on the client's actual documents: scanned receipts, inconsistent formatting from three different point-of-sale systems, handwritten annotations, in ways that would have meant a staff member manually re-checking every extraction anyway, defeating the purpose of automating the work. Claude handled the same messy real-world batch with a meaningfully lower error rate, which is the only comparison that mattered for that client's $120,000-a-year bookkeeping workload.
That is the pattern we see repeatedly. Clean benchmark, close scores. Real Australian business documents, a clear gap. We would rather build our reputation on the second test than the first, because that is the one our clients are actually paying us to get right.
What this means for a client evaluating us
If you are comparing consultancies and one of them is pitching whatever model topped last week's leaderboard, ask them the questions we ask ourselves: how does it perform on your actual documents, not a public test set, who is accountable when it gets something wrong, and what does the data handling agreement actually say. Those answers matter more than a benchmark score, and they are the reason Automata AI builds on Claude, backed by Anthropic's enterprise agreements, rather than chasing the open-weight model of the month.



