Claude, GPT, Kimi, Llama: every few months a new open-weight model tops a benchmark leaderboard and a wall of vendor charts lands with it. When Kimi K3 launched, its maker published numbers putting it ahead of Claude Opus 4.8 and GPT-5.5 on agentic benchmark suites. Independent testers then ran their own pass over the same model and found a 51 per cent hallucination rate on a factual category that never showed up in the vendor's charts at all. That gap isn't evidence of dishonesty. Vendors publish the evaluations they chose to run, on the tasks they chose to test. The distance between those published numbers and what your business actually asks the model to do is a gap only you can close, and closing it takes a proper evaluation harness, not a glance at a leaderboard before you sign off on a model.
This matters most right now because the incentive to switch off Claude and onto a cheaper open-weight model has never been stronger. Every launch comes with a chart showing the new model beating Claude on cost, or on a benchmark suite, or both. Some of those claims hold up under your own testing. Plenty don't, and the only way to know which is which before a client sees the output is to build the test yourself.
Why the public benchmark misses your risk
Public benchmarks measure general capability against public data. The risk your business is actually carrying is narrower than that, and it rarely overlaps with what a vendor chose to publish.
Does the model invent a clause number when summarising a contract
Does it state a figure that isn't present in the source document
Does it answer confidently when the honest response is that the information isn't available
Does it get Australian specifics right, including GST treatment, state jurisdiction and local terminology
A model can score well on SWE-bench and fail all four of those, and nothing on the launch page will warn you in advance. A leaderboard ranking is a poor substitute for testing against your own documents, your own workflow, and your own definition of an acceptable answer. This is doubly true for anything touching Australian regulatory or tax detail, where a benchmark built on US or global data has no reason to have been tested on local rules at all.
A harness you can build in a week
You don't need an evaluation platform or a data science team to get a working answer. You need a repeatable test set and a scoring rule you're willing to stand behind.
Collect 100 to 200 real documents from your own workflow, de-identified where required
Write the correct answer for each one, or the correct refusal where the source doesn't contain the answer
Include 20 to 30 adversarial cases where the answer is genuinely absent, because unsupported answers are exactly where hallucination shows up
Score with exact match where you can, use a second model as grader where the output is prose, and have a human spot-check 10 per cent of the results
Run the harness against every model you're considering, not just the one you're leaning toward. Run it again every time you change a prompt or upgrade a version, because a harness result from three months ago tells you nothing about the model sitting in production today. Store the results with dates attached. That record is what lets you justify a model change with evidence instead of a hunch, and it's the first thing worth showing a client, a board, or an auditor who asks why you trust the system.
What a pass actually looks like
Picture a Sydney bookkeeping firm testing whether a new open-weight model can summarise supplier contracts alongside Claude. The harness throws 150 real contracts at both models, including 25 where the renewal clause has been deliberately removed. Claude flags all 25 as missing information. The open-weight model invents a plausible-sounding renewal date for 14 of them. That single adversarial batch tells the firm more about which model to trust in front of a client than any public benchmark score would, and it took an afternoon to build, not a research team.
Common ways teams get this wrong
Building a harness is easy to do badly. The failure modes we see most often in Australian businesses setting one up for the first time are predictable, and worth checking against before you trust your own results.
Testing only on tidy, well-formatted documents that don't resemble the messy scans and half-complete forms your workflow actually produces
Letting the same person who wrote the prompt also grade the output, which quietly biases every close call toward a pass
Running the harness once at launch and never touching it again, even after the prompt changes or the model gets upgraded under the hood
Treating a strong public benchmark score as validation on its own, when it's marketing copy from the vendor that built the model
Fix those four and the harness earns its keep. Skip them and you've built a checkbox exercise that looks rigorous without actually reducing risk.
Set the threshold before you see the numbers
The order matters here. Decide what an acceptable failure rate looks like before you run the test, not after you've seen how the model performed. For most Australian business workflows we treat an unsupported-claim rate above 2 per cent as unfit for anything client-facing without human review, and anything above 5 per cent as unfit for use at all, full stop. Pick your own numbers if two and five don't suit your risk tolerance, but pick them in writing and pick them first. A threshold set after the results are in isn't a threshold, it's a justification dressed up as one.
What this actually costs
Building a harness like this is a few days of someone's time, call it $3,000 to $6,000 depending on how messy your source documents are and how much de-identification they need before anyone outside your business looks at them. The cost of skipping it isn't zero, it's just deferred. It shows up the day a hallucinated figure lands inside a client deliverable, a board paper, or a number reported to a regulator, and by then the cost is a lot higher than a harness would ever have been. Every serious AI deployment we build for Australian businesses out of Sydney has one of these harnesses attached, whichever model ends up doing the work underneath.
Make evaluation part of how you choose a model, not an afterthought
None of this is an argument against open-weight models. Some of them are genuinely strong, and some will beat Claude on a given benchmark for a given task. The argument is against trusting a vendor's chart as a substitute for your own test. Claude, Kimi, GPT, Llama: run all of them through the same harness, on the same documents, against the same threshold, and let the numbers you generated yourself make the call. That's the only benchmark result that actually describes the risk your business is carrying, and it's the only one worth putting in front of a client or a regulator.
If you want help building an evaluation set for your own workflow, or want a second opinion on a harness you've already started, get in touch and we'll walk through it together.



