Picking an AI model used to mean choosing between three or four names you already knew. As at September 2026 it means choosing from a field big enough that nobody reads all of it. That is not more choice in any useful sense. It is a search problem with a price attached, and the price lands on whoever does the searching.
How many open-weight AI models are there in 2026?
One widely used public tracker lists 224 canonical open-weight models as at September 2026, inside a broader index of 398 models spanning the major labs and inference providers. Counts differ between trackers because each sets its own bar for what counts as a distinct release, so treat any single figure as an order of magnitude rather than a census. The direction is the part that matters for a buyer: the field is now far larger than any one business can assess before it has to decide.
The benchmarks multiply the problem
Models are not scored on one yardstick. The same public trackers rank on GPQA Diamond, SWE-Bench Verified, MMLU-Pro, AIME, LiveCodeBench and MMMU, among others, and each measures something different. A model can lead on code generation and sit mid-field on document reasoning. So the comparison space is not the model count. It is the model count multiplied by the number of scores you decide to care about, which is how a two-week evaluation becomes a two-month one.
A coding score tells you nothing about how a model handles a scanned supplier invoice
Reasoning benchmarks run on curated problems, not on your own messy inputs
Point releases land every four to eight weeks, so an evaluation more than a quarter old is already stale
Trackers disagree on which releases count, so two leaderboards can order the same two models differently
What the comparison actually costs
Here is what we see when a mid-market business in Sydney or Melbourne runs this evaluation properly and without help. The figures below are our own estimates for a typical build rather than survey data, and they assume internal staff time costed at a fully loaded rate.
| Stage | What it involves | Estimate (AUD) |
|---|---|---|
| Reading the methodology | Working out which published scores map to your use case | $1,500 to $2,500 |
| Building a test set | Assembling 50 to 100 of your own real examples with known answers | $2,000 to $4,000 |
| Running the bake-off | Two or three shortlisted models scored against that test set | $3,000 to $6,000 |
| Switching after a wrong pick | Re-prompting, rework and fresh governance sign-off | $8,000 to $20,000 |
The last row is the one that hurts. A model chosen off a chart rather than off your own documents tends to fail quietly: acceptable on the demo, wrong often enough in production that people quietly stop using it. By then the integration is built, the training is done, and the cost of changing course is no longer a licence fee.
The three checks we run instead
We do not run forty benchmarks for a client. We run three checks, because these are the three that map to money.
Does it get your own real examples right, scored against a set you hold back rather than one the model has likely seen
What does it cost per 1,000 requests at the volume you actually expect, not the volume in the pricing page example
What happens when it is wrong, how often that happens, and who carries the consequence when it does
Why cost per 1,000 requests, not cost per token
Token prices compare badly across models because they hide how many tokens each one spends on the same job. A cheaper model that reasons at length can cost more per finished task than a dearer one that answers directly. Pricing the unit of work you care about, a processed invoice or a drafted response, makes the comparison honest. Add the retries you will need, because those are real requests and they appear on the bill.
What not to conclude from this
Three misreadings worth heading off.
Not that benchmarks are worthless. They are a reasonable first screen for ruling models out, and a poor basis for ruling one in
Not that a bigger open-weight field means open weights always win. For most Australian businesses the shortlist still ends up with Claude on it, usually at the top
Not that you should skip evaluation. The argument here is for a smaller evaluation you actually finish, not for trusting a vendor slide
Where to start this week
Pick one process. Write down 50 real examples from it, with the answers you would accept. That single artefact settles a model choice faster than any leaderboard, and it keeps working when the next release lands six weeks from now. If the process touches personal information, note where that data would sit under each option, because the Privacy Act question is much easier to answer before the build than after it.
We have written before on reading benchmark claims sceptically, building your own hallucination eval harness, and what a leaderboard-topping score should and should not change. If you would rather not build the filter yourself, our assessment runs it against your own processes, the ROI calculator sizes the payoff first, and you can book a time to talk it through.



