Blog

224 Open Models: The Real Cost of Leaderboard Shopping

September 2026 · 6 min read · ROI & Business Case

Notebook sketch: a wide grid of model boxes narrowing through a funnel past three ticks to one shortlisted model
← Back to all posts

Picking an AI model used to mean choosing between three or four names you already knew. As at September 2026 it means choosing from a field big enough that nobody reads all of it. That is not more choice in any useful sense. It is a search problem with a price attached, and the price lands on whoever does the searching.

How many open-weight AI models are there in 2026?

One widely used public tracker lists 224 canonical open-weight models as at September 2026, inside a broader index of 398 models spanning the major labs and inference providers. Counts differ between trackers because each sets its own bar for what counts as a distinct release, so treat any single figure as an order of magnitude rather than a census. The direction is the part that matters for a buyer: the field is now far larger than any one business can assess before it has to decide.

The benchmarks multiply the problem

Models are not scored on one yardstick. The same public trackers rank on GPQA Diamond, SWE-Bench Verified, MMLU-Pro, AIME, LiveCodeBench and MMMU, among others, and each measures something different. A model can lead on code generation and sit mid-field on document reasoning. So the comparison space is not the model count. It is the model count multiplied by the number of scores you decide to care about, which is how a two-week evaluation becomes a two-month one.

  • A coding score tells you nothing about how a model handles a scanned supplier invoice

  • Reasoning benchmarks run on curated problems, not on your own messy inputs

  • Point releases land every four to eight weeks, so an evaluation more than a quarter old is already stale

  • Trackers disagree on which releases count, so two leaderboards can order the same two models differently

What the comparison actually costs

Here is what we see when a mid-market business in Sydney or Melbourne runs this evaluation properly and without help. The figures below are our own estimates for a typical build rather than survey data, and they assume internal staff time costed at a fully loaded rate.

Automata AI estimates of what a do-it-yourself model evaluation costs a mid-market Australian business
StageWhat it involvesEstimate (AUD)
Reading the methodologyWorking out which published scores map to your use case$1,500 to $2,500
Building a test setAssembling 50 to 100 of your own real examples with known answers$2,000 to $4,000
Running the bake-offTwo or three shortlisted models scored against that test set$3,000 to $6,000
Switching after a wrong pickRe-prompting, rework and fresh governance sign-off$8,000 to $20,000

The last row is the one that hurts. A model chosen off a chart rather than off your own documents tends to fail quietly: acceptable on the demo, wrong often enough in production that people quietly stop using it. By then the integration is built, the training is done, and the cost of changing course is no longer a licence fee.

The three checks we run instead

We do not run forty benchmarks for a client. We run three checks, because these are the three that map to money.

  • Does it get your own real examples right, scored against a set you hold back rather than one the model has likely seen

  • What does it cost per 1,000 requests at the volume you actually expect, not the volume in the pricing page example

  • What happens when it is wrong, how often that happens, and who carries the consequence when it does

Why cost per 1,000 requests, not cost per token

Token prices compare badly across models because they hide how many tokens each one spends on the same job. A cheaper model that reasons at length can cost more per finished task than a dearer one that answers directly. Pricing the unit of work you care about, a processed invoice or a drafted response, makes the comparison honest. Add the retries you will need, because those are real requests and they appear on the bill.

What not to conclude from this

Three misreadings worth heading off.

  • Not that benchmarks are worthless. They are a reasonable first screen for ruling models out, and a poor basis for ruling one in

  • Not that a bigger open-weight field means open weights always win. For most Australian businesses the shortlist still ends up with Claude on it, usually at the top

  • Not that you should skip evaluation. The argument here is for a smaller evaluation you actually finish, not for trusting a vendor slide

Where to start this week

Pick one process. Write down 50 real examples from it, with the answers you would accept. That single artefact settles a model choice faster than any leaderboard, and it keeps working when the next release lands six weeks from now. If the process touches personal information, note where that data would sit under each option, because the Privacy Act question is much easier to answer before the build than after it.

We have written before on reading benchmark claims sceptically, building your own hallucination eval harness, and what a leaderboard-topping score should and should not change. If you would rather not build the filter yourself, our assessment runs it against your own processes, the ROI calculator sizes the payoff first, and you can book a time to talk it through.

FAQ

Frequently asked questions

How many open-weight AI models are available in 2026?

One widely used public tracker listed 224 canonical open-weight models as at September 2026, within a broader index of 398 models, though counts vary because each tracker sets its own rule for what counts as a distinct release.

Which AI benchmark should I trust when comparing models?

None of them on its own. Benchmarks such as SWE-Bench Verified or GPQA Diamond measure narrow capabilities on curated problems, so use them to shortlist candidates and then score those candidates against your own real examples.

How often do AI model leaderboards change?

Point releases arrive every four to eight weeks across the major labs, which means a leaderboard position more than a quarter old should be read as history rather than as current guidance for a buying decision.

Is a higher benchmark score worth paying more for?

Only if the extra score shows up on your own workload. Price the unit of work rather than the token, include the retries you will need, and compare finished-task cost across two or three candidates.

Can a small business evaluate AI models without an engineering team?

Yes, provided the evaluation stays small. Fifty real examples with agreed correct answers, run through two or three candidates, will settle most decisions without reading a single benchmark methodology paper.

Ready to move from AI pilot to production?

We help mid-market Australian businesses deploy AI automations that actually reach production and deliver measurable ROI.