Blog

What OpenAI's New AI Scorecard Gets Right, and How to Apply It to Claude

August 2026 · 7 min read · ROI & Business Case

A hand-drawn gauge with a needle pointing toward a rising line ending in a terracotta dot, representing cost-per-outcome improving as AI usage scales
← Back to all posts

OpenAI's leadership recently proposed a new way to judge whether AI spending is actually working. Instead of counting seats, users, or raw token totals, the pitch is to measure useful intelligence per dollar: does the AI complete meaningful work, what does each successful task actually cost once retries and review are factored in, is the result dependable, and does the economics improve as usage scales.

It's a solid framework, and it doesn't belong to any one vendor. It applies to Claude exactly as well as it applies to ChatGPT, and for an Australian business owner trying to work out whether an AI project is worth the spend, it's a test worth borrowing regardless of which model sits behind the work.

Why token price is the wrong number to watch

Most vendor comparisons, and most internal budget conversations about AI, still centre on price per token or price per seat. That's the wrong unit to anchor on. The real cost of an AI workflow is the cost of a successfully completed task: the model's fee, plus every retry, plus the time a person spends reviewing or fixing the output before it's usable.

A cheaper model that needs constant supervision isn't a cheaper workflow. If a fractionally-priced model produces an invoice-matching result that a finance team has to re-check line by line, the true cost per completed task can end up higher than a pricier model that gets it right the first time. Businesses running early Claude pilots in Sydney and Melbourne have found this out the hard way: the headline price of the API call was never the number that actually mattered to the bottom line.

A practical checklist for Australian business owners

Applied to a small or mid-size Australian business already running Claude, or weighing it up, the four questions in OpenAI's scorecard translate into a straightforward pre-engagement checklist. Before signing off on any AI automation project, it's worth asking:

  • What specific, completed piece of work is this AI meant to produce, not just what capability does it demonstrate in a demo?

  • What does it actually cost per successful outcome, including the time your team spends checking or fixing the result?

  • How dependable is that outcome without someone checking it every single time?

  • Does the economics improve as usage grows, or does the review overhead grow just as fast as the volume?

None of these questions require a data science team to answer them. They require someone in the business, often the owner, to sit down with whoever is running the AI project and ask for the actual numbers, not the demo.

Where this fits into how AI projects should actually be scoped

This is close to how a well-run AI engagement should already be structured, and it's the reasoning behind scoping most AI automation work in three separate stages rather than one open-ended build.

Stage one is a fixed-cost audit, typically $3,500 to $6,000 AUD, to work out what's actually worth automating in the business and what it should cost per outcome once it's running. Stage two is a defined build phase, priced against that audited scope rather than an open-ended hourly rate, so the business knows what it is paying before the work starts. Stage three is an ongoing retainer, and it only starts once the workflow is live and the cost-per-outcome is a known number rather than an assumption.

That sequencing matters because it forces the ROI conversation to happen before the money is spent, not after. A business that skips the audit and goes straight to a build is, in effect, running OpenAI's scorecard blind: no baseline for what a successful task should cost, no way to tell if the retries are eating the margin, and no clean way to check whether the return is actually scaling.

A worked example

Take a mid-size Brisbane logistics business processing supplier invoices by hand. An AI pilot that automates data entry might look impressive on a demo screen, extracting fields from a PDF in seconds. But if the extracted data still needs a bookkeeper to check every line before it hits the accounting system, the true cost per invoice processed may barely beat the manual process. Run the same pilot through the four-question checklist first and the answer becomes clear well before a five-figure build is committed.

What this means for a Claude investment

OpenAI arriving at this conclusion from the vendor side is a useful confirmation for any Australian business owner who has been quietly sceptical of AI ROI claims that only ever cite raw capability or price per token. The scorecard is right, and the businesses already getting real value from Claude in Melbourne, Brisbane, and Sydney are mostly using some version of it, whether they've put a name to it or not.

The practical takeaway isn't to switch vendors or chase the newest benchmark. It's to insist on the same four questions before committing budget to any AI project, Claude included: what task gets completed, what it actually costs once the review overhead is counted, how dependable it is unsupervised, and whether the economics hold up as usage grows. A framework only earns its keep once it changes what gets measured before the invoice arrives, rather than after the budget is already spent.

If you're weighing up an AI project and want the cost-per-outcome question answered before you commit budget, book a session to talk through what a fixed-cost audit would look like for your business.

Ready to move from AI pilot to production?

We help mid-market Australian businesses deploy AI automations that actually reach production and deliver measurable ROI.