Blog

Claude ROI Tracking: Usage Is Not Value, So Measure Both

October 2026 · ROI & Business Case

Hand-drawn balance scale weighing a bar chart of Claude usage against terracotta coins of business value
← Back to all posts

On 16 September 2026, OpenAI added analytics to the ChatGPT Admin Console that put usage, cost, task insights and outcome metrics in one place. It is a sensible release, and it names a problem every Claude customer has too: the board asks what the AI spend returned, and the only thing on hand is a chart of active users.

We run Claude rollouts for Australian businesses, so this is the question we get in month three of almost every engagement. The short version of our answer: a usage dashboard tells you whether people turned up. It does not tell you whether the business got anything. You need both numbers, and the second one never comes out of an admin console on its own.

What did OpenAI actually ship?

According to OpenAI's announcement, the new console views cover ChatGPT Work and Codex:

  • A Usage view showing active users, credits and tokens.

  • An Insights task classifier that groups messages into use cases.

  • Task-detail breakdowns by model, reasoning level and speed, meant to guide training.

  • A plugin leaderboard and a Skills view.

  • A Codex Outcomes view tracking Codex's share of merged commits and lines of code against review time, defects and rework.

  • An Admin plugin and API for custom dashboards and leadership decks.

OpenAI also published a worked example, labelled illustrative: a 20-seller sales team saving 3 hours per account brief produces $207,000 of annual capacity value against a $60,000 first-year cost, or 245% ROI. Customer proof points sit beside it, including 1Password reporting 553% ROI through Codex.

How do you measure Claude ROI in an Australian business?

You measure Claude ROI by pairing two records: what the admin and usage analytics say people did with Claude, and a baseline you took before the rollout of how long the same work used to take and how often it went wrong. The return is the gap between those two, priced at a loaded hourly rate and discounted for hours that were saved but never put to other work.

The usage half is the easy half. We covered how to read it in Claude usage analytics for team adoption. The baseline half is where most pilots fall over, because nobody timed the old process before changing it. If you have not started yet, time ten real examples of each task this week. If you already started, a short time-ledger method can rebuild a usable baseline from people still doing the task the old way.

The same calculation with Australian numbers

Here is OpenAI's format applied to an illustrative mid-market case: a Brisbane commercial builder with 12 estimators using Claude to prepare tender comparisons. Every figure below is illustrative, not a client result.

Illustrative first-year Claude ROI for a 12-person estimating team, in AUD
LineWorkingValue
Hours saved per person2.5 hours a week over 46 working weeks115 hours
Gross capacity value12 people x 115 hours x $85 loaded rate$117,300
Realisation discountHalf the saved hours go to billable or tender work$58,650
First-year costSeats, setup, training and review time$38,000
Headline ROI($117,300 - $38,000) / $38,000209%
Realised ROI($58,650 - $38,000) / $38,00054%

Both ROI rows are arithmetically correct. Only one of them survives a CFO. The 209% figure assumes every saved hour turns into value, which is the assumption behind most vendor examples, OpenAI's included. The 54% figure asks where the hours went. In this case the builder submitted more tenders with the same team, and that is what makes half the hours real.

Three things a dashboard cannot tell you

  • Whether the work was any good. Faster tender comparisons that miss an exclusion clause are a cost, so track error and rework rates beside time.

  • Where the hours went. Saved time that dissolves into the week is not a return. Name the work that absorbed it.

  • What you would have spent anyway. If a contractor or a new hire was the alternative, the avoided cost belongs in the model. If nothing was, it does not.

The first point deserves its own tracking. We set out the measures in AI ROI beyond hours saved: error rates, cycle times and customer scores sit beside time, because a regulator or a client will judge the output, not the speed. For an APRA-regulated firm, a rework rate is often the more persuasive number.

What not to conclude from the comparison

It would be easy to read OpenAI's release as a feature gap to score. That misses the point. A task classifier and an outcomes view are useful, and Claude's admin analytics cover similar ground on usage and spend. Neither product knows your loaded rates, your baseline or what your people did with the time. That part is your work, whichever assistant you run.

So if you are a Claude customer, there is nothing to migrate and nothing to wait for. Pick two workflows, take a baseline, keep a ledger for six weeks and run the table above with your own figures. Report the realised number, and show the headline one next to it so nobody thinks you hid it.

Where to start

Our ROI calculator runs the same build-up with your team size and rates. If you would rather work through it with someone who has done it for other Australian teams, book a short conversation and bring one workflow you suspect is paying off. We will help you prove it or drop it.

FAQ

Frequently asked questions

Does Claude have an ROI dashboard?

Claude's admin analytics show usage and spend across your team, which covers the cost side and the adoption side. The value side still needs a baseline and loaded rates that only your business can supply.

What is a good ROI for a Claude rollout?

There is no universal benchmark. A realised first-year return that is clearly positive after discounting unredeployed hours is a sound result, and it matters more than a large headline percentage built on gross hours.

How long before you can measure Claude ROI?

Around six weeks of ledger data on one or two workflows is usually enough for a first defensible number, provided you timed the old process before the rollout or can reconstruct it.

Is hours saved the right metric for AI ROI?

Hours saved is a starting point only. Pair it with error and rework rates and cycle times, then discount for hours that were saved but not redeployed into billable or revenue work.

Ready to move from AI pilot to production?

We help mid-market Australian businesses deploy AI automations that actually reach production and deliver measurable ROI.