Opus 4.5 was the model that made a lot of Australian businesses stop treating Claude as a writing tool and start treating it as something that could run a process. The headline capability shifts were in sustained multi-step work, tool use, and holding a large amount of context without losing the thread. If you are deciding which model to point at which job, that framing is more useful than any benchmark table.
What actually changed for business work
The improvements that mattered commercially were not about writing better prose. They were about reliability across long chains of work, which is precisely what separates a demo from something you let near a real process.
Longer coherent task chains: fewer cases where step seven forgets what step two decided
More dependable tool use, meaning it calls the right system with the right arguments rather than confidently improvising
Better handling of large context, so you can hand it a full contract or a year of transactions rather than excerpts
More willingness to say it does not know, which sounds minor and is the single most valuable trait in a business setting
That last point deserves more weight than it usually gets. A model that invents a plausible answer under uncertainty is unusable for anything with consequences, because you cannot tell the invented answers from the correct ones without checking every single one. Reliability under uncertainty is what makes delegation possible.
Where it helps and where it is overkill
The mistake we see most often is running everything on the largest available model because it feels safer. It is not safer, it is just slower and more expensive. Model choice should follow the shape of the task, not the anxiety of the person configuring it.
High-judgement, low-volume work: contract review, proposal drafting, analysing a messy dataset. Worth the strongest model available
High-volume, low-judgement work: classifying enquiries, extracting fields, tagging records. A smaller, cheaper model usually matches it
Long agentic chains where a mid-sequence error is expensive: worth the stronger model purely for the reduced failure rate
Anything customer-facing without a review step: the model choice matters less than the fact that you removed the review step
A useful test is to run the same fifty real examples through two models and compare the output side by side. Most teams discover the cheaper model is indistinguishable on the bulk of their work and clearly worse on a narrow slice. That slice is the only place the expensive model is earning anything.
The cost side
Running a frontier model on every task in a business adds up faster than people expect, because the cost is per token rather than per seat. A mid-sized Australian firm putting genuine daily volume through a top-tier model can find itself at $2,000 to $6,000 a month without anything unusual happening.
Splitting traffic so the expensive model only handles work that needs it routinely cuts that by more than half, and nobody notices the difference in output quality. This is the single highest-return piece of housekeeping in most AI deployments, and it is almost always left undone because nobody owns the bill. If your AI spend has no named owner, it will keep growing until someone in finance notices.
Model versions move faster than your processes
Opus 4.5 is no longer the newest model in the line, and by the time you have embedded any model into a workflow there will be a newer one. That is an argument for building processes that name the job rather than the model: your invoice-triage workflow should specify what good output looks like, not which version produced it.
Teams that hard-code a model name into fifty scripts spend a weekend on every upgrade. Teams that route through a single configurable setting spend ten minutes. The difference costs nothing to get right at the start and is tedious to retrofit, which is the usual shape of infrastructure decisions.
How to decide without a benchmark obsession
Public benchmarks tell you almost nothing about whether a model will do your job well, because your job is not in the benchmark. The businesses that choose well build a small private evaluation: twenty or thirty real tasks with known good answers, run against any model they are considering.
It takes a day to build and it answers the question permanently, including for every future model release. Compared with reading comparison articles written by people who have never seen your data, it is a bargain.
What not to conclude
A better model does not fix a badly defined process, and it does not remove the need for a human to check work that carries consequences. If your current results are inconsistent, upgrading the model is the most expensive way to find out that the problem was your inputs.
Treat model selection as a cost and reliability decision, not a status one. The right question is never "are we on the best model", it is "is any task running on a model that is wrong for it in either direction". In most Sydney businesses we look at, the answer is yes in both directions at once.
If you want a read on which of your workloads are on the wrong model, book a short call and we will look at where the spend and the errors actually sit.



