Picture two cost reports landing on a Melbourne operations manager's desk in September 2026. One says the open-weight model is a third of the price of Claude. The other, built by a different engineer, says the gap is barely worth the migration. Both used the same models. The difference was how many tokens each engineer decided to send.
That gap is the story of the year for anyone running language models at volume. Through 2025, the default instinct was to give the model more: bigger context windows, longer prompts, the top 20 retrieved documents on every call just in case. Practitioners now describe 2026 as the year that discipline flipped. Token minimisation, deliberately trimming what gets sent to a model, has become the standard pattern for teams running open-weight models at scale.
What is token minimisation in LLM deployment?
Token minimisation is the practice of sending a language model only the context it needs to finish a task, and nothing more. It means trimming retrieved documents to the smallest set that answers the question, caching instructions that never change instead of resending them, and treating every token in a prompt as a cost line. Because both price and latency scale with what you send, a leaner prompt is cheaper and faster on any model, open-weight or managed.
The shift matters most for self-hosted and open-weight deployments because there the cost of a bloated prompt is paid twice: once in GPU time and again in slower responses that tie up capacity other requests are waiting for. On a managed API the cost shows up on the invoice. On your own hardware it shows up as a server you need to buy sooner.
The habits that replaced 'give it everything'
None of these are exotic. They are the kind of changes an engineer can make in a week, and they apply whichever model sits underneath:
Retrieval trimmed to the smallest set of chunks that actually answers the question, rather than returning the top 20 to be safe.
Stable system prompts cached and reused across calls instead of resending the same boilerplate instructions on every request.
Old conversation turns summarised or dropped once they stop being relevant, instead of dragging the full history into every call.
Output length capped for tasks that only need a label, a number or a short answer.
The second habit is exactly what Claude's prompt caching and its cheaper cache-read pricing are built to reward, which is why Claude users have been nudged toward this discipline for a while. We have written up what caching taught us inside Claude Code and four token budgeting patterns for Claude agents if you want the implementation detail.
Why token minimisation breaks most cost comparisons
Here is the mistake we see most often in Australian businesses weighing Claude against an open-weight alternative. The open-weight model gets run with a minimised, hand-tuned prompt, because someone had to work hard to get it performing. Claude gets run with the bloated prompt carried over from an early prototype, because it already worked and nobody touched it.
That comparison is not measuring the models. It is measuring engineering effort, and it consistently overstates the open-weight model's cost advantage. The more work went into one side's prompt, the more lopsided the result.
An honest side-by-side has to hold the plumbing constant. The table below is the checklist we use before any number goes in front of a client.
| Variable | Common unfair setup | Fair setup |
|---|---|---|
| Retrieval strategy | Tuned top 3 chunks on one side, top 20 on the other | Same retriever, same chunk size, same count |
| System prompt | Trimmed on one side, prototype-era on the other | Same structure, both minimised |
| Caching | Enabled on one side only | Enabled wherever the platform supports it |
| Success criterion | Different definitions of a finished task | One written definition applied to both |
| Volume assumption | Peak month on one side, average on the other | Same monthly call volume |
What an honest benchmark costs, and what it usually shows
Rebuilding the prompt architecture for both sides of a comparison is real work. For a mid-sized workload, a properly optimised evaluation typically takes our team 15 to 25 hours. That is not free, and it is fair to ask whether it is worth it.
The answer is usually yes, because the resulting cost delta is routinely smaller than the unoptimised comparison suggested, sometimes by more than $1,000 a month at moderate volume. Over a year that is $12,000 or more of apparent saving that was never going to materialise, set against a migration that carries its own rework, retraining and new failure modes. Getting the benchmark right before switching is almost always cheaper than finding out after.
It also protects you from a quieter problem. Benchmarks published by model vendors are run by teams with every incentive to tune their own side hardest. The same logic applies there, which is why we suggest a buyer's filter for open-weight benchmark claims before any headline number moves your budget.
Where minimisation goes too far
Trimming has a failure mode, and it is worth naming. Cut retrieval too hard and the model stops seeing the paragraph that held the answer. It then produces something plausible and wrong, which costs far more than the tokens you saved. For work that touches customer records, financial figures or anything reportable under the Privacy Act, that trade is rarely worth it.
The fix is to minimise against a measured success rate, not against the token count alone. Keep a small set of real test cases, trim until accuracy starts to drop, then step back one notch. That turns token minimisation from a cost-cutting exercise into an engineering one, and it gives you a number you can defend to a finance team.
If someone has shown you a cost comparison between Claude and an open-weight model, ask whether both sides were optimised the same way before you trust the number. Our ROI calculator is a reasonable first pass, and if you want the comparison rebuilt properly, book a session with us and bring the prompts you are running today.



