Blog

Token Minimisation for Open-Weight LLM Deployment

September 2026 · 6 min read · Technical

Five long lines of text passing through a funnel and emerging as one short highlighted block
← Back to all posts

Picture two cost reports landing on a Melbourne operations manager's desk in September 2026. One says the open-weight model is a third of the price of Claude. The other, built by a different engineer, says the gap is barely worth the migration. Both used the same models. The difference was how many tokens each engineer decided to send.

That gap is the story of the year for anyone running language models at volume. Through 2025, the default instinct was to give the model more: bigger context windows, longer prompts, the top 20 retrieved documents on every call just in case. Practitioners now describe 2026 as the year that discipline flipped. Token minimisation, deliberately trimming what gets sent to a model, has become the standard pattern for teams running open-weight models at scale.

What is token minimisation in LLM deployment?

Token minimisation is the practice of sending a language model only the context it needs to finish a task, and nothing more. It means trimming retrieved documents to the smallest set that answers the question, caching instructions that never change instead of resending them, and treating every token in a prompt as a cost line. Because both price and latency scale with what you send, a leaner prompt is cheaper and faster on any model, open-weight or managed.

The shift matters most for self-hosted and open-weight deployments because there the cost of a bloated prompt is paid twice: once in GPU time and again in slower responses that tie up capacity other requests are waiting for. On a managed API the cost shows up on the invoice. On your own hardware it shows up as a server you need to buy sooner.

The habits that replaced 'give it everything'

None of these are exotic. They are the kind of changes an engineer can make in a week, and they apply whichever model sits underneath:

  • Retrieval trimmed to the smallest set of chunks that actually answers the question, rather than returning the top 20 to be safe.

  • Stable system prompts cached and reused across calls instead of resending the same boilerplate instructions on every request.

  • Old conversation turns summarised or dropped once they stop being relevant, instead of dragging the full history into every call.

  • Output length capped for tasks that only need a label, a number or a short answer.

The second habit is exactly what Claude's prompt caching and its cheaper cache-read pricing are built to reward, which is why Claude users have been nudged toward this discipline for a while. We have written up what caching taught us inside Claude Code and four token budgeting patterns for Claude agents if you want the implementation detail.

Why token minimisation breaks most cost comparisons

Here is the mistake we see most often in Australian businesses weighing Claude against an open-weight alternative. The open-weight model gets run with a minimised, hand-tuned prompt, because someone had to work hard to get it performing. Claude gets run with the bloated prompt carried over from an early prototype, because it already worked and nobody touched it.

That comparison is not measuring the models. It is measuring engineering effort, and it consistently overstates the open-weight model's cost advantage. The more work went into one side's prompt, the more lopsided the result.

An honest side-by-side has to hold the plumbing constant. The table below is the checklist we use before any number goes in front of a client.

Variables to hold constant in a Claude versus open-weight cost comparison
VariableCommon unfair setupFair setup
Retrieval strategyTuned top 3 chunks on one side, top 20 on the otherSame retriever, same chunk size, same count
System promptTrimmed on one side, prototype-era on the otherSame structure, both minimised
CachingEnabled on one side onlyEnabled wherever the platform supports it
Success criterionDifferent definitions of a finished taskOne written definition applied to both
Volume assumptionPeak month on one side, average on the otherSame monthly call volume

What an honest benchmark costs, and what it usually shows

Rebuilding the prompt architecture for both sides of a comparison is real work. For a mid-sized workload, a properly optimised evaluation typically takes our team 15 to 25 hours. That is not free, and it is fair to ask whether it is worth it.

The answer is usually yes, because the resulting cost delta is routinely smaller than the unoptimised comparison suggested, sometimes by more than $1,000 a month at moderate volume. Over a year that is $12,000 or more of apparent saving that was never going to materialise, set against a migration that carries its own rework, retraining and new failure modes. Getting the benchmark right before switching is almost always cheaper than finding out after.

It also protects you from a quieter problem. Benchmarks published by model vendors are run by teams with every incentive to tune their own side hardest. The same logic applies there, which is why we suggest a buyer's filter for open-weight benchmark claims before any headline number moves your budget.

Where minimisation goes too far

Trimming has a failure mode, and it is worth naming. Cut retrieval too hard and the model stops seeing the paragraph that held the answer. It then produces something plausible and wrong, which costs far more than the tokens you saved. For work that touches customer records, financial figures or anything reportable under the Privacy Act, that trade is rarely worth it.

The fix is to minimise against a measured success rate, not against the token count alone. Keep a small set of real test cases, trim until accuracy starts to drop, then step back one notch. That turns token minimisation from a cost-cutting exercise into an engineering one, and it gives you a number you can defend to a finance team.

If someone has shown you a cost comparison between Claude and an open-weight model, ask whether both sides were optimised the same way before you trust the number. Our ROI calculator is a reasonable first pass, and if you want the comparison rebuilt properly, book a session with us and bring the prompts you are running today.

FAQ

Frequently asked questions

Does sending more context to an LLM improve answer quality?

Only up to a point. Past the chunks that actually hold the answer, extra context adds cost and latency and can distract the model, so accuracy often plateaus or drops while the bill keeps climbing.

How does prompt caching reduce token costs?

Caching stores a stable prefix such as a system prompt so repeat calls read it at a lower rate instead of paying full price to resend it. It rewards prompts that keep fixed instructions at the start and variable content at the end.

Why are Claude versus open-weight cost comparisons often wrong?

They frequently compare a hand-tuned, minimised prompt on the open-weight side with an unoptimised prototype prompt on the Claude side, so the result measures engineering effort rather than the models themselves.

How much retrieval context should a RAG system send per query?

There is no fixed number. Start with the smallest set of chunks that answers a sample of real questions correctly, then measure accuracy as you add or remove chunks and settle just above the point where it drops.

Ready to move from AI pilot to production?

We help mid-market Australian businesses deploy AI automations that actually reach production and deliver measurable ROI.