Most small businesses running AI in 2026 are overpaying for about a third of their workload. Not because their provider is expensive, but because they send simple jobs, like tagging an email or reading a line item off a standard invoice, to a model built for far harder work.
The reason this has become fixable is that small open-weight models got good. Models such as Qwen3-8B, GLM-4-9B-0414 and Llama 3.1 8B Instruct have closed enough of the capability gap that they now handle a meaningful share of tasks that used to be routed to much larger models, at a fraction of the compute cost.
Can an 8B model replace a 70B model for small business tasks?
For narrow, repetitive, low-stakes tasks, yes. An 8-billion-parameter model can reliably sort inbound emails, tag support tickets or extract fields from a consistent document template, and it does so far more cheaply than a 70B-class model. For anything needing multi-step reasoning, tool use or judgment with real consequences, no. The practical question is not which model wins a leaderboard but which of your tasks are small enough for a small model.
Part of what changed is deployment. Quantisation methods like GPTQ, AWQ and GGUF have matured to the point where an 8B model runs comfortably on a single consumer-grade GPU, something that was impractical eighteen months ago. If you want the mechanics, our guide to quantisation and fitting models on small GPUs covers them.
A task-by-task routing check
The table below is the rough sorting we use when mapping a client's workload. It is a starting point, not a verdict, and every row should be tested on your own data.
| Task | Small 8B model | Keep on Claude | Why |
|---|---|---|---|
| Sorting inbound emails into queues | Yes | No | Narrow labels, cheap to correct |
| Tagging support tickets | Yes | No | Repetitive, low-stakes |
| Flagging invoices for review | Yes | No | A human reviews the flag anyway |
| Line items from a standard supplier template | Yes | No | Consistent format, checkable output |
| Drafting a client-facing proposal | No | Yes | Judgment and tone carry real risk |
| Handling a complex customer complaint | No | Yes | Multi-step reasoning, reputational stakes |
| Multi-step agent work with tools | No | Yes | Small models are unreliable at tool use |
The rule of thumb underneath the table is simple: route a task to a small model only if getting it wrong costs less than the model saves. In practice that rules out most customer-facing and financial work.
What the saving actually looks like
We have restructured a handful of client AI stacks this year to route simple, high-volume tasks to a cheap small model and keep everything judgment-heavy on Claude. For a business processing a few thousand simple classification calls a month, this typically cuts that slice of the AI bill from several hundred dollars to under $50 a month.
That is a real saving, and it is also a narrow one. It only applies to the part of the workload that was paying for capability it never used. The parts of the workflow that genuinely need reasoning stay on Claude, where the quality bar matters, and their cost does not change.
So the honest pitch for a Brisbane or Sydney business with a lean AI budget is not that small models slash your bill. It is that they remove waste from one slice of it, and finding that slice is worth an afternoon.
Three traps when moving to a small model
Assuming the saving applies everywhere. It applies to the boring, repetitive, low-stakes third of most workflows, not the whole stack.
Forgetting the running cost. A self-hosted 8B model still needs someone to deploy, patch and monitor it. If your volume is modest, a small managed model may be the simpler route.
Skipping the test set. Build 50 to 100 real examples of each task before switching, and compare error rates, not just price.
On the second point, it is worth remembering that the small-model idea is not exclusive to open weights. Claude's own smaller tier exists for the same reason, and our write-ups on when Haiku beats Opus on cost and where small models win in production cover that route. For some teams, a managed small model plus one vendor is cheaper overall than running a second stack.
Keeping one stack manageable
Routing work across two models adds a moving part. The routing logic has to be written down, monitored and owned by someone, or it drifts. We generally suggest a single router with a clear rule per task type and a fallback to Claude when the small model's confidence is low. That keeps the one model, many jobs discipline intact while still sending the easy work somewhere cheaper.
If you want us to map which parts of your workflow are candidates for a small model and which need to stay on Claude, try the ROI calculator for a first estimate, or book a session with us and bring a month of your AI usage.



