Forty thousand dollars. That is roughly what a multi-GPU inference server costs to provision in Australia, and until recently it was the entry price for running a capable 27-billion-parameter model on your own hardware. The Qwen3.8-27B build released on 2 September 2026 claims to bring that down to a single consumer 24GB card worth $2,500 to $4,000.
The reason is not a bigger or smaller model. It is a different way of doing attention, the step where a transformer works out which earlier words matter to the next one. That is worth understanding properly, because it changes which jobs get cheaper and which do not.
What is linear attention in Qwen3.8-27B?
Linear attention is a replacement for the standard attention mechanism that dense transformers have used since 2017. Standard attention compares every token with every other token, so the work grows with the square of the input length. Linear attention keeps a running summary instead, so compute grows roughly in step with length. For Qwen3.8-27B, released under Apache 2.0 in September 2026, that is what lets inference fit on one 24GB GPU.
In plain terms: double the length of a document under standard attention and the attention work roughly quadruples. Under linear attention it roughly doubles. On a 2,000-word email that difference is invisible. On a 60-page lease, a quarter of call-centre transcripts or a data room of contracts, it is the difference between a job that fits on a desktop card and one that needs a server rack.
Why the 24GB number matters, and what it hides
A 24GB card in the RTX 4090 class is hardware an Australian development shop in Sydney or Melbourne may already have under a desk. Fitting a 27B model on it moves local inference from a capital project into a line item. We have already costed the dense Qwen3.8-27B against a Claude API bill, and that analysis assumed a heavier hardware footprint than this build needs.
Two caveats before anyone orders a card. First, a 27-billion-parameter model stored at full 16-bit precision needs far more than 24GB for its weights alone, so single-card claims generally rely on compressed weights. Check which precision the claim was measured at and test quality at that precision on your own documents. Our explainer on quantisation and fitting big models on small GPUs covers what you typically give up. Second, linear attention is an approximation of the full comparison. It trades some precision on exact long-range recall for speed and memory, so needle-in-a-haystack tasks deserve their own test.
| Cost line | Multi-GPU server | Single 24GB card |
|---|---|---|
| Upfront hardware | $40,000 or more | $2,500 to $4,000 |
| Monthly DevOps time (small deployment) | $1,500 to $3,000 | $1,500 to $3,000 |
| Vendor support line | None for the open model | None for the open model |
| Licence cost (Apache 2.0) | $0 | $0 |
| Long-document cost curve | Grows with the square of length | Grows roughly in line with length |
Look at the second row. The hardware saving is dramatic, but the labour does not shrink with the card. Someone still has to patch, monitor and secure the machine, and at $1,500 to $3,000 a month that person quickly outweighs a one-off $4,000 purchase. We break that line down in our piece on DevOps labour for self-hosted models.
Jobs this architecture actually suits
The linear cost curve rewards long inputs and patient schedules. The best fits are batch jobs that run overnight, where nobody is waiting on the answer:
Nightly lease or contract clause extraction for a property manager with hundreds of agreements on file
Overnight clean-up of call transcripts for a contact centre before the QA team starts
Bulk PDF summarisation of supplier catalogues, tender packs or policy manuals where latency is irrelevant
What these share is a tolerance for the occasional miss. A clause that is extracted wrongly gets caught when a person reads the summary. A transcript with a tidy-up error does not reach a customer. Live chat, anything customer-facing, and anything where one wrong answer costs a client are a different category, and there we still default to Claude.
The licence is the quiet win
Apache 2.0 is genuinely permissive: no royalty and no field-of-use restriction. That is a cleaner footing than Llama 4's community licence, and it means a Brisbane software business could build a product feature on this model without negotiating terms. It does not come with a support contract. If the linear attention build misbehaves on an edge case in your data, there is no vendor to call and the fix is yours to find.
A worked comparison for a mid-sized firm
Take a Sydney property manager processing around 400 lease documents a month for clause extraction. The documents are long, the job can run overnight, and a person checks every summary anyway. On a single card, the first-year spend is roughly $4,000 in hardware plus $18,000 to $36,000 in DevOps time. The same job on Claude's managed API has no hardware, no patching and a support path, and for most firms below roughly 2 million tokens a day it comes out ahead on total cost of ownership. Our token-volume break-even maths shows how to find your own number.
The honest read: this architecture makes self-hosting plausible for more businesses, not advisable for most. If your volume is high, your documents are long, and you already employ someone who runs infrastructure, a pilot is worth doing. If none of those are true, the cheaper card does not change the answer. Our services page describes how we scope that comparison, and the ROI calculator gives a quick first pass.
If you are weighing a local Qwen3.8-27B pilot against staying on Claude, book a short call and we will run the cost comparison on your actual workload.



