GPT-OSS-20B is one of the few open-weight models a small technical team can test on hardware it could plausibly rent this afternoon. One well-specced GPU is enough. That makes it a good candidate for an honest trial, and also an easy one to get wrong, because the part that works on day one is not the part that decides whether it survives in production. This guide, written in October 2026, covers the hardware, the decisions to settle before you deploy, and the cost pattern that catches Australian teams out.
Can GPT-OSS-20B run on a single GPU?
Yes. A single 24GB GPU, such as an RTX 4090-class card or an equivalent cloud instance, is typically enough to run GPT-OSS-20B when the model is quantised to 8-bit or 4-bit. Running it at full precision needs closer to 48GB of memory, which usually means a larger card or a second one. For most trials, a quantised model on one 24GB card is the sensible starting point.
Renting that class of GPU from an Australian availability zone costs somewhere between $1.80 and $4.50 an hour, depending on the provider and the card. The hourly figure looks small. It stops looking small when the instance is left running after the team goes home.
Four decisions to make before you deploy
Downloading weights and starting a server takes an afternoon. These four questions take longer, and skipping them is how a trial turns into an unowned system.
Inference server.
Quantisation.
Data location.
After-hours monitoring.
If the first question is still open, our comparison of Ollama, vLLM and LocalAI as open-model runtimes walks through how to choose one for a small Australian team.
The three steps that decide whether it lasts
Getting the model to answer a prompt is the easy part. Three pieces of engineering separate a demo from a tool people can rely on.
Set a hard token budget per request
A local model has no invoice per call, so nothing stops a runaway prompt from tying up the card. A fixed ceiling on tokens per request keeps one user from starving everyone else and keeps latency predictable.
Build a fallback to a hosted API
One GPU is one point of failure. When it is unavailable, requests need somewhere to go. Routing them to a hosted model such as Claude means a crashed process becomes a slower afternoon, not an outage. Build this path before launch, and test it by switching the GPU off.
Load-test with realistic concurrency
A Melbourne engineering team running GPT-OSS-20B as a coding assistant for internal tooling found that a single GPU handled roughly 15 to 20 concurrent lightweight requests before latency became noticeable. That ceiling is fine for a development team. It matters a great deal if the plan is to put the model in front of customers. Run the test with your own request sizes before anyone in the business depends on the result.
What it costs to keep one GPU switched on
A dedicated, always-on GPU instance in Australia running GPT-OSS-20B typically costs $900 to $2,500 a month, depending on the card and provider. That is before the engineering time to maintain it, which we have costed separately in our piece on the DevOps labour that self-hosting calculators miss.
The important property of that number is that it is fixed. The bill is the same whether the model handles ten requests a day or ten thousand. API pricing works the opposite way: you pay for what you use and nothing when it is quiet. The table below shows what a fixed bill does to the cost of each request. The per-request figures are illustrative arithmetic on the lower $900 monthly figure, over a 30-day month, and exclude engineering time.
| Requests per day | Requests per month | GPU cost per request |
|---|---|---|
| 10 | 300 | $3.00 |
| 100 | 3,000 | $0.30 |
| 1,000 | 30,000 | $0.03 |
| 10,000 | 300,000 | $0.003 |
On-demand rental has its own version of the same problem. An instance left idle for a 14-hour night costs between $25 and $63 at the hourly rates above, for no work done. Falling prices change the sums at the margin, as we covered in GPU spot prices and the self-hosting break-even, but they do not change the shape: steady, high volume favours the fixed bill, and everything else favours paying per use.
Signs a single-GPU deployment is the wrong fit
Free weights are appealing, and the appeal can get ahead of the numbers. A single-GPU deployment is probably the wrong choice when any of these describe your workload.
Traffic is spiky, so the GPU sits idle, and still billed, for most of the day.
The use case is customer-facing and needs uptime the team cannot commit to monitoring.
Nobody has the time to own security patching on an ongoing basis.
The workload would run comfortably within a managed platform's existing token pricing.
For spiky or low-volume workloads, a managed API will usually come out cheaper once engineering time is priced in honestly. For many Australian mid-market teams that means running the work on Claude and keeping a self-hosted model for the narrow cases where data location or steady volume justify it.
What to do next
Run the trial. A quantised GPT-OSS-20B on one rented card for a week will teach a team more than a month of reading. Go in with the token budget, the fallback and the load test planned, and measure your real request volume while it runs. Then put that volume through our ROI calculator before committing to an always-on instance.
If you would like a second opinion on whether your workload belongs on a GPU you maintain or on Claude, book a short call with us. We will look at your usage pattern and tell you plainly which way the numbers point.



