Blog

Deploying GPT-OSS-20B on a Single GPU: An Australian Guide

October 2026 · Technical

Hand-drawn single GPU card with one terracotta fan and a power lead, showing a model deployed on one card
← Back to all posts

GPT-OSS-20B is one of the few open-weight models a small technical team can test on hardware it could plausibly rent this afternoon. One well-specced GPU is enough. That makes it a good candidate for an honest trial, and also an easy one to get wrong, because the part that works on day one is not the part that decides whether it survives in production. This guide, written in October 2026, covers the hardware, the decisions to settle before you deploy, and the cost pattern that catches Australian teams out.

Can GPT-OSS-20B run on a single GPU?

Yes. A single 24GB GPU, such as an RTX 4090-class card or an equivalent cloud instance, is typically enough to run GPT-OSS-20B when the model is quantised to 8-bit or 4-bit. Running it at full precision needs closer to 48GB of memory, which usually means a larger card or a second one. For most trials, a quantised model on one 24GB card is the sensible starting point.

Renting that class of GPU from an Australian availability zone costs somewhere between $1.80 and $4.50 an hour, depending on the provider and the card. The hourly figure looks small. It stops looking small when the instance is left running after the team goes home.

Four decisions to make before you deploy

Downloading weights and starting a server takes an afternoon. These four questions take longer, and skipping them is how a trial turns into an unowned system.

  • Inference server.

  • Quantisation.

  • Data location.

  • After-hours monitoring.

If the first question is still open, our comparison of Ollama, vLLM and LocalAI as open-model runtimes walks through how to choose one for a small Australian team.

The three steps that decide whether it lasts

Getting the model to answer a prompt is the easy part. Three pieces of engineering separate a demo from a tool people can rely on.

Set a hard token budget per request

A local model has no invoice per call, so nothing stops a runaway prompt from tying up the card. A fixed ceiling on tokens per request keeps one user from starving everyone else and keeps latency predictable.

Build a fallback to a hosted API

One GPU is one point of failure. When it is unavailable, requests need somewhere to go. Routing them to a hosted model such as Claude means a crashed process becomes a slower afternoon, not an outage. Build this path before launch, and test it by switching the GPU off.

Load-test with realistic concurrency

A Melbourne engineering team running GPT-OSS-20B as a coding assistant for internal tooling found that a single GPU handled roughly 15 to 20 concurrent lightweight requests before latency became noticeable. That ceiling is fine for a development team. It matters a great deal if the plan is to put the model in front of customers. Run the test with your own request sizes before anyone in the business depends on the result.

What it costs to keep one GPU switched on

A dedicated, always-on GPU instance in Australia running GPT-OSS-20B typically costs $900 to $2,500 a month, depending on the card and provider. That is before the engineering time to maintain it, which we have costed separately in our piece on the DevOps labour that self-hosting calculators miss.

The important property of that number is that it is fixed. The bill is the same whether the model handles ten requests a day or ten thousand. API pricing works the opposite way: you pay for what you use and nothing when it is quiet. The table below shows what a fixed bill does to the cost of each request. The per-request figures are illustrative arithmetic on the lower $900 monthly figure, over a 30-day month, and exclude engineering time.

Illustrative cost per request on a $900 a month dedicated GPU, by daily volume
Requests per dayRequests per monthGPU cost per request
10300$3.00
1003,000$0.30
1,00030,000$0.03
10,000300,000$0.003

On-demand rental has its own version of the same problem. An instance left idle for a 14-hour night costs between $25 and $63 at the hourly rates above, for no work done. Falling prices change the sums at the margin, as we covered in GPU spot prices and the self-hosting break-even, but they do not change the shape: steady, high volume favours the fixed bill, and everything else favours paying per use.

Signs a single-GPU deployment is the wrong fit

Free weights are appealing, and the appeal can get ahead of the numbers. A single-GPU deployment is probably the wrong choice when any of these describe your workload.

  • Traffic is spiky, so the GPU sits idle, and still billed, for most of the day.

  • The use case is customer-facing and needs uptime the team cannot commit to monitoring.

  • Nobody has the time to own security patching on an ongoing basis.

  • The workload would run comfortably within a managed platform's existing token pricing.

For spiky or low-volume workloads, a managed API will usually come out cheaper once engineering time is priced in honestly. For many Australian mid-market teams that means running the work on Claude and keeping a self-hosted model for the narrow cases where data location or steady volume justify it.

What to do next

Run the trial. A quantised GPT-OSS-20B on one rented card for a week will teach a team more than a month of reading. Go in with the token budget, the fallback and the load test planned, and measure your real request volume while it runs. Then put that volume through our ROI calculator before committing to an always-on instance.

If you would like a second opinion on whether your workload belongs on a GPU you maintain or on Claude, book a short call with us. We will look at your usage pattern and tell you plainly which way the numbers point.

FAQ

Frequently asked questions

How much VRAM does GPT-OSS-20B need?

A single 24GB GPU is typically enough to run GPT-OSS-20B with 8-bit or 4-bit quantisation. Full-precision inference needs closer to 48GB, which usually means a larger card.

Which inference servers support GPT-OSS-20B?

vLLM, TGI and llama.cpp all support GPT-OSS models. Each makes different trade-offs, so choose based on your concurrency needs and how much operational work your team can take on.

How many users can one GPU running GPT-OSS-20B handle?

One Melbourne engineering team using it as an internal coding assistant saw roughly 15 to 20 concurrent lightweight requests before latency became noticeable. Load-test with your own request sizes before relying on that figure.

How much does it cost to run GPT-OSS-20B in Australia?

Renting a suitable GPU in an Australian availability zone costs about $1.80 to $4.50 an hour. A dedicated always-on instance typically runs $900 to $2,500 a month, before engineering time.

Is GPT-OSS-20B cheaper than using a hosted API?

Only for steady, high-volume workloads. The GPU bill is fixed whatever the usage, so for spiky or low-volume work a managed API such as Claude usually costs less once engineering time is included.

Ready to move from AI pilot to production?

We help mid-market Australian businesses deploy AI automations that actually reach production and deliver measurable ROI.