OpenAI, working with Cerebras, previewed Ultrafast this week: a new tier for GPT-5.6 Sol that runs up to 14 times faster than standard, hitting around 750 output tokens per second. Same underlying intelligence as standard Sol, just delivered much faster. OpenAI is pointing it at incident response, customer service, financial market analysis, and e-commerce. Access is limited to a small group of customers for now, expanding as capacity grows.
It is a genuinely impressive infrastructure achievement. It is also worth being precise about which businesses that speed actually helps, because the honest answer is not most of them.
Where speed is the real constraint
There is a real category of workload where output speed is the bottleneck: live customer-facing chat where a two-second delay feels broken, real-time trading signals, or an agent loop where every millisecond compounds across thousands of calls. If that is your workload, a 14x speed tier is a legitimate reason to pay attention, and you should test it properly.
Where speed isn't the constraint, which is most businesses
For the large majority of AI use cases we see in Australian SMBs, speed was never the limiting factor. Processing invoices overnight, drafting a follow-up email, summarising a contract, triaging leads before a morning call: none of these need a response in under a second. What actually limits value in those workflows is accuracy and judgement, not throughput. A wrong answer delivered instantly is still a wrong answer, and now you have less time to catch it before it does something.
Before paying a premium for speed, the questions worth asking are simple:
Does the task have a human reviewing the output anyway? If so, shaving seconds off generation does not change your cycle time, review does.
Is the current model's accuracy already the bottleneck, not its speed? Faster wrong answers compound the problem, they do not solve it.
What is the actual cost delta, and does the speed gain translate into a business outcome or just a nicer demo?
Put a number on it
Say a mid-size Australian business runs a customer-service workflow with a person checking every AI-drafted reply before it goes out. Move to a tier that generates the draft 14 times faster and you might save a few seconds per reply, but the human review is still the gate, so your real cycle time barely moves. Now suppose that faster model is a shade less accurate and one in twenty replies needs a correction it would not have before. On 1,000 replies a month that is 50 extra corrections, and at ten minutes each on a $60-an-hour operator, roughly $500 a month spent buying speed you could not use in the first place. Speed is only cheap when it was actually your constraint.
Where Claude keeps investing
Ultrafast is a real capability aimed at a real, if narrow, category of use case. For most Australian businesses the honest answer is that inference speed stopped being the constraint a while ago. Where Claude keeps investing, context handling in tools like Claude Tag, agent reliability, and verification loops, tends to matter more for the workflows that actually move a small business's bottom line: fewer errors, less rework, and more of the process running unsupervised with confidence. For anything touching money or regulated data under the Privacy Act, that reliability is worth more than tokens per second.
The counterargument worth stating
To be fair to OpenAI, speed can unlock things that were not possible before, not just make existing tasks marginally quicker. A model fast enough to hold a natural back-and-forth voice conversation, or to run an agent loop that would have timed out at slower speeds, is doing something new, not just something faster. If Ultrafast lets you build a workflow you genuinely could not run before, that is a real gain and worth the premium. The mistake is assuming that because speed is impressive it must be valuable to you specifically.
So the honest position is not speed does not matter. It is that speed matters for a specific, identifiable set of workloads, and most Australian small businesses do not run those workloads yet. Work out which camp you are in before you pay for the fast tier, because buying a solution to a constraint you do not have is one of the most common and least visible ways to waste an AI budget.
If you want a straight read on whether speed, accuracy, or something else entirely is your real bottleneck, book a session and we will look at your actual workflow before recommending anything.



