Claude doesn't transcribe audio. It never has, and going by OpenAI's latest release, that isn't about to change. On 28 July 2026, OpenAI launched two new speech-to-text models, GPT Transcribe and GPT Live Transcribe, aimed squarely at the transcription market Whisper used to own. If you run a business in Sydney, Melbourne or anywhere else in Australia and you're wondering whether Claude needs to catch up, the short answer is no. It doesn't need to, because transcription and reasoning are two different jobs, and stapling them into one model is not the smart move it sounds like.
This is a competitor announcement, so it's worth being straight about what OpenAI actually shipped before making the architecture argument. No spin, just the numbers.
What OpenAI Actually Shipped
GPT Transcribe is built for async, batch-style transcription, the kind you run over a completed recording rather than a live call. GPT Live Transcribe is the low-latency sibling, designed for streaming audio as it happens. Both sit alongside, not on top of, Whisper.
Here's what changed with the launch:
GPT Transcribe is priced at US$0.0045 per minute, down from US$0.006 per minute for the previous gpt-4o-transcribe model.
GPT Live Transcribe is priced at US$0.017 per minute, matching the existing gpt-realtime-whisper rate.
On the Common Voice benchmark across 22 languages, word error rate reportedly dropped from roughly 40% to roughly 19% compared with whisper-1.
Neither model ships with word-level timestamps, SRT or VTT subtitle export, speaker diarization, or English translation at launch.
That's a genuine improvement in accuracy and a genuine cut in price. Nobody running high-volume transcription should ignore it. It's also, notably, a narrower release than it might sound like: two models that do one thing, better and cheaper than before.
The Gaps Worth Planning Around
The missing features matter more than they might look at first glance. No word-level timestamps means you can't easily jump to the moment in a two-hour meeting where a decision was made. No SRT or VTT export means you can't drop a transcript straight into video captioning without an extra conversion step. No speaker diarization means a transcript of a three-person sales call comes back as one undifferentiated block of text, with no indication of who said what. No English translation means multilingual teams still need a separate step for non-English audio.
None of that is a knock on the release, it's a batch and streaming transcription model, not a full meeting-intelligence platform, and it was never positioned as one. But it does mean a business building a real workflow around it needs to plan for those gaps rather than assume they're covered. If speaker attribution matters for your use case, whether that's a client call where you need to know who committed to what, or a compliance record where you need to show who said what and when, that has to be solved either upstream, through a diarization-capable capture tool, or downstream, by structuring how the audio is captured in the first place so Claude has something workable to reason over. Claude can make excellent use of a well-structured transcript. It can't invent structure that was never captured.
This is exactly the kind of gap that gets missed when a business buys a transcription tool expecting it to be a complete solution rather than one component in a larger system. The fix isn't more features bolted onto the transcription model. It's building the workflow properly around it from the outset.
Why Claude Isn't Trying to Win This Race
Transcription is a tightly scoped task. Audio goes in, text comes out, and the model is judged almost entirely on word error rate. Reasoning is a different task entirely. It's not about guessing which phonemes were spoken, it's about understanding what the words mean, what matters in them, and what should happen next.
Claude's job in a business workflow was never to compete on transcription accuracy. It's to read a transcript once it exists and do something useful with it: summarise a call, pull out the action items, draft a follow-up email, or update a customer record. Bolting a transcription head onto a general-purpose reasoning model doesn't make the reasoning sharper, and tuning a transcription model harder on word error rate doesn't make it a competent business analyst. These are two different optimisation targets, and trying to serve both with one model usually means compromising on each.
That split isn't a workaround. It's the same principle behind most well-built software: components with a single responsibility that you can swap out independently, rather than one monolith trying to be everything at once. A dedicated transcription API optimised for accuracy, paired with a reasoning model optimised for judgement, is a more sensible architecture than a single model straining to do both jobs at once.
The Pattern We Already Build for AU Businesses
This isn't theoretical. It's the same pattern Automata AI already builds for Australian clients running meeting-note and call-summary workflows, and it holds regardless of which vendor supplies the transcription step.
Call or meeting audio, a sales call, a site visit, a customer support conversation, is captured and sent to a dedicated transcription API, whether that's OpenAI's new models or another provider entirely.
The raw transcript is handed to Claude, which extracts the structured pieces a business actually needs: a summary, the decisions made, the action items, and any follow-up dates.
Claude drafts the next step, whether that's a CRM update, a follow-up email, an internal note, or a task assignment for the right person.
A human reviews and approves before anything goes out, because a transcription slip or a misread instruction is cheap to catch at that point and expensive to catch later.
Swap the transcription vendor, from Whisper to GPT Transcribe to whatever comes next, and none of that touches the Claude layer. That's the entire point of keeping the two apart: the reasoning layer stays stable while the transcription layer keeps getting cheaper and more accurate underneath it.
The Real Bottleneck Isn't Transcription
With GPT Transcribe now priced at under half a US cent per minute, the cost of turning speech into text has more or less stopped being something a business needs to worry about. An hour of meeting audio costs a fraction of a dollar to transcribe, and even converted to Australian dollars that's well under A$1 an hour. For a Sydney business handling hundreds of client calls a month, that per-minute rate is not the expense line worth optimising anymore.
The bottleneck is what happens to the text afterward, and that's a reasoning and integration problem, not a transcription problem. This is where the real cost and the real value actually sit. A properly built workflow that turns a raw transcript into a CRM entry, a compliant record, or a drafted follow-up typically runs from around A$3,500 for a single, well-defined workflow, up to the low five figures for something that touches multiple systems and needs more careful handling. That build is what pays for itself. The transcription rate underneath it barely moves the needle.
What This Means If You're Weighing Up a Build
If you're an Australian business owner comparing AI vendors right now, the useful question isn't "does this model transcribe as well as OpenAI's." It's "what happens to the text once it exists." Businesses that need to keep proper records under the Privacy Act, or that operate in sectors watched by ASIC or APRA, need the reasoning layer to be reliable and auditable at least as much as they need the transcript to be word-perfect.
A general-purpose model built for careful, structured reasoning, paired with a purpose-built transcription tool doing the one job it's optimised for, is a more sensible architecture than waiting for a single vendor to be excellent at both. It's also easier to maintain, cheaper to swap out, and easier to explain to an auditor or a client who wants to know how a decision got made.
OpenAI's release is a real improvement for anyone running high-volume transcription, and it's worth using if that's your bottleneck. For most Australian businesses, it isn't. The interesting problem was never turning speech into text. It's turning text into action, a summary that becomes a CRM update, a call that becomes a follow-up, a meeting that becomes a task list. If you want help building that layer, get in touch and we'll walk through what a Claude-based workflow could look like for your business.



