A caller rings a bank to dispute a charge. The voice on the line sounds natural, switches to Vietnamese when she does, and never pauses awkwardly. Then it fails to lodge the dispute. According to Google's own published numbers, that second half is where its newest voice model still struggles, and it is the half an Australian business should read first.
What Google launched on 15 September 2026
Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its newest live-dialogue voice models. The claims, from Google's announcement, are strong on conversation quality:
First place on Artificial Analysis' Speech-to-Speech Quality Index, with a score of 82.6.
97.7% on Big Bench Audio and second place in the Speech Agent Arena for the base Live model.
68.6% on tau-Voice agentic task completion, and 35.1% on Sierra's tau-Voice banking benchmark.
Automatic switching across 97 languages mid-conversation, near-real-time visual context, and tool or API calls that run in the background without stopping the conversation.
An Extended Thinking variant that reasons while it speaks, with verbal cues such as acknowledging it is working on a multi-step task.
It is available through the Gemini API, inside Google Workspace (Docs Live, Gmail Live and Keep Live), in Search Live and the Gemini app, and for enterprise voice-agent builders through the Gemini Live API and Gemini Enterprise Agent Platform.
Is Gemini 3.8 Live better than Claude for voice agents?
On live speech itself, Google's published results put Gemini 3.8 Live ahead, and Claude does not currently lead on native speech-to-speech. But a voice agent for an Australian business is judged on whether the job gets done: the booking made, the claim lodged, the record updated correctly. Google's own task-completion scores, 68.6% overall and 35.1% on banking tasks, show that part is still far from solved by voice quality alone.
Read the 35.1% figure slowly. On a benchmark built around banking tasks, the model completed roughly one in three. That is not a criticism unique to Google; agentic banking tasks are hard for every model. It is a reminder that the headline quality score and the number that decides whether you can put the agent on your phone line are different numbers.
Where each approach fits
Claude-based voice agents are usually built as a pipeline: speech-to-text in, Claude for reasoning and tool calls, text-to-speech out, often on a platform such as Vapi or Retell. That adds some latency compared with a native speech-to-speech model, and in exchange gives you a text layer you can log, test and audit. Our comparison of voice AI platforms for business phone lines covers the platform choice.
| If your priority is | Lean towards | Why |
|---|---|---|
| Most natural-sounding conversation | Gemini 3.8 Live | Leads Google's cited speech quality rankings |
| Many languages on one line | Gemini 3.8 Live | Switches across 97 languages automatically |
| Staff already live in Google Workspace | Gemini 3.8 Live | Built into Docs, Gmail and Keep |
| Multi-step actions in your own systems | Claude pipeline | Tool calls through MCP connectors you already run |
| Full transcript and audit trail | Claude pipeline | Every turn exists as text you can review |
| One model across voice, chat and back office | Claude pipeline | Same prompts, tools and tests everywhere |
What this means for Australian businesses
For a Sydney or Melbourne business, the 97-language switching is genuinely attractive. A medical centre or real estate agency serving customers who speak Mandarin, Cantonese, Vietnamese, Arabic or Hindi could handle first contact without a separate line for each. That is a real advantage and worth testing.
The obligations do not change with the model. Calls that capture personal information fall under the Privacy Act, and in financial services a voice agent that gives or implies advice raises ASIC questions long before it raises technical ones. For APRA-regulated firms, a voice agent that acts on accounts is a material system with the oversight that implies.
Sizing the opportunity is simple arithmetic. If a clinic misses 30 calls a week and one in five would have become a $150 appointment, that is about $46,800 a year in lost bookings. The first question is whether an agent can reliably book those appointments in your system, not how natural it sounds while failing to.
How we would test it
Write 30 real call scenarios from your own logs, including the awkward ones.
Score each on whether the task was completed correctly in the back-office system, not on how the call sounded.
Run the same scenarios through a native voice model and a Claude pipeline, and compare completion rates side by side.
Check where transcripts and audio are stored, and for how long, before any live customer call.
For background, we have compared dictation with task completion, looked at OpenAI's full-duplex voice architecture, and set out where spoken interfaces fit in 2026. If you are choosing a voice-agent stack now, our AI assessment will score your own call scenarios, or book a time with us to talk it through.



