Google shipped intelligent dictation in the Gemini app for macOS on 25 August 2026. You speak into any window on the desktop and the text lands at your cursor already cleaned up: filler words stripped, mid-sentence corrections applied, formatting handled. On its own terms it works well, and if you dictate a lot you will feel the difference on day one.
The reason to write about it is not the feature. It is the category confusion around it. Australian business owners are being asked to pick an "AI assistant" while the term covers two very different things: tools that change how words get into a document, and tools that take a job description and return finished work. Those are not competing versions of the same product. They solve different halves of the same hour.
Dictation changes the input, not the outcome
Voice-to-text moves you from typing at roughly 40 words a minute to speaking at 130. That is a genuine saving on the drafting step. What it does not touch is everything that happens after the words exist.
Take a quote for a fit-out job. Dictating the covering note takes two minutes instead of six. The rest of the task is unchanged: pull last quarter's rates, check the margin, confirm the subcontractor is still current on insurance, format the document, log it against the opportunity, and set a reminder to chase it. The typing was never the expensive part.
Dictation shortens composition. It does not gather the inputs.
It produces clean text. It does not check whether the text is right.
It works one document at a time. It has no view of the twenty other documents that share the same numbers.
It stops at the cursor. Filing, sending, logging and following up all stay with you.
None of that is a criticism of the feature. It is a description of its scope. The mistake is buying it as though it closes the loop.
What task completion actually looks like
Claude Cowork and Claude Code sit on the other side of that line. You describe an outcome, and the model does the multi-step work to reach it, using your files and your connected systems along the way. The comparison is not "which one transcribes better". It is "which one is still working after step one".
The jobs that suit this pattern share a shape. They are repetitive, they touch more than one system, and a competent person could do them with clear instructions but would rather not.
Read a folder of 40 supplier invoices, pull the totals and due dates, flag anything outside agreed payment terms, and produce a one-page summary for the bookkeeper.
Turn a call transcript into a draft quote, a CRM record and a follow-up email, all populated from the same set of facts.
Take last month's job records and write the variance commentary a director actually reads, with the exceptions named rather than averaged away.
Check a batch of tender documents against a compliance checklist and list only the clauses that fail it.
The output of a dictation tool is text. The output of the second pattern is a decision you can act on, plus the artefacts that decision needs. Both save time. Only one of them takes work off the list.
The maths for an Australian small business
Worked example, and treat it as illustrative rather than a promise. An office manager on A$85,000 costs roughly A$55 an hour once superannuation, leave and overheads are counted. Say six hours a week goes to document handling: quotes, supplier admin, reports, the weekly summary nobody wants to write.
Faster input might reclaim 30 to 45 minutes of that. Call it A$1,700 a year in recovered time, which comfortably justifies a subscription and needs no project to implement. Moving two of those six hours into an agent that completes the task is worth closer to A$5,700 a year for the same role, and it compounds because the work is done the same way every time rather than differently depending on who is busy.
The catch is that the second number is not free. It needs the task written down properly, a check on the output for the first few weeks, and someone who owns the result. A setup engagement of A$3,500 pays for itself inside a year on that example, but only if the task chosen is one that actually repeats. Automating something you do twice a quarter is a hobby.
What not to conclude from this
Three things this argument does not say.
Dictation is not a lesser product. For anyone managing RSI, a disability, or simply thinking better out loud than on a keyboard, it is the more important of the two features.
Task completion is not always the right call. Anything with real consequences and no easy check still wants a person on it, and the review time can exceed the time saved.
This is not a verdict on Google versus Anthropic. Both will ship the other side of the line eventually. The question is what your business needs this quarter, not who wins the category.
There is also a data question that sits underneath both. Dictation means audio leaves the machine; task completion means your files do. Under the Privacy Act, that is your obligation regardless of which vendor logo is on the tool, so know where the processing happens and what is retained before you roll either one across a team.
The question to ask before you buy either
Pick one task that costs you real hours. Write down every step, from the trigger to the point where it is finished and filed. Then ask any vendor which step their tool completes. Most demos answer for step one. The value is in steps three through seven.
If the honest answer is that the tool speeds up typing, buy it for that and be happy. If the task has six steps and you want five of them handled, you are shopping for something else, and the setup work matters more than the model.
We do this work with Australian businesses from Sydney: pick the task, prove it on real files, then decide whether it is worth building. If you want a second opinion on which of your processes is the right first candidate, book a short call.



