Blog

Claude vs Google's New Gemini 3.6 Flash: What the Efficiency and Cyber Claims Actually Mean

August 2026 · 6 min read · AI Strategy

Two hand-drawn gauge dials side by side, one with a terracotta-filled efficiency wedge, compared under a magnifying glass, in flat ink-line notebook style
← Back to all posts

Google dropped a new Gemini Flash lineup on 21 July 2026: Gemini 3.6 Flash, 3.5 Flash-Lite, and a specialised 3.5 Flash Cyber model built to work alongside Google's CodeMender security agent. The headline numbers are all about efficiency. Google says 3.6 Flash cuts output token usage by 17% compared with 3.5 Flash, with reductions of up to 65% on some coding benchmarks, priced at $1.50 per million input tokens and $7.50 per million output tokens.

If you're an Australian business owner weighing Claude against Gemini for an agent build, the efficiency pitch deserves a closer look than a launch page gives it. We build automation on Claude for clients across Sydney, Melbourne and Brisbane, so we're not pretending to be neutral here. But the questions below are worth asking of any vendor's numbers, Claude's included, before you let a benchmark decide anything.

It's also a fast-moving picture. New Flash, Lite and Cyber variants landing in the same week means most comparison articles you'll find are already stale by the time you read them, built on a press release rather than a real workload. That's exactly why the test needs to be your own task, not someone else's benchmark chart.

Cheaper tokens are not the same as a cheaper task

Google's framing measures token efficiency and raw benchmark scores. What it doesn't measure is the cost of a completed, correct piece of work, once retries, human review and the time spent catching an agent's mistake before it reaches a customer are added into the total. That's close to the exact argument OpenAI's own leadership has been making this month, about measuring "useful intelligence per dollar" rather than per-token pricing. It's the right question to ask of any model on the market right now, not a competitor talking point.

A model that's cheaper per token but needs more oversight, more retries, or produces work a person has to redo isn't actually cheaper. It just moves the cost from the invoice line to the hours your team spends checking the output. For a small business running lean, without a dedicated review layer, that hidden cost is usually the bigger one. It just doesn't show up on the pricing page.

This is where the comparison gets genuinely useful. Run the same real task, an invoice reconciliation, a customer email draft, a piece of code that needs to ship, through both models and time the whole loop: draft, check, fix, approve. The per-token price rarely predicts which one finishes first with less babysitting.

The security angle is the more interesting claim

The 3.5 Flash Cyber variant, paired with CodeMender, is Google's answer to a real gap. Most AI coding tools generate code faster than anyone can review it for security issues by hand. Anthropic has been building in the same direction, publishing its own Zero Trust for Agents framework and detailing how its security team defends a codebase where Claude authors most of the merged code.

The difference worth watching isn't which vendor ships a security-labelled model first. It's whether security review is built into the workflow by default, or whether your team has to bolt on a separate, specialised model and remember to switch it on every time. For a business handling client data under the Privacy Act, or reporting obligations that touch APRA or AUSTRAC-adjacent processes, that default matters more than a benchmark score on a launch blog.

What to actually compare before you switch

Before switching, or choosing, a model for an agent build, we'd suggest comparing three things properly rather than taking any vendor's launch page at face value:

  • The end-to-end cost of a successfully completed task, not the advertised price per million tokens.

  • How much unsupervised work the model can be trusted with before something needs a human check.

  • Whether security review is a default part of the workflow, or something your team has to add on top and remember to use.

Google's new Flash family is a real product update and worth testing if you're already building on Gemini. It isn't, on its own, a reason to switch away from a model your team trusts and has already built workflows around. That decision should rest on the same measure now being pointed at across the industry: work completed, cost included, per dollar spent.

What this looks like for an Australian business

When we run a model and workflow audit for a client, a fixed $3,500 AUD engagement, we're not comparing marketing claims side by side. We're timing the actual task from prompt through to a piece of work a person would sign off on, then pricing in the review step most vendor benchmarks leave out entirely. That's usually where the real gap between Claude and a cheaper-looking competitor shows up, or quietly disappears once the review time is counted.

For most Australian SMBs, the decision isn't which model wins a coding benchmark this quarter. It's which one your team can run with the least supervision, on the tasks that actually matter to the business, without creating a new security review job nobody signed up for and nobody has time to do properly.

There's also a switching cost that rarely makes it into a vendor comparison. If your team has already built prompts, agent workflows and internal guardrails around Claude, ripping that out for a 17% token saving on paper can cost more in rebuild hours than it saves in the first year. Worth testing the new Gemini Flash models on a low-stakes task first, side by side with what you already run, before touching anything customer-facing.

If you're weighing Claude against Gemini 3.6 Flash, or any other model, for an agent build and want a straight answer rather than a sales pitch, get in touch with Automata AI and we'll help you work out the real cost of the task, not just the cost of the token.

Ready to move from AI pilot to production?

We help mid-market Australian businesses deploy AI automations that actually reach production and deliver measurable ROI.