On 1 September 2026 Google turned on agentic video understanding across Gemini 3.7 Flash, Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. The published claim is large: up to 88% less token consumption, up to 66% lower cost and up to 7% better quality against static frame-rate video processing.
Before anyone files this next to the last two comparisons we wrote, the subject is different. Those were about video generation, which is the business of producing a clip that did not exist before. This is video understanding: handing a model footage you already hold and asking it questions. Different capability, different buyers, different bill. If the generation question is the one you actually have, our comparison of Claude and Gemini Omni on agentic reliability versus video generation answers it and this post will not.
What is agentic video understanding?
Agentic video understanding is a processing mode in which a model's reasoning is paired with native video tools so it can dynamically search, scan and inspect target segments across visual frames, audio and transcripts, rather than ingesting the whole file at a fixed frames-per-second rate. Google describes it as similar in spirit to the agentic vision it already offers for images, and says it makes possible sub-second moment retrieval, more accurate anomaly detection and precise counting.
The comparison point is static processing, where the model takes the video in at a fixed rate, one frame per second by default and adjustable through the API. On a short clip that is fine. On a multi-hour recording it forces a choice nobody enjoys making.
Where the numbers come from
Google states the efficiency gains are most pronounced on long-form video, described as running from ten-minute how-to guides through ninety-minute lectures and multi-hour recordings, because that is exactly where static processing makes developers choose between high token costs and techniques that drop critical details. The 88% token reduction and the 7% accuracy lift are quoted for Gemini 3.7 Flash specifically, which Google positions as the best quality-to-cost tradeoff of the three supported models.
The feature is available for video uploads and for YouTube videos through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.
It is switched on by setting the API configuration to agentic, not by changing model.
Every headline figure is an up-to figure, published by the vendor, measured against the vendor's own static baseline.
None of the published numbers are a comparison against another provider's video handling.
| Claim | What Google states |
|---|---|
| Token consumption | Down by up to 88% against static frame-rate processing |
| Cost | Down by up to 66% |
| Quality | Up by up to 7% |
| Models supported | Gemini 3.7 Flash, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite |
| Where it runs | Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform |
| Best tradeoff | Gemini 3.7 Flash, per Google's own positioning |
What Claude builders should actually do about it
Most of the automation work we do for Sydney and Melbourne businesses is not video-bound. It is documents, email, tickets, catalogues, spreadsheets and phone notes. When video does turn up it is rarely a stream of short clips. It is a small number of long recordings that nobody has time to watch: site inspections, franchise compliance footage, training libraries, council and board meetings, contact-centre screen captures.
If that is your shape of problem, a capability aimed squarely at long-form comprehension is worth testing, and it is worth testing honestly. Our position has not moved: Claude stays the default for the reasoning, the tool use and the orchestration around the task. Whatever reads the footage can be a tool that agent calls.
Count the footage first. Hours per month, average length, and how many questions per hour of video you actually need answered.
Write the questions down before you test anything, because retrieval accuracy only means something against questions that matter to the business.
Measure cost per correct answer, not cost per minute of footage processed.
Decide where the video sits. A second provider in the stack is a contract, a data flow and a Privacy Act question, not just an API key.
On our engagements a proper bake-off against a real video corpus runs roughly $8,000 to $15,000 of work: assembling a labelled sample, writing the questions, and running the same set through each candidate. That is inexpensive next to committing a production pipeline to the wrong choice and discovering it in month four.
What not to conclude from an 88% figure
Up to is carrying weight in every one of those numbers. They describe a best case on the workload the feature was built for, measured against the same vendor's previous approach. A token reduction is also not automatically a cost reduction for you, because your bill depends on how many questions you ask per hour of footage, not on how efficiently one hour is read.
The other thing the announcement does not address is what happens around the model. Moment retrieval that is usually right still needs a human review path for the times it is not, and anomaly detection that flags a safety incident needs somebody whose job it is to act on the flag. That work is the same whichever provider reads the video.
A sensible split for an Australian team
Keep one agent that owns the conversation and the business rules, and treat video comprehension as a capability it calls. That way the decision you are making is reversible: if a better option ships in six months, you change a tool, not an architecture. It also keeps your evals in one place, which is the part teams regret skipping.
We help Australian businesses make these calls with numbers rather than vendor blog posts. See our consulting services, run the ROI calculator on the hours the footage is costing you today, or read Google's own write-up of agentic video in Gemini.



