Blog

Claude vs Gemini: Agentic Video Understanding

September 2026 · 7 min read · Technical

Line drawing of a filmstrip with a magnifier inspecting one segment of it
← Back to all posts

On 1 September 2026 Google turned on agentic video understanding across Gemini 3.7 Flash, Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. The published claim is large: up to 88% less token consumption, up to 66% lower cost and up to 7% better quality against static frame-rate video processing.

Before anyone files this next to the last two comparisons we wrote, the subject is different. Those were about video generation, which is the business of producing a clip that did not exist before. This is video understanding: handing a model footage you already hold and asking it questions. Different capability, different buyers, different bill. If the generation question is the one you actually have, our comparison of Claude and Gemini Omni on agentic reliability versus video generation answers it and this post will not.

What is agentic video understanding?

Agentic video understanding is a processing mode in which a model's reasoning is paired with native video tools so it can dynamically search, scan and inspect target segments across visual frames, audio and transcripts, rather than ingesting the whole file at a fixed frames-per-second rate. Google describes it as similar in spirit to the agentic vision it already offers for images, and says it makes possible sub-second moment retrieval, more accurate anomaly detection and precise counting.

The comparison point is static processing, where the model takes the video in at a fixed rate, one frame per second by default and adjustable through the API. On a short clip that is fine. On a multi-hour recording it forces a choice nobody enjoys making.

Where the numbers come from

Google states the efficiency gains are most pronounced on long-form video, described as running from ten-minute how-to guides through ninety-minute lectures and multi-hour recordings, because that is exactly where static processing makes developers choose between high token costs and techniques that drop critical details. The 88% token reduction and the 7% accuracy lift are quoted for Gemini 3.7 Flash specifically, which Google positions as the best quality-to-cost tradeoff of the three supported models.

  • The feature is available for video uploads and for YouTube videos through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.

  • It is switched on by setting the API configuration to agentic, not by changing model.

  • Every headline figure is an up-to figure, published by the vendor, measured against the vendor's own static baseline.

  • None of the published numbers are a comparison against another provider's video handling.

What Google published about agentic video understanding, as at September 2026
ClaimWhat Google states
Token consumptionDown by up to 88% against static frame-rate processing
CostDown by up to 66%
QualityUp by up to 7%
Models supportedGemini 3.7 Flash, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite
Where it runsGemini API in Google AI Studio and the Gemini Enterprise Agent Platform
Best tradeoffGemini 3.7 Flash, per Google's own positioning

What Claude builders should actually do about it

Most of the automation work we do for Sydney and Melbourne businesses is not video-bound. It is documents, email, tickets, catalogues, spreadsheets and phone notes. When video does turn up it is rarely a stream of short clips. It is a small number of long recordings that nobody has time to watch: site inspections, franchise compliance footage, training libraries, council and board meetings, contact-centre screen captures.

If that is your shape of problem, a capability aimed squarely at long-form comprehension is worth testing, and it is worth testing honestly. Our position has not moved: Claude stays the default for the reasoning, the tool use and the orchestration around the task. Whatever reads the footage can be a tool that agent calls.

  • Count the footage first. Hours per month, average length, and how many questions per hour of video you actually need answered.

  • Write the questions down before you test anything, because retrieval accuracy only means something against questions that matter to the business.

  • Measure cost per correct answer, not cost per minute of footage processed.

  • Decide where the video sits. A second provider in the stack is a contract, a data flow and a Privacy Act question, not just an API key.

On our engagements a proper bake-off against a real video corpus runs roughly $8,000 to $15,000 of work: assembling a labelled sample, writing the questions, and running the same set through each candidate. That is inexpensive next to committing a production pipeline to the wrong choice and discovering it in month four.

What not to conclude from an 88% figure

Up to is carrying weight in every one of those numbers. They describe a best case on the workload the feature was built for, measured against the same vendor's previous approach. A token reduction is also not automatically a cost reduction for you, because your bill depends on how many questions you ask per hour of footage, not on how efficiently one hour is read.

The other thing the announcement does not address is what happens around the model. Moment retrieval that is usually right still needs a human review path for the times it is not, and anomaly detection that flags a safety incident needs somebody whose job it is to act on the flag. That work is the same whichever provider reads the video.

A sensible split for an Australian team

Keep one agent that owns the conversation and the business rules, and treat video comprehension as a capability it calls. That way the decision you are making is reversible: if a better option ships in six months, you change a tool, not an architecture. It also keeps your evals in one place, which is the part teams regret skipping.

We help Australian businesses make these calls with numbers rather than vendor blog posts. See our consulting services, run the ROI calculator on the hours the footage is costing you today, or read Google's own write-up of agentic video in Gemini.

FAQ

Frequently asked questions

What is agentic video understanding in Gemini?

It is a processing mode that pairs the model's reasoning with native video tools so it can search, scan and inspect specific segments across frames, audio and transcripts instead of reading the file at a fixed frame rate.

How is video understanding different from video generation?

Video generation produces a clip that did not previously exist. Video understanding takes footage you already have and answers questions about it, such as finding a moment, counting objects or spotting an anomaly.

Which Gemini models support agentic video?

Google lists Gemini 3.7 Flash, Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, and positions Gemini 3.7 Flash as offering the best quality-to-cost tradeoff of the three supported models.

How do you turn agentic video processing on?

Google says you set your API configuration to agentic in Google AI Studio or the Gemini Enterprise Agent Platform. It applies to uploaded video and to YouTube videos rather than requiring a different model.

Does agentic video processing reduce costs by 66% for everyone?

No. The published figures are up-to numbers measured against static frame-rate processing on long-form video, so shorter clips and question-heavy workloads will see considerably less benefit than the headline suggests.

Ready to move from AI pilot to production?

We help mid-market Australian businesses deploy AI automations that actually reach production and deliver measurable ROI.