Vision-capable models have quietly become the most commercially useful development for a specific set of Australian businesses: the ones whose work produces photographs and paper. Inspections, claims, compliance checks and document processing all involve someone looking at an image and writing down what they saw, which is now partly automatable.
What vision models are actually good at
Describing what is visible, extracting text and structure from documents, and checking an image against a written standard. They are reliable at reading and transcription and much less reliable at judgement calls that depend on domain expertise.
Reading scanned invoices, delivery dockets and forms into structured data
Producing a first-draft written description from inspection photographs
Flagging whether an image appears to show a listed condition, for a human to confirm
Sorting and labelling large photo sets so someone can find things later
The pattern in all four is that the model prepares work and a person decides. That boundary is where the technology is dependable, and moving it is where projects go wrong.
Where the money is
Document processing, unglamorously. Accounts payable is the clearest case: invoices arrive as PDFs and images in wildly inconsistent formats, and someone types their contents into a system every day.
Inspection reporting is the second. A field inspector photographs a site, and the report gets written that evening or several days later from memory and thumbnails. Generating a structured first draft at the point of capture removes the worst part of the job and improves accuracy simply by being closer to the event.
What it is worth
An operation processing 600 supplier invoices a month, at four minutes each of manual entry and checking, spends about 40 hours a month on it. At a fully loaded $45 an hour that is roughly $21,600 a year for a task with no judgement in it whatsoever.
Automating extraction and leaving the exceptions to a person typically removes most of that while catching more errors, because a machine reading every line item is more thorough than a human reading the total.
Accuracy needs a threshold, not a hope
The right design routes confident extractions straight through and sends uncertain ones to a person. A system that extracts everything and asks for review on everything has saved nobody anything.
Set the threshold using your own documents, and measure what falls out. Most businesses find that a large share of their invoices are from a handful of regular suppliers in consistent formats, which is exactly the easy majority worth automating.
Open weights in the multimodal category
Vision-capable open models have closed much of the gap on document extraction, which matters because this is high-volume work where per-token cost compounds. It is the clearest case for a cheaper model doing the bulk and a stronger one handling exceptions.
For sensitive imagery there is a second argument. Site photographs and identity documents can be processed inside your own environment, which is easier to explain to a client than sending them to a third party, whatever the contractual protections say.
Privacy obligations around images
Photographs routinely capture more than intended: bystanders, vehicle plates, documents on a desk, the inside of someone's home. Under the Privacy Act and the Australian Privacy Principles that is personal information you are collecting, storing and now processing.
Decide retention and access before you scale up. An inspection business accumulating images for years without a deletion policy is building a liability, and the fix is far cheaper at the start than after the fourth year of storage.
Where these projects fail
Almost always by aiming at judgement too early. A model asked to determine whether structural cracking is significant, or whether a defect breaches a code, will produce a confident answer that is sometimes wrong in ways a layperson cannot detect.
Keep it on description and extraction, and let the qualified person make the call. The commercial value is in removing the typing, not in replacing the expertise, and the professional liability sits with the person signing the report either way.
A sensible first month
Pick one document type with high volume and low variation. Run it in parallel with your existing process, compare the outputs, and count how often a person had to intervene and why.
Those intervention reasons become your routing rules. After a month you will know precisely which share of the work is safely automatable, which is a far better basis for a decision than a vendor accuracy claim.
What not to conclude
Vision capability does not mean the model understands what it is looking at in any professional sense. It describes appearances well and infers significance poorly, and the difference matters enormously in inspection, insurance and compliance work.
Nor does high accuracy on a demo set predict performance on your documents. Every business has its own worst-case: the supplier who sends photographs of paper invoices, the handwritten site note. Test on the ugly examples, because those are the ones that determine whether anyone keeps using it.
If your team is retyping documents or writing reports from photographs, book a short call and we will look at what is safely automatable.



