Public sector interest in open-weight models is driven by something private businesses rarely have to weigh: an obligation to explain decisions to the public, and a procurement process that treats vendor dependency as a risk in its own right. Those two pressures push agencies toward models they can run and inspect themselves, sometimes past the point where it makes practical sense.
Why agencies look at open weights at all
Not cost, usually. The drivers are control and defensibility, both of which matter more in a setting where a decision may be reviewed years later by someone with statutory powers.
Processing can be kept inside infrastructure the agency controls
The model can be frozen, so a decision made in 2026 can be reproduced in 2029
No vendor-side change alters behaviour without notice mid-programme
Nothing about the arrangement depends on a commercial relationship continuing
That third point is the one that comes up in review. A hosted model that improves quietly is a feature commercially and a problem when you have to demonstrate what the system did at a particular moment.
Where the argument overreaches
Self-hosting solves reproducibility and it introduces an operational burden that agencies are frequently not staffed for. Running inference infrastructure properly means capacity planning, patching, monitoring and someone accountable when it fails at four in the morning.
For a mid-sized council or department that is a standing commitment of $150,000 to $400,000 a year once you count the people, and it competes directly with the service delivery the programme was meant to improve. Plenty of agencies would be better served by a hosted arrangement with strong contractual terms.
What actually needs answering in procurement
Most agency AI procurement questions are not about model quality. They are about where processing occurs, what is retained, who can access it, how decisions are logged, and what happens to the data when the contract ends.
A vendor who answers those crisply will get further than one with better benchmark numbers. If you are selling into this market, the questionnaire is the product conversation, not a formality afterwards.
Transparency obligations are the real constraint
Administrative decisions affecting people generally need to be explainable, and freedom of information regimes mean the working can be requested. A system that produces an outcome nobody can account for is a problem regardless of how accurate it is on average.
The practical implication is that AI in a public sector context belongs in preparation and analysis rather than in the decision itself. Drafting an assessment for an officer to review is defensible; automating the determination is a different proposition entirely.
Where the value actually is
The highest-return public sector uses are unglamorous and internal: summarising submissions, drafting correspondence, triaging enquiries, checking documents against a checklist, making a decade of records searchable.
A council receiving 400 submissions on a planning proposal spends weeks reading them. Producing a themed summary with every claim traceable to a source submission does not make the decision, it makes the human work tractable, and it is reviewable in a way an automated determination is not.
Procurement reality in Australia
Panel arrangements and standing offers shape what agencies can buy and how quickly. Where a relevant panel exists, being on it matters more than any capability argument; where one has lapsed or does not cover a category, agencies fall back to their own procurement thresholds.
Either way, check the current status directly rather than relying on what was true last year. These arrangements change, and building a bid around a panel that has since expired wastes a quarter.
Start with a bounded internal pilot
The pattern that survives scrutiny is narrow, internal, and human-reviewed: one team, one document type, a defined period, and a written evaluation at the end covering accuracy, time saved and what went wrong.
That evaluation is the asset. It is what makes the second, larger programme approvable, and it is considerably more persuasive to a risk committee than a vendor case study from another jurisdiction.
What not to conclude
Open weights are not automatically the compliant choice, and hosted models are not automatically disqualified. The obligations are about control, transparency and accountability, and a well-governed hosted arrangement can satisfy all three while a poorly run internal deployment satisfies none.
Nor should any of this replace the officer. The consistent lesson from Australian public sector deployments is that AI earns its place preparing work for a human decision-maker, and creates problems the moment it is positioned as the decision-maker itself.
If your agency is scoping a pilot and needs it to survive a risk review, book a short call and we will look at what is defensible and what is not.



