Most vendor risk questionnaires have a line that reads something like "describe how your AI model is tested for safety". Until now, every answer to it has come from the vendor's own team. On 18 September 2026, the company behind Claude changed that answer.
Anthropic announced a partnership with Accenture, through its AI arm Faculty, to run independent, embedded evaluation of Claude and of future frontier models. You can read the announcement itself. This post covers what it means for an Australian organisation choosing an AI platform, and what it does not mean.
What is embedded evaluation of Claude?
Embedded evaluation means outside evaluators work inside the AI company instead of visiting it. Under the arrangement announced in September 2026, Accenture's evaluators get access similar to an employee's. They can watch Claude models take shape during training, review the decisions behind deployment calls, and talk directly to the staff who build the models. Both companies expect to invest at least $1 billion over five years building this capability.
The contrast is with the occasional spot-check, where an auditor is engaged for a fixed window, sees what it is shown, and leaves. An evaluator who is present throughout can catch things a quarterly visit cannot, and can report incidents independently of the vendor's communications team.
| Approach | Who does the checking | What they can see | Main weakness |
|---|---|---|---|
| Self-reported | The vendor's own team | Everything | You are trusting the summary |
| Spot-check audit | Outside auditor, occasionally | What is shown in the audit window | Misses what happens between visits |
| Embedded evaluation | Outside evaluators, inside the company | Training, deployment decisions, staff | New practice, norms still forming |
Why it matters for vendor due diligence
Most AI safety claims today are self-reported. A vendor publishes a model card, runs its own red-teaming, and asks you to trust the write-up. That is not worthless, but it is the weakest form of evidence a procurement team accepts in any other category. Nobody takes a supplier's word on its own financial accounts.
Embedded evaluation is a different kind of evidence, not a bigger pile of the same kind. So it gives you a new question to ask every AI vendor on your shortlist, not only Claude: do you have independent evaluators inside the building, and can they report without your sign-off?
We have written before about the vendor safety questions every AU business should be asking and about Claude's safety incident reports as a trust signal. This announcement adds one more line to that list, and it is the hardest one for a vendor to fake.
Where Australian buyers will feel it first
Anthropic has been explicit that funding and access norms for this kind of evaluation do not yet exist as an industry standard. It is building the practice and inviting scrutiny of how it is done. For Australian organisations, the effect shows up in procurement before it shows up anywhere else.
Financial services. Firms regulated by APRA and ASIC already have to show how they manage risk from material service providers. "Checked by someone other than the vendor" is an easier story to tell a risk committee than a model card.
Health and government-adjacent work. Tender panels increasingly ask whether safety claims are independently verified, on top of whether the model performs well.
Mid-market firms selling into either. Your customers' questionnaires flow down to you. If your product runs on an AI model, their question about model assurance becomes your question.
Expect that question to appear in more Australian government and enterprise AI tenders over the next 12 months. It also sits comfortably beside the guardrails in the Voluntary AI Safety Standard, which ask organisations to understand and test the AI systems they deploy.
What not to conclude from this
Three cautions, because an announcement like this is easy to over-read.
It does not make your deployment safe. Independent evaluation looks at the model and at how the vendor decides to release it. It says nothing about your prompts, your data permissions, or whether a staff member pastes client records into the wrong tool. Those controls are still yours.
It is not exclusive. Anthropic frames the arrangement as non-exclusive. Other evaluators and other AI labs are expected to adopt similar arrangements over time. Treat it as a moving bar across the market, not a permanent Claude-only feature.
It is early. The capability is being built over five years. Ask what has been evaluated so far and what has been published, and write down the answer with a date beside it.
A worked example for a procurement team
Take an illustrative case. A Sydney insurance broker with 150 staff is about to commit $180,000 a year to a Claude rollout across claims and client service. The risk committee wants one page on model assurance before it signs.
Before September 2026, that page would have listed the vendor's published safety documents and stopped. Now it can carry three rows: what the vendor says about its own testing, who outside the vendor checks that testing and with what access, and what the broker itself controls. The second row used to be blank for every vendor. A committee can tell the difference between a blank row and a filled one in about ten seconds.
The same one-page format works for comparing vendors. Ask each the same three questions and put the answers side by side. Where a vendor can only fill the first row, you have learned something useful about how much of its safety story rests on trust.
What to do this quarter
Add one line to your AI vendor questionnaire: who independently evaluates your models, what access do they have, and can they report without your approval? Then check your own half of the page. Permissions, logging and staff guidance are the parts no outside evaluator covers.
At Automata AI we apply Claude to work that costs Australian mid-market businesses time and money, and we take deployments past the pilot stage into production. If governance and vendor evaluation questions are coming up in your procurement conversations, book a conversation with us before you lock in a platform. You can also see how we structure rollouts on our services page.



