OpenAI published an account this week of pausing access to one of its own long-running models after internal testing turned up a failure mode its standard evaluation suite hadn't caught. The team's stated conclusion was direct: pre-deployment testing built for models that answer a single prompt and stop isn't enough for models that act autonomously over long stretches. Persistence gives a model more opportunities to take an unwanted action before anyone notices, and ordinary spot-check testing wasn't built to catch that kind of drift.
The fix, according to OpenAI's own writeup, was to build new evaluations directly from the failures observed, strengthen alignment for long-horizon behaviour, add monitoring across the full sequence of actions rather than judging only the final output, and restore access gradually with more visibility and manual controls for users. None of that response is unusual on its own. What's useful here is that a major AI lab described the whole episode in public rather than quietly patching it and moving on.
Why this matters even if your business has never touched OpenAI's models
This is the same category of disclosure Anthropic has built its safety practice around for Claude: published system cards for each model release, a public Responsible Scaling Policy that sets out what testing has to happen before a more capable model ships, and ongoing red-teaming research aimed at finding a failure mode before a customer does. A second major lab describing an identical problem, in public, using the same vocabulary, is a signal worth reading regardless of which vendor's models sit inside your stack today.
Long-running, lightly supervised agents are becoming a normal part of how Australian businesses use AI: processing inbound leads overnight, reconciling invoices while the office is closed, drafting and queuing customer replies before anyone reviews them. Every vendor building tools like this is working through the same open problem, which is how you know an agent hasn't gone wrong when nobody is watching it in real time. OpenAI just showed its working. That's worth noticing.
The vendor diligence checklist worth running on every AI vendor you use
For a Sydney business handing an agent longer-running, less supervised tasks, whether that's claims processing, sales pipeline management, or unattended overnight workflows, the practical takeaway is a short list of questions to put to any vendor, including the one you already use.
Does the vendor publish what testing happens before a new model or agent capability ships, and what specifically that testing looks for?
Is there a documented, repeatable process for pausing or rolling back a model if a new failure mode turns up after release?
Does monitoring cover the full sequence of actions a long-running agent takes, or only the final output it hands back to you?
What visibility and manual control does your business actually retain once an agent is running unattended overnight or over a weekend?
If something goes wrong, how quickly will you find out, and from whom, rather than from an unhappy customer?
What skipping this checklist actually costs
An AI agent making one uncaught mistake overnight, in a workflow nobody was watching, can turn into a $40,000 to $100,000 cleanup once you count the error itself, the investigation, and the client conversation that follows. That range isn't hypothetical. It sits in a similar order of magnitude to the incident-response costs a business needs to be able to document under the Privacy Act, and for regulated entities, under APRA's CPS 230 operational risk standard. Getting the diligence checklist wrong isn't an abstract governance exercise. It's a line item, and a conversation with a regulator or a client that's much harder to have after the fact than before it.
What to actually do this quarter
None of this is a reason to avoid long-running agents. The productivity case for them is real and growing across Melbourne, Brisbane and Sydney businesses alike. It's a reason to build the same questions into how your business evaluates any AI vendor, existing or new, on a fixed schedule rather than only after something breaks.
Ask your current AI vendor, in writing, for their pre-deployment testing and rollback process. A vague answer is itself useful data.
Set a rule that no agent runs fully unattended on a business-critical workflow without a defined escalation path for when something looks wrong.
Review which of your current automations touch personal data covered by the Privacy Act, or fall inside APRA's operational risk scope, and confirm monitoring actually covers the full action sequence rather than just the final output.
Put a quarterly vendor review on the calendar now, rather than waiting for an incident to force one later.
A vendor that publishes its failures and fixes in public, the way OpenAI did here and the way Anthropic does with Claude's system cards, is giving your business more to work with than one that says nothing at all. If you'd like a second set of eyes on how your current automations hold up against this checklist, Automata AI runs a short AI vendor and governance review for Sydney businesses. You can book a session through the link below and we'll walk through the checklist against what you've actually got running.
Ready to talk it through? Book a brainstorm session and we'll go through your current AI vendors against the checklist above.



