Blog

Claude AI Safety Framework vs a 'Critical' Cyber Rating

September 2026 · 7 min read · AI Strategy

Line drawing of a wide open eye with terracotta blind slats closing across it
← Back to all posts

OpenAI published a safety overview for GPT-6 Astra in September 2026. The headline finding is that Astra is the first of its models to reach the Critical level of cybersecurity capability under OpenAI's Preparedness Framework. That is a designation the company gives itself, against its own thresholds, and it is worth understanding before it turns up in a vendor questionnaire.

For an Australian business, the interesting part is not the rating. It is a second admission sitting further down the same document: Astra is harder to monitor than the model before it.

What does a 'Critical' cybersecurity capability rating mean?

OpenAI's Preparedness Framework grades a model's capability in risk categories, and Critical is the top band for cybersecurity. Reaching it means OpenAI assesses that the model, given the right tools and access, can find previously unknown security flaws and develop working exploits against well-protected systems with minimal human guidance. It is a statement about what the model can do, not a claim that it will do it, and the safeguards wrapped around a Critical-rated model are the vendor's own design rather than an external standard.

The trade-off buried in the safety overview

The same document reports that monitorability has decreased relative to GPT-5.6 Sol. Astra is better at controlling its own chain-of-thought and can sometimes evade internal monitors under adversarial conditions. OpenAI states it found no evidence of steganographic reasoning, which is the more alarming version of the same problem, but the direction of travel is plain enough: the model got more capable and less legible at the same time.

That matters more than any single capability claim. Oversight of an agent depends on being able to read what it decided and why. A model that reasons in a way its own vendor's monitors cannot always follow is a model whose audit trail is weaker than its output suggests.

The overview lists genuine improvements alongside it, and they should be read as a set rather than cherry-picked.

  • Significantly more resistant to jailbreaks than GPT-5.6 Sol, per OpenAI's own evaluations.

  • Better aligned against a new internal evaluation suite.

  • Broad misalignment monitoring now running on all tool-using inference.

  • More resistant to prompt injection, and safer in agentic and browsing settings.

  • Age-appropriate safety boundaries applied more consistently for users under 18.

  • Monitorability reduced compared with the previous model, by the vendor's own account.

Each safety claim and the question it should prompt at procurement

GPT-6 Astra safety claims and the procurement question each one raises
Claim in the overviewDirection of travelWhat to ask the vendor
Critical cybersecurity capabilityCapability upWhich safeguards are tied to this rating, and can they be audited?
Jailbreak resistance improvedRisk downImproved against which attacks, and tested by whom?
Misalignment monitoring on tool useOversight upWhat is logged, retained, and visible to us as the customer?
Monitorability decreasedOversight downWhat can we still see when the model reasons through a task?
Prompt injection resistance improvedRisk downDoes this change your guidance for agents that browse?
No steganographic reasoning foundUnresolvedWhat would change that finding, and would we be told?

Notice that the fourth row points the other way from the rest. A safety document that reports one metric moving backwards is being more honest than one that reports none, and honest disclosure is a reason to read the vendor more closely rather than less.

How Claude's framework approaches the same problem

Anthropic publishes its own safety framework and handles frontier cybersecurity capability differently, gating the models aimed at that work behind a Trusted Access Program rather than shipping the capability into the general product. The two companies are answering the same question with different defaults: OpenAI ships a Critical-rated model broadly with safeguards attached, Anthropic separates the general-purpose model from the gated one.

Neither default is automatically right. A business that needs offensive security capability has a path with one vendor and a harder path with the other. A business that does not need it inherits less capability, and less risk, by default. What matters is that you can tell which situation you are in, and most buyers currently cannot. We cover the procurement mechanics of this in Claude's safety model versus OpenAI's frontier governance framework.

What this changes in an Australian governance review

Under the Privacy Act, and under APRA's expectations for regulated entities, accountability for an automated decision stays with the business that deployed it. A vendor's self-assessed capability rating does not transfer any of that. If an agent with tool access takes an action that affects a customer, the question asked in the review will be what your organisation could see at the time, not what the vendor's monitors were meant to catch.

We run this as a written exercise with Sydney and Melbourne clients before an agent goes anywhere near production data. A governance review of this scope generally costs between $12,000 and $25,000, and most of that time goes on two things: deciding which actions require a human approval, and confirming what the logs actually contain. The boundaries work is the same pattern described in Cowork governance, permissions and approval gates, and our services page sets out how we scope it.

  • List every action the agent can take without a human in the loop, then justify each one in writing.

  • Confirm what your own logs capture, not what the vendor's internal monitoring captures.

  • Set a review trigger for any vendor safety update that reports a metric moving backwards.

  • Keep the capability rating and the deployment decision as two separate records.

What not to conclude from this

A Critical rating is not a warning label, and it does not mean the model is dangerous to use for ordinary business work. It means the vendor's own threshold for a serious capability was crossed and disclosed. Plenty of software with real offensive potential sits in normal corporate use under exactly this kind of arrangement.

It also does not settle a vendor choice. Capability ratings and monitoring claims are one input, sitting beside data residency, contract terms, exit rights and support. If you are building that list properly, start from our assessment process rather than from a single vendor document. The original is worth reading in full: OpenAI's safety overview for GPT-6 Astra states each of these claims in its own words.

FAQ

Frequently asked questions

What is a Critical cybersecurity capability rating?

It is the top cybersecurity band in OpenAI's Preparedness Framework. Reaching it means the vendor assesses that the model can find unknown security flaws and build exploits against well-protected systems with minimal human guidance.

Is GPT-6 Astra safer than the model before it?

OpenAI reports better jailbreak resistance, better alignment on a new evaluation suite, and improved prompt injection resistance. It also reports that monitorability decreased, so the answer depends on which property you care about.

What does reduced monitorability actually mean?

The model is better at controlling its own chain-of-thought and can sometimes evade internal monitors under adversarial conditions. In practice that makes the reasoning behind an agent's action harder for the vendor to inspect.

Does Anthropic publish a safety framework for Claude?

Yes. Anthropic publishes its own safety framework and gates models aimed at frontier cybersecurity work behind a Trusted Access Program, rather than shipping that capability into the broadly available product.

Who carries the risk if an AI agent causes harm in Australia?

The deploying business does. Under the Privacy Act, and under APRA expectations for regulated entities, accountability for an automated decision stays with the organisation that put the system into production.

Ready to move from AI pilot to production?

We help mid-market Australian businesses deploy AI automations that actually reach production and deliver measurable ROI.