OpenAI launched GPT-6 Astra on 3 September 2026, and its launch post does something unusual: it puts Claude Fable 5.1, Claude Fable 5 and Claude Opus 5 straight into its own benchmark tables. If you are picking a model to build agents on, that saves you assembling the comparison yourself. It also means you are reading a scoreboard kept by one of the players.
Below is what the tables actually say, what the categories predict about real work, and the part an Australian business has to test for itself regardless of who published the numbers.
How does GPT-6 Astra compare to Claude Fable 5.1?
On OpenAI's own published benchmark tables, GPT-6 Astra leads on most computer-use and cybersecurity evaluations, while Claude models stay competitive or ahead on parts of the coding set and on alignment-related measures. OpenAI describes Astra as its most intelligent and aligned model and claims leading results across computer use, browsing, software engineering, cybersecurity and professional work. Every one of those figures is self-reported by OpenAI, measured on its chosen evaluations, and none of them has been independently reproduced at the time of writing in September 2026.
Four benchmark families and what each one predicts
The tables group into four categories, and they predict very different things about the work an agent will do for you. A model that wins one category can be the wrong choice if your work sits in another.
Benchmark families in OpenAI's launch tables and the work each one stands in for
| Family | Evaluations named | Reported leader | What it predicts |
|---|---|---|---|
| Coding | Terminal-Bench, DeepSWE, FrontierCode | Mixed, Claude competitive | Multi-step work in a real repository |
| Computer use | OSWorld, ScreenSpot-Pro, Agents' Last Exam | Astra on most | Driving software with no API |
| Science | FrontierMath, GPQA Diamond | Astra on the headline scores | Hard reasoning on closed problems |
| Cybersecurity | ExploitBench, ExploitGym, SRE-Bench | Astra on most | Offensive and operational security tasks |
Three headline claims sit above those tables: 98 per cent on FrontierMath Tier 4, 99.9 per cent on ARC-AGI-3, and 100 per cent on ExploitBench. Read a perfect score as information about the benchmark rather than the model. Once an evaluation is saturated it has stopped separating anything, and the next comparison has to be made somewhere else.
Why the source of the scoreboard changes how you read it
Vendor-published comparisons are not worthless. They are a claim made in public that competitors can contest, which is more accountability than a sales deck offers. But the vendor chooses which evaluations appear, which versions of each competitor are tested, and which configuration each model runs in, and those choices do a lot of work.
The evaluation set is selected by the vendor, so absent categories are a result too.
Competitor models are run by the party with an interest in the outcome, not by their makers.
Effort and configuration settings change results substantially and are rarely matched across entrants.
A benchmark measures a task someone designed, which is never quite the task your business runs.
Astra's tables report lower misalignment and circumvention rates for Claude than for OpenAI's own previous model, which is a useful signal precisely because it does not flatter the publisher.
That last point is the one to hold onto. When a vendor's own numbers favour a competitor on a measure, that measure is worth more than the ones where the vendor wins. It is the same reasoning we apply in Claude Code versus GPT-5.6 for AU dev teams.
What actually decides an agent stack for a business
Astra is available through ChatGPT Plus, Pro, Business and Enterprise, the OpenAI API, Microsoft Azure and AWS Bedrock. That breadth of distribution matters more to some buyers than any benchmark, because it decides whether the model can sit inside infrastructure you already run under contracts you have already signed.
For most Australian mid-market businesses the decision comes down to four things, and only one of them appears on a scoreboard.
Whether the model reliably finishes your work, measured on your own tasks over a fortnight.
Whether it runs where your data is allowed to be, under agreements your board already accepts.
Whether your team can operate it, review its output, and correct it without specialist help.
Whether it holds up under load, which is an operational question rather than a capability one.
We run a two-week bake-off on a client's real work rather than on published evaluations, typically for between $9,000 and $18,000 for a Sydney team, because a model that wins a benchmark and loses on your document formats has lost. Reliability under real load is a separate matter again, covered in the operational cost of cheap AI. If you would rather scope that yourself first, start with our assessment and the approach on our services page.
What this comparison does not settle
It does not settle cost, which moves independently of capability and has to be modelled on your own token profile. It does not settle safety posture or governance, which deserve their own review rather than a line in a benchmark table. And it does not settle model choice for a specific job, because Claude Opus 5 and Claude Fable 5 are still in these tables for a reason: the strongest model is not always the right one for a given task.
What it does give you is a structure. Decide which of the four families your work actually lives in, test the shortlist on that work, and treat every published figure as a starting hypothesis. OpenAI's GPT-6 Astra launch post carries the full tables, and Claude Opus 5 on frontier performance covers the other side of the shortlist.



