Blog

Claude Fable 5.1 vs GPT-6 Astra: What to Weigh Up

September 2026 · 7 min read · AI Strategy

Line drawing of a tipped balance beam weighing a stack of blocks against one disc
← Back to all posts

OpenAI launched GPT-6 Astra on 3 September 2026, and its launch post does something unusual: it puts Claude Fable 5.1, Claude Fable 5 and Claude Opus 5 straight into its own benchmark tables. If you are picking a model to build agents on, that saves you assembling the comparison yourself. It also means you are reading a scoreboard kept by one of the players.

Below is what the tables actually say, what the categories predict about real work, and the part an Australian business has to test for itself regardless of who published the numbers.

How does GPT-6 Astra compare to Claude Fable 5.1?

On OpenAI's own published benchmark tables, GPT-6 Astra leads on most computer-use and cybersecurity evaluations, while Claude models stay competitive or ahead on parts of the coding set and on alignment-related measures. OpenAI describes Astra as its most intelligent and aligned model and claims leading results across computer use, browsing, software engineering, cybersecurity and professional work. Every one of those figures is self-reported by OpenAI, measured on its chosen evaluations, and none of them has been independently reproduced at the time of writing in September 2026.

Four benchmark families and what each one predicts

The tables group into four categories, and they predict very different things about the work an agent will do for you. A model that wins one category can be the wrong choice if your work sits in another.

Benchmark families in OpenAI's launch tables and the work each one stands in for

Benchmark families compared across Claude and GPT-6 Astra in OpenAI's own tables
FamilyEvaluations namedReported leaderWhat it predicts
CodingTerminal-Bench, DeepSWE, FrontierCodeMixed, Claude competitiveMulti-step work in a real repository
Computer useOSWorld, ScreenSpot-Pro, Agents' Last ExamAstra on mostDriving software with no API
ScienceFrontierMath, GPQA DiamondAstra on the headline scoresHard reasoning on closed problems
CybersecurityExploitBench, ExploitGym, SRE-BenchAstra on mostOffensive and operational security tasks

Three headline claims sit above those tables: 98 per cent on FrontierMath Tier 4, 99.9 per cent on ARC-AGI-3, and 100 per cent on ExploitBench. Read a perfect score as information about the benchmark rather than the model. Once an evaluation is saturated it has stopped separating anything, and the next comparison has to be made somewhere else.

Why the source of the scoreboard changes how you read it

Vendor-published comparisons are not worthless. They are a claim made in public that competitors can contest, which is more accountability than a sales deck offers. But the vendor chooses which evaluations appear, which versions of each competitor are tested, and which configuration each model runs in, and those choices do a lot of work.

  • The evaluation set is selected by the vendor, so absent categories are a result too.

  • Competitor models are run by the party with an interest in the outcome, not by their makers.

  • Effort and configuration settings change results substantially and are rarely matched across entrants.

  • A benchmark measures a task someone designed, which is never quite the task your business runs.

  • Astra's tables report lower misalignment and circumvention rates for Claude than for OpenAI's own previous model, which is a useful signal precisely because it does not flatter the publisher.

That last point is the one to hold onto. When a vendor's own numbers favour a competitor on a measure, that measure is worth more than the ones where the vendor wins. It is the same reasoning we apply in Claude Code versus GPT-5.6 for AU dev teams.

What actually decides an agent stack for a business

Astra is available through ChatGPT Plus, Pro, Business and Enterprise, the OpenAI API, Microsoft Azure and AWS Bedrock. That breadth of distribution matters more to some buyers than any benchmark, because it decides whether the model can sit inside infrastructure you already run under contracts you have already signed.

For most Australian mid-market businesses the decision comes down to four things, and only one of them appears on a scoreboard.

  • Whether the model reliably finishes your work, measured on your own tasks over a fortnight.

  • Whether it runs where your data is allowed to be, under agreements your board already accepts.

  • Whether your team can operate it, review its output, and correct it without specialist help.

  • Whether it holds up under load, which is an operational question rather than a capability one.

We run a two-week bake-off on a client's real work rather than on published evaluations, typically for between $9,000 and $18,000 for a Sydney team, because a model that wins a benchmark and loses on your document formats has lost. Reliability under real load is a separate matter again, covered in the operational cost of cheap AI. If you would rather scope that yourself first, start with our assessment and the approach on our services page.

What this comparison does not settle

It does not settle cost, which moves independently of capability and has to be modelled on your own token profile. It does not settle safety posture or governance, which deserve their own review rather than a line in a benchmark table. And it does not settle model choice for a specific job, because Claude Opus 5 and Claude Fable 5 are still in these tables for a reason: the strongest model is not always the right one for a given task.

What it does give you is a structure. Decide which of the four families your work actually lives in, test the shortlist on that work, and treat every published figure as a starting hypothesis. OpenAI's GPT-6 Astra launch post carries the full tables, and Claude Opus 5 on frontier performance covers the other side of the shortlist.

FAQ

Frequently asked questions

Does GPT-6 Astra beat Claude Fable 5.1 on benchmarks?

On OpenAI's own tables Astra leads most computer-use and cybersecurity evaluations, while Claude models remain competitive or ahead on parts of the coding set and on alignment-related measures. The figures are self-reported by OpenAI.

When was GPT-6 Astra released?

OpenAI launched GPT-6 Astra on 3 September 2026, describing it as its most intelligent and aligned model, with claimed leading results across computer use, browsing, software engineering, cybersecurity and professional work.

Which Claude models appear in OpenAI's comparison tables?

Claude Fable 5.1, Claude Fable 5 and Claude Opus 5 all appear, alongside Gemini 3.8 Flash, across coding, computer use, science and cybersecurity benchmark families in the GPT-6 Astra launch post.

Where can businesses access GPT-6 Astra?

OpenAI lists availability through ChatGPT Plus, Pro, Business and Enterprise plans, the OpenAI API, Microsoft Azure and AWS Bedrock, which matters for buyers with existing cloud agreements in place.

Should a benchmark score decide which model a business builds on?

No. Benchmarks measure designed tasks, not your workload. Test a shortlist on your own work over a fortnight, and weigh data location, team operability and reliability under load alongside any published score.

Ready to move from AI pilot to production?

We help mid-market Australian businesses deploy AI automations that actually reach production and deliver measurable ROI.