Cost Calculator

Work out what your AI models should cost

Enter what you spend now, choose where the work moves, and see the difference. 17 models across 6 providers, priced from published USD list prices.

Calculate your savings

Enter what you spend today, then choose where the work moves. All figures are USD.

A document, a system prompt and a few tool definitions in context.

US$0.1305 per task on this profile

US$0.0522 per task on this profile

60%
0%100%

Same volume of work, repriced. US$0.0835 per task 36.0% cheaper than Claude Opus 5 alone.

US$0.0835
Blended cost per task
36.0%
Spend reduction
US$9,000
Saved per month
US$108,000
Saved per year

Modelled on published list prices and the task profile above. The Cost view reprices the same tasks; the Tasks view holds your budget fixed. Actual results depend on your workload.

Get your team's estimate →

Free 30-minute session. We will look at your actual usage, not a model of it.

Methodology

Every figure traces back to a published price list

There is no proprietary benchmark behind these numbers and no rate we have negotiated on your behalf. Each model carries its vendor's own published per-million-token prices for input, cached input and output, last verified on 22 August 2026. Cost per task applies those rates to a task profile you can see and change.

Cache writes are left out deliberately. In a steady-state production workload the same prefix is written once and read thousands of times, so the write cost amortises to almost nothing. A workload that never repeats its context will cost more than this model shows, which is itself a finding worth acting on.

The Cost view reprices the same volume of work. The Tasks view holds your budget fixed and shows how much more work fits inside it. Both read from the same blended rate, so they are two views of one calculation rather than two calculations. What they cannot know is your actual traffic mix, which is where a model of a bill and a real bill start to diverge.

The landscape

Capability and cost do not move together

Plotted on published prices and published scores, the models do not form a straight line. Several sit well below the trend, delivering most of the capability of the tier above them at a fraction of the cost. That gap is what the calculator above is pricing.

50%55%60%65%70%75%$0.005$0.010$0.025$0.050$0.10$0.25Cost per task, USD — logarithmic scaleCursorBench 3.2 scoreGPT-5.6 Luna (OpenAI) — US$0.00582 per task, 61.1%GPT-5.6 LunaGemini 3.7 Flash (Google) — US$0.0196 per task, 61.6%Gemini 3.7 FlashGemini 3.6 Flash (Google) — US$0.0196 per task, 53.5%Gemini 3.6 FlashGLM 5.2 (Z.ai) — US$0.0313 per task, 55.0%GLM 5.2Grok 4.6 (xAI) — US$0.0465 per task, 70.8%Grok 4.6Claude Sonnet 5 (Anthropic) — US$0.0522 per task, 61.5%Sonnet 5GPT-5.6 Terra (OpenAI) — US$0.0582 per task, 64.9%GPT-5.6 TerraKimi K3 (Moonshot) — US$0.0783 per task, 60.8%Kimi K3GPT-5.6 Sol (OpenAI) — US$0.1044 per task, 67.2%GPT-5.6 SolClaude Opus 5 (Anthropic) — US$0.1305 per task, 70.0%Opus 5Claude Opus 4.8 (Anthropic) — US$0.1305 per task, 62.3%Opus 4.8GPT-5.5 (OpenAI) — US$0.1455 per task, 58.4%GPT-5.5Claude Fable 5 (Anthropic) — US$0.2610 per task, 70.5%Fable 5

Cost per task is derived from each vendor's published list prices on the standard automation profile below. Scores are published CursorBench 3.2 results, highest-effort variant of each model, from cursor.com/evals. Cursor evaluates agents on ambiguous, multi-file tasks drawn from real sessions. Four models in the table below carry no published score, so they are priced but not plotted: GLM 4.7, Claude Haiku 4.5, Grok 4.5, Gemini 3.1 Pro Preview.

Published list prices, all providers

Published API list prices in USD per million tokens, with cost per task derived on the standard automation profile.
ModelProviderInputCached inOutputCost / taskScore
GPT-5.6 LunaOpenAI$0.20$0.020$1.20US$0.0058261.1%
GLM 4.7Z.ai$0.60$0.11$2.20US$0.0143
Gemini 3.7 Flash*Google$0.75$0.075$3.75US$0.019661.6%
Gemini 3.6 Flash*Google$0.75$0.075$3.75US$0.019653.5%
Claude Haiku 4.5Anthropic$1.00$0.10$5.00US$0.0261
GLM 5.2Z.ai$1.40$0.26$4.40US$0.031355.0%
Grok 4.5xAI$2.00$0.30$6.00US$0.0423
Grok 4.6xAI$2.00$0.50$6.00US$0.046570.8%
Claude Sonnet 5Anthropic$2.00$0.20$10.00US$0.052261.5%
GPT-5.6 TerraOpenAI$2.00$0.20$12.00US$0.058264.9%
Gemini 3.1 Pro PreviewGoogle$2.00$0.20$12.00US$0.0582
Kimi K3Moonshot$3.00$0.30$15.00US$0.078360.8%
GPT-5.6 SolOpenAI$4.00$0.40$20.00US$0.104467.2%
Claude Opus 5Anthropic$5.00$0.50$25.00US$0.130570.0%
Claude Opus 4.8Anthropic$5.00$0.50$25.00US$0.130562.3%
GPT-5.5OpenAI$5.00$0.50$30.00US$0.145558.4%
Claude Fable 5Anthropic$10.00$1.00$50.00US$0.261070.5%

Input, cached input and output are USD per million tokens. Cost per task applies those rates to the standard automation profile: 30,000 input tokens at a 70% cache hit rate and 3,000 output tokens.

Promotional pricing published through 31 December 2026. Standard rates resume 1 January 2027.

Priced in context tiers. The rate shown is the tier below 200K tokens; longer contexts are charged at roughly double.

Sources: platform.claude.com, developers.openai.com, ai.google.dev, docs.x.ai, platform.kimi.ai, docs.z.ai.

Where the savings come from

Three changes account for most of it

Route by difficulty

The largest single lever

Most production systems send every request to whichever model was chosen on day one. Classification, extraction and routing rarely need a frontier model; multi-step reasoning does. Splitting a workload across two or three models by difficulty is usually the difference between a bill that scales and one that does not.

Cache the repeated context

Cheapest change to make

System prompts, tool definitions and reference documents are re-sent on every call. Cached input costs roughly a tenth of standard input across most providers, so a stable prefix pays for itself after one or two reads. Move the volatile parts of a prompt to the end and the cache stops being invalidated on every request.

Batch what is not urgent

Half price, asynchronously

Overnight enrichment, backfills, document classification and report generation do not need an answer in two seconds. Several providers publish an asynchronous batch tier at half the standard input and output rate. The work is identical; only the latency guarantee changes.

The Australian view

A USD bill lands on an AUD profit and loss

Every provider on this page invoices in US dollars, so a model bill that looks flat in USD moves with the exchange rate in your accounts. At meaningful volume that is a real budgeting problem: a ten per cent currency move is a ten per cent cost move you did not cause and cannot control. Whether GST applies to overseas digital services depends on your registration and how the supply is treated, which is a question for your accountant rather than an assumption to build a budget on.

For regulated clients there is a second constraint that outranks price. Under APRA CPS 230 the material service providers behind a critical operation have to be identified and managed, and where inference physically runs is part of that conversation. Some providers offer regional or US-pinned inference at a premium; the cheapest model on the table is not always one you can put a customer record through. Model choice is a control decision as much as a cost decision.

If the question you are actually asking is what a manual process costs before any of this applies, the AI automation ROI calculator answers that one in AUD. If you already have engineers building against these APIs, our Claude Code consulting work covers the routing and caching patterns directly, and the engagement tiers explain how a larger piece of work is scoped.

FAQ

Frequently asked questions

Where do the prices in this calculator come from?

Every rate is the vendor's own published API list price, read from their pricing documentation and last verified on 22 August 2026. Nothing here is a negotiated rate, a reseller margin or an estimate. Sources are linked under the comparison table so you can check any figure yourself.

How is cost per task calculated?

By applying each model's published per-million-token rates to a stated task profile. The default profile assumes 30,000 input tokens at a 70 per cent cache hit rate and 3,000 output tokens, which is typical of an automation step carrying a system prompt, tool definitions and a document. You can switch profiles in the calculator, and the assumption is shown on screen rather than buried in a footnote.

Does switching models mean losing quality?

Sometimes, which is why the chart plots published benchmark scores against cost rather than showing cost alone. The useful question is not which model is best overall but which model is sufficient for each class of work in your system. A task that only classifies an email does not get better on a frontier model; a task that reasons across a contract usually does.

Why are the prices in USD when we are an Australian business?

Because every major model provider bills in USD. Converting to AUD on this page would bake in an exchange rate that is wrong the following week. Your actual AUD cost is the figures here multiplied by whatever rate your card or invoice settles at, which is worth modelling separately if model spend is a material line item.

How accurate is this against a real invoice?

It is a model, not a bill. It will be close when your workload resembles the selected profile and your cache hit rate is roughly what you set. It will be optimistic if your prompts change on every request so nothing caches, and pessimistic if you already batch heavily. Treat the output as a direction and a rough magnitude, then check it against a month of real usage data.

Can Automata AI do this analysis on our actual usage?

Yes. We pull your real token usage by endpoint and by task type, work out where the spend actually sits, and implement the routing, caching and batching changes rather than handing over a recommendation. Most of the saving in practice comes from three or four high-volume paths, not from an across-the-board model swap.

Next step

Price it on your actual usage, not a profile

Book a free 30-minute call. We'll look at where your token spend really sits and what routing, caching and batching are worth against your own traffic.