Blog

AI Pilot to Production: Claude's Four Key Decisions

September 2026 · 6 min read · AI Strategy

A small test box with a wavy line connected by an arrow to a larger structured box divided into tiers, showing a pilot moving into production
← Back to all posts

Only 23% of C-suite leaders say their organisation is getting sustained, enterprise-wide impact from AI. That figure comes from Accenture's July 2026 Pulse of Change report, and it anchors a new joint guide from Claude and Accenture on moving AI from pilot to production. A second number from the same guide is more useful to an Australian business: according to Accenture's September 2026 Tokenomics research, 42% of organisations have no single owner accountable for AI costs and outcomes.

The guide is written for CIOs at large enterprises. Its logic applies just as well to a $20 million Brisbane distributor or a Melbourne professional services firm, and in some ways applies more, because smaller firms have less slack to absorb a pilot that goes nowhere. The full guide is on Claude's blog.

Why do most AI pilots fail to reach production?

Most AI pilots fail to reach production because the pilot was never a fair test of production. Pilots run with hand-picked, enthusiastic teams, protected budgets and a tightly defined scope. Production has reluctant users, competing budgets and messy edge cases. A pilot that succeeds under those protected conditions tells you the model can do the task. It does not tell you the business can run it.

The guide's answer is to make seven decisions in a fixed order, split across three phases: before the pilot, during it, and once in production. Four of those decisions do most of the work, and they are the ones worth copying.

1. Define the job before you pick a tool

The guide asks for a four-part definition of the job the AI will do: the user, the task, the output, and a measurable quality threshold. Most pilots we are asked to rescue skipped the fourth part. Everyone agreed the output looked good, and nobody wrote down what good meant.

An illustrative example for an Australian insurance broker:

  • User: account managers handling renewals for commercial clients.

  • Task: compare the incumbent policy with up to four renewal quotes.

  • Output: a one-page comparison flagging changes in cover, excess and exclusions.

  • Quality threshold: every changed exclusion caught on a 50-file test set, with no more than one false flag per ten files.

With that last line written down before the pilot starts, the go or no-go decision at the end becomes arithmetic rather than a debate. We made the same argument in our four gates for prototype to production.

2. Cost the whole thing before the pilot starts

The guide recommends a lightweight total cost of ownership model built before the pilot, not after. That sounds obvious until you notice how rarely it happens. A pilot bill of a few hundred dollars a month hides the real run-rate costs: concurrency, growing context, monitoring, and the staff time spent reviewing output.

A useful minimum for a mid-sized Australian firm is a one-page model with setup cost, monthly model spend at full volume, review hours at full volume, and a named owner. If the pilot would cost $15,000 and the production run rate is $6,000 a month plus two days a week of review time, that is the number the business is actually approving. Our piece on where pilot costs balloon in production walks through the usual surprises, and our ROI calculator gives a quick first pass.

3. Match human review to the risk of the output

This is the most practical idea in the guide: a four-tier oversight model that sets how much human review each output gets, based on what happens if it is wrong.

Four-tier oversight model from the Claude and Accenture guide, with Australian examples
TierHuman involvementExample output
AutomatedNo routine review; monitored in aggregateTagging inbound emails by topic
SampledA person checks a regular sampleDrafting internal meeting summaries
ReviewedA person checks every output before it is usedCustomer-facing quotes and renewal letters
AdvisoryAI suggests; a person decides and owns the callCredit, claims or advice covered by ASIC or APRA obligations

The mistake we see most is putting everything in the reviewed tier during a pilot, then being unable to afford that review at scale. Deciding the tier per output type up front lets you model review cost honestly. Our guide to placing approval checkpoints goes deeper on where the human sits.

4. Name who owns each transition decision

The guide ends with a transition blueprint: what has to be decided, when, and who owns each decision. This is where the 42% statistic bites. When responsibility for AI is shared between IT and finance, nobody owns the trade-off between quality and cost, so nobody makes it, and the pilot drifts.

In an Australian business turning over $5 million to $50 million, the right owner is rarely the IT lead. It is usually the operations or department head whose team does the work, with finance consulted on cost. ABC Legal's rollout is a good example of what changes once ownership is clear, covered in our write-up of its 50-agent deployment.

Where the guide is thinner

Two gaps for local readers. The guide says little about data handling, and for any Australian business moving customer information through a model, Privacy Act obligations belong in the pre-pilot phase, not as a production afterthought. And the statistics are survey results from large enterprises, so treat 23% and 42% as direction rather than as benchmarks for your firm.

If you have a pilot that worked but has not moved, our AI readiness assessment is built around these four decisions, or you can book a call and talk through where it stalled.

FAQ

Frequently asked questions

What percentage of AI pilots reach production?

There is no single reliable figure, but Accenture's July 2026 Pulse of Change report found only 23% of C-suite leaders see sustained, enterprise-wide AI impact, which suggests most pilots do not scale.

What is a four-tier oversight model for AI?

It sorts AI outputs into automated, sampled, reviewed and advisory tiers, so the amount of human checking each output receives matches the harm it could cause if it were wrong.

Who should own an AI program in a mid-sized business?

Usually the head of the team whose work the AI changes, with finance consulted on cost. Accenture found 42% of organisations have no single accountable owner, which stalls cost and quality decisions.

When should you build a total cost of ownership model for AI?

Before the pilot starts. A pilot's small bill hides production costs such as concurrency, growing context, monitoring and human review time, which only a model built up front will surface.

How long should an AI pilot run?

Long enough to test against the quality threshold you defined at the start, often four to eight weeks. Duration matters less than having a written pass mark before day one.

Ready to move from AI pilot to production?

We help mid-market Australian businesses deploy AI automations that actually reach production and deliver measurable ROI.