Blog

Claude Code CI Scaling: What 25x Job Growth Broke

September 2026 · 6 min read · Technical

An exponential curve rising past three short horizontal steps, each shorter than the last, showing shrinking runway from quick fixes
← Back to all posts

A service that had been fine for years ran out of headroom in 70 days. The fix for that bought 29 days. The fix after that bought less than one. That is the short version of a Claude engineering write-up on what happened to Anthropic's CI once Claude was writing most of the code, and it is the most useful thing published this month for any team planning to scale Claude Code.

What happened to Anthropic's CI

Over six months, CI job volume rose 25 times. Claude was by then authoring around 80% of the code and doing much of the pull request review, so every agent-driven change triggered builds and tests at a pace no human team produces. The piece that buckled was test impact analysis, the service that decides which tests need to run for a given change so the whole suite does not run on every commit.

That service has two halves. A listener records test results as they come in. A selector uses that history to decide which tests to run on each pull request. Both were designed for human commit rates. Neither was designed for a codebase where agents open changes around the clock. The full write-up, by engineer Sachin Malhotra, is on Claude's engineering blog.

Why does agentic coding put so much load on CI?

Because CI cost scales with the number of changes, not the number of engineers. A human developer might push a handful of commits a day. An agent working through a backlog can open, revise and re-run changes continuously, and each revision triggers builds and test selection. When an agent also reviews pull requests and requests fixes, the loop tightens further. Headcount stays flat while CI demand multiplies.

This is the mistake teams make when they budget for Claude Code. They model seat costs and token costs, then discover the real bill lands in CI minutes, runner capacity and whichever internal service sits in the critical path of every merge.

Three patches, each worth less than the last

The team did what most teams would do first: patch the existing design. The pattern of diminishing returns is the lesson.

Runway bought by each quick fix to Anthropic's test selection service before redesign
Fix triedRunway it boughtWhy it stopped working
Move to a bigger machineAbout 70 daysVertical scaling has a ceiling; load kept compounding
Shard the serviceAbout 29 daysShards still shared the same growth curve
Restart the service dailyLess than a dayA stopgap that treated the symptom, not the design
Stateless redesign on an in-memory data storeScales horizontallyBuilt for the growth curve rather than against it

Each patch was reasonable in isolation. Together they show what happens when load grows exponentially and fixes are linear: every fix is smaller than the last in calendar time, even if it took the same engineering effort.

How the redesign got built in three weeks

The replacement moved test impact analysis to a stateless architecture that scales horizontally, backed by an in-memory data store. It was built in three weeks, with Claude doing much of the implementation autonomously once it had a long-running monitoring session through Claude Tag. We have written before about how Anthropic uses Claude Tag internally, and this is a strong example of the pattern: the agent watches the service, sees what breaks, and fixes it in small increments.

That only worked because the service was instrumented well enough for Claude to diagnose itself. An agent cannot fix what it cannot observe. Logs, metrics and clear error messages stop being nice-to-haves and become the interface the agent works through.

The two design rules the write-up lands on

  • Plan for the exponential. Once agents write and review most code, give version-zero designs 10 to 20 times headroom rather than the usual two or three.

  • Instrument for the agent. Build services so Claude can read their state, diagnose faults and ship small fixes, not just so a human on call can.

What this means for a Sydney or Melbourne engineering team

Most Australian teams are not at Anthropic's scale, and do not need a stateless test-selection service. The growth curve still applies. A useful exercise before widening Claude Code access is to take your current monthly CI spend and multiply it. A team spending $4,000 a month on runners and build minutes is looking at $40,000 at 10 times the load and $100,000 at 25 times. Even a fraction of that growth is a budget conversation worth having before, not after.

  • Find the one internal service every merge depends on: test selection, a shared staging database, a licence server. That is where you will break first.

  • Measure your flaky test rate now. Agents re-run failures relentlessly, and a 3% flake rate turns into a queue problem at agent volume.

  • Decide how review scales. If Claude is reviewing pull requests too, human review has to move to risk-based sampling.

  • For APRA-regulated teams, check that change-management evidence still holds when an agent opens and reviews most changes.

On the flaky test point, our walkthrough on using Claude Code for flaky-test hunts is a practical starting place. On review, see how to stop reviewing Claude Code output line by line. On the security side of the same shift, Anthropic's team has also described how it secures a pipeline where Claude writes most of the code.

Where the lesson does not transfer

Two cautions. First, the 25 times figure describes Anthropic, a company whose engineers are unusually quick to hand work to agents. Your growth multiple as at September 2026 could be three times or 30; measure it over a month rather than borrowing theirs. Second, the redesign succeeded partly because the team could hand a long-running agent a well-instrumented service. If your services are poorly observed, the first project is observability, not autonomy.

If you are planning a wider Claude Code rollout and want the CI and review costs modelled up front, our services cover exactly that scoping, or book a time to talk it through.

FAQ

Frequently asked questions

How much did Claude Code increase CI load at Anthropic?

CI job volume rose 25 times over six months as Claude came to author around 80% of the code and handle much of the pull request review, according to Anthropic's own engineering write-up.

What is test impact analysis?

It is a service that records test results and uses that history to select which tests need to run for each code change, so the full test suite does not run on every commit.

Why did scaling up the server not fix the CI problem?

A bigger machine only buys a fixed amount of extra capacity, while agent-driven load was compounding. The upgrade bought about 70 days before the service fell behind again.

How much headroom should systems have when agents write most of the code?

Anthropic's engineers suggest designing early versions for 10 to 20 times current load, well above the two or three times headroom teams usually allow for human-driven growth.

Ready to move from AI pilot to production?

We help mid-market Australian businesses deploy AI automations that actually reach production and deliver measurable ROI.