A service that had been fine for years ran out of headroom in 70 days. The fix for that bought 29 days. The fix after that bought less than one. That is the short version of a Claude engineering write-up on what happened to Anthropic's CI once Claude was writing most of the code, and it is the most useful thing published this month for any team planning to scale Claude Code.
What happened to Anthropic's CI
Over six months, CI job volume rose 25 times. Claude was by then authoring around 80% of the code and doing much of the pull request review, so every agent-driven change triggered builds and tests at a pace no human team produces. The piece that buckled was test impact analysis, the service that decides which tests need to run for a given change so the whole suite does not run on every commit.
That service has two halves. A listener records test results as they come in. A selector uses that history to decide which tests to run on each pull request. Both were designed for human commit rates. Neither was designed for a codebase where agents open changes around the clock. The full write-up, by engineer Sachin Malhotra, is on Claude's engineering blog.
Why does agentic coding put so much load on CI?
Because CI cost scales with the number of changes, not the number of engineers. A human developer might push a handful of commits a day. An agent working through a backlog can open, revise and re-run changes continuously, and each revision triggers builds and test selection. When an agent also reviews pull requests and requests fixes, the loop tightens further. Headcount stays flat while CI demand multiplies.
This is the mistake teams make when they budget for Claude Code. They model seat costs and token costs, then discover the real bill lands in CI minutes, runner capacity and whichever internal service sits in the critical path of every merge.
Three patches, each worth less than the last
The team did what most teams would do first: patch the existing design. The pattern of diminishing returns is the lesson.
| Fix tried | Runway it bought | Why it stopped working |
|---|---|---|
| Move to a bigger machine | About 70 days | Vertical scaling has a ceiling; load kept compounding |
| Shard the service | About 29 days | Shards still shared the same growth curve |
| Restart the service daily | Less than a day | A stopgap that treated the symptom, not the design |
| Stateless redesign on an in-memory data store | Scales horizontally | Built for the growth curve rather than against it |
Each patch was reasonable in isolation. Together they show what happens when load grows exponentially and fixes are linear: every fix is smaller than the last in calendar time, even if it took the same engineering effort.
How the redesign got built in three weeks
The replacement moved test impact analysis to a stateless architecture that scales horizontally, backed by an in-memory data store. It was built in three weeks, with Claude doing much of the implementation autonomously once it had a long-running monitoring session through Claude Tag. We have written before about how Anthropic uses Claude Tag internally, and this is a strong example of the pattern: the agent watches the service, sees what breaks, and fixes it in small increments.
That only worked because the service was instrumented well enough for Claude to diagnose itself. An agent cannot fix what it cannot observe. Logs, metrics and clear error messages stop being nice-to-haves and become the interface the agent works through.
The two design rules the write-up lands on
Plan for the exponential. Once agents write and review most code, give version-zero designs 10 to 20 times headroom rather than the usual two or three.
Instrument for the agent. Build services so Claude can read their state, diagnose faults and ship small fixes, not just so a human on call can.
What this means for a Sydney or Melbourne engineering team
Most Australian teams are not at Anthropic's scale, and do not need a stateless test-selection service. The growth curve still applies. A useful exercise before widening Claude Code access is to take your current monthly CI spend and multiply it. A team spending $4,000 a month on runners and build minutes is looking at $40,000 at 10 times the load and $100,000 at 25 times. Even a fraction of that growth is a budget conversation worth having before, not after.
Find the one internal service every merge depends on: test selection, a shared staging database, a licence server. That is where you will break first.
Measure your flaky test rate now. Agents re-run failures relentlessly, and a 3% flake rate turns into a queue problem at agent volume.
Decide how review scales. If Claude is reviewing pull requests too, human review has to move to risk-based sampling.
For APRA-regulated teams, check that change-management evidence still holds when an agent opens and reviews most changes.
On the flaky test point, our walkthrough on using Claude Code for flaky-test hunts is a practical starting place. On review, see how to stop reviewing Claude Code output line by line. On the security side of the same shift, Anthropic's team has also described how it secures a pipeline where Claude writes most of the code.
Where the lesson does not transfer
Two cautions. First, the 25 times figure describes Anthropic, a company whose engineers are unusually quick to hand work to agents. Your growth multiple as at September 2026 could be three times or 30; measure it over a month rather than borrowing theirs. Second, the redesign succeeded partly because the team could hand a long-running agent a well-instrumented service. If your services are poorly observed, the first project is observability, not autonomy.
If you are planning a wider Claude Code rollout and want the CI and review costs modelled up front, our services cover exactly that scoping, or book a time to talk it through.



