OpenAI has published internal numbers on how its own research team uses coding agents day to day. By mid-August 2026 the median OpenAI researcher was running more than $600 a day of agent inference, and the research organisation logged 3.1 agent-workdays for every one human workday. That is not a demo. That is a frontier lab saying, in its own data, that agentic coding has become default infrastructure.
For Australian businesses already running Claude Code, or weighing whether to start, the useful part is not whose tools these were. It is the shape of the curve. The question stops being whether to use a coding agent and becomes how far behind you are if you have not started. You can read OpenAI's write-up for the full picture.
What does OpenAI's agent usage data actually show?
Three things. Median daily agent spend per researcher climbed from occasional use in January 2026 to more than $600 a day by August. The ratio of agent-workdays to human workdays passed three to one. And the kind of work handed to agents widened, from research and infrastructure code out to monitoring, technical support and higher-level planning. All three figures are OpenAI's own, self-reported and not independently checked, so read them as direction rather than as a measured benchmark.
The trend line matters more than any single number. Usage did not plateau, and the tasks kept getting harder. OpenAI also reports that internal support channels, where researchers used to ask colleagues for help, have seen traffic fall. People ask the agent instead.
| Measure | January 2026 | August 2026 |
|---|---|---|
| Median researcher daily agent spend | Occasional use | More than $600 per day |
| Agent-workdays per human workday | Below one to one | 3.1 to one |
| Researchers running four or more concurrent agents | Rare | Rising sharply |
| Dominant task type | Research and infrastructure code | Monitoring, support, planning |
Why this matters if you are not a frontier lab
OpenAI's researchers are an extreme case: elite engineers with effectively unlimited compute. Most Sydney and Melbourne businesses will never run agents at that intensity and do not need to. But the direction of travel is the same one we see on client work. The businesses getting the most out of Claude Code are not the ones with the biggest budgets. They are the ones that handed over a real task early, then widened the scope once the pattern held.
Agent usage compounds. Teams that start with one well-scoped, low-risk job build the habit and the trust needed to hand over bigger ones later.
The bottleneck shifts rather than disappearing. Once agents clear the mechanical work, judgement and prioritisation become the limiting factor.
This data is a proof point, not a template. One Claude Code workflow that saves a bookkeeper two hours a week pays for itself long before frontier-lab scale.
Waiting for a safe-enough moment usually means starting later than a competitor who did not wait.
A first workflow of this kind is usually a $6,000 to $15,000 piece of work for an Australian SMB, most of it scoping and review rather than build. Our ROI calculator will give you a rough shape for your own numbers.
The support-channel signal is the one to watch
Buried in the usage data is a detail with more operational meaning than the spend figure: traffic to internal help channels fell as agent use rose. People stopped queuing for a colleague's attention and asked the agent first. In a small Australian business that same pattern shows up as fewer interruptions to whoever happens to be the one person who knows how something works.
That is a real productivity gain and a real risk at once. The gain is obvious. The risk is that undocumented knowledge stops being asked about out loud, which is often the only reason it ever gets written down. Teams that come through this well tend to treat agent transcripts as documentation and review them, rather than letting the answers evaporate after each session.
What not to read into the $600 figure
It is not a budget recommendation. A researcher at a frontier lab burning that much a day is paying for exploratory work with an enormous expected payoff, on infrastructure the lab already owns. It is also OpenAI's own number, with no stated currency conversion and no disclosed baseline. Treating it as a benchmark for what your team should spend would be reading a lab's research budget as a business operating cost.
The transferable part is the ratio and the widening task mix, not the dollar amount. If you want the comparison that actually applies to a business team rather than a research org, our write-up on getting more out of each Claude Code session is closer to the ground.
Where Claude fits
We work Claude-first because reliability matters more to a business owner than benchmark position. Handing a real process to an agent is a trust decision, and predictability is what makes it survivable. If a lab that builds its own models is restructuring how its people work around agentic coding, that is a reasonable signal for any Australian business still treating AI as a chatbot rather than something you delegate to. Our Claude Code playbook for AU startups sets out the first few moves, and our consulting services cover running the rollout properly. If a short conversation would be more use than another article, book a time.



