Anthropic had to rewrite its own engineering hiring test three times because its own model kept passing it. That is the short version of a post by Tristan Hume, who leads Anthropic's performance optimisation team, published on 21 January 2026 and available on Anthropic's engineering blog. It reads like a hiring story. For a business deciding how much work to hand to Claude Code, it is really a story about checkpoints.
What happened to the take-home test
The original take-home was a performance optimisation problem on a simulated accelerator. It ran from late 2023 and helped Anthropic hire dozens of engineers. It worked well until Claude Opus 4 matched or beat the best human candidates inside the time limit.
The second version shortened the time and asked for more depth. It held for months, until Claude Opus 4.5 solved that too, including finding a non-obvious trick mid-session once it was nudged. The third version uses a deliberately unusual, tightly constrained instruction set, in the spirit of Zachtronics-style programming puzzle games, with no built-in debugging tools. Building your own tooling is now part of what the test measures. So far that version has held, and Anthropic has released the original as an open challenge for candidates who want to try beating Claude's best recorded score.
How should a business evaluate Claude Code on its own work?
Evaluate Claude Code on tasks that look like your real work, not on tidy puzzles it has effectively seen before. Anthropic's experience shows that well-specified, self-contained problems are exactly where current models match strong engineers. The useful signal sits in long-horizon tasks, unfamiliar constraints, and work where judgment about what to build matters as much as writing the code.
That gives a simple rule for Australian teams deciding what Claude Code can do unsupervised. If a task is well-defined, self-contained and has a clear right answer, expect Claude Code to do it well, and check it with automated tests rather than a person reading every line. If a task is long, ambiguous, or depends on context that lives in someone's head, keep a human checkpoint in the loop.
| Task shape | Example | Suggested checkpoint |
|---|---|---|
| Well-specified and self-contained | Add input validation to an existing form handler | Automated tests plus a quick diff skim |
| Well-specified but long | Migrate 40 API endpoints to a new client library | Tests plus checkpoints at agreed milestones |
| Unfamiliar constraints | Change code in a legacy system with no test suite | Build the harness first, then review each change |
| Judgment-heavy | Decide how to restructure billing logic | Human designs, Claude implements, human reviews |
| Regulated output | Changes touching data covered by APRA or the Privacy Act | Named reviewer sign-off every time |
Three lessons that transfer from hiring to delivery
The three iterations of Anthropic's test map neatly onto how a business should design its own checks for AI-assisted work:
Your test will go stale. A checkpoint that caught problems with last year's model may pass everything this year. Review your acceptance criteria when you upgrade models, not just when something breaks.
Length and ambiguity are where the signal lives. Short, clean tasks tell you little. Give Claude Code a realistic multi-step job from your backlog and watch where it needs steering.
Tooling is part of the job. The version that held asked candidates to build their own debugging tools. The equivalent for a business is making sure Claude Code can run your tests, linters and build, so its work proves itself.
That last point is the one most teams skip. We have written about a Claude Code skill that proves agent work is done and about moving away from reviewing output line by line. Both come back to the same idea: verification should be built into the workflow, not bolted on at the end by a tired reviewer.
A worked example for a small team
Take a Melbourne software business with four developers, each costing roughly $150K a year fully loaded. If a senior developer spends a day a week reviewing AI-generated changes line by line, that is about $30,000 a year of senior time on review alone. Moving well-specified work to a test-gated checkpoint and keeping human review for judgment-heavy changes can plausibly halve that, without lowering the bar on the work that matters.
The trap is doing the reverse: trusting Claude Code on long, ambiguous tasks because it aced a few tidy ones. Anthropic's own test shows that tidy tasks are the easy part. Before rolling Claude Code across a team, run a short evaluation on five to ten real tickets, graded the way our test plan for Claude apps describes.
Careful with the takeaway
This is not evidence that developers are no longer needed. Anthropic still hires engineers; it changed what the test measures. The skills that now separate candidates, such as building tools, working under odd constraints and knowing which problem to solve, are the same skills your team needs to supervise Claude Code well. As at September 2026, the businesses getting the most from it are the ones that invested in that supervision.
Our services include setting up these checkpoints for teams adopting Claude Code. If you want help working out which of your tasks are safe to hand over, book a time with us and bring a few recent tickets.



