Blog

Why Rakuten Trusts Claude Fable 5 to Run Agents Unattended Overnight

August 2026 · 6 min read · AI Strategy

Flat line illustration of a city skyline at night with one terracotta-lit window showing a small robot working at a desk under a crescent moon
← Back to all posts

Claude has crossed a threshold that matters more to Australian business owners than any benchmark score: it can now be trusted to run overnight, unsupervised, without someone standing by to catch its mistakes. That's the headline buried in Rakuten's own account of testing Claude Fable 5, and it's worth unpacking properly, because the implications reach well past one large Japanese company's engineering team.

Rakuten has been testing Claude models since September 2024. Since March 2025 the company has used Claude Code to ship production software, and it has built custom Claude Managed Agents across product, sales, marketing and finance, rolling each one out company-wide within a week of the underlying feature launching. Those agents are wired into Slack, Microsoft Teams and Rakuten's internal task systems, not sitting in a sandbox somewhere waiting to be demoed. Yusuke Kaji, Rakuten's General Manager of AI for Business, describes testing each new model release as a "new quest": he sets it a stretch task the way a good manager would set one for a person, then watches how it copes when the first step doesn't go to plan.

For an Australian business, overnight is not a marginal sliver of the day. It's close to half of it, and it's the half where almost nothing gets watched in real time. An agent that can be trusted to work through those hours without supervision isn't a small productivity tweak. It changes what a lean team, the kind most Australian SMBs run with, can realistically take on without adding headcount.

The problem earlier models couldn't solve

Before Fable 5, letting an agent loose on a multi-hour task without supervision was, in Kaji's own words, a gamble. Get the first step right and the run went fine. Get it wrong, and the model would burn hours heading in the wrong direction, sometimes without ever noticing it had gone off course at all. The issue wasn't raw intelligence. Earlier models could reason well enough in isolation. What was missing was self-verification: nothing in the loop was checking the work as it happened, so a small early mistake had hours to compound before anyone, human or model, caught it.

  • It re-checks its own assumptions mid-task, correcting a wrong turn instead of committing to it for hours.

  • It returns to first principles at each step, re-validating against the original goal without being prompted to.

  • Its judgement on ambiguous calls lines up with the team's own, a quality Kaji calls "taste alignment".

The upshot, in Kaji's own words, is that the model "understands its mistake before I point it out at 2am or 3am, so that I can sleep." That's not a marketing line dressed up for a blog post. It's a working description of what unattended reliability actually looks like once you strip the hype away: not a model that never errs, but one that catches itself.

Why this matters more than the benchmark

Rakuten's agents are now closing issues roughly ten times faster across every domain the company has deployed them in. But Kaji is careful to draw a line here: adding more agents doesn't add judgement. The organisation's progress still depends on people closing the loop at the right moments, and nobody at Rakuten is claiming the agents run the business unsupervised end to end. The gain isn't that humans are removed. It's that the handful of moments where a human genuinely needs to weigh in get isolated from the hundreds of moments where they don't.

That distinction is the one Australian business owners should sit with before getting excited about running agents overnight. The genuine question isn't whether a model is fast, and it isn't whether it can pass a benchmark. It's whether it will notice its own mistake before that mistake compounds into something a person has to clean up the next morning, whether the task is processing invoices, triaging support tickets, reconciling transactions, or watching systems while the office is closed.

What this is worth to an Australian business

Here's the number that makes this concrete. A support or operations hire to cover overnight monitoring typically costs an Australian business $70,000 to $90,000 AUD a year in salary alone, and that person still has to sleep at some point, take leave, and get sick. That's not an argument for replacing the role outright. It's an argument for testing, on a bounded and well-defined task, whether the right agent, set up with clear limits on what it's allowed to act on, can do the noticing itself and escalate only the small number of decisions each night that genuinely need a human.

  • Does the task have a genuine multi-hour arc, or is it a short job you're overcomplicating by calling it "autonomous"?

  • Can you define clear boundaries for what the agent may act on without approval, and what it must escalate instead?

  • Is there a way to audit what happened overnight before the team walks in, rather than just trusting it went fine?

  • Have you actually tested it on a real stretch task, the way Kaji tests new model releases, rather than a scripted demo that only ever shows the happy path?

None of this requires enterprise scale or a dedicated AI team. It requires the same discipline Rakuten applied at a much larger size: set the agent a real task, watch closely what it does when the first step goes wrong, and only widen its remit once it's earned that trust. Most Australian small and mid-sized businesses have at least one overnight or after-hours process sitting idle for exactly this kind of pilot, whether that's a support inbox, an invoice queue, or a monitoring dashboard nobody wants to be woken up for.

If you're weighing up whether an agent could take the night shift in your business, that's worth a proper conversation before you build anything. Automata AI is a Sydney-based Claude consultancy that helps Australian businesses work out what's genuinely worth automating unattended, and what still needs a person watching. Book a time on our contact page and we'll talk through where your business actually sits on that line.

Ready to move from AI pilot to production?

We help mid-market Australian businesses deploy AI automations that actually reach production and deliver measurable ROI.