Anthropic's own data team has been running an interesting experiment: taking the same governed data setup that gets their data scientists roughly 95% accuracy in Claude Code, and exposing it to anyone at the company through Slack, via Claude Tag. It is a small idea with a large payoff, and it maps onto how most businesses actually get stuck on data.
The problem this solves
Every business hits the same wall eventually: the people who can answer a data question, how many signups yesterday, which product is trending down, what is the refund rate this month, are a small handful of analysts, and everyone else either waits in a queue or guesses. Anthropic's data team built their accuracy foundation around three things:
A governed semantic layer, so revenue means the same thing every time someone asks about it.
A set of skill files encoding the team's actual analytical conventions, not generic best practice.
An evaluation suite to keep accuracy measurable over time, not just it felt right that one time.
That accuracy work happened in Claude Code, where the data engineers actually build things. The interesting shift is what came next: getting that same governed foundation into Slack, where the rest of the company already lives, via Claude Tag.
Deployment is a different problem to accuracy
The team's own framing is worth sitting with: getting an agent accurate and getting it usefully deployed to non-analysts turned out to be genuinely different problems. Once a data agent is answering questions for people who are not analysts, new questions show up that do not matter in a controlled evaluation: who actually has access and through what surface, whether someone can see numbers they should not, whether the answer is based on data from an hour ago or a week ago, and how anyone finds out when it gets something wrong. Accuracy is a lab problem. Distribution, permissions and freshness are field problems, and they are where most rollouts quietly fail.
Why this matters for AU businesses
Most Australian SMBs do not have a dedicated data team, which is exactly why this pattern is worth attention rather than skipping past as not relevant, we are too small. The version that scales down: a handful of well-defined metrics (revenue, pipeline, churn), a Claude Skill that encodes how your business actually defines them, and access through the tool your team already uses daily. Businesses we have seen do this well typically start with three to five metrics, not thirty. A tightly governed answer to what is our cash position beats a sprawling dashboard nobody opens. Getting that first slice right, on a Slack-sized budget rather than a full BI-platform budget, is usually a matter of weeks and a few thousand dollars, around $4,000 in setup, not a data-team hire.
The takeaway
The accuracy work and the deployment work are two separate jobs, and most AI rollouts only budget for the first. If you already have a working Claude Skill or agent that is accurate in testing but nobody outside your ops team actually uses it, distribution is probably the missing half. Fix that and the tool starts earning its keep, because the value was never in the model being right in a test; it was in the right answer reaching the person who needed it, in the place they already work.
How to start small
If this pattern sounds useful, the worst move is to try to boil the ocean with thirty metrics and every data source at once. Start narrow and prove it. Pick the three questions your team actually asks most often, write down exactly how each metric is defined so there is one true answer, and encode those definitions in a Claude Skill. Wire it to the one or two data sources those questions depend on, then put it in the Slack channel where people already ask. Give it a fortnight, watch which answers people trust and which they double-check, and only then widen the scope.
Done this way, the cost stays small and the risk stays contained. You are not betting on a platform migration, you are testing whether a governed answer in the place people already work saves them from waiting on an analyst. If it does, you expand one metric at a time. If it does not, you have spent a few thousand dollars finding out, not a data-team salary, and that is a cheap lesson by any reasonable measure.
If you want help turning an accurate internal AI tool into one your whole team actually reaches for, book a session and we will start with the metrics that matter most.



