Blog

Prompt Injection Needs Two Doors Open: The Cowork Safety Model in Plain English

August 2026 · 6 min read · Technical

Prompt Injection Needs Two Doors Open: The Cowork Safety Model in Plain English
← Back to all posts

Most explanations of prompt injection describe the attack and then stop, which leaves you worried but no better defended. Anthropic's Cowork safety documentation does something more practical. It names the two conditions an attack needs, and both have to be true at the same moment. Claude must be able to read information from outside your trusted boundary, and Claude must be able to take an action that could compromise you.

Close either door and the attack has nowhere to go. That is the entire governance model in one sentence, and it beats any checklist of warnings because it tells you what to actually change.

What an injection looks like in practice

Malicious instructions sit inside content Claude reads while doing a legitimate job. A line buried in a supplier PDF. Hidden text on a web page nobody scrolls to. A paragraph in an email forwarded from outside the business. Claude reads it as part of the task, and the instruction tries to redirect the work toward something you did not ask for.

Anthropic runs three defences against this, and they are real:

  • Model level training, so Claude recognises and refuses malicious instructions.

  • Content classifiers that scan untrusted content as it enters Claude's context.

  • Action screening in Auto mode, where each action is reviewed before it runs. A blocked action sends Claude looking for a safer path or back to you with a question.

None of those are a reason to skip the two door rule. Defence in depth means the layers are independent, not that the top layer excuses the others.

Closing door one: what Claude can read

On desktop, Claude reaches only the folders you have connected. That is a decision you make rather than a default you inherit, which makes it the cheapest control in the whole model.

  • Connect a working folder, not a whole drive. If a folder holds client contracts, staff records or anything covered by the Privacy Act, connect it for the specific task and disconnect afterwards.

  • Be deliberate about which browser tabs are open when using the Chrome side panel. The session can see the pages you are on, including pages behind a login.

  • Treat any document originating outside your business as untrusted content, because that is exactly what it is. A quote from a supplier is untrusted content that happens to be useful.

The mistake here is connecting a top level folder once, because it is convenient, and never revisiting it. Six months later that folder holds three unrelated clients and a payroll export.

Closing door two: what Claude can do

Write tools carry the risk. Reading your inbox is a different proposition to sending from it, and Anthropic's own architecture draws that line explicitly.

  • Switch to manual approval for tasks touching sensitive files, accounts or sites, and for anything difficult to undo such as sending messages or making purchases.

  • For scheduled work, start with low risk summarising before automating anything consequential, review the output after each run, and pause tasks you are not actively using.

  • Keep sensitive data out of unattended jobs entirely, rather than relying on a review step to catch it after the fact.

A task that only reads cannot hurt you through a write it has no ability to make. If you find yourself nervous about a job, check whether it needs write tools at all. Often the nervousness is about a capability that was never in scope.

The line Australian businesses need in writing

Anthropic is unusually direct about accountability. You remain responsible for every action Claude takes, including published content, financial transactions, data changes and respecting the terms of third party websites.

For a firm operating under APRA or ASIC obligations that is not a footnote, it is the compliance position, and it will be quoted back to you in a review. The practical response is small: a one page addendum to your existing IT policy covering which folders are connected, which approval mode applies to which task type, and which tasks may run unattended.

That page costs almost nothing to write and settles most internal objections before they become a project delay. We include one in our $3,500 Cowork setup because it is the artefact clients get asked for first, usually by someone who has never opened the product.

What good looks like after a month

A team with the two door model working does not talk about prompt injection much, because the answer to every version of the question is the same. What can it read, and what can it do. If both answers are narrow for a given job, the job is safe to automate. If both are broad, a human stays in the loop.

The attack needs both doors. Your policy only has to shut one of them, reliably, on the jobs that matter. That is a far easier standard to hold than trying to anticipate every hostile paragraph on the internet.

If you want that addendum drafted against how your team actually works rather than a template, we can help. Start at /contact.

Ready to move from AI pilot to production?

We help mid-market Australian businesses deploy AI automations that actually reach production and deliver measurable ROI.