OpenAI launched Astra for Law on 17 September 2026: a version of GPT-6 Astra wired up to a legal search index covering more than 230 million URLs, including case law from the Free Law Project and CourtListener, plus legal-analysis instructions built for law firms and legal-tech vendors. Harvey and Legora are already building on it via API. If you're an Australian law firm weighing Claude against this, here's what's actually worth paying attention to, and what isn't.
What Astra for Law Actually Is
It's a single-purpose configuration: GPT-6 Astra plus a legal search index plus legal-analysis instructions, aimed squarely at legal research accuracy. OpenAI claims a jump from 38.7% to 54.0% correctness over GPT-6 Astra with plain web search on Vals AI's Legal Research Bench, alongside 24% more reference cases found and up to 54% more relevant passages retrieved. There's also a Trusted Access Program for eligible firms with Zero Data Retention on the API, ChatGPT Enterprise usage excluded from human review by default, and named governance and custom-build work with Latham & Watkins, Sullivan & Cromwell, Ropes & Gray and Cooley.
Is Claude Actually Worse at Legal Research Than Astra for Law?
OpenAI's own launch material includes a comparison where Astra for Law reportedly matched precedents more accurately than Claude Fable 5.1 on a litigation and transactional memo task, including a claim that Claude cited a reversed holding in one example and found no matching case in another. That's OpenAI's own test, on its own launch page, comparing its purpose-built legal tool against a general-purpose Claude configuration, not an independently run or third-party-verified benchmark. Worth knowing about, not worth treating as settled fact, and not a reason on its own to switch platforms without running your own test on your own matters.
Where the Real Comparison Sits
Astra for Law is a research tool bolted onto a general model. Claude's advantage for a law firm is less about winning a single case-lookup benchmark and more about what happens around the research: Claude Code and Claude Skills let a firm build an actual matter workflow, drafting, redlining, client comms, matter-management integration, agent handoffs to specific team members, rather than a search box with better citations. The Astra launch also brought 26 new ecosystem plugins including iManage, Intapp, DeepJudge and Thomson Reuters HighQ, so OpenAI is clearly chasing the same workflow ground Claude has been building on for longer.
Legal research accuracy on case law: worth testing directly against your own matters, not taking either vendor's word for it.
Workflow depth: agentic drafting, matter-specific automation, and integration into how a firm actually practises, not just retrieval.
Governance and data handling: Zero Data Retention and human-review exclusion are meaningful for a firm handling privileged material, and worth comparing feature-for-feature against Claude's own enterprise controls.
Vendor lock-in: a legal-search-specific tool is narrower than a general agentic platform a firm can extend into other parts of the practice.
Astra for Law versus a Claude-based legal workflow
| Dimension | Astra for Law | Claude-based workflow |
|---|---|---|
| Core design | Legal search index bolted onto GPT-6 Astra | General agentic platform (Claude Code, Skills, MCP) |
| Best fit | Case-law lookup and citation accuracy | End-to-end matter workflows and drafting |
| Ecosystem | 26 new legal-tech plugins at launch | Growing MCP connector and skills ecosystem |
| Independent verification of the launch benchmark | Not yet available | Not applicable — worth your own test either way |
This isn't the first time a competitor's launch has framed itself against Claude, and it won't be the last. We track these moves as they land on our blog, specifically because a single vendor comparison on a vendor's own launch page is rarely the full picture, and AU firms deserve a clearer read on what actually changes for their day-to-day practice than either company's marketing will give them.
What an AU Firm Should Actually Do About This
Run a small, real test before deciding anything. Pick five or six of your own past matters with known outcomes, run the same research question through both platforms, and check the citations by hand rather than trusting either vendor's benchmark. For an Australian mid-sized firm, the practical question usually isn't which model is marginally more accurate on US case law, since a meaningful share of AU legal work leans on Australian case law and legislation neither tool was purpose-built around. The more useful question is which platform your team will actually adopt into daily matter work, since a tool nobody opens delivers $0 of the promised time saving no matter what its benchmark says.
If you'd like help running that comparison properly, on your own matters rather than a vendor's demo set, that's exactly the kind of evaluation work we do for AU professional services firms. Have a look at our services or get in touch to talk through what a fair, firm-specific test would look like.



