OpenAI launches a legal AI platform
Harvey and Legora can build on Astra for Law, while planned software integrations are not yet live.
By Saeed Ezzati7 min read
The audio edition
Listen to this newsletter
0:003:54
Read transcript
OpenAI is moving further into professional software with Astra for Law, a legal AI platform for research, drafting advice, and custom applications. The initial rollout is deliberately narrow: access is limited to selected law firms, with OpenAI citing protections for confidential client work. Astra for Law combines GPT-6 Astra with an index of U.S. case law, statutes, regulations, and other legal materials. It also adds specialized instructions for legal analysis and writing. OpenAI says that combination is intended to support lawyers, but the company has not presented independently verified performance results. So this is a product and platform launch, not evidence that the system can independently handle legal judgment. The audience extends beyond firms. OpenAI says legal AI companies Harvey and Legora will be able to build products on Astra for Law. It also announced planned integrations with Relativity, Clio, Intapp, and Thomson Reuters. Those connections are not demonstrated as live products yet, which makes the ecosystem a future direction rather than a current capability. Sullivan and Cromwell, Ropes and Gray, Cooley, Latham and Watkins, and Wachtell Lipton helped test and develop applications. That gives OpenAI named design partners, while leaving the eventual reach of the selected-firm program unclear. The important boundary is that legal source material, firm access, and a broader software network are being assembled under one product—but deployment, confidentiality controls, and real-world reliability still have to be established in practice. That same push into structured professional workflows appears in Anthropic’s beta redesign of Claude Code Projects. A coordinator can take a larger software goal, divide it among parallel cloud sessions, review the results, and assemble the changes while the developer directs the project from a main conversation. Each worker gets its own repository branch and copy, so it can run tests or open pull requests independently. Shared project memory carries requirements, decisions, and status across threads. But this is not an escape from engineering basics: overlapping changes can still create merge conflicts. The sessions run in the cloud, and several workers draw on the subscriber’s existing Claude usage, so complex projects can hit limits faster. Local tools and private-network access are planned, not available yet. The efficiency question is also central to Google-led research called Dream-RSI. It records an agent’s search history in a discovery tree, then replays those past outcomes to test alternative exploration policies without rerunning every candidate and evaluator. The researchers reported up to 162 times fewer discovery-agent calls than SimpleTES on a Lasso path-discovery task, plus smaller gains on GPU-kernel benchmarks. But that headline is not a like-for-like runtime comparison: it measures calls against a specific baseline and model setup. The code is public, and the open question is whether replay-selected policies retain their advantage through repeated live searches and different kinds of work. And a reminder from OpenAI researcher Noam Brown: more agents are not a universal scaling law. Teams can reduce waiting time when one model’s reasoning would take too long. In OpenAI evaluations, four agents were roughly twice as fast as one on some tasks, while 16 still improved results with diminishing efficiency. Brown said parallelism suits divisible work such as mathematics or reviewing many sources better than tasks needing one unified context, like writing a novel. OpenAI also lacks reliable evidence on very large fleets because the experiments are expensive. Across today’s launches and research, the practical watchpoint is clear: AI is moving from single answers toward coordinated systems, but the winners will be determined by deployment boundaries, compute costs, shared context, and measured reliability—not by agent count alone.



