The APEX benchmark – built by Mercor, with tasks authored by BigLaw-experienced lawyers and advised by Cass Sunstein – is the most rigorous test of whether AI can perform real legal work. The answer is more specific than vendors or skeptics suggest.
OpenAI's Astra for Law and Anthropic's Claude for Legal work with many of the same legal vendors, but OpenAI builds the legal knowledge into the harness while Anthropic ships it as skills and connectors you can edit, and neither vendor compared itself to the other's real setup.