Legora, a legal AI platform used by law firms for contract review, litigation drafting, and deal work, has published a new evaluation called BAR, its Benchmark for Agentic Reasoning. The scores come from running frontier models inside Legora’s own production software rather than a standalone test rig built purely for grading. That distinction matters for buyers: most legal AI benchmarks run in simplified environments, not the tool a customer would actually touch on a Tuesday.
Legora argues that separation understates real performance, since public benchmark cases eventually leak into training data and inflate scores on tasks a model has effectively already seen. Its response is a private pool of roughly 5,161 legal cases spanning 28 practice areas, built with law firm partners and Legora’s own legal engineers, then checked by a licensed lawyer before use. BAR itself is a curated slice of that pool, weighted so its mix of legal specialties, work formats, and difficulty tiers mirrors the full corpus. None of it is public, so no outside researcher can reproduce a single result.
Each case pairs a prompt with a folder of source documents and a scoring rubric built from weighted, pass or fail criteria. A lawyer sets what a passing answer must contain, the model produces the deliverable inside Legora’s harness, and a language model acting as judge checks that output against the rubric three separate times per case. Missing a high-importance item costs far more than missing a minor one, which Legora says tracks how a partner grades a junior associate’s draft.
Legora ran the evaluation across models from Anthropic, OpenAI, and SpaceXAI. On short tasks, roughly a few hours of work, the models scored close together. The gap widens on the longest assignments, week-plus matters requiring sustained judgment, where four models pulled clear of the seven-model average: Sonnet 5, Opus 4.8, Grok 4.5, and Fable 5. Grok 4.5 also posted the fastest median time per case while still scoring above average on quality, and it joined Sonnet 5 and Opus 4.8 in pairing high quality with below-average cost.
Citation reliability split differently. OpenAI’s models led on what Legora calls Cited Answers, meaning more of their claims carried a citation at all. Fable 5 and Opus 4.8 led on Grounding, meaning their citations more consistently matched the documents the model actually used. Legora separately reports that harness changes alone, with no underlying model swap, lifted output quality roughly 5 percent for models already live in production between June and July 2026.
Here is the catch buyers should sit with. Legora sells the harness it is scoring. Every result in BAR reflects a model wrapped in Legora’s own prompts, retrieval system, and workflow tooling, not the model’s raw legal reasoning in isolation. A 5 percent harness-driven quality gain is a real engineering result, but it is also a number Legora fully controls, produced by a judge model Legora does not name, on cases nobody outside the company can inspect. None of that makes the scores false. It means BAR answers “how good is this model once Legora builds the scaffolding,” a different question from “how good is this model,” and procurement teams should not treat the two as interchangeable.
Legal teams evaluating AI vendors should ask Legora, and any competitor publishing a similar benchmark, whether outside auditors have reviewed the case set and whether the same models rank the same way outside the vendor’s own harness before citing a BAR score as portable proof of model quality.
Legora published these findings, including its June-to-July 2026 harness comparison, in its own report titled “The Legora Benchmark for Agentic Reasoning” at legora.com/bar.