A legal research benchmark marks the best agent fully correct on 43% of questions

A preprint posted to arXiv on 30 September introduces Legal Research Bench, a set of 413 open-ended questions about United States law written by expert lawyers. Each comes with a gold answer, the supporting authorities and a binary grading rubric. Thirteen frontier models were run in a harness with web search, case-law search, page parsing and retrieval tools.

The grading is the interesting part. The authors use what they call all-pass grading with source verification: a response counts as correct only if every required criterion is satisfied and the authorities it cites actually verify. They also validated their automated judge against practising attorneys. On that measure the strongest model tested, Claude Opus 4.8, is “fully correct on 42.9% of questions”, and scores fall further on questions that require reconciling conflicting authorities.

The finding most worth repeating is about effort. The paper reports that across models, “more turns, tool calls, and inference cost do not predict higher accuracy”, which cuts against the assumption that letting an agent grind for longer fixes reliability.

Caveats: this is a preprint and has not been peer reviewed, author affiliations are not listed on the abstract page, and a thirteen-model line-up dates quickly. Why it matters: partial credit flatters these systems, and a lawyer cannot use an answer that is mostly right.

Source: the paper on arXiv.


Related