AI Briefing
KO

Harvey Releases Legal Agent Benchmark

·2026.05.07 09:00

Key point

Harvey has released LAB, an open-source benchmark for legal agents.

1 / 2

Details

Harvey has released LAB (Legal Agent Benchmark), an open-source benchmark for legal work. The goal is to measure, based on real-world work standards, how far agents can perform lawyer tasks and at what point human review becomes necessary. This allows law firms to more clearly judge AI investment ROI and how work should be divided, while also giving model providers and researchers a common standard for comparing performance.

LAB is designed to reflect the actual workflows of real law firms.

  • Instructions: A short work assignment a partner gives to an associate
  • Environment: A client matter mixing contracts, emails, and templates
  • Output: A reviewable legal work product
  • Verification: Grading based on expert rubrics

The instructions for each task are short, averaging around 50 words, and the environment includes both core and peripheral files, requiring context to be pieced together across multiple documents. The initial version consists of over 1,200 tasks, 24 legal practice areas, and more than 75,000 expert-written criteria.

Evaluation adopts all-pass grading. A task is only counted as complete if it passes all criteria—facts, citations, amounts, deadlines, severity, and recommendations—and getting only some issues right is treated as incomplete rather than earning partial credit.

As an example, an M&A task presented involves reviewing change-of-control clauses in a hypothetical $458 million acquisition, organizing risks, and writing a memo containing per-contract analysis and recommended responses. Harvey noted that existing evaluations such as LegalBench, CUAD, LEXam, and BigLaw Bench have mostly remained focused on short-form reasoning, and explained that a common standard is needed to measure long, multi-step legal work.

Just as SWE-Bench Pro, SWE-Bench Verified, and Terminal-Bench 2.0 marked a turning point in the coding field, the view is that such benchmarks could similarly serve as a signal of agent performance in the legal field. This version has been released without a leaderboard for now, and Harvey plans to continue expanding the benchmark by publishing baseline results and normalization rules together with research partners going forward.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.