AI Briefing
KO

Legal Agent Benchmark (LAB) Early Results Analysis

·2026.05.26 09:00

Key point

Even the latest frontier models complete complex legal tasks perfectly less than 10% of the time.

1 / 2

Details

Legal Agent Benchmark (LAB) is an open-source benchmark designed to evaluate the performance of agents performing complex legal tasks. Analysis using a strict all-pass standard, which counts a task as passed only if all requirement criteria are met, revealed three major trends.

First, frontier models' ability to perform legal tasks is still at an early stage. The evaluated models completed tasks perfectly from start to finish in less than 10% of cases overall.

Second, model intelligence is unevenly distributed across legal specialty areas. A jagged intelligence phenomenon appears, where models excel in certain areas but not in others, and the leading model varies by legal practice area.

Third, achieving top-tier performance requires significant cost and time. Running the top-ranked models on the benchmark incurs a cost of about $50 per task and a latency of over 20 minutes.

All-pass scores for major models are as follows:

  • Claude Opus 4.7: 7.1%
  • Sonnet 4.6: 5.4%
  • Opus 4.6: 4.2%
  • GPT-5.5: 2.1%
  • Gemini 3.5 Flash: 0.8%

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.