Wake-sleep Algorithm Boosts Long-Horizon Legal Agent Performance to 15.7% All-Pass Rate
Key point
Applying the wake-sleep algorithm to long-horizon legal agents increased the all-pass rate from 2.9% to 15.7% and halved the cost per task through selective memory retrieval.
Details
Researchers evaluated a variant of the wake-sleep algorithm on long-horizon legal agents using the Legal Agent Benchmark (LAB). The process involves an online 'wake' phase where agents complete tasks and an offline 'sleep' phase where a review model generates and merges reusable lessons into the agent's memory. Across 10 cycles and 196 tasks in Corporate M&A and Capital Markets, the wake-sleep agent achieved a 15.7% all-pass rate compared to 2.9% for a baseline agent with no memory. The agent also demonstrated improved generalization, passing over 10% more rubric criteria on familiar matters and nearly 10% more on new matters. To address the increased cost and latency associated with more thorough work (indicated by ~2.5x more tool calls), the team implemented lesson retrieval using Jev, which injected only relevant memories into the context. This optimization halved the cost per task while maintaining the same average quality.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.