LLMs corrupt documents when delegated editing tasks
Key point
On the DELEGATE-52 benchmark, even frontier models showed significant drops in document fidelity during long delegated editing sessions.
Details
DELEGATE-52 is a benchmark that measures how well document fidelity is maintained when users delegate long document editing tasks to LLMs.
It includes in-depth document editing tasks across 52 specialized domains, and the simulation consists of 20 consecutive delegation tasks.
Across experiments with 19 LLMs, even frontier models such as Gemini 3.1 Pro, Claude 4.6 Opus, and GPT 5.4 corrupted an average of 25% of document content by the end of long workflows.
Document corruption was not a matter of occasional typos, but rather a pattern where documents quietly broke down as long interactions accumulated. Performance degraded further as documents grew larger, interactions grew longer, and distractor files increased, and simple agentic tool use alone did not improve this.
The following examples illustrate the sharp degradation observed:
- Linux Kernel Architecture document: fidelity dropped from 79% → 49% → 48% → 48% across iterations with Gemini 3.1 Pro
- 12-Shaft Twill Diamond document: dropped from 100% → 40% → 27% → 34% with Claude 4.6 Opus
- ActionBoy Palm Tree document: dropped from 100% → 31% → 15% → 6% with GPT-5.2
Public resources include microsoft/DELEGATE52 and datasets/microsoft/DELEGATE52.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.