AI Briefing
KOSign in

DoGBench: First User-Facing Documentation Benchmark Shows No Local Model Scores Above 50%

·2026.10.02 08:10

Key point

The highest composite score on the 117-item held-out split was 47.3, achieved by Qwen3.8 Max with OpenCode, highlighting significant gaps in AI agents' ability to maintain expert-level documentation.

Details

DoGBench introduces the first benchmark specifically evaluating AI agents on their ability to write and maintain user-facing software documentation in response to real repository events. The benchmark assesses whether agents can recognize when documentation needs updating and produce changes that meet expert review standards, using rubrics validated with project maintainers.

Benchmark Structure and Results

The dataset contains 292 tasks drawn from real open-source projects such as Helm, PostHog, Mautic, and Doc Detective. Of these, 205 require a documentation change and 87 require leaving documentation unchanged. The primary evaluation uses a random, stratified 117-item held-out split (82 requiring updates, 35 requiring no change).

Key performance metrics from the evaluation of seven model-and-harness lanes:

  • Highest composite score: 47.3 out of 100, achieved by Qwen3.8 Max with OpenCode.
  • Highest P0-clean delivery rate: 39.0%, achieved by GPT-5.6 Sol with Codex (delivering a critical-failure-free patch for 32 of 82 update tasks).
  • Cloud agents: Three cloud-agent systems were also evaluated on the same split, with the highest composite score among all systems reported as 54.8.

Common Failure Modes

Agents struggle with both the decision to update and the quality of the update. A broader audit of 1,267 submissions revealed:

  • 45.5% had a task-completion gap (missing prerequisites, steps, or verification).
  • 36.6% contained technical inaccuracies.
  • 32.5% omitted part of the central concept or reference information.
  • 6.1% contained fabricated content (invented classes, flags, or endpoints).

Agents often document internal refactors that require no user guidance or overlook necessary updates because they fail to find existing pages, missing the fact that the absence of coverage is the problem to solve.

Evaluation Methodology

DoGBench reports update decisions and patch quality separately using a combined score based on the harmonic mean of delivered patch quality and abstention recall. A critical failure (P0 criterion) caps a patch’s score at 60 out of 100, ensuring that strengths in secondary criteria cannot average away major defects. The benchmark emphasizes that expert review remains necessary, as agents lack the contextual understanding of reader workflows and project-specific nuances.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.