AI Briefing
KO

AI Self-Improvement Benchmark Released

·2026.08.28 05:13

Key point

Scale AI has released HarnessOpt-Bench, a benchmark for measuring the self-improvement capabilities of AI agents.

Details

The Scale AI team introduced HarnessOpt-Bench to measure how well AI can improve other AI agents. This benchmark isolates the evaluation targets outside the sandbox, structurally preventing optimization agents from stealing answers or cheating.

Key Findings

  • Impact of Model Selection: When using the same coding harness, Claude Opus 5 achieved the highest performance in 3 out of 4 tasks.
  • Impact of Harness Selection: Even with the same model, OpenCode yielded better results in 11 out of 20 model-task combinations compared to its own harness.
  • Core Insight: The performance improvement from changing models was 1.8 times greater than the effect of changing harnesses.

This study was validated through 111 runs across 5 frontier models and 4 downstream tasks, and the related paper and code have been made public.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.