AI Briefing
KO

UiPath Reveals Performance Gap Between Claude Models Based on Agent Task Length

·2026.09.22 21:49

Key point

Analysis shows that as agent tasks become longer, Claude Haiku's success rate drops sharply, widening the gap with Sonnet.

Details

UiPath's engineering team analyzed the performance difference between Claude Haiku 4.5 and Sonnet 4.6 using their own agent task benchmark. While the score difference between the two models was only 6 points (73.3% vs 79.6%) on SWE-bench Verified, a significant gap of 43.6 percentage points (42.0% vs 85.6%) occurred in actual agent tasks.

Performance Changes by Task Length

When task complexity was categorized by the number of commands required by Sonnet, Haiku's performance degradation became pronounced.

  • 1-2 commands: Haiku 89.5%, Sonnet 89.5% (tie)
  • 3-5 commands: Haiku 70.7%, Sonnet 98.3%
  • 11-20 commands: Haiku 46.2%, Sonnet 90.8%
  • 41 or more: Haiku 21.2%, Sonnet 78.8%

For every doubling of task length, Haiku's success probability drops by half, whereas Sonnet decreased by only 11 points across the entire range.

Limitations of Benchmarks

SWE-bench provides an environment where errors can be corrected after test execution, but UiPath's internal checks are hidden, causing early-stage errors to lead to final failure. This suggests that while Haiku is competitive enough for short tasks, the necessity for higher-tier models increases sharply in multi-step agent tasks.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.