AI Agents' Real-World Performance
·2026.04.15 02:21
Key point
In a benchmark of 153 web tasks, even the best model achieved only 33.3%.
Details
ClawBench is a benchmark that evaluates how well AI agents perform everyday online tasks in real web environments.
Using 153 tasks and 144 live websites, it examines actual service interaction ability rather than simple static tests.
The results are fairly sobering. The best model's success rate was 33.3%, showing that current agents are not yet capable of reliably completing real-world-level web tasks.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.