Introducing the 'Agent Execution Tax' Concept
Key point
This presents the concept of an 'Agent Execution Tax,' where the actual cost per successful task matters more than token pricing, along with benchmark results by model.
Details
When measuring browser agent performance, considering only simple token pricing can be misleading. A new metric called the 'Agent Execution Tax'—referring to wasted inference cost caused by retries—has been proposed.
In a WebVoyager benchmark test of 4 models, a certain model recorded an execution tax (ratio of wasted inference to productive inference) of 22.9%. This suggests that the model with the cheapest cost per token could have a cost per successful task that is 2.3x higher.
The key model comparison results are as follows:
- MiniMax M2.5: Cost per successful task is 2.3x cheaper compared to Gemini.
- GLM-5: Showed the highest performance with 57.1% accuracy, with strength in structured data processing.
- Kimi K2.5: Recorded 0% parse retries out of 852 calls (compared to 18.6% for Gemini).
The analysis suggests that the reason recent open-weight models are performing well on agent benchmarks is not simply because the models have become more intelligent, but because reliability per call has improved.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.