AI Briefing
KO

Armature Releases Data on 17k Agent Runs

·2026.09.12 01:45

Key point

Analysis of 17,000 agent runs reveals that Claude Code rarely performs web searches and that tool selection discrepancies among coding agents are frequent.

Details

Armature has released data analyzing tool selection patterns of coding agents by measuring 17,000 agent sessions under conditions similar to real-world environments. This analysis targeted various repositories and personas (vibe coders, junior/senior engineers).

Key Findings

  • Differences in Web Search Frequency: Claude Code rarely performs web searches, whereas Codex uses them frequently, and Cursor shows moderate frequency.
  • Tool Selection Discrepancies: Coding agents disagreed more often than they agreed with each other.
  • Gap Between Mention and Selection: Key players in specific categories such as LangChain, Supabase, Netlify, PayPal, and Adyen were frequently mentioned but not actually selected.
  • Impact of Context: Modifying repository context can completely change the agent's tool selection results.

All execution traces are available via the Armature blog, allowing developers to gain a deeper understanding of agent tool selection logic.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.