HarnessTax Study: Same Model Costs Nearly 2x More Depending on Harness
Key point
The HarnessTax study reveals that even with the same AI model, costs can differ by up to 2x depending on the harness choice, with the simple open-source harness Pi demonstrating competitive performance.
Details
The HarnessTax study evaluated combinations of 7 models and 3 harnesses (Claude Code, Codex CLI, Pi) on SWE-bench Lite and Terminal-Bench 2.0. The results showed that while harness choice had a negligible impact on task success rate (within ±2–5%), it significantly affected cost. Specifically, Claude Code incurred approximately 2.0x the cost of Pi and 1.6x the cost of Codex.
Regarding cost differences for the same model, Claude Fable 5 cost $1.33 on Claude Code and $0.67 on Pi, with Pi achieving a similar success rate (96.7% vs 96.7%) at roughly half the cost. Additionally, on Terminal-Bench 2.0, GPT-5.6 Sol achieved a higher success rate (83.3% vs 78.9%) with the Pi harness ($0.42) compared to the Codex harness ($0.76), while operating at 'about half the cost' according to the original text. Pi reached the Pareto frontier using only 4 tools—read, write, edit, and bash—and the fact that Claude Code's initial context is more than 10x that of Pi was identified as the primary cause of the cost difference.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.