AI Briefing
KO

Metrics for the AI Era

·2026.07.17 19:00

Key point

OpenAI has proposed a way to measure AI investment returns using four metrics centered on 'useful intelligence per dollar.'

Details

Executives are asking how to get more value from AI spending. OpenAI proposes evaluating AI economics around task completion rather than cost per token.

1. How much useful work is actually getting done?

Measuring real outcomes—solved customer problems, shipped code, reviewed contracts—is the starting point. Better models can handle more complex tasks, freeing up time for people to focus on judgment and creativity.

2. What's the real cost per successful task?

A lower price per token doesn't mean a lower cost per outcome. If an advanced model gets the answer right on the first try, it reduces the cost of retries, review, and correction, which can result in a lower final cost. To calculate this, divide total cost (compute + human review + rework) by the number of successful tasks.

GPT-5.6, which OpenAI launched last week, consists of three tiers: Sol (top-tier, highest performance), Terra (balanced performance and cost), and Luna (fastest and cheapest). Teams can choose based on the task's economics. On the DeepSWE v1.1 benchmark, GPT-5.6 Sol achieved 72.7% while cutting API costs by 36.2% compared to other major models.

3. How often does AI perform the task correctly?

AI adoption deepens in stages. It progresses from drafting → finding context and connecting tools → actually performing the task, with each stage requiring higher reliability. Teams can track this with three metrics:

  • Usable: The output meets quality standards as presented
  • Needs correction: Requires a retry or human edit
  • Escalated: Requires human intervention to complete

Clear boundaries are essential for reliability. Data access permissions, authority to make system changes, and points for review and approval should all be defined in advance. ChatGPT Work is built on enterprise security, privacy, compliance, and workspace management.

4. As scale increases, does more work get done per AI dollar?

Track the same workflow over time, recording completed tasks, total cost, and cost per task. If completed tasks grow faster than total cost while quality is maintained or improved, each AI dollar is producing more value.

Compute is at the center of this equation. Better models, efficient inference, purpose-built hardware, high utilization, smart routing, and strong product design all improve compute returns. From a human perspective, these advantages are experienced as better answers, faster results, fewer corrections, more reliable products, and lower task costs.

Benefits compound. Better infrastructure → accelerated research → better, more efficient models → improved products → increased adoption and revenue growth → continued investment in next-generation research and safety.

OpenAI brings all of this together into a single intelligent platform through ChatGPT, ChatGPT Work, Codex, and the API. When one layer improves, every product and customer can benefit.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.