AI Briefing
KO

AI Gateway Production Metrics

·2026.05.14 06:00

Key point

In Vercel AI Gateway production traffic, Anthropic led in spend while Google led in usage.

1 / 2

Details

An analysis of production traffic from 200,000+ teams over the past 7 months by Vercel AI Gateway showed that model competition plays out more clearly in real-world operations than in benchmarks. As of April, spend share was Anthropic 61%, Google 21%, OpenAI 12%, but token usage was the opposite: Google 38%, Anthropic 26%, OpenAI 13%, xAI 10%.

High-volume workloads moved across an average of 30+ models and were not tied to a specific lab. Open-source models are also gradually being adopted, but actual customers mix multiple models by use case rather than relying on a single model.

  • Premium reasoning calls concentrated on Claude Opus, while low-cost, high-speed calls concentrated on Gemini Flash.
  • OpenAI's spend share tripled in April compared to March following the launch of GPT-5.4/5.5.
  • Google's spend share rose from 8% → 21%, driven by the spread of Gemini Flash.

The split between cost and usage was also clear across workloads. Personal assistants accounted for 20% of cost and 40% of tokens, coding agents for 22% of cost and 20% of tokens, back-office agents for 6% of cost and 15% of tokens, and app generation for 7% of cost and 11% of tokens.

  • B2C generates volume through low-cost, high-volume calls, while B2B generates spend through more expensive calls.
  • Cost per token for B2B is roughly 2x that of B2C.
  • Agentic workloads account for 59% of all tokens and have doubled in 6 months.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.