AI Briefing
KO

Why token prices are falling but AI bills are rising

·2026.07.21 09:00

Key point

Per-token prices have plummeted, but actual AI costs have risen instead due to a surge in usage driven by agent architectures.

1 / 2

Details

According to Ramp data, between January 2025 and April 2026, enterprise token consumption increased 1,000%, while spending increased only 500%. Usage growth is outpacing the rate of price decline by a factor of two.

There are three main reasons AI costs have risen.

  • Change in the unit of consumption: In 2024, the basic pattern was one prompt → one response, but by 2026, an agent handles a single request by triggering dozens of model calls.
  • Context bloat: RAG pipelines inject documents into every prompt, conversation history is resent with every turn, and reasoning models consume tokens for the thinking process itself.
  • Background inference: Monitoring agents, compliance checks, and other processes that run continuously without any human request keep consuming tokens.

This is a textbook case of the Jevons paradox. As inference costs got cheaper, everyone adopted token-heavy approaches because they produce better results.

On top of this, supply-side pressure is also mounting. H100 rental prices bottomed out in October 2025, then surged 40% by March 2026, and GPU capacity through September is already fully booked.

Proposed countermeasures include matching models to the maturity of the use case, or, as Tesla does, setting a $200 per-employee weekly cap that requires manager approval to exceed.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.