Opus 4.7's New Tokenizer: What's the Real Cost
Key point
Opus 4.7's new tokenizer increased real-world costs by 12-27%.
Details
Anthropic added a new tokenizer to Claude Opus 4.7 but kept the input $5/M, output $25/M pricing unchanged. However, the same text is now counted as more native tokens, which can change the actual billed amount. OpenRouter compared over 1 million text-only, non-cancelled requests from a switcher cohort whose primary model changed from 4.6 to 4.7.
The baseline was OpenRouter's own tokenizer, QuadChars, which counts 4 printable ASCII characters as 1 token and treats each non-ASCII character as a separate token, isolating only the difference from the native count. This approach separates out the effect of the tokenizer change alone, not changes in model content.
- < 2K: native token ratio increased by about 45%
- 2K–10K: increased by about 42%
- 10K–25K: increased by about 34%
- 25K–50K and 50K–128K: increased by about 32%
- 128K+: increased by about 33%
For longer prompts, prompt caching absorbed a significant portion of the increased tokens. Above 25K, most of the increase went into the cache, and in the 128K+ range, 93% of the additional tokens were handled by the cache. Conversely, requests under 2K had a low cache hit rate, so the buffering effect was minimal.
Response length also changed. Under 2K, 4.7's median completion was 62% shorter, but at 10K and above, it was 13-30% longer. As a result, the actual average cost rose by 12-27% in the 2K and above range, with per-bracket figures of 2K–10K +27.2%, 10K–25K +25.2%, 25K–50K +21.3%, 50K–128K +11.9%, and 128K+ +15.3%. Under 2K was actually cheaper, at -1.6%.
In other words, the impact of the new tokenizer varied greatly depending on the workload. Short requests became cheaper thanks to shorter responses, while real-world prompts of medium length or longer showed a clear increase in billed amounts even after the cache absorbed part of the increase.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.