AI Briefing
KO

Analysis of the 'Inference Gap' Between Claude's Reduced Reasoning Tokens and Benchmarks

·2026.09.18 09:00

Key point

Analysis reveals that actual reasoning token usage in Claude models is significantly lower than benchmark levels, creating an 'Inference Gap' and prompting calls for transparency and regulatory discussions.

1 / 7

Details

From the perspective that inference effort may be the most critical resource for AI model performance, an analysis of reasoning token usage in Claude models revealed a severe gap (Inference Gap) between benchmarks and actual production environments.

Decline in Reasoning Token Usage

According to the analysis, 39.2% of calls performed no thinking at all. Additionally, thinking token usage decreased in August compared to July, dropping by 21.9% on a turn-weighted basis and 35.1% on a session-weighted basis. Notably, there was a trend where the proportion of thinking tokens decreased as input tokens increased.

Gap Between Benchmarks and Actual Usage Environments

The average number of thinking tokens in Anthropic's benchmarks at xhigh and max effort levels ranges from 24.4K to 32.1K. However, 95% of the actual dataset fell short of 18.2K thinking tokens, which is the minimum among the top 5 calls in the benchmarks, based on 48 calls (long-horizon task). Best-of-N or parallel processing cannot succeed with candidate generation lacking sufficient inference compute, which can lead to intermittent performance degradation or unstable agent behavior.

Transparency and Regulatory Discussions

To bridge the gap between model potential and actual reliability, it is essential to disclose transparency regarding actual reasoning token counts, maximum continuous reasoning length, and the zero-thinking ratio. Performance loss due to resources limited by providers can expand into issues of provider liability, marking a time when discussions related to AI safety and regulation are necessary.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.