Inception Releases Ultra-Fast LLM 'Mercury 2.5' Generating 770 Tokens Per Second
·2026.09.24 07:16
Key point
Although its intelligence index is below average, it offers approximately 7x faster output speed and lower cost compared to models in the same price range.
Details
Inception has released its proprietary LLM Mercury 2.5. According to Artificial Analysis, this model generates 770.4 output tokens per second, a significantly faster speed compared to the median (108.6 t/s) of reasoning models in the same price range.
Performance and Cost
- Speed: Generating 770.4 tokens per second ranks it 2nd among 175 models, but its Time to First Token (TTFT) is 2.91 seconds, slightly slower than the median (2.22 seconds).
- Intelligence: The Artificial Analysis Intelligence Index v4.3.2 score is 12 points, lower than the median for the same price range (13 points), placing it 91st among 175 models.
- Cost: Priced at $0.25 / 1M tokens for input and $0.75 / 1M tokens for output. The blended price with a 90% cache discount is $0.14 / 1M tokens.
- Conciseness: During the Intelligence Index evaluation, it generated a total of 35M output tokens, showing very concise responses compared to the median for the same price range (85M).
Technical Specifications and Deployment
- Architecture: It is a reasoning model (supporting extended thinking) with a 260k token context window. It supports only text input/output and does not support multimodal (image input).
- Deployment: Currently accessible only via the Inception API, with weights and parameters remaining private.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.