AI Briefing
KO

Inception Labs Releases Mercury 2.5 with 40% Intelligence Boost and 1,107 Tokens Per Second Generation

·2026.09.08 09:00

Key point

Inception Labs has released Mercury 2.5, which increases intelligence by 40% and generates 1,107 tokens per second.

1 / 4

Details

Inception Labs has launched Mercury 2.5, its latest diffusion language model (dLLM). Trained as the largest diffusion language model to date, it achieves a 40% improvement in intelligence over previous versions, delivering performance comparable to cost-efficient frontier models such as GPT-5.6 Luna and Gemini 3.5 Flash-Light.

Key Performance and Specifications

  • Speed: Generates 1,107 tokens per second on NVIDIA GPUs
  • Intelligence: 40% increase over Mercury 2
  • Context: Supports up to 260K tokens
  • Pricing: Input $0.04/M, Output $0.15/M (based on launch promotion)
  • Features: Supports parallel tool calls and parallel reasoning

Production Use Cases

  • Search Agents: Handles complex search pipelines including query rewriting, reranking, and summarization within a single interaction
  • Voice Agents: Model response latency reduced to under 170ms during live call processing by OpenCall
  • Coding Sub-agents: Augment Code reduced context compaction latency from 150 seconds to 27 seconds, cutting costs by 90%

Additional Announcements and Accessibility

  • Mercury Voice (Preview): A dLLM optimized for voice agents with a time-to-first-token (TTFT) of under 170ms
  • Mercury Router (Preview): Analyzes prompts to route to the optimal model based on quality, speed, and cost
  • Access: Available via the Inception Labs API, Baseten, and OpenRouter, with 100 million free tokens provided initially

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.