AI Briefing
KOSign in

Microsoft introduces Microsoft-Decision-1, a low-latency model for structured decision-making

·2026.10.10 03:35

Key point

The new model, available in Microsoft Foundry, delivers calibrated probability scores for routing and classification tasks at 35 times the speed of GPT-6 Sol.

1 / 2

Details

Microsoft has released Microsoft-Decision-1, a new AI model designed specifically for fast, structured decision-making tasks such as routing, classification, prioritization, and workflow control. Unlike general-purpose LLMs, this model is purpose-built to deliver immediate, actionable structured outputs with high performance and low latency. It is currently available in Microsoft Foundry and will soon be accessible via OpenRouter.

Performance and Benchmarks

Microsoft-Decision-1 was evaluated across 36 benchmarks comprising nearly 150,000 questions that were kept blind from the training data. The results highlight significant advantages in speed and accuracy:

  • Speed: The model achieved a P50 latency 35 times faster than GPT-6 Sol and 4.5 times faster than the runner-up, Quyet-1.0-Large.
  • Accuracy: It recorded the highest accuracy across the benchmark suite, demonstrating strong generalization across routing, ranking, long-context, multilingual, and out-of-distribution tasks.
  • Robustness: In perturbation tests involving paraphrased options or shuffled choices, the model changed its decision in only 1.3% of cases, with zero flips when options were reversed or shuffled.

Technical Architecture and Capabilities

The model was post-trained on Qwen3.5-9B to enable fast, single-pass decision scoring. It provides calibrated probability scores for each answer option, supporting yes/no, multiple-choice, and rating formats. This calibration allows applications to determine when to act, defer, or request human review based on confidence levels.

Key features include:

  • Rubric-based grading: Ability to grade AI responses and agent actions against defined criteria.
  • Safety screening: Tested on 5,250 requests across 11 benchmarks, the model successfully refused harmful content while maintaining high utility.
  • API Structure: Decisions are made through a simple structured API call, integrating easily into existing agents and workflows.

Internal Use Cases and Pricing

Microsoft teams have already deployed the model for various internal operations:

  • XBOX Research: Processed over 10,000 pieces of feedback, achieving quality competitive with GPT-6 Sol while running 14 times faster and costing 200 times less.
  • Copilot Team: Used for quality control of chat and agentic responses, finding it competitive with GPT5.6 Luna but 100 times faster.
  • Microsoft Discovery: Implemented for adaptive replanning in scientific experiments, resulting in 46 times more consistent scoring and nearly 4 times faster replanning speeds compared to LLM-based approaches.

Pricing: Input tokens cost $0.042 USD per million tokens, while output tokens are free.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.