EmpirioLabs Releases Aplomb 1: Open-Weights 5.3B Decision Model with 1M Context and Multimodal Input
Key point
The model scores 44.86 on Decision Index 0.2.1, ranking #1 among 4B models, and processes 1M-token documents in ~3 seconds via API.
Details
EmpirioLabs has released Aplomb 1, an open-weights 5.3B decision model built on Qwen3.5-4B with the audio encoder from Qwen3-Omni-30B-Audio. The model supports 1M token context and accepts text, image, video, and audio inputs in a single request. It is designed to make decisions with probabilistic outputs for tool arguments, allowing agents to act on high-confidence calls while routing uncertain ones to larger models.
Performance and Benchmarks
On the Decision Index 0.2.1, Aplomb 1 scores 44.86, ranking #1 among 4B models on the published board. It achieves the top score among models up to 5.3B on 8 of 38 benchmarks, including MMLA, GPQA Diamond, and BBH.
API Capabilities and Latency
While the open weights are available, the 1M-token speed and decision features are currently exclusive to the EmpirioLabs API. Key API metrics include:
- Latency: Processes a 1M-token document in ~3 seconds (compared to ~111 seconds for a full read in standard mode). Short queries take ~15 ms model time (~111 ms end-to-end).
- Pricing: $0.00 per 1M input tokens, with free output and Zero Data Retention (ZDR) by default.
- Compatibility: Supports OpenAI, Anthropic, and Gemini formats alongside the native Decisions API.
Technical Details
The model expands the context window from 262K to 1M tokens and includes a custom decision head. It runs in bf16 on ~12 GB of GPU memory using the reference script. The weights are licensed under the EmpirioLabs License, free for research, evaluation, personal use, and internal use by companies with under $1M in annual revenue. Training data included public train splits of WinoGrande and ContractNLI.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.